Compositions for target primed reverse transcription

Modified 3' UTR sequences in nucleic acid molecules improve the efficiency and specificity of TPRT by enhancing the interaction with R2 retroelement polypeptides, enabling precise genetic integration into target genomic locations.

WO2026055129A1PCT designated stage Publication Date: 2026-03-12ADDITION THERAPEUTICS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing methods for target primed reverse transcription (TPRT) face inefficiencies and challenges in inserting exogenous nucleic acid sequences into a subject genome using non-long terminal repeat (non-LTR) retrotransposons.

Method used

Development of non-naturally occurring nucleic acid molecules with modified 3' UTR sequences, featuring specific mutations and modifications such as pseudouridine, to enhance the efficiency of TPRT by improving the interaction with R2 retroelement polypeptides for targeted nucleic acid insertion.

Benefits of technology

The modified 3' UTR sequences significantly enhance the efficiency and specificity of TPRT, allowing for precise integration of cargo nucleic acid sequences into target genomic locations, such as rDNA, thereby improving the overall effectiveness of genetic modification processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044479_12032026_PF_FP_ABST
    Figure US2025044479_12032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided herein are nucleic acid constructs and systems for target primed reverse transcription gene insertion. The constructs and systems may comprise a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence comprising a mutation relative to a wild-type R2 3' UTR sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Attomey Docket No. 67098-712601COMPOSITIONS FOR TARGET PRIMED REVERSE TRANSCRIPTIONCROSS REFERENCE

[0001] This application claims the benefit of U.S. Provisional Pat. App. No. 63 / 690,254, filedSeptember 3, 2024, which is incorporated by reference herein in its entirety.SEQUENCE LISTING

[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on August 19, 2025, is named 67098-712_601_SL.xml and is 97,389 bytes in size.BACKGROUND

[0003] Target primed reverse transcription (TPRT) inserts a nucleic acid sequence into a subject genome using non-long terminal repeat (non-LTR) retrotransposons. The insertion of an exogenous nucleic acid sequence into a subject via TPRT faces various difficulties and drawbacks. Additional methods and systems are needed to improve the efficiency and other characteristics of TPRT methods and systems.SUMMARY

[0004] In one aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 1. In some aspects, the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1.

[0005] In some aspects, the 3’ UTR sequence comprises a nucleotide sequence ACCCGGAA. In some aspects, the 3’ UTR sequence comprises a nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97).

[0006] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 1. In some aspects, the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1. In some aspects, the 3’ UTR sequence further comprises a first nucleotide sequence ACCCGGAA and a second nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97).Attomey Docket No. 67098-712601

[0007] In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 12, 18, 23, 24, 50, or 76 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 12, or 18 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 5, 23, 24, 50, and 76 of SEQ ID NO: 1.

[0008] In some aspects, the 3’ UTR sequence comprises at least two mutations. In some aspects, the 3’ UTR sequence comprises two mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence comprises at least three mutations. In some aspects, the 3’ UTR sequence comprises three mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 4 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 12 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 18 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 5 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 23 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 24 of SEQ ID NO: 1, (iv) a fourth mutation located at a position corresponding nucleotide position 50 of SEQ ID NO: 1, and (v) a fifth mutation at a position corresponding nucleotide position 76 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 7 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 27 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 81 of SEQ ID NO: 1, and (iv) a fourth mutation located at a position corresponding nucleotide position 85 of SEQ ID NO: 1.Attomey Docket No. 67098-712601

[0009] In some aspects, the 3’ UTR sequence has at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1.

[0010] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to a stem-loop region of the wild-type R2 3’ UTR sequence. In some aspects, the wild-type R2 3’ UTR sequence is selected from the group consisting of a wildtype Geospiza fortis (GeFo) R2 3’ UTR sequence, a wild-type Zonotrichia albicollis (ZoAl) R2 3’ UTR sequence, a wild-type Taeniopygia guttata (TaGu) R2 3’ UTR sequence, and a wild-type Tinamus guttatus (TiGu) R2 3’ UTR sequence.

[0011] In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.

[0012] In some aspects, the wild-type R2 3’ UTR sequence comprises at least two stem-loop regions. In some aspects, the wild-type R2 3’ UTR sequence comprises at least three stem-loop regions. In some aspects, the wild-type R2 3’ UTR sequence comprises at least four stem-loop regions.

[0013] In some aspects, the 3’ UTR sequence comprises at least two mutations relative to the wild-type R2 3’ UTR sequence. In some aspects, the at least two mutations are both located in regions of the 3’ UTR sequence corresponding to stem-loop regions of the wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence comprises two mutations located in a region of the 3’ UTR sequence corresponding to a same stem-loop region of the wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence comprises two mutations located in regions of the 3’ UTR sequence corresponding to two discrete stem-loop regions of the wild-type R2 3’ UTR sequence.

[0014] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2Attomey Docket No. 67098-7126013’ UTR sequence. In some aspects, the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence.

[0015] In some aspects, the wild-type R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-31 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.

[0016] In some aspects, the 3’ UTR sequence comprises an additional mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence. In some aspects, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence. In some aspects, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

[0017] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence.

[0018] In some aspects, the wild-type R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding toAttomey Docket No. 67098-712601 nucleotide position 50 of SEQ ID NO: 1. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.

[0019] In some aspects, the 3’ UTR sequence comprises an additional mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence. In some aspects, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

[0020] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

[0021] In some aspects, the wild-type R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.

[0022] In some aspects, the 3’ UTR sequence comprises an additional mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

[0023] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to a stem-loop region of the wild-type R2 3’ UTR sequence. InAttomey Docket No. 67098-712601 some aspects, the mutation is located in a region that does not correspond to a position at the bottom of a stem-loop in the wild-type R2 3’ UTR sequence.

[0024] In some aspects, the wild-type R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4. In some aspects, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence.

[0025] In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4-30 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1. In some aspects, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. In some aspects, the mutation is located in a region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence. In some aspects, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 69-91 of SEQ ID NO: 1. In some aspects, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1.

[0026] In some aspects, the 3’ UTR sequence comprises a nucleotide sequence ACCCGGAA. In some aspects, the 3’ UTR sequence comprises a nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97).

[0027] In some aspects, the mutation is a purine to pyrimidine mutation or a pyrimidine to purine mutation. In some aspects, the mutation is a purine to pyrimidine mutation. In some aspects, the mutation is an adenine to uracil mutation. In some aspects, the mutation is a guanine to uracil mutation. In some aspects, the mutation is an adenine to cytosine mutation. In some aspects, the mutation is a guanine to cytosine mutation. In some aspects, the mutation is a pyrimidine to purine mutation. In some aspects, the mutation is a uracil to adenine mutation. In some aspects, the mutation is a uracil to guanine mutation. In some aspects, the mutation is a cytosine to adenine mutation. In some aspects, the mutation is a cytosine to guanine mutation. In some aspects, the mutation is a purine to purine mutation or a pyrimidine to pyrimidine mutation. In some aspects, the mutation is a purine to purine mutation. In some aspects, the mutation is an adenine to guanine mutation. In some aspects, the mutation is a guanine to adenine mutation. InAttomey Docket No. 67098-712601 some aspects, the mutation is a pyrimidine to pyrimidine mutation. In some aspects, the mutation is a cytosine to uracil mutation. In some aspects, the mutation is a uracil to cytosine mutation.

[0028] In some aspects, the 3’ UTR sequence comprises a sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 5-90. In some aspects, the 3’ UTR sequence comprises a sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of any one of SEQ ID NOs: 5- 90. In some aspects, the 3’ UTR sequence comprises a sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 5, 11, 6, or 9. In some aspects, the 3’ UTR sequence comprises any one of SEQ ID NOs: 5-90. In some aspects, the 3’ UTR sequence comprises any one of SEQ ID NOs: 5-12. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 5. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 6. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 7. In some aspects, the 3’ UTR sequence comprises SEQ ID NO 5. In some aspects, the 3’ UTR sequence comprises SEQ ID NO 11. In some aspects, the 3’ UTR sequence comprises SEQ ID NO 9.

[0029] In some aspects, the 3’ UTR sequence comprises no more than 200 nucleotides.

[0030] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 5, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 5. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 5.

[0031] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 11, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1- 35, 47-56, or 65-102 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 11. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 11.Attomey Docket No. 67098-712601

[0032] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 6, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 6. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 6.

[0033] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 9, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1. In some aspects, the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 9. In some aspects, the 3’ UTR sequence comprises SEQ ID NO: 9.

[0034] In some aspects, the non-naturally occurring nucleic acid comprises a modified uridine. In some aspects, the modified uridine is a pseudouridine or a N1 -methylpseudouridine. In some aspects, the modified uridine is a pseudouridine. In some aspects, the modified uridine is a Nl- methylpseudouridine. In some aspects, the non-naturally occurring nucleic acid comprises one or more modified uridines but not unmodified uridines.

[0035] In some aspects, the non-naturally occurring nucleic acid further comprises a cargo sequence. In some aspects, the cargo sequence is at least 500 nucleotides in length. In some aspects, the cargo sequence encodes a protein. In some aspects, the cargo sequence comprises a therapeutic gene or a diagnostic gene.

[0036] In some aspects, the non-naturally occurring nucleic acid further comprises a promoter sequence. In some aspects, the promoter is an RNAP II promoter or a T7 RNA polymerase promoter. In some aspects, the non-naturally occurring nucleic acid further comprises a terminator sequence. In some aspects, the non-naturally occurring nucleic acid further comprises a 5’ UTR sequence on a 5’ side of the cargo sequence. In some aspects, the non-naturally occurring nucleic acid further comprises a stability element configured to protect the non- naturally occurring nucleic acid from degradation. In some aspects, the stability element is selected from K3, K4, eK5, MALATl mm, MALATl hs, MALATl_hs-core, PAN KSHV, dENE os, HSL, U7, K9, KI 6, KEMCV, or KBDV1. In some aspects, the non-naturally occurring nucleic acid further comprises a self-cleaving ribozyme motif. In some aspects, theAttomey Docket No. 67098-712601 self-cleaving ribozyme motif comprises a hepatitis delta virus fold. In some aspects, the non- naturally occurring nucleic acid further comprises a 3’ tail region comprising an adenine-rich region. In some aspects, the adenine-rich region comprises a polyadenosine sequence. In some aspects, the polyadenosine sequence has a length of from 22 to 41 nucleotides. In some aspects, the 3’ tail region comprises one or more non-adenine nucleobases. In some aspects, the adenine- rich region has no more than 10%, 15%, 20%, or 25% non-adenine nucleobases. In some aspects, the non-naturally occurring nucleic acid further comprises a complementarity region that is complementary to a region in a 28 S rRNA gene. In some aspects, the complementarity region comprises a sequence of UAGC or TAGC.

[0037] In another aspect, the present disclosure provides a cell comprising the non-naturally occurring nucleic acid molecule. In some aspects, the cell is a eukaryotic cell. In some aspects, the cell is an epithelial cell.

[0038] In another aspect, the present disclosure provides a system for modifying a target nucleic acid in a cell of a subject, the system comprising: (a) the non-naturally occurring nucleic acid molecule; and (b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide. In some aspects, the R2 retroelement polypeptide binds to the 3’ UTR sequence of the non-naturally occurring nucleic acid molecule.

[0039] In some aspects, the R2 retroelement polypeptide is selected from the group consisting of a Geospiza fortis (GeFo) R2 retroelement polypeptide, a Zonotrichia albicollis (ZoAl) R2 retroelement polypeptide, a Taeniopygia guttata (TaGu) R2 retroelement polypeptide, and a Tinamus guttatus (TiGu) R2 retroelement polypeptide. In some aspects, the R2 retroelement polypeptide and the 3’ UTR sequence are derived from different species. In some aspects, the R2 retroelement polypeptide and the 3’ UTR sequence are derived from a same species. In some aspects, the R2 retroelement polypeptide comprises a Geospiza fortis (GeFo) R2 retroelement polypeptide. In some aspects, the R2 retroelement polypeptide comprises a Taeniopygia guttata (TaGu) R2 retroelement polypeptide. In some aspects, the system further comprises a target nucleic acid molecule. In some aspects, the target nucleic acid molecule comprises DNA. In some aspects, the target nucleic acid molecule comprises genomic DNA. In some aspects, the target nucleic acid molecule comprises a rDNA sequence. In some aspects, the target nucleic acid molecule comprises a 28S rRNA sequence.

[0040] In another aspect, the present disclosure provides a cell comprising the system. In some aspects, the cell is a eukaryotic cell. In some aspects, the cell is an epithelial cell.

[0041] In another aspect, the present disclosure provides a method of modifying a target nucleic acid in a cell of a subject, the method comprising providing the cell with: (a) the non-naturally occurring nucleic acid molecule; and (b) an R2 retroelement polypeptide or a second nucleic acidAttomey Docket No. 67098-712601 encoding the R2 retroelement polypeptide. In some aspects, the R2 retroelement polypeptide inserts a cargo nucleic acid sequence of the non-naturally occurring nucleic acid molecule, or a reverse complement thereof, into the target nucleic acid.

[0042] In some aspects, the R2 retroelement polypeptide is selected from the group consisting of a Geospiza fortis (GeFo) R2 retroelement polypeptide, a Zonotrichia albicollis (ZoAl) R2 retroelement polypeptide, a Taeniopygia guttata (TaGu) R2 retroelement polypeptide, and a Tinamus guttatus (TiGu) R2 retroelement polypeptide. In some aspects, the R2 retroelement polypeptide binds to the 3’ UTR sequence of the non-naturally occurring nucleic acid molecule. In some aspects, the target nucleic acid is in a cell. In some aspects, the cell is a eukaryotic cell. In some aspects, the cell is an epithelial cell. In some aspects, the R2 retroelement polypeptide inserts the cargo nucleic acid sequence or the reverse complement thereof into a rDNA sequence of the target nucleic acid. In some aspects, the R2 retroelement polypeptide inserts the cargo nucleic acid sequence or the reverse complement thereof into a 28S rDNA sequence of the target nucleic acid molecule.INCORPORATION BY REFERENCE

[0043] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Various features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG.” herein), of which:

[0045] FIG. 1A depicts a first structural conformation of a GeFo 3’ UTR sequence (SEQ ID NO 1). FIG. IB depicts a second structural conformation of a GeFo 3’ UTR sequence (SEQ ID NO 1). FIG. 1C depicts a first structural conformation of a TaGu 3’ UTR sequence (SEQ ID NO 2). FIG. ID depicts a second structural conformation of a TaGu 3’ UTR sequence (SEQ ID NO 2).

[0046] FIG. 2 shows a graph characterizing 3’ UTR sequences based on structural similarity of their respective minimal free energy (MFE) folds to that of the wild-type GeFo 3’ UTR and the ensemble diversity of RNA folding.Attomey Docket No. 67098-712601

[0047] FIG. 3 shows a graph of selected 3’ UTR sequences for screening for TPRT gene insertion and categorization based on structural similarity of their respective MFE folds to that of the wild-type GeFo 3’ UTR and the ensemble diversity of RNA folding.

[0048] FIG. 4A shows a design of an example construct described herein for TPRT gene insertion. FIG, 4B shows a design of an example construct with specific elements in the 5’ and 3’ regions for a TPRT gene insertion assay.

[0049] FIG. 5 shows data indicating performance of constructs tested in a TPRT gene insertion screen after 24 hours in BNL cells, Huh7 cells, and RPE1 cells.

[0050] FIG. 6 shows data indicating performance of constructs tested in a TPRT gene insertion screen after 48 hours in BNL cells, Huh7 cells, and RPE1 cells.

[0051] FIG. 7A shows data indicating performance of top-performing constructs relative to a control N1 -methylpseudouridine construct comprising the GeFo sequence (SEQ ID NO: 1), tested in a TPRT gene insertion screen after 48 hours in BNL cells, Huh7 cells, and RPE1 cells.

[0052] FIG. 7B shows data indicating performance of top-performing constructs relative to a cognate construct (with the same uridine modification or lack thereof) comprising the GeFo sequence (SEQ ID NO: 1), tested in a TPRT gene insertion screen after 48 hours in BNL cells, Huh7 cells, and RPE1 cells.

[0053] FIG. 8A shows data indicating correlation between construct performance and the base position of a mutation in the variant construct relative to the wild-type GeFo 3’ UTR sequence (SEQ ID NO: 1) and the “do-not-touch” regions within the 3’ UTR sequence that are important for performance. FIG. 8B depicts the structure of the wild-type GeFo 3’ UTR (SEQ ID NO: 1) and the “do-not-touch” regions within the 3’ UTR sequence that are important for performance.

[0054] FIG. 9A shows data from the top performing 3 ’UTR variant sequences indicating correlation between construct performance and the base position of a mutation in the variant construct relative to the wild-type GeFo 3’ UTR sequence (SEQ ID NO: 1) and the “do-not- touch” regions within the 3’ UTR sequence that are important for performance. FIG. 9B shows data from the bottom performing 3 ’UTR variant sequences indicating correlation between construct performance and the base position of a mutation in the variant construct relative to the wild-type GeFo 3’ UTR sequence (SEQ ID NO: 1) and the “do-not-touch” regions within the 3’ UTR sequence that are important for performance.

[0055] FIG. 10 shows performance data of constructs tested in a TPRT gene insertion screen, grouped by: Nl-methyl pseudouridine-containing construct, pseudouridine-containing construct, or uridine-containing construct.Attomey Docket No. 67098-712601

[0056] FIG. 11A shows performance data of constructs tested in a TPRT gene insertion screen, expressed as geometric mean across 3 cell lines tested. The x-axis has number labels ordered from 1-268 for corresponding constructs shown in FIG. 11C. FIG. 11B shows performance data of constructs tested in a TPR gene insertion screen for each of the 3 cell lines. The x-axis has number labels ordered from 1-268 for corresponding constructs shown in FIG. 11C. FIG. 11C depicts the name of the 3’ UTR variant sequence in the construct, the type of construct (Nl-methylpseudouridine-containing construct labeled as “ml -psi”, pseudouridine- containing construct labeled as “psi”, and uridine-containing construct labeled as “U”), and the corresponding number label in the x-axis of both FIG. 11A and FIG. 11B.

[0057] FIG. 12A shows performance data of constructs tested in a TPRT gene insertion screen, grouped by: Nl-methyl pseudouridine-containing construct, pseudouridine-containing construct, or uridine-containing construct. The x-axis has number labels ordered from 1-90 for corresponding constructs shown in FIG. 12B.

[0058] FIG. 13A shows performance data of Nl-methyl pseudouridine-containing constructs tested in a TPRT gene insertion screen for each of the three cell lines. The x-axis has number labels ordered from 1-90 for corresponding constructs shown in FIG. 13B. FIG. 13B depicts the name of the 3’ UTR variant sequence in the construct and the corresponding number label in the x-axis of FIG. 13A

[0059] FIG. 14A shows performance data of pseudouridine-containing constructs tested in a TPRT gene insertion screen for each of the three cell lines. The x-axis has number labels ordered from 1-90 for corresponding constructs shown in FIG. 14B. FIG. 14B depicts the name of the 3’ UTR variant sequence in the construct and the corresponding number label in the x-axis of FIG.14A

[0060] FIG. 15A shows performance data of uridine-containing constructs tested in a TPRT gene insertion screen for each of the three cell lines. The x-axis has number labels ordered from 1-90 for corresponding constructs shown in FIG. 15B. FIG. 15B depicts the name of the 3’ UTR variant sequence in the construct and the corresponding number label in the x-axis of FIG. 15A.

[0061] FIG. 16 shows performance of constructs in a TPRT gene insertion screen, grouped by categorization based on their MFE folds.DETAILED DESCRIPTION

[0062] While various embodiments of the disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in theAttomey Docket No. 67098-712601 art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed.Definitions

[0063] Derived from: As used herein, the term “derived from” refers to a nucleic acid or protein sequence that is isolated from or obtained from a specific source, such as a non-long terminal repeat (non-LTR) retrotransposon. The term includes native sequences isolated from or obtained from a specific source. The term also includes man-made variants of sequences from the original source that have the same or similar functional properties, e.g., the variant can comprise a nucleic or amino acid sequence that has been modified from the original source to have improved functional properties compared to the original source molecule.

[0064] If an RNA sequence is recited using deoxyribonucleotides, any thymidines (“T”s) can be replaced with uridines (“U”s) or uridine analogs (e.g., N1 -methylpseudouridine or pseudouridine) to convert the DNA sequence to an RNA sequence.

[0065] Encode: As used herein, the term “encode” refers broadly to any process whereby the information in a polymeric macromolecule is used to direct the production of a second molecule that is different from the first. The second molecule may have a chemical structure that is different from the chemical nature of the first molecule.

[0066] Gene Insertion Construct: As used herein, the term “Gene Insertion Construct”, or GIC, refers to an RNA construct which comprises the cargo nucleic acid sequence or a template sequence for a reverse transcriptase (RT) protein.

[0067] Gene-Insertion System: As used herein, the term “Gene-Insertion System” or “GIS,” is a system of components (modules) which may be used to insert a genetic sequence (transgene) into a location of a subject genome via reverse transcription, including TPRT.

[0068] Liposome: As used herein, “liposome” generally refers to a vesicle composed of lipids (e.g., amphiphilic lipids) arranged in one or more spherical bilayers or bilayers.

[0069] Mutation: As used herein in the context of one or more mutations relative to reference nucleic acid sequence, the term “mutation” refers to a change in nucleotide sequence relative to the reference sequence.

[0070] Sequence Identity: The “percent sequence identity” between a reference nucleic acid sequence and a query nucleic acid sequence (i.e., the nucleic acid sequence being analyzed to determine whether it is within a particular percent sequence identity with the reference nucleic acid sequence) is determined by optimally aligning the sequences using the Needleman-Wunsch alignment algorithm (with match / mismatch scores of 2,-3, a gap existence penalty of 5, and a gap extension penalty of 2) and comparing the aligned nucleic acids. The number of exact matchesAttomey Docket No. 67098-712601 divided by the total number of nucleotides in the alignment (which corresponds with the number of nucleotides in the reference sequence plus any gaps in the reference sequence when aligned with the query sequence) is determined and expressed as a percentage. This is the percent sequence identity between the query nucleic acid sequence and the reference nucleic acid sequence (i.e., percent sequence identity = (# of exact matches) / (total # of nucleotides in the alignment)* 100). An alignment using the Needleman-Wunsch alignment algorithm (with match / mismatch scores of 2,-3, a gap existence penalty of 5, and a gap extension penalty of 2) can be generated using the “Global Align” BLAST program available at https: / / blast.ncbi.nlm.nih.gov / Blast.cgi. A nucleotide “position” of a sequence refers to the position that corresponds with the specified nucleotide position when the query sequence is globally aligned to the reference sequence as specified above using the Needleman-Wunsch algorithm.

[0071] Target Cell: As used herein, the phrase “targeted cells” refers to any one or more cells of interest. The cells may be found in vitro, in vivo, in situ or in the tissue or organ of an organism. The organism may be an animal, preferably a mammal, more preferably a human and most preferably a patient.

[0072] Target Primed Reverse Transcription: As used herein, the term “target primed reverse transcription” refers to any process where a reverse transcriptase uses a genome-embedded nicked DNA 3’ end at the target site as the primer to initiate cDNA synthesis.Overview

[0073] Recognized herein is a need for improved methods for editing nucleic acid molecules, for example, for gene therapy. Many current gene editing technologies are hindered by their inability to make large gene insertions or inserting a large cargo nucleic acid sequence into a target nucleic acid molecule. Furthermore, current gene editing technologies encounter challenges of unintended genetic alterations that can be deleterious to the target cell genome or cellular physiology. Another challenge for gene therapy is making gene insertions in nondividing cells, such as neurons, since many gene editing technologies only work well during DNA replication in dividing cells. Finally, many current gene editing technologies are hindered by low gene editing efficiencies.

[0074] The present disclosure provides compositions and methods that can enable efficient insertion of a cargo nucleic acid sequence, such as a transgene encoding a therapeutic protein or a non-protein regulatory element, into a target nucleic acid molecule. The compositions and methods described herein can provide the advantages of enabling large cargo nucleic acid insertions into a target nucleic acid molecule with high efficiencies and enabling insertion intoAttomey Docket No. 67098-712601“safe harbor” sites that do not cause deleterious or undesirable alterations to the target cell. Furthermore, compositions and methods described herein can enable gene insertions in postmitotic cells, overcoming challenges associated with other gene editing technologies. The compositions and methods described herein can provide benefits for many different therapeutic applications, for example, to treat genetic diseases or other disorders associated with gene mutations or dysfunctional proteins.

[0075] In some aspects, the compositions and methods described herein make use of target primed reverse transcription (TPRT) using sequences derived from a retroelement, such as a nonlong terminal repeat (non-LTR) retrotransposon. A system for carrying TPRT can comprise a retroelement polypeptide, such as a retrotransposase, and a nucleic acid comprising a cargo nucleic acid sequence, a 5’ UTR sequence located 5’ of the cargo nucleic acid sequence, and a 3’ UTR sequence located 3’ of the cargo nucleic acid sequence. The 5’ UTR and / or the 3’ UTR sequence of the nucleic acid can interact with the retroelement polypeptide. In some cases, the 3’ UTR sequence of the nucleic acid enables the retroelement polypeptide to carry out TPRT of the cargo nucleic acid sequence and insert the cargo nucleic sequence into a target nucleic acid molecule.

[0076] In some aspects, the retroelement polypeptide comprises an endonuclease domain and a reverse transcriptase domain. Without being bound by theory, it is currently thought that the endonuclease domain of the retroelement polypeptide cleaves the bottom strand of the target nucleic acid molecule, which provides a 3’ hydroxyl end that serves as a primer for reverse transcription of the cargo nucleic acid sequence by the reverse transcriptase domain of the retroelement polypeptide. Following first strand synthesis to produce cDNA, the endonuclease domain or a host endonuclease cleaves the opposite (e.g., the top strand) of the genomic DNA. The nick in the top strand produces another 3’ hydroxyl end that serves as a primer for second strand cDNA synthesis. It is currently unknown if second strand DNA synthesis is performed by the retroelement polypeptide or by a cellular polymerase. The nick is then repaired, resulting in integration of the double-stranded cDNA into the target site in the genomic DNA.

[0077] In some aspects, the current disclosure provides a nucleic acid comprising a 3’ UTR sequence for inserting a cargo nucleic acid sequence into a target nucleic acid molecule. The 3’ UTR sequence can interact with a retroelement polypeptide (e.g., an R2 retroelement polypeptide). In some cases, the 3’ UTR sequence of the nucleic acid enables the retroelement polypeptide to carry out TPRT or insertion of the cargo nucleic acid sequence or a reverse complement thereof into a target nucleic acid molecule. Altering the 3’ UTR sequence can affect the efficiency of TPRT or insertion of the cargo nucleic acid sequence or reverse complement thereof into the target nucleic acid molecule. Certain regions of the 3’ UTR sequence can beAttomey Docket No. 67098-712601 important to insertion activity and mutations within those regions can reduce or inhibit insertion by a retroelement polypeptide. In some embodiments, the 3’ UTR sequence of the nucleic acid comprises a mutation relative to a wild-type R2 3’ UTR sequence. While certain mutations (such as those within certain important or invariant regions) can reduce or inhibit insertion, other mutations can enhance insertion by a retroelement polypeptide. The present disclosure provides 3’ UTR sequences that can enable high efficiency insertion of a cargo nucleic acid sequence to a target nucleic acid molecule.I. Compositions

[0078] In some aspects, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a cargo nucleic acid sequence for insertion into a target nucleic acid molecule. The non-naturally occurring nucleic acid molecule can further comprise 5’ region that is located 5’ of the cargo nucleic acid sequence and a 3’ region that is located 3’ of the cargo nucleic acid sequence. Examples of such a nucleic acid construct or gene insertion construct are shown in FIGs. 4A and B. FIG. 4A depicts a nucleic acid construct comprising a 5’ region that is located 5’ of the cargo nucleic acid sequence and a 3’ region that is located 3’ of the cargo nucleic acid sequence. In FIG. 4A, the 5’ region comprises a 5’ module comprising the 5’ UTR configured to interact with an R2 polypeptide, and the 3’ region comprises a 3’ module comprising the 3’ UTR configured to interact with an R2 polypeptide. In this example, the 5’ region further comprises a promoter. In this example, the 3’ region further comprise a tail. In FIG. 4B, the 5’ region comprises a 5’ module or 5’ UTR sequence comprising HDV-gRZ (SEQ ID NO: 91), a promoter, and the Neo3-Kozak sequence. In this example, the 3’ region comprises a minpAS sequence (SEQ ID NO: 95), a 3’ UTR sequence, and a tail region. The 5’ region can further comprise a pseudo-cap. In some embodiments, the 3’ region comprises a 3’ UTR sequence derived from a retroelement (e.g., an R2 retroelement). In some embodiments, the 3’ UTR sequence has at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least97%, at least 98%, or at least 99% sequence identity to a wild-type 3’ UTR sequence. In some embodiments, the 3’ region comprises a 3’ UTR sequence comprising a mutation relative to a wild-type R2 3’ UTR sequence. The 3’ UTR sequence can interact with a R2 retroelement polypeptide. In some embodiments, the 3’ UTR sequence interacts with or binds to a cognate R2 retroelement polypeptide derived from the same species. In some embodiments, the 3’ UTR sequence interacts with or binds to a non-cognate R2 retroelement polypeptide derived from a different species as the 3’ UTR sequence.Attomey Docket No. 67098-7126013 ’ UTR sequence

[0079] In some aspects, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence. The 3’ UTR sequence can be in a 3’ region located 3’ of a cargo nucleic acid sequence for insertion into a target nucleic acid molecule, as described elsewhere herein. The 3’ UTR sequence can have least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a wildtype R2 3’ UTR sequence. The wild-type R2 3’ UTR sequence can be a wild-type Geospiza fortis (GeFo) R2 3’ UTR sequence, a wild-type Zonotrichia albicollis (ZoAl) R2 3’ UTR sequence, a wild-type Taeniopygia guttata (TaGu) R2 3’ UTR sequence, or a wild-type Tinamus guttatus (TiGu) R2 3’ UTR sequence. The wild-type R2 3’ UTR sequence can be selected from SEQ ID NOs: 1-4. In some aspects, the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence.

[0080] The 3’ UTR sequence can interact with or bind to an R2 retroelement polypeptide, for example, a GeFo R2 polypeptide, a ZoAl R2 polypeptide, a TaGu R2 polypeptide, or a TiGu R2 polypeptide. In some embodiments, the 3’ UTR sequence interacts with or bind to a cognate R2 retroelement polypeptide derived from the same species. In some embodiments, the 3’ UTR sequence interacts with or binds to a non-cognate R2 retroelement polypeptide derived from a different species as the 3’ UTR sequence.

[0081] In an aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 5, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to SEQ ID NO: 5. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 5. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at leastAttomey Docket No. 67098-71260198%, at least 99% sequence, or 100% sequence identity to a continuous region of at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 5.

[0082] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 11, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1- 35, 47-56, or 65-102 of SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to SEQ ID NO: 11. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 11. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to a continuous region of at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 11.

[0083] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 6, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1- 35, 47-56, or 65-102 of SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to SEQ ID NO: 6. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 6. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, atAttomey Docket No. 67098-712601 least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to a continuous region of at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 6.

[0084] In another aspect, the present disclosure provides a non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 9, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1- 35, 47-56, or 65-102 of SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to SEQ ID NO: 9. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 9. In some embodiments, the 3’ UTR sequence has at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence, or 100% sequence identity to a continuous region of at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, or at least 100 nucleotides of SEQ ID NO: 9.Invariant sequence

[0085] In some embodiments, the 3’ UTR sequence comprises an invariant sequence. In some cases, the invariant sequence is important for binding to or interacting with an R2 retroelement polypeptide. In some cases, the invariant sequence is important for insertion of the 3’ UTR sequence or a cargo nucleic acid sequence of the non-naturally occurring nucleic acid molecule, or a reverse complement thereof, in a target nucleic acid molecule.

[0086] In some embodiments, the invariant sequence comprises a nucleotide sequence ACCCGGAA. In some embodiments, the invariant sequence comprises a nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97). In some embodiments, the invariant sequence comprises ACGGG. The 3’ UTR sequence can comprise a first nucleotide sequence ACCCGGAA and a second nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97). The invariant sequence can be in a stem-loop region of the 3’ UTR sequence. In some embodiments, the invariant sequence comprises a loop region of a stem-loop region. For example, as shown in FIG. 8B, the secondAttomey Docket No. 67098-712601 stem loop region in order from 5’ to 3’ comprises the nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97). The loop region of the second stem loop region comprises nucleotide sequence ACGGG. In FIG. 8B, the third stem loop region in order from 5’ to 3’ comprises the nucleotide sequence ACCCGGAA, which is in the loop region of the third stem loop region. The invariant sequence can be part of a wild-type R2 3’ UTR sequence.

[0087] A 3’ UTR sequence comprising an invariant sequence described herein can have increased binding to the R2 retroelement polypeptide as compared to an otherwise identical 3’ UTR sequence comprising a mutation in the invariant sequence. A 3’ UTR sequence comprising an invariant sequence described herein can have at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 2 fold, at least 3 fold, at least 5 fold, at least 10 fold, at least 15 fold, at least 20 fold, at least 30 fold, at least 50 fold, or at least 100 fold higher binding affinity compared to an otherwise identical 3’ UTR sequence comprising a mutation in the invariant sequence. A 3’ UTR sequence comprising an invariant sequence described herein can result in increased insertion efficiency of a cargo nucleic acid sequence in the non-naturally occurring nucleic acid molecule as compared to an otherwise identical nucleic acid construct comprising a mutation in the invariant sequence. A 3’ UTR sequence comprising an invariant sequence described herein can improve insertion efficiency of a cargo nucleic acid sequence by least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 2 fold, at least 3 fold, at least 5 fold, at least 10 fold, at least 15 fold, at least 20 fold, at least 30 fold, at least 50 fold, or at least 100 fold compared to an otherwise identical 3’ UTR sequence comprising a mutation in the invariant sequence.Mutations in 3 ’ UTR sequence

[0088] In some aspects, the 3’ UTR sequence of the non-naturally occurring nucleic acid molecule comprises a mutation relative to a wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence of the non-naturally occurring nucleic acid molecule comprises one or more invariant sequences described herein and a mutation relative to a wild-type R2 3’ UTR sequence. In some aspects, the 3’ UTR sequence has at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1, and comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1.

[0089] The 3’ UTR sequence of the non-naturally occurring nucleic acid molecule further comprises one or more invariant sequences described herein. In some embodiments, the 3’ UTR sequence further comprises a nucleotide sequence ACCCGGAA. In some embodiments, the 3’Attomey Docket No. 67098-712601UTR sequence further comprise a nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97). In some embodiments, the 3’ UTR sequence further comprise a nucleotide sequence ACGGG. In some embodiments, the 3’ UTR sequence further comprises a first nucleotide sequence ACCCGGAA and a second nucleotide sequence ACGGG. In some embodiments, the 3’ UTR sequence further comprises a first nucleotide sequence ACCCGGAA and a second nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97).

[0090] In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1. In some embodiments, the one or more invariant sequences correspond to nucleotides positions 36-46, and the mutation is located outside the region corresponding to the invariant sequence. In some embodiments, the one or more invariant sequences correspond to nucleotides positions 39-43, and the mutation is located outside the region corresponding to the invariant sequence. In some embodiments, the one or more invariant sequences correspond to nucleotides positions 57-64, and the mutation is located outside the region corresponding to the invariant sequence. In some embodiments, the one or more invariant sequences correspond to nucleotides positions 57-64, and the mutation is located outside the region corresponding to the invariant sequence. In some embodiments, the one or more invariant sequences correspond to nucleotides positions 39-43 and 57-64, and the mutation is located outside the region corresponding to the invariant sequence. In some embodiments, the one or more invariant sequences correspond to nucleotides positions 36-46 and 57-64, and the mutation is located outside the region corresponding to the invariant sequence.

[0091] In some embodiments, the mutation is further located at a position of the 3’ UTR sequence corresponding to a stem-loop region of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 12, 18, 23, 24, 50, or 76 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 12, or 18 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 5, 23, 24, 50, and 76 of SEQ ID NO: 1.Attomey Docket No. 67098-712601

[0092] In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-31 of SEQ ID NO: 1. The mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1. The mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1.

[0093] In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 4 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 5 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 7 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 12 of SEQ ID NO: 1. In some embodiments, the mutation is e located at a position of the 3’ UTR sequence corresponding to nucleotide position 18 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 23 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 24 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 27 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 50 of SEQ ID NO: 1. The mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 76 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 81 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 85 of SEQ ID NO: 1.

[0094] In some embodiments, the 3’ UTR sequence comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 mutationsAttomey Docket No. 67098-712601 relative to the wild-type R2 3’ UTR sequence. In some embodiments, each mutation of the at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 mutations in the 3’ UTR sequence is located is located at position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1- 35, 47-56, or 65-102 of SEQ ID NO: 1. In some embodiments, each mutation in the 3’ UTR sequence is located at position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1.

[0095] In some embodiments, the 3’ UTR sequence comprises at least two mutations. The 3’ UTR sequence can comprise two mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. The 3’ UTR sequence can comprise two mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 3-31 of SEQ ID NO: 1. The 3’ UTR sequence can comprise two mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1. The 3’ UTR sequence can comprise two mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 68-92 of SEQ ID NO: 1.

[0096] In some embodiments, the 3’ UTR sequence comprises at least three mutations. The 3’ UTR sequence can comprise three mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1. The 3’ UTR sequence can comprise three mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 3-31 of SEQ ID NO: 1. The 3’ UTR sequence can comprise three mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1. The 3’ UTR sequence can comprise three mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 68-92 of SEQ ID NO: 1.

[0097] In some embodiments, the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 4 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 12 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 18 of SEQ ID NO: 1.

[0098] In some embodiments, the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 5 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 23 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 24 of SEQ ID NO: 1, (iv) a fourthAttomey Docket No. 67098-712601 mutation located at a position corresponding nucleotide position 50 of SEQ ID NO: 1, and (v) a fifth mutation at a position corresponding nucleotide position 76 of SEQ ID NO: 1.

[0099] In some embodiments, the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 7 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 27 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 81 of SEQ ID NO: 1, and (iv) a fourth mutation located at a position corresponding nucleotide position 85 of SEQ ID NO: 1.

[0100] In certain aspects, the 3’ UTR sequence has at least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a wildtype R2 3’ UTR sequence, and comprises a mutation relative to the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to a stem-loop region of the wild-type R2 3’ UTR sequence. The wild-type type R2 3’ UTR sequence can comprise at least two, at least three, at least four, at least five, or at least six stem-loop regions. The wild-type R2 3’ UTR sequence can be selected from the group consisting of a wildtype Geospiza fortis (GeFo) R2 3’ UTR sequence, a wild-type Zonotrichia albicollis (ZoAl) R2 3’ UTR sequence, a wild-type Taeniopygia guttata (TaGu) R2 3’ UTR sequence, or a wild-type Tinamus guttatus (TiGu) R2 3’ UTR sequence. In some embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. In other embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2. In further embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3. In further embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.

[0101] FIGs. 1A-1D show examples of structural conformations of R2 3’ UTR sequences. FIGs. 1A and IB depict two different structural conformation of a wild-type Geospiza fortis (GeFo) R2 3’ UTR sequence (SEQ ID NO: 1). In the conformation depicted in FIG. 1A, the RNA structure comprises four stem-loop regions 101, 102, 103, and 104. In the conformation depicted in FIG. 1A, the first stem-loop region 101 comprises nucleotide positions 3-31 of SEQ ID NO: 1, the second stem-loop region 102 comprises nucleotide positions 32-50 of SEQ ID NO: 1, the third stem-loop region 103 comprises nucleotide positions 53-67 of SEQ ID NO: 1, and the fourth stem-loop region 104 comprises nucleotide positions 68-92 of SEQ ID NO: 1.

[0102] FIGs. 1C and ID depict two different conformations of a wild-type Taeniopygia guttata (TaGu) R2 3’ UTR sequence (SEQ ID NO: 2). In the conformation depicted in FIG. 1C, the RNA structure comprises four stem-loop regions 111, 112, 113, and 114. In the conformation depicted in FIG. 1C, the first stem-loop region 111 comprises nucleotide positions 3-31 of SEQAttomey Docket No. 67098-712601ID NO: 2, the second stem-loop region 112 comprises nucleotide positions 32-50 of SEQ ID NO: 2, the third stem-loop region 113 comprises nucleotide positions 53-67 of SEQ ID NO: 1, and the fourth stem-loop region 114 comprises nucleotide positions 68-92 of SEQ ID NO: 1.

[0103] In some embodiments, the 3’ UTR sequence comprises at least two mutations relative to the wild-type R2 3’ UTR sequence. The at least two mutations can both be located in regions of the 3’ UTR sequence corresponding to stem-loop regions of the wild-type R2 3’ UTR sequence. In some embodiments, the 3’ UTR sequence comprises two mutations located in a region of the 3’ UTR sequence corresponding to a same stem-loop region of the wild-type R2 3’ UTR sequence. In some embodiments, the 3’ UTR sequence comprises two mutations located in regions of the 3’ UTR sequence corresponding to two discrete stem-loop regions of the wild-type R2 3’ UTR sequence.

[0104] In some embodiments, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. The 3’ UTR can comprise a mutation relative to the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence (e.g., 101 of FIG. 1 A or 111 of FIG. 1C). The wild-type R2 3’ UTR sequence can be or comprise any one of SEQ ID NOs: 1-4.

[0105] In some embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. The mutation can located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-31 of SEQ ID NO: 1. In some embodiments, mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1.

[0106] The 3’ UTR sequence can comprise an additional mutation relative to the wild-type R2 3’ UTR sequence. In some embodiments, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence. In other embodiments, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence. The additional mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1. In further embodiments, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence, for example, at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1.Attomey Docket No. 67098-712601

[0107] In some embodiments, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. The 3’ UTR can comprise a mutation relative to the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence (e.g., 102 of FIG. 1 A or 112 of FIG. 1C). The wild-type R2 3’ UTR sequence can be or comprise any one of SEQ ID NOs: 1- 4.

[0108] In some embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. The mutation can located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1. In some embodiments, mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 50 of SEQ ID NO: 1.

[0109] The 3’ UTR sequence can comprise an additional mutation relative to the wild-type R2 3’ UTR sequence. In some embodiments, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence. The additional mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 32-35 or 47- 50 of SEQ ID NO: 1. In other embodiments, the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence, for example, at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1.

[0110] In some embodiments, the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. The 3’ UTR can comprise a mutation relative to the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence (e.g., 104 of FIG. 1A or 114 of FIG. 1C). The wild-type R2 3’ UTR sequence can be or comprise any one of SEQ ID NOs: 1-4. [OHl] In some embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1. The mutation can located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1.

[0112] The 3’ UTR sequence can comprise an additional mutation relative to the wild-type R2 3’ UTR sequence. In some embodiments, the additional mutation is located in the region ofAttorney Docket No. 67098-712601 the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

[0113] In some embodiments, the 3’ UTR sequence further comprises a nucleotide sequence ACCCGGAA. In some embodiments, the 3’ UTR sequence further comprise a nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97). In some embodiments, the 3’ UTR sequence further comprises a first nucleotide sequence ACCCGGAA and a second nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97). In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1.

[0114] In some embodiments, the 3’ UTR sequence comprises a mutation located in a region of the 3’ UTR sequence corresponding to a stem -loop region of the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region that does not correspond to a position at the bottom of a stem-loop in the wild-type R2 3’ UTR sequence. The wild-type R2 3’ UTR sequence can be or comprise any one of SEQ ID NOs: 1-4.

[0115] For example, in some embodiments, the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4-30, 33-35, 47-49, or 69-91 of SEQ ID NO: 1. In some embodiments, the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 76, 81, or 85 of SEQ ID NO: 1.

[0116] The wild-type R2 3’ UTR sequence can comprise, in order from 5’ to 3’, a first stemloop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region. In some embodiments, the mutation is located in a region of the 3’ UTR sequence corresponding to the first stem-loop region but does not correspond to a position at the bottom of the first stemloop. For example, as shown in FIG. 1A, the mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4- 30 of SEQ ID NO: 1. In some embodiments, the mutation is located a position corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1.

[0117] In some embodiments, the mutation is located in a region of the 3’ UTR sequence corresponding to the second stem-loop region but does not correspond to a position at the bottom of the second stem-loop. For example, as shown in FIG. 1A, the mutation can be located at aAttomey Docket No. 67098-712601 position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 33-49 of SEQ ID NO: 1. In some embodiments, the mutation is located a position corresponding to a nucleotide position selected from any one of nucleotide positions 33- 35 or 47-49 of SEQ ID NO: 1.

[0118] In some embodiments, the mutation is located in a region of the 3’ UTR sequence corresponding to the third stem-loop region but does not correspond to a position at the bottom of the third stem-loop. For example, as shown in FIG. 1A, the mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 54-66 of SEQ ID NO: 1. In some embodiments, the mutation is located a position corresponding to a nucleotide position selected from any one of nucleotide positions 54- 56, 65, or 66 of SEQ ID NO: 1. The third stem-loop region may be a “low-propensity” structure. The third stem-loop region may be about equally likely to be unstructured as be in the stem-loop structure.

[0119] In some embodiments, the mutation is located in a region of the 3’ UTR sequence corresponding to the fourth stem-loop region but does not correspond to a position at the bottom of the fourth stem-loop. For example, as shown in FIG. 1A, the mutation can be located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1. In some embodiments, the mutation is located a position corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1.

[0120] In some embodiments, the 3’ end region of the 3’ UTR sequence comprises a complementarity region that is complementary to a target site in the genome. In some embodiments, the complementarity region is complementary to a region of a rRNA gene. In some embodiments, the rRNA gene is a 28S rRNA gene. In some embodiments, the complementarity region has a sequence of UAGC or TAGC. In some embodiments, the complementarity region has as length from 4 to 10 nucleotides, from 4 to 15 nucleotides, from 4 to 20 nucleotides, from 4 to 25 nucleotides, or from 4 to 30 nucleotides. In some embodiments, the complementarity region has a length up to 4 nucleotides, up to 10 nucleotides, up to 15 nucleotides, up to 20 nucleotides, up to 25 nucleotides, or up to 30 nucleotides. In some embodiments, the complementarity region has a length of at least 3 nucleotides, at least 4 nucleotides, at least 6 nucleotides, at least 8 nucleotides, at least 10 nucleotides, at least 15 nucleotides, at least 20 nucleotides, or at least 25 nucleotides.

[0121] Any of the mutations in the 3’ UTR sequence described herein can be a purine to pyrimidine mutation, a pyrimidine to purine mutation, a purine to purine mutation, or a pyrimidine to pyrimidine mutation. In some embodiments, a mutation described herein is aAttomey Docket No. 67098-712601 purine to pyrimidine mutation. For example, the mutation can be an adenine to uracil mutation, a guanine to uracil mutation, an adenine to cytosine mutation, or a guanine to cytosine mutation. In some embodiments, a mutation described herein is a pyrimidine to purine mutation. For example, the mutation can be a uracil to adenine mutation, a uracil to guanine mutation, a cytosine to adenine mutation, or a cytosine to guanine mutation. In some embodiments, a mutation described herein is a purine to purine mutation. For example, the mutation can be an adenine to guanine mutation or a guanine to adenine mutation. In some embodiments, a mutation described herein is a pyrimidine to pyrimidine mutation. For example, the mutation can be a cytosine to uracil mutation or a uracil to cytosine mutation. In some embodiments, the 3’ UTR sequences any one of the mutations listed in the fourth column of Table 2 relative to SEQ ID NO: 1. The 3’ UTR sequence can comprise any one of the mutations selected from the group consisting of: G4A, A12C, U18G, A5U, G23A, A24G, U50C, A76U, U7A, U27C, U81G, A85G, C47U, U51C, C53A, A65G, C6U, U10G, U26A, U55A, A12C, A15U, U16C, A76G, A97C, A24G, G87A, U7A, A24G, and A79G relative to SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence comprises any one of the mutations selected from the group consisting of: G4A, A12C, and U18G relative to SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence comprises mutations G4A, A12C, and U18G relative to SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence comprises any one of the mutations selected from the group consisting of: A5U, G23 A, A24G, U50C, and A76U relative to SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence comprises mutations A5U, G23A, A24G, U50C, and A76U relative to SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence comprises any one of the mutations selected from the group consisting of: U7A, U27C, U81G, and A85G relative to SEQ ID NO: 1. In some embodiments, the 3’ UTR sequence comprises mutations U7A, U27C, U81G, and A85G relative to SEQ ID NO: 1.

[0122] In some embodiments, the 3’ UTR sequence comprises a sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 5- 90. The 3’ UTR sequence can comprise any one of SEQ ID NOs: 5-90. In some embodiments, the 3’ UTR sequence comprises any one of SEQ ID NOs: 5-12. In some aspects, the 3’ UTR sequence comprises a sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to any one of SEQ ID NOs: 5, 11, 6, or 9. In some embodiments, the 3’ UTR sequence comprises SEQ ID NO: 5. In some embodiments, the 3’ UTR sequence comprises SEQ ID NO: 6. In some embodiments, the 3’ UTR sequence comprises SEQ ID NO: 7. In some embodiments, the 3’ UTR sequence comprises SEQ ID NO 5. In some aspects, the 3’ UTRAttorney Docket No. 67098-712601 sequence comprises SEQ ID NO 11. In some embodiments, the 3’ UTR sequence comprises SEQID NO 9.

[0123] In some embodiments, the 3’ UTR sequence comprises a structure having a structural similarity of from 0.9 to 1.0, from 0.95 to 1.0, from 0.7 to 0.9, from 0.82 to 0.88, from 0.7 to 0.8, from 0.5 to 0.7, from 0.6 to 0.7, from 0.65 to 0.7, from 0.55 to 0.64, from 0.5 to 0.65, or from 0.4 to 0.55 to the MFE fold of SEQ ID NO: 1, as calculated by the restricted Damerau-Levenshtein distance between the predicted MFE fold of the 3 'UTR sequence to the MFE fold of SEQ ID NO: 1.

[0124] In some embodiments, the 3’ UTR sequence comprises a structure having ensemble diversity of from 0 to 6, from 0 to 8, from 6 to 10, from 15 to 18, from 19 to 35, or from 22 to 33, as calculated by RNAlib / RNAfold. Ensemble diversity may be calculated by reading single RNA sequences, computing minimum free energy (MFE) structures, and printing the result together with the corresponding MFE structure in dot-bracket notation.

[0125] In some embodiments, the 3’ UTR sequence comprises an MFE-match constrained structure, as determined in Example 1. The MFE-match constrained structure can have a structural similarity of from 0.95 to 1.0 to the MFE fold of SEQ ID NO: 1 and an ensemble diversity of from 0 to 6, as shown in FIG. 3. In some embodiments, the 3’ UTR sequence comprises an MFE-match diverse structure, as determined in Example 1. The MFE-match diverse structure can have a structural similarity of from 0.95 to 1.0 to the MFE fold of SEQ ID NO: 1 and an ensemble diversity of from 22 to 33, as shown in FIG. 3. In some embodiments, the 3’ UTR sequence comprises a ZoAl vicinity, as determined by Example 1. The ZoAl vicinity structure can have a structural similarity of from 0.82 to 0.88 to the MFE fold of SEQ ID NO: 1 and an ensemble diversity of from 6 to 10, as shown in FIG. 3. In some embodiments, the 3’ UTR sequence comprises a TiGu vicinity structure, as determined by Example 1. The TiGu vicinity structure can have a structural similarity of from 0.7 to 0.8 to the MFE fold of SEQ ID NO: 1 and an ensemble diversity of from 15 to 18, as shown in FIG. 3. In some embodiments, the 3’ UTR sequence comprises a TaGu vicinity structure, as determined by Example 1. The TaGu vicinity structure can have a structural similarity of from 0.5 to 0.65 to the MFE fold of SEQ ID NO: 1 and an ensemble diversity of from 0 to 8, as shown in FIG. 3. In some embodiments, the 3’ UTR sequence comprises an MFE-mismatch constrained structure, as determined by Example 1. In some embodiments, the 3’ UTR sequences comprise an MFE- match diverse structure, as determined by Example 1.

[0126] A 3’ UTR sequence comprising a mutation described herein can have increased binding to the R2 retroelement polypeptide as compared to the wild-type R2 3’ UTR sequence. A 3’ UTR sequence comprising a mutation described herein can have at least 10%, at least 20%, atAttomey Docket No. 67098-712601 least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 2 fold, at least 3 fold, at least 5 fold, at least 10 fold, at least 15 fold, at least 20 fold, at least 30 fold, at least 50 fold, or at least 100 fold higher binding affinity compared to the wild-type R2 3’ UTR sequence.

[0127] In some embodiments, any one of the nucleic acid constructs herein comprising a 3’ UTR sequence comprising a mutation relative to the wild-type R2 3’ UTR sequence is more efficiently inserted into a genome by target primed reverse transcription as compared to an otherwise identical nucleic acid construct comprising the wild-type R2 3’ UTR sequence. In some embodiments, insertion efficiency is assessed by determining a level of gene expression from the template sequence following insertion into a target nucleic acid of a cell. A 3’ UTR sequence comprising a mutation described herein can improve insertion efficiency of a cargo nucleic acid sequence by least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 2 fold, at least 3 fold, at least 5 fold, at least 10 fold, at least 15 fold, at least 20 fold, at least 30 fold, at least 50 fold, or at least 100 fold compared to the wild-type R2 3’ UTR sequence.Cargo nucleic acid sequence

[0128] In some aspects, the non-naturally occurring nucleic acid described herein further comprises a cargo nucleic acid sequence. The cargo nucleic acid sequence can be a heterologous cargo sequence. As used herein, the term “heterologous” refers to any polynucleotide or polypeptide sequence that is not naturally occurring in a host cell or organism or is inserted in a location not naturally occurring in the host cell or organism. The cargo nucleic acid sequence can be a cargo sequence that is not naturally occurring with a 3’ UTR sequence in the non-naturally occurring nucleic acid.

[0129] In some embodiments, the cargo nucleic sequence can be or comprise a template sequence for TPRT by the retroelement polypeptide. The cargo nucleic acid sequence can have a length of at least about 200 bases. The template sequence may have a length of at least about 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, or 2000 bases. The template may have a length of longer than about 2000 bases.

[0130] In some embodiments, the cargo nucleic acid sequence encodes a protein. In some embodiments, the cargo nucleic acid sequence comprises a transgene. In some embodiments, the transgene is a therapeutically active gene or a diagnostic gene. The transgene may comprise PKU, OTC, MMUT, ATP7B, F8, F9, or SERPINA1. The template sequence may comprise one or more non-native transgenes that rescue loss of function in a human disease or confer beneficial function.Attomey Docket No. 67098-712601

[0131] The cargo nucleic acid sequence can further comprise a promoter sequence, e.g., an RNAP II Promoter or a T7 RNA polymerase promoter. The cargo nucleic acid sequence can further comprise a terminator, e.g., a RNAP I terminator.Other nucleic acid elements

[0132] In some aspects, the non-naturally occurring nucleic acid comprises one or more additional nucleic acid elements.

[0133] In some aspects, the non-naturally occurring nucleic acid comprises a 3’ region that is located 3’ of the cargo nucleic acid sequence that further comprises one or more additional elements in addition to the 3’ UTR sequence having least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a wild-type R23’ UTR sequence, described elsewhere herein. In some embodiments, the 3’ region comprises a reverse transcriptase translation stop codon, a second 3’ untranslated region (3’ UTR), a tail region, or a combination thereof. The second 3 ’ UTR sequence can be a UTR designed for stability of the non-naturally occurring nucleic acid molecule or specificity of insertion into a target nucleic acid molecule. In the example shown in FIG. 4B, the nucleic acid construct comprises a first 3’ UTR and a minpAS sequence (SEQ ID NO: 95).

[0134] In some aspects, the non-naturally occurring nucleic acid further comprises a tail region. The tail region can be in the 3’ region located 3’ of the cargo nucleic acid sequence, as shown in FIGs. 4A and 4B. In some embodiments, the tail region comprises an adenine-rich region of at least 22 nucleotides. In some embodiments, at least 70% of the nucleobases in the adenine-rich region are adenine.

[0135] In some embodiments, the tail region further comprises a stability element. In some embodiments, the stability element is 5’ to the adenine-rich region. In some embodiments, the stability element is 3’ to the adenine-rich region.

[0136] In some embodiments, the adenine-rich region is between 22 and 41 nucleotides in length. In some embodiments, the adenine-rich region is 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, or 41 nucleotides in length. In some embodiments, the adenine-rich region is 22, 32, or 41 nucleotides in length. In some embodiments, the adenine-rich region is 22 nucleotides in length. In some embodiments, the adenine-rich region is 32 nucleotides in length. In some embodiments, the adenine-rich region is 41 nucleotides in length.

[0137] In some embodiments, the adenine-rich region has between 0% and 30% non-adenine nucleobases. In some embodiments, the adenine-rich region has 0% to 5% non-adenineAttomey Docket No. 67098-712601 nucleobases. In some embodiments, the adenine-rich region has 5% to 10% non-adenine nucleobases. In some embodiments, the adenine-rich region has 10% to 15% non-adenine nucleobases. In some embodiments, the adenine-rich region has 15% to 20% non-adenine nucleobases. In some embodiments, the adenine-rich region has 20% to 25% non-adenine nucleobases. In some embodiments, the adenine-rich region has 25% to 30% non-adenine nucleobases. In some embodiments, the adenine-rich region has 0% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than 10% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than 20% non-adenine nucleobases. In some embodiments, the adenine-rich region has no more than 30% non-adenine nucleobases. In some embodiments, the adenine-rich region has 10% non-adenine nucleobases. In some embodiments, the adenine-rich region has 20% non-adenine nucleobases. In some embodiments, the adenine-rich region has 30% non-adenine nucleobases. In some embodiments, the non- adenine nucleobases comprise any one of cytosine, guanine, uridine, or a modified uridine. In some embodiments, the modified uridine is pseudouridine or N1 -methylpseudouridine.

[0138] In some embodiments, the three, 3 ’-most terminal nucleotides of the adenine-rich region each comprise adenine, e.g., the sequence of the three terminal nucleotides of the adenine- rich region is AAA, where A is a nucleotide comprising adenine. In some embodiments, 3 ’-most terminal nucleotide of the adenine-rich region does not comprise adenine. In some embodiments, the sequence of the three 3 ’-most terminal nucleotides of the adenine-rich is AAX, where A is a nucleotide comprising adenine and X is a nucleotide comprising a non-adenine nucleobase. In some embodiments, the two 3 ’-most terminal nucleotides of the adenine region do not comprise adenine. In some embodiments, the sequence of the three 3 ’-most terminal nucleotides of the adenine-rich is AXX, where A is a nucleotide comprising adenine and X is a nucleotide comprising a non-adenine nucleobase.

[0139] In some embodiments, the stability element in the tail region is selected from K3, K4, eK5, MALATl mm, MALATl hs, MALAT1 Jis-core, PAN KSHV, dENE os, HSL, U7, K9, K16, KEMCV, or KBDVl.

[0140] In some embodiments, the non-naturally occurring nucleic acid molecule further comprises a promoter sequence or a terminator sequence. The promoter sequence can be an RNAP II promoter or a T7 RNA polymerase promoter. The terminator sequence can be an RNA polymerase terminator sequence.

[0141] In some aspects, the non-naturally occurring nucleic acid further comprises a 5’ region that is located 5’ of the cargo nucleic acid sequence. In some embodiments, the 5’ region further comprises a sequence derived from a native retroelement 5’ region, a 5’ UTR sequence, an rRNA sequence, a ribozyme sequence, a folding motif sequence, a Kozak sequence or anAttomey Docket No. 67098-712601 internal ribosome entry site, a non-native translation start codon, a 5’ cap, or any combination thereof. The ribozyme sequence can be an HDV ribozyme (e.g., HDV_ac2, HDV gul, HDV_gu5b, HDV gu6. HDV_gu5b_NP2), a TriCasA ribozyme, an L8 ribozyme (e.g., L8_gu6) an SL28 ribozyme, or a. native cognate or semi-cognate ribozyme, or modified variants thereof. In the example shown in FIG. 4B, the 5’ region can comprise HDV-gRZ (SEQ ID NO: 91) and the Neo3-Kozak sequence. In some embodiments, the 5’ region comprises a stability element.

[0142] In some embodiments, the 5’ region comprises a 5’ UTR sequence that interacts with or binds to a cognate R2 retroelement polypeptide derived from the same species. In some embodiments, the 3’ UTR sequence interacts with or binds to a non-cognate R2 retroelement polypeptide derived from a different species as the 3’ UTR sequence. In some embodiments, the 5’ UTR sequence has least 70%, at least 75%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to a wild-type R2 5’ UTR sequence. In some embodiments, the 5’ UTR sequence is or comprises a wild-type R2 5’ UTR sequence. The wildtype R2 5’ UTR sequence can be a wild-type Geospiza fortis (GeFo) R2 5’ UTR sequence, a wild-type Zonotrichia albicollis (ZoAl) R2 5’ UTR sequence, a wild-type Taeniopygia guttata (TaGu) R2 5’ UTR sequence, or a wild-type Tinamus guttatus (TiGu) R2 5’ UTR sequence.Modified Uridines

[0143] In some embodiments, the non-naturally occurring nucleic acid molecule may comprise one or more native or non-native additions or modifications.

[0144] In some embodiments, the non-naturally occurring nucleic acid molecule comprises RNA. The non-naturally occurring nucleic acid molecule can comprise one or more modified uridine (U) nucleosides. RNAs containing unmodified uridines can activate the innate immune response and can be less stable in cells. Modified uridines can reduce the innate immune response in a host organism when cells are transfected the non-naturally occurring nucleic acid molecule. Modified uridines can also increase RNA stability.

[0145] In some embodiments, the non-naturally occurring nucleic acid molecule comprises one or more modified uridine (U) nucleosides, selected from the group consisting of Nl-methyl- pseudouridine (NlmYU), pseudouridine (YU), 5 -methyluridine (5meU), 5-methyoxyuridine (5moU), and or a combination thereof. In some embodiments, the non-naturally occurring nucleic acid molecule comprises Nl-methyl-pseudouri dine (NlmYU). In some embodiments, the non-naturally occurring nucleic acid molecule comprises pseudouridine (YU). In some embodiments, the non-naturally occurring nucleic acid molecule comprises a mixture orAttomey Docket No. 67098-712601 combination of unmodified uridines and modified uridines selected from the group consisting of NlmYU, YU, 5meU, and 5moU.

[0146] In some embodiments, the modified uridines are distributed throughout the non- naturally occurring nucleic acid molecule. Ion some embodiments, all the uridines comprise the same modified uridine (e.g., all the uridines are Nl-methyl-pseudouridine (NlmYU) or all the modified uridines are pseudouridine (YU)).

[0147] In some embodiments, the non-naturally occurring nucleic acid molecule comprising modified uridines is not cleavable by a ribozyme. In some embodiments, a non-naturally occurring nucleic acid molecule comprising the modified uridines Nl-methyl-pseudouridine (NlmYU) or pseudouridine (YU) is not cleavable by a ribozyme. In some embodiments, a non- naturally occurring nucleic acid molecule comprising a modified uridine increases the efficiency of insertion of the cargo nucleic acid sequence into a target nucleic acid molecule compared to template RNA comprising an unmodified uridine. In some embodiments, cellular toxicity is decreased when the non-naturally occurring nucleic acid molecule comprises a modified uridine.II. Systems

[0148] In some aspects, the present disclosure provides systems for modifying a target nucleic acid in a cell. The systems are gene-insertion systems, which comprise any one of the nucleic acid constructs described herein and a partnered retroelement polypeptide or second nucleic acid encoding the partnered retroelement polypeptide. The nucleic acid construct in the gene-insertion system, also referred to as a gene insertion construct, can comprise a cargo nucleic acid sequence, and the gene-insertion system components can cause insertion of a cargo nucleic acid sequence or a reverse complement thereof into a target nucleic acid molecule.

[0149] In some embodiments, the partnered retroelement polypeptide or second nucleic acid encoding the partnered retroelement polypeptide comprises or encodes at least one retroelement protein. In some embodiments, the retroelement polypeptide can be a non-LTR retroelement polypeptide containing a TPRT-competent reverse transcriptase (RT) or strand-nicking endonuclease activity that is active when assayed for RT primer extension or in vitro TPRT. The retroelement polypeptide can be site-specific.

[0150] The retroelement polypeptide can be derived from a non-LTR retroelement, for example of the RLE-type or APE-type or Penelope-type. An RLE-type non-LTR retrotransposon may be from any one of many clades, including but not limited to R2, R4, CRE, Genie, HERO, NeSL. An APE-type non-LTR retrotransposon may be from any one of many clades, including but not limited to I, Rl, LI, Txl, CR1, Rexl, Jockey, L2, Tad, RTE, RTEX, ingi, Vingi, TR AS, SARI’, or any combination thereof. In some embodiments, gene-insertion system componentsAttomey Docket No. 67098-712601 may be derived from retroelements that insert into rDNA, i.e., the so-called R elements, such as retroelements of the R1 or R2 clades. In some embodiments, the R2 clade retroelement may have canonical R2 retroelement insertion site specificity or may be derived from an R8 or R9 retroelement in the larger R2 clade that have changed target sequence relative to the canonical R2 retroelements or may be derived from R2NS retroelements that appear to have lost target site specificity. In some embodiments, the retroelement proteins may be R2 retroelement proteins, or an R2 / R8 / R9 domain architecture of non-LTR RT proteins, or a naturally occurring protein or protein complex.

[0151] Gene-insertion system components may be derived from portions or domains of retroelements found in any species, including those of distant evolutionary relation to the subject. For example, suitable retroelements from which gene-insertion system components may be derived may comprise those found in birds (e.g., Zonotrichia albicoUis. Taeniopygia giilala, Tinamus giiUalus. and Geospiza fords), fish (e.g., Pungitis pungilis, Oryzias talipes, Danio rerio, Oryzias melasligma, Petromyzon marinus, Salmo iriila, Salmo salar, or Gasterosteus aciilealiis), insects (e.g., Drosophila mercalorum. Drosophila melanogasler, Nasonia vitripennis, Tribolium caslaneum. Drosophila simulans. Apis cerana, and Bombyx mori), crustaceans (e.g., Lepidurus couesii, and Triops cancriformis), other invertebrates (e.g., Limulus polyphemus, Hydra magnipapillata, or Adineta vaga), chordates (e.g., Ciona intestinalis) including mammals, and any combination thereof.

[0152] In some embodiments, the retroelement polypeptide is derived from a same species as a 3’ UTR sequence in the nucleic acid construct comprising the cargo nucleic acid sequence. In some embodiments, the retroelement polypeptide is derived from a different species as a 3’ UTR sequence in the nucleic acid construct comprising the cargo nucleic acid sequence. In some embodiments, the retroelement polypeptide is derived from a same species as a 5’ UTR sequence in the nucleic acid construct comprising the cargo nucleic acid sequence. In some embodiments, the retroelement polypeptide is derived from a different species as a 5’ UTR sequence in the nucleic acid construct comprising the cargo nucleic acid sequence. The retroelement polypeptide can be a GeFo R2 polypeptide, a ZoAl R2 polypeptide, a TaGu R2 polypeptide, or a TiGu R2 polypeptide.

[0153] In some embodiments, the system further comprises the target nucleic acid molecule. The target nucleic acid molecule can comprise DNA. The target nucleic molecule can comprise double-stranded DNA. In some embodiments, the target nucleic acid molecule comprises a plasmid. In some embodiments, the target nucleic acid molecule comprises genomic DNA. The target nucleic acid can comprise a rDNA sequence. In some embodiments, the target nucleic acid comprises a 28 rRNA sequence.Attomey Docket No. 67098-712601

[0154] In some embodiments, the present disclosure provides a cell comprising a composition or system descried herein. The cell can be a eukaryotic cell. In some embodiments, the cell is a mammalian cell or a human cell. For example, the cell may be a liver cell, a hepatocyte, a hepatic stellate cell, a retinal cell, a Kupffer cell, an endothelial cell, a liver sinusoidal endothelial cell, an epithelial cell, or a neuron. The cell can be a liver epithelial cell (e.g., BNL). The cell can be a hepatocellular carcinoma cell (e.g, Huh7). The cell can be a retinal pigment endothelial cell (e.g., RPE-1). In some embodiments, the cell is a dividing cell. In some embodiments, the cell is a post-mitotic cell. In some embodiments, the cell is a cancerous cell. In some embodiments, the cell is a non-cancerous cell.III. Methods of Use

[0155] In some aspects, the present disclosure provides methods for modifying a target nucleic acid molecule. A method for modifying a target nucleic acid molecule can comprise inserting a cargo nucleic acid sequence into a target nucleic acid molecule.

[0156] In some embodiments, the template sequence is inserted at one or more target insertion sites. In some embodiments, the cargo nucleic acid sequence or a reverse complement thereof is inserted into a eukaryotic genome. In some cases, the insertion into the target nucleic acid molecule is site-specific. The insertion can occur at a “safe harbor site,” such as in rRNA. In some embodiments, the insertion occurs in the 28 S rRNA gene. The method can leverage an R2 retroelement polypeptide to support RNA-mediated transgene insertion, like Target Primed Reverse Transcription (TPRT)-initiated transgene insertion. TPRT may be used to modify a target nucleic acid in a mammalian cell rDNA, using a directly introduced RNA template comprising the cargo nucleic acid sequence. The systems and methods may involve other species’ genomes as targets for TPRT -mediated transgene insertion, or for non-genomic targets.

[0157] In some embodiments, the method comprises modifying a target nucleic acid molecule in a cell. The methods may comprise introducing an effective amount of at least one composition or gene-insertion system described elsewhere herein to the cell. The cell can be a eukaryotic cell. In some embodiments, the cell is a mammalian cell or a human cell. For example, the cell may be a liver cell, a hepatocyte, a hepatic stellate cell, a retinal cell, a Kupffer cell, an endothelial cell, a liver sinusoidal endothelial cell, an epithelial cell, or a neuron. The cell can be a liver epithelial cell (e.g., BNL). The cell can be a hepatocellular carcinoma cell (e.g, Huh7). The cell can be a retinal pigment endothelial cell (e.g., RPE-1). In some embodiments, the cell is a cancerous cell. In some embodiments, the cell is a non-cancerous cell. In some embodiments, the cell is actively proliferating or expanding. In some embodiments, the cell isAttomey Docket No. 67098-712601 progressing through the cell cycle. In some embodiments, the cell comprises transcriptionally active or replicationally active rDNA. In some embodiments, the cell is a post-mitotic cell.

[0158] In some embodiments, the gene-insertion system comprises delivery of a transgene to a cell via delivery of the gene-insertion system described herein.

[0159] Introduction of the gene-insertion system to cells may involve standard methods, such as lipid-enabled transfection, electroporation, or other methods. The gene-insertion system may be encapsulated in one or more delivery vehicles. Delivery vehicles may facilitate in vivo or in vitro transfection of subject cells by protecting gene-insertion system components from degradation in the extracellular environment, facilitating uptake by subject cells, enhancing endosomal escape, or any combination thereof. Delivery vehicles may be lipid-based (e.g., lipid nanoparticles (LNPs), liposomes, and micelles) or non-lipid-based (e.g., virus like particles (VLPs) and polymeric delivery particles).

[0160] For example, the gene-insertion system may be introduced to cells via a lipid nanoparticle (LNP). LNPs possess an exterior lipid layer including a hydrophilic exterior surface that is exposed to the non-LNP environment, a non-aqueous or an aqueous interior space (i.e., micelle-like and vesicle-like LNPs respectively), and at least one hydrophobic inter-membrane space. LNP membranes may be non-lamellar or lamellar, with 1, 2, 3, 4, 5 or more than 5 layers. LNPs may be solid or semi-solid. In some embodiments, at least one cargo or a payload (such as the gene-insertion system) is comprised in the interior space, the inter membrane space, on the exterior surface, or any combination thereof of the LNP.

[0161] In some embodiments, the LNPs may comprise an ionizable (cationic) lipid, a phospholipid, cholesterol, and a polymer-conjugated lipid. Cholesterol promotes membrane fusion and aids in LNP stability. In some embodiments, the cholesterol is a sterol. Phospholipids aid in endosomal escape and provide structure to the LNP bilayer. The phospholipid may be a non-cationic lipid. Polymer-conjugated lipids reduce LNP aggregation and “protect” the LNP from non-specific endocytosis by immune cells. Polymer-conjugated lipids comprise PEG-lipids. Ionizable (cationic) lipids enhance endosomal escape and complex with a negatively charged cargo (such as polynucleotides of the gene-insertion system). Cationic lipids comprise ionizable cationic lipids.

[0162] In some embodiments, the LNPs have a selective preference for delivery to liver cells. Following injection into circulation, the LNPs may shed the PEG component. The LNPs then bind to apolipoprotein E (ApoE), and then be taken into liver cells, including hepatocytes, via the LDL receptor. Thus, the LNPs facilitate hepato-specific delivery of cargo contained in the LNPs.

[0163] In some embodiments, the method comprises inserting a transgene into a target nucleic acid molecule in a cell and functionally expressing the transgene in the cell. The methodAttorney Docket No. 67098-712601 may employ various techniques to assess successful gene introduction. For example, In Vivo Imaging System (IVIS) imaging facilitates real-time, non-invasive monitoring of gene expression. IVIS also facilitates evaluation of biological expression patterns of the introduced gene in live animals. Further, non-secreted genes may be evaluated by immunohistochemistry. Similarly, secreted genes may be evaluated by microplate assays. These microplate assays may test blood samples.

[0164] The compositions, systems, or methods described herein may be utilized in treating a subject in need thereof. A composition or system described herein can be used to treat a disease or disorder. The methods can be used to treat a disease associated with a defective or mutated gene in a subject, such as but not limited to diseases caused by single- gene defects (monogenic disorders), such as Sickle cell anemia, Severe Combined Immunodeficiency (ADA-SCID / X- SCID), Cystic fibrosis, Hemophilia, Duchenne muscular dystrophy, Huntington’s disease, Parkinson’s, Hypercholesterolemia, Alpha-1 antitrypsin, Chronic granulomatous disease, Fanconi Anemia and Gaucher Disease. In some embodiments, the methods can be used to treat spinal muscular atrophy and inherited retinal dystrophy. In some embodiments, the methods can be used to treat polygenic disorders, such as but not limited to heart disease, cancer, diabetes, schizophrenia, Parkinson’s disease and Alzheimer’s disease. In some embodiments, the methods can be used to treat infectious diseases, such as HIV.

[0165] For example, in patients with hemophilia A, the payload can encode a wild-type factor VIII protein, in patients with hemophilia B, the payload can encode a wild-type factor IX protein. In some embodiments, the payload can encode a wild-type p53 gene in a subject with a defective p53 gene to help prevent tumor growth.

[0166] Representative examples of diseases or conditions that can be treated by the methods of the disclosure are shown in Table 1.Table 1. Representative diseases and conditions that can be treated by the methods of the disclosure.Attorney Docket No. 67098-712601Attorney Docket No. 67098-712601Attorney Docket No. 67098-712601

[0167] In some aspects, the method is an in vivo method. In some embodiments, the method is an ex vivo method.

[0168] In some embodiments, the methods comprise administering an effective dose of a pharmaceutical composition of the disclosure to a patient in need of treatment. The pharmaceutical composition can be administered via any suitable method that results in targeted integration of the payload sequence into one or more cells of the subject. In some embodiments, the pharmaceutical composition is administered intravenously, intramuscularly, subcutaneously, intraocularly, intraretinally, within the CNS or other neural tissue, or intranasally.

[0169] In some embodiments, the cell is removed from the subject or patient before being transfected ex vivo with mRNA encoding a retroelement polypeptide and a nucleic acid construct of the disclosure. Following ex vivo transfection, correct insertion of the heterologous polynucleotide comprising the payload sequence can be determined, for example by amplifyingAttomey Docket No. 67098-712601 sequences at the 5’ and / or 3’ insertion junctions, and / or amplifying the payload sequence. A correctly targeted insertion can also be determined by sequencing the genomic target site. Expression of the payload sequence can also be determined, for example, by detecting expression of a product encoded by the payload sequence, such as a protein or regulatory RNA. After correct integration and / or expression of tire payload sequence is determined, the correctly targeted cells can be administered to the subject (autologous therapy).EXAMPLESExample 1: Screening 3’ modules comprising variant 3’UTR sequences for improved gene insertion efficiency

[0170] In this example, variants of a 3’ module of a template RNA molecule of a TPRT gene insertion system were designed and screened for levels of gene insertion activity in a TPRT gene insertion assay.Design and structural characterization of 3 ’ module variants

[0171] To design the test constructs for the 3’ module variant screen, 100,000 variants of a GeFo 3’ UTR sequence, each comprising 1-5% mutations relative to the wild-type GeFo 3’ UTR sequence (SEQ ID NO: 1), were generated in a non-Markovian process and structurally characterized. The ensemble diversity of RNA folding and structural similarity of the MFE fold to that of the wild-type GeFo 3’ UTR (conformation 1, depicted in FIG. 1) were calculated.

[0172] For screening in the TPRT gene insertion assay, variants were selected that exhibit 1) propensity to fold into conformation 1 + low ensemble diversity 2) propensity to fold into conformation 1 + high ensemble diversity, 3) propensity to fold into alternative conformations + low ensemble diversity, and 4) propensity to fold into alternative conformations + high ensemble diversity, as depicted in FIGs. 2 and 3. As shown in FIG. 3, the selected variants were categorized into the following categories based on their propensity to fold into conformation 1 or propensity to fold into alternative conformations and their ensemble diversity: MFE-match constrained, MFE-match diverse, ZoAl vicinity, TiGu vicinity, TaGu vicinity, MFE-mismatch constrained, and MFE-mismatch diverse.Screening 3 ’ module variants in TPRT gene insertion assay in 3 cell lines

[0173] The selected 3’ UTR variant sequences were used as the 3’ module in the test constructs of template RNA molecules for screening in TPRT gene insertion assays in 3 different cell lines. The test constructs were designed with the structure depicted in FIG. 4B: HDV-gRZ (SEQ ID NO: 91) - promoter - Neo3-Kozak - cargo - minpAS (SEQ ID NO: 95) - 3’ UTR - tail. In each test construct, the promoter was pCMV (SEQ ID NO: 92) and the cargo comprised reporter gene secNL_intron002_pos03 (SEQ ID NO: 94). The tail was A22 (SEQ ID NO: 96).Attomey Docket No. 67098-712601The 3’ UTR sequences selected for screening and their mutations relative to the wild-type GeFo 3’ are shown in Table 2. In addition to the variant 3’ UTR sequences, the GeFo control, the wildtype TaGu 3’ UTR sequence (SEQ ID NO: 2), the wild-type ZoAl 3’ UTR sequence (SEQ ID NO: 3), and the wild-type TiGu 3’ UTR sequence (SEQ ID NO: 4) were also tested in the screen.

[0174] For the screen, plasmids were generated encoding the reporter RNA downstream of a T7 promoter and upstream of a Bbsi cut site. Plasmids were then digested, purified, and used as a template for in vitro transcription to synthesize RNA. The in vitro transcribed RNA was purified using an oligo dT resin.

[0175] A first set of RNA template constructs was generated with unmodified uridine. A second set of RNA template constructs was generated with the uridines substituted with pseudouridines. A third set of RNA template construct was generated with the uridines substituted with N1 -methylpseudouridines.

[0176] The three sets of RNA template constructs were then screened in TPRT gene insertion assays to measure corresponding levels of gene insertion activity by a Taeniopygia guttata (TaGu) R2 polypeptide in BNL cells, Huh7 cells, and RPE1 cells. The relative gene insertion activity resulting from each test construct was then calculated as a relative geometric mean activity over the control N1 -methylpseudouridine construct comprising the GeFo sequence.

[0177] To assess gene insertion activity, the RNAs were complexed with Lipofectamine Messenger Max to form lipoplexes that were then transfected into cells in 384-well plates using either forward or reverse transfection. At designated readout times (usually 24, 48, 72 hours etc.), the number of GFP positive cells were counted using the Opera Phenix Plus High-Content Screening System (Revvity). The supematant / cell media was completely removed, and fresh media was applied to the cells. The extracted supernatant was used to assay for secreted Nano Luciferase using the Nano-GLO Luciferase Assay System (Promega). Background signal was subtracted, and gene insertion activity was reported as RLU (Relative Luminescent Units) per transfected cell. The relative gene insertion activity over the N1 -methylpseudouridine construct comprising the GeFo sequence was calculated. The performance of the test constructs is illustrated in FIGs. 5-16.

[0178] Tables 3a-5c show the relative gene insertion activity resulting from the RNA template constructs after 24 hours. Tables 3a-3c show the gene insertion activity of the Nl- methylpseudouridine, pseudouridine, and uridine RNA template constructs, respectively, in BNL cells. The N1 -methylpseudouridine constructs comprising a 3’ module comprising a 3’ UTR sequence selected from SEQ ID NOs: 5 or 6 resulted in higher gene insertion activity compared to the control N1 -methylpseudouridine construct comprising the GeFo sequence (SEQ ID NO: 1). The pseudouridine constructs comprising a 3’ module comprising a 3’ UTR sequence selectedAttomey Docket No. 67098-712601 from SEQ ID NOs: 5, 2, 10, 9, 6, 7, 18, 27, 13, or 12 also resulted in higher gene insertion activity compared to the control N1 -methylpseudouridine construct comprising the GeFo sequence. Of all the constructs tested, the pseudouridine constructs comprising SEQ ID NOs: 5, 2, 10, 9, and 6 resulted in the highest gene insertion activities (all greater than 1.5 fold that of the control construct) in BNL cells.

[0179] Tables 4a-4c show the gene insertion activity of the N1 -methylpseudouridine, pseudouridine, and uridine RNA template constructs, respectively, in Huh7 cells. The Nl- methylpseudouridine constructs comprising a 3’ module comprising a 3’ UTR sequence selected from SEQ ID NOs: 6, 5, 8, 7, or 9 resulted in higher gene insertion activity compared to the control N1 -methylpseudouridine construct comprising the GeFo sequence (SEQ ID NO: 1). The pseudouridine constructs comprising a 3’ module comprising a 3’ UTR sequence selected from SEQ ID NOs: 6, 5, 8, 7, 13, 9, 10, 14, 11, 2, 19, 12, 16, 18, or 15 also resulted in higher gene insertion activity compared to the control N1 -methylpseudouridine construct comprising the GeFo sequence. Of all the constructs tested, the N1 -methylpseudouridine constructs comprising SEQ ID NOs: 6, 5, and 8, and the pseudouridine constructs comprising SEQ ID NOs: 6, 5, 8, 7, and 13 resulted in the highest gene insertion activities (all greater than 1.5 fold that of the control construct) in Huh7 cells.

[0180] Tables 5a-5c show the gene insertion activity of the N1 -methylpseudouridine, pseudouridine, and uridine RNA template constructs, respectively, in RPE1 cells. The Nl- methylpseudouridine constructs comprising a 3’ module comprising a 3’ UTR sequence selected from SEQ ID NO: 6, 5, 7, 8, 9, 12, 21, 14, 10, 19, 27, 20, and 11 resulted in higher gene insertion activity compared to the control N1 -methylpseudouridine construct comprising the GeFo sequence (SEQ ID NO: 1). The pseudouridine constructs comprising a 3’ module comprising a 3’ UTR sequence selected from SEQ ID NOs: 5, 6, 8, 7, 13, 14, 10, 9, 3, or 2 also resulted in higher gene insertion activity compared to the control N1 -methylpseudouridine construct comprising the GeFo sequence. Of all the constructs tested, the N1 -methylpseudouridine constructs comprising SEQ ID NOs: 6, 5, 7, 8, and 9, and the pseudouridine constructs comprising SEQ ID NOs: 5 and 6 resulted in the highest gene insertion activities (all greater than 1.5 fold that of the control construct) in RPE1 cells.

[0181] FIGs. 5-11 depict the results from Tables 3a-5c. FIGs 11 A-l 1C show the performance of all 268 constructs as geometric mean across the 3 cell lines tested. FIG. 10 shows the performance of the constructs, indicated by their respective structural category. FIG. 5 shows the performance of the constructs in each structural category after 24 hours. FIG. 6 show the performance of the constructs after 48 hours. FIGs. 5 and 6 show that the constructs with the highest gene insertion activity in BNL, Huh7, and RPE1 cells fall in the category ofAttomey Docket No. 67098-712601TiGu vicinity. Certain constructs falling in the MFE-match-diverse and TaGu vicinity also demonstrated strong performance. FIGs. 12A-15B show the performance of the N1 -methyl pseudouridine-containing constructs, the pseudouridine-containing constructs, and the uridine- containing constructs. Overall, the N1 -methylpseudouridine constructs and pseudouridine constructs demonstrated much higher performance than the uridine constructs with the same sequences in all three cell lines. On average, the highest performing 3’ UTR variant sequences across all three cell lines were SEQ ID NOs: 5-12

[0182] FIGs. 7A and 7B further illustrate the performance of the top performing 3 ’ UTR variant sequences (SEQ ID NOs: 5-12) in the three cell lines. FIG. 7A illustrates the performance relative to the control N1 -methylpseudouridine construct comprising the GeFo sequence (SEQ ID NO: 1), indicating that N1 -methylpseudouridine constructs perform particularly well in RPE1 cells, while pseudouridine constructs perform better in Huh7 and BNL cells. FIG. 7B illustrates the performance of the constructs relative to the cognate construct comprising the GeFo sequence (SEQ ID NO: 1). FIG. 7B indicates that the top two performing 3’ UTR variant sequences, GeFo98_var004829 (SEQ ID NO: 5) and GeFo98_var065243 (SEQ ID NO: 6) are consistently better relative to their cognate wild-type GeFo controls in the context of N1 -methylpseudouridine constructs, pseudouridine constructs, and uridine constructs in all three cell-lines.Correlation of mutation position in 3 ’ UTR variant sequence and variant performance

[0183] The results from the TPRT assays were then analyzed for correlation of mutation position in the 3’ UTR variant sequence and the variant’s performance. FIG. 8A depicts the correlation between the construct performance and the base position of a mutation in the variant construct relative to the wild-type GeFo 3’ UTR sequence (SEQ ID NO: 1). FIGs. 8A and 8B indicate regions 801 and 802 within the 3’ UTR sequence, where mutations from the wild-type GeFo 3’ UTR sequence led to poor construct performance. These regions correspond to nucleotide positions 36-46 (SEQ ID NO: 97, UGGACGGGCCA) and 57-64 (ACCCGGAA) of the 3’ UTR sequence.

[0184] FIG. 9A indicates the top performing 3 ’UTR variant sequences across the three cell lines and their associated mutations, while FIG. 9B indicates the bottom 20 performing 3’ variant UTR sequences across the three cell lines and their associated mutations. Notably, none of the top performing 3’ UTR variant sequences had any mutations in the sequences UGGACGGGCCA (SEQ ID NO: 97) and ACCCGGAA, corresponding to nucleotide positions 36-46 and 57-64, respectively. The worst performing 3’ UTR variant sequences had a number of mutations within nucleotide positions 36-46 and 57-64 relative to the wild-type GeFo 3’ UTR sequence (SEQ IDAttomey Docket No. 67098-712601NO: 1). These results further support the importance of the sequences UGGACGGGCCA (SEQ ID NO: 97) and ACCCGGAA, corresponding to nucleotide positions 36-46 and 57-64, respectively, for variant performance in TPRT gene insertion.

[0185] Examining the top performing 3’ UTR variant sequences further provides insight into mutation positions that are associated with strong performance in TPRT gene insertion. For instance, the top three performing constructs comprised mutations selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, and 85 relative to the wild-type GeFo 3’ UTR sequence (SEQ ID NO: 1).ttorney Docket No.67098-712601 able 2-3’ UTR variant sequences screenedttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601 able 3a - Gene insertion activity over GeFo control in BNL cells after 24 hours (Nl-methyl pseudouridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttomey Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 3b - Gene insertion activity over GeFo control in BNL cells after 24 hours (Pseudouridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 3c - Gene insertion activity over GeFo control in BNL cells after 24 hours (Uridine)-70-ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 4a - Gene insertion activity over GeFo control in Huh7 cells after 24 hours (Nl-methyl pseudouridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601 able 4b - Gene insertion activity over GeFo control in Huh7 cells after 24 hours (Pseudouridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 4c - Gene insertion activity over GeFo control in Huh7 cells after 24 hours (Uridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 5a - Gene insertion activity over GeFo control in RPE1 cells after 24 hours (Nl-methyl pseudouridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 5b - Gene insertion activity over GeFo control in RPE1 cells after 24 hours (Pseudouridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601 able 5c - Gene insertion activity over GeFo control in RPE1 cells after 24 hours (Uridine)ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601ttorney Docket No. 67098-712601able 6 - Other sequencesttorney Docket No. 67098-712601ttorney Docket No. 67098-712601

Claims

Attorney Docket No. 67098-712601CLAIMSWHAT IS CLAIMED IS:

1. A non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1.

2. The non-naturally occurring nucleic acid molecule of claim 1, wherein the 3’ UTR sequence comprises a nucleotide sequence ACCCGGAA.

3. The non-naturally occurring nucleic acid molecule of claim 1 or 2, wherein the 3’ UTR sequence comprises a nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97).

4. A non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 1, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the 3’ UTR sequence further comprises a first nucleotide sequence ACCCGGAA and a second nucleotide sequence UGGACGGGCCA (SEQ ID NO: 97}.

5. The non-naturally occurring nucleic acid molecule of any one of claims 1-4, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1.

6. The non-naturally occurring nucleic acid molecule of any one of claims 1-5, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1.

7. The non-naturally occurring nucleic acid molecule of any one of claims 1-6, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 12, 18, 23, 24, 50, or 76 of SEQ ID NO: 1.

8. The non-naturally occurring nucleic acid molecule of any one of claims 1-7, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 12, or 18 of SEQ ID NO: 1.

9. The non-naturally occurring nucleic acid molecule of any one of claims 1-7, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to aAttorney Docket No. 67098-712601 nucleotide position selected from any one of nucleotide positions 5, 23, 24, 50, or 76 of SEQ ID NO: 1.

10. The non-naturally occurring nucleic acid molecule of any one of claims 1-9, wherein the 3’ UTR sequence comprises at least two mutations.

11. The non-naturally occurring nucleic acid molecule of any one of claims 1-10, wherein the 3’ UTR sequence comprises two mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1.

12. The non-naturally occurring nucleic acid molecule of any one of claims 1-11, wherein the 3’ UTR sequence comprises at least three mutations.

13. The non-naturally occurring nucleic acid molecule of any one of claims 1-12, wherein the 3’ UTR sequence comprises three mutations located at positions corresponding to nucleotide positions selected from nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1.

14. The non-naturally occurring nucleic acid molecule of any one of claims 1-13, wherein the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 4 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 12 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 18 of SEQ ID NO: 1.

15. The non-naturally occurring nucleic acid molecule of any one of claims 1-13, wherein the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 5 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 23 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 24 of SEQ ID NO: 1, (iv) a fourth mutation located at a position corresponding nucleotide position 50 of SEQ ID NO: 1, and (v) a fifth mutation at a position corresponding nucleotide position 76 of SEQ ID NO: 1.

16. The non-naturally occurring nucleic acid molecule of any one of claims 1-13, wherein the 3’ UTR sequence comprises (i) a first mutation located at a position corresponding to nucleotide position 7 of SEQ ID NO: 1, (ii) a second mutation located at a position corresponding to nucleotide position 27 of SEQ ID NO: 1, and (iii) a third mutation located at a position corresponding nucleotide position 81 of SEQ ID NO: 1, and (iv) a fourth mutation located at a position corresponding nucleotide position 85 of SEQ ID NO: 1.

17. The non-naturally occurring nucleic acid molecule of any one of claims 1-16, wherein the 3’ UTR sequence has at least 90%, at least 91%, at least 92%, at least 93%, at leastAttorney Docket No. 67098-71260194%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 1.

18. A non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence, wherein the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to a stem-loop region of the wild-type R2 3’ UTR sequence, wherein the wild-type R2 3’ UTR sequence is selected from the group consisting of a wild-type Geospiza fortis (GeFo) R2 3’ UTR sequence, a wild-type Zonotrichia albicollis (ZoAl) R2 3’ UTR sequence, a wild-type Taeniopygia guttata (TaGu) R2 3’ UTR sequence, and a wild-type Tinamus guttatus (TiGu) R2 3’ UTR sequence.

19. The non-naturally occurring nucleic acid molecule of claim 18, wherein the wildtype R2 3’ UTR sequence comprises SEQ ID NO: 1.

20. The non-naturally occurring nucleic acid molecule of claim 18 or 19, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-35, 47-50, or 68-92 of SEQ ID NO: 1.

21. The non-naturally occurring nucleic acid molecule of claim 20, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, 27, 50, 76, 81, or 85 of SEQ ID NO: 1.

22. The non-naturally occurring nucleic acid molecule of claim 18, wherein the wildtype R2 3’ UTR sequence comprises SEQ ID NO: 2.

23. The non-naturally occurring nucleic acid molecule of claim 18, wherein the wildtype R2 3’ UTR sequence comprises SEQ ID NO: 3.

24. The non-naturally occurring nucleic acid molecule of claim 18, wherein the wildtype R2 3’ UTR sequence comprises SEQ ID NO: 4.

25. The non-naturally occurring nucleic acid molecule of any one of claims 18-24, wherein the wild-type R2 3’ UTR sequence comprises at least two stem-loop regions.

26. The non-naturally occurring nucleic acid molecule of any one of claims 18-25, wherein the wild-type R2 3’ UTR sequence comprises at least three stem-loop regions.

27. The non-naturally occurring nucleic acid molecule of any one of claims 18-26, wherein the wild-type R2 3’ UTR sequence comprises at least four stem-loop regions.

28. The non-naturally occurring nucleic acid molecule of any one of claims 18-27, wherein the 3’ UTR sequence comprises at least two mutations relative to the wild-type R2 3’ UTR sequence.Attorney Docket No. 67098-71260129. The non-naturally occurring nucleic acid molecule of claim 28, wherein the at least two mutations are both located in regions of the 3’ UTR sequence corresponding to stemloop regions of the wild-type R2 3’ UTR sequence.

30. The non-naturally occurring nucleic acid molecule of any one of claims 18-29, wherein the 3’ UTR sequence comprises two mutations located in a region of the 3’ UTR sequence corresponding to a same stem-loop region of the wild-type R2 3’ UTR sequence.

31. The non-naturally occurring nucleic acid molecule of any one of claims 18-29, wherein the 3’ UTR sequence comprises two mutations located in regions of the 3’ UTR sequence corresponding to two discrete stem-loop regions of the wild-type R2 3’ UTR sequence.

32. A non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence, wherein the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence, wherein the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stemloop region, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence.

33. The non-naturally occurring nucleic acid molecule of claim 32, wherein the wildtype R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4.

34. The non-naturally occurring nucleic acid molecule of claim 32 or 33, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1.

35. The non-naturally occurring nucleic acid molecule of claim 34, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 3-31 of SEQ ID NO: 1.

36. The non-naturally occurring nucleic acid molecule of claim 34 or 35, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1.

37. The non-naturally occurring nucleic acid molecule of claim 32 or 33, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2.

38. The non-naturally occurring nucleic acid molecule of claim 32 or 33, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3.

39. The non-naturally occurring nucleic acid molecule of claim 32 or 33, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.Attorney Docket No. 67098-71260140. The non-naturally occurring nucleic acid molecule of any one of claims 32-39, wherein the 3’ UTR sequence comprises an additional mutation relative to the wild-type R2 3’ UTR sequence.

41. The non-naturally occurring nucleic acid molecule of claim 40, wherein the additional mutation is located in the region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence.

42. The non-naturally occurring nucleic acid molecule of claim 40, wherein the additional mutation is located in the region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence.

43. The non-naturally occurring nucleic acid molecule of claim 40, wherein the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

44. A non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence, wherein the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence, wherein the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stemloop region, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence.

45. The non-naturally occurring nucleic acid molecule of claim 44, wherein the wildtype R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4.

46. The non-naturally occurring nucleic acid molecule of claim 44 or 45, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1.

47. The non-naturally occurring nucleic acid molecule of claim 46, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 32-35 or 47-50 of SEQ ID NO: 1.

48. The non-naturally occurring nucleic acid molecule of any one of claim 46 or 47, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to nucleotide position 50 of SEQ ID NO: 1.

49. The non-naturally occurring nucleic acid molecule of claim 44 or 45, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2.

50. The non-naturally occurring nucleic acid molecule of claim 44 or 45, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3.Attorney Docket No. 67098-71260151. The non-naturally occurring nucleic acid molecule of claim 44 or 45, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.

52. The non-naturally occurring nucleic acid molecule of any one of claims 44-51, wherein the 3’ UTR sequence comprises an additional mutation relative to the wild-type R2 3’ UTR sequence.

53. The non-naturally occurring nucleic acid molecule of claim 52, wherein the additional mutation is located in the region of the 3’ UTR sequence corresponding to the second stem-loop region of the wild-type R2 3’ UTR sequence.

54. The non-naturally occurring nucleic acid molecule of claim 52, wherein the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

55. A non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence, wherein the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence, wherein the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stemloop region, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

56. The non-naturally occurring nucleic acid molecule of claim 55, wherein the wildtype R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4.

57. The non-naturally occurring nucleic acid molecule of claim 55 or 56, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1.

58. The non-naturally occurring nucleic acid molecule of claim 57, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 68-92 of SEQ ID NO: 1.

59. The non-naturally occurring nucleic acid molecule of claim 57 or 58, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1.

60. The non-naturally occurring nucleic acid molecule of claim 55 or 56, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 2.

61. The non-naturally occurring nucleic acid molecule of claim 55 or 56, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 3.

62. The non-naturally occurring nucleic acid molecule of claim 55 or 56, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 4.Attorney Docket No. 67098-71260163. The non-naturally occurring nucleic acid molecule of any one of claims 55-62, wherein the 3’ UTR sequence comprises an additional mutation relative to the wild-type R2 3’ UTR sequence.

64. The non-naturally occurring nucleic acid molecule of claim 63, wherein the additional mutation is located in the region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

65. A non-naturally occurring nucleic acid molecule comprising a 3’ UTR sequence having at least 85% sequence identity to a wild-type R2 3’ UTR sequence, wherein the 3’ UTR sequence comprises a mutation relative to the wild-type R2 3’ UTR sequence, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to a stem-loop region of the wild-type R2 3’ UTR sequence, and wherein the mutation is located in a region that does not correspond to a position at the bottom of a stem-loop in the wild-type R2 3’ UTR sequence.

66. The non-naturally occurring nucleic acid molecule of claim 65, wherein the wildtype R2 3’ UTR sequence comprises any one of SEQ ID NOs: 1-4.

67. The non-naturally occurring nucleic acid molecule of claim 65 or 66, wherein the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the first stem-loop region of the wild-type R2 3’ UTR sequence.

68. The non-naturally occurring nucleic acid molecule of claim 66 or 67, wherein the wild-type R2 3’ UTR sequence comprises SEQ ID NO: 1, and wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4-30 of SEQ ID NO: 1.

69. The non-naturally occurring nucleic acid molecule of claim 68, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 4, 5, 7, 12, 18, 23, 24, or 27 of SEQ ID NO: 1.

70. The non-naturally occurring nucleic acid molecule of claim 65 or 66, wherein the wild-type R2 3’ UTR sequence comprises, in order from 5’ to 3’, a first stem-loop region, a second stem-loop region, a third stem-loop region, and a fourth stem-loop region, wherein the mutation is located in a region of the 3’ UTR sequence corresponding to the fourth stem-loop region of the wild-type R2 3’ UTR sequence.

71. The non-naturally occurring nucleic acid molecule of claim 70, wherein the wildtype R2 3’ UTR sequence comprises SEQ ID NO: 1, wherein the mutation is located at a positionAttorney Docket No. 67098-712601 of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 69-91 of SEQ ID NO: 1.

72. The non-naturally occurring nucleic acid molecule of claim 71, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 76, 81, or 85 of SEQ ID NO: 1.

73. The non-naturally occurring nucleic acid molecule of any one of claims 18-72, wherein the 3’ UTR sequence comprises a nucleotide sequence ACCCGGAA.

74. The non-naturally occurring nucleic acid molecule of any one of claims 18-73, wherein the 3’ UTR sequence comprises a nucleotide sequence UGGACGGGCCA (SEQ ID75. The non-naturally occurring nucleic acid molecule of any one of claims 1-74, wherein the mutation is a purine to pyrimidine mutation or a pyrimidine to purine mutation.

76. The non-naturally occurring nucleic acid molecule of any one of claims 1-73, wherein the mutation is a purine to pyrimidine mutation.

77. The non-naturally occurring nucleic acid molecule of any one of claims 1-76, wherein the mutation is an adenine to uracil mutation.

78. The non-naturally occurring nucleic acid molecule of any one of claims 1-76, wherein the mutation is a guanine to uracil mutation.

79. The non-naturally occurring nucleic acid molecule of any one of claims 1-76, wherein the mutation is an adenine to cytosine mutation.

80. The non-naturally occurring nucleic acid molecule of any one of claims 1-76, wherein the mutation is a guanine to cytosine mutation.

81. The non-naturally occurring nucleic acid molecule of any one of claims 1-73, wherein the mutation is a pyrimidine to purine mutation.

82. The non-naturally occurring nucleic acid molecule of any one of claims 1-73 or 81, wherein the mutation is a uracil to adenine mutation.

83. The non-naturally occurring nucleic acid molecule of any one of claims 1-73 or 81, wherein the mutation is a uracil to guanine mutation.

84. The non-naturally occurring nucleic acid molecule of any one of claims 1-73 or 81, wherein the mutation is a cytosine to adenine mutation.

85. The non-naturally occurring nucleic acid molecule of any one of claims 1-73 or 81, wherein the mutation is a cytosine to guanine mutation.

86. The non-naturally occurring nucleic acid molecule of any one of claims 1-72, wherein the mutation is a purine to purine mutation or a pyrimidine to pyrimidine mutation.Attorney Docket No. 67098-71260187. The non-naturally occurring nucleic acid molecule of any one of claims 1-72 or 86, wherein the mutation is a purine to purine mutation.

88. The non-naturally occurring nucleic acid molecule of any one of claims 1-72, 86, or 87, wherein the mutation is an adenine to guanine mutation.

89. The non-naturally occurring nucleic acid molecule of any one of claims 1-72, 86, or 87, wherein the mutation is a guanine to adenine mutation.

90. The non-naturally occurring nucleic acid molecule of any one of claims 1-72 or 86, wherein the mutation is a pyrimidine to pyrimidine mutation.

91. The non-naturally occurring nucleic acid molecule of any one of claims 1-72, 86, or 90, wherein the mutation is a cytosine to uracil mutation.

92. The non-naturally occurring nucleic acid molecule of any one of claims 1-72, 86, or 90, wherein the mutation is a uracil to cytosine mutation.

93. The non-naturally occurring nucleic acid molecule of any one of claims 1-92, wherein the 3’ UTR sequence comprises a sequence having at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 5-90.

94. The non-naturally occurring nucleic acid molecule of any one of claims 1-93, wherein the 3’ UTR sequence comprises any one of SEQ ID NOs: 5-90.

95. The non-naturally occurring nucleic acid molecule of any one of claims 1-94, wherein the 3’ UTR sequence comprises any one of SEQ ID NOs: 5-12.

96. The non-naturally occurring nucleic acid molecule of any one of claims 1-95, wherein the 3’ UTR sequence comprises SEQ ID NO: 5.

97. The non-naturally occurring nucleic acid molecule of any one of claims 1-95, wherein the 3’ UTR sequence comprises SEQ ID NO: 6.

98. The non-naturally occurring nucleic acid molecule of any one of claims 1-95, wherein the 3’ UTR sequence comprises SEQ ID NO: 7.

99. The non-naturally occurring nucleic acid molecule of any one of claims 1-95, wherein the 3’ UTR sequence comprises SEQ ID NO: 11.

100. The non-naturally occurring nucleic acid molecule of any one of claims 1-95, wherein the 3’ UTR sequence comprises SEQ ID NO: 9.

101. The non-naturally occurring nucleic acid molecule of any one of claims 1-100, wherein the 3’ UTR sequence comprises no more than 200 nucleotides.

102. A non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 5, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1,Attorney Docket No. 67098-712601 wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1.

103. The non-naturally occurring nucleic acid molecule of claim 102, wherein the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:5.

104. The non-naturally occurring nucleic acid molecule of claim 102 or 103, wherein the 3’ UTR sequence comprises SEQ ID NO: 5.

105. A non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 11, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1.

106. The non-naturally occurring nucleic acid molecule of claim 105, wherein the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 11.

107. The non-naturally occurring nucleic acid molecule of claim 105 or 106, wherein the 3’ UTR sequence comprises SEQ ID NO: 11.

108. A non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 6, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1, wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1.

109. The non-naturally occurring nucleic acid molecule of claim 108, wherein the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO:6110. The non-naturally occurring nucleic acid molecule of claim 108 or 109, wherein the 3’ UTR sequence comprises SEQ ID NO: 6.

111. A non-naturally occurring nucleic acid molecule comprising a 3' UTR sequence having at least 85% sequence identity to SEQ ID NO: 9, wherein the 3’ UTR sequence comprises a mutation relative to a wild-type R2 3’ UTR sequence comprising SEQ ID NO: 1,Attorney Docket No. 67098-712601 wherein the mutation is located at a position of the 3’ UTR sequence corresponding to a nucleotide position selected from any one of nucleotide positions 1-35, 47-56, or 65-102 of SEQ ID NO: 1.

112. The non-naturally occurring nucleic acid molecule of claim 111, wherein the 3’ UTR sequence has at least 90%, at least 95%, or at least 99% sequence identity to SEQ ID NO: 9.

113. The non-naturally occurring nucleic acid molecule of claim 111 or 112, wherein the 3’ UTR sequence comprises SEQ ID NO: 9.

114. The non-naturally occurring nucleic acid molecule of any one of claims 1-111, wherein the non-naturally occurring nucleic acid comprises a modified uridine.

115. The non-naturally occurring nucleic acid molecule of claim 102, wherein the modified uridine is a pseudouridine or a N1 -methylpseudouridine.

116. The non-naturally occurring nucleic acid molecule of claim 102 or 115, wherein the modified uridine is a pseudouridine.

117. The non-naturally occurring nucleic acid molecule of claim 102 or 115, wherein the modified uridine is a N1 -methylpseudouridine.

118. The non-naturally occurring nucleic acid molecule of any one of claims 1-117, wherein the non-naturally occurring nucleic acid comprises one or more modified uridines but not unmodified uridines.

119. The non-naturally occurring nucleic acid molecule of any one of claims 1-118, wherein the non-naturally occurring nucleic acid further comprises a cargo sequence.

120. The non-naturally occurring nucleic acid molecule of claim 119, wherein the cargo sequence is at least 500 nucleotides in length.

121. The non-naturally occurring nucleic acid molecule of claim 119 or 120, wherein the cargo sequence encodes a protein.

122. The nucleic acid construct of any one of claims 119-121, wherein the cargo sequence comprises a therapeutic gene or a diagnostic gene.

123. The non-naturally occurring nucleic acid molecule of any one of claims 1-122, further comprising a promoter sequence.

124. The non-naturally occurring nucleic acid molecule of claim 123, wherein the promoter is an RNAP II promoter or a T7 RNA polymerase promoter.

125. The non-naturally occurring nucleic acid molecule of any one of claims 1-124, further comprising a terminator sequence.

126. The non-naturally occurring nucleic acid molecule of any one of claims 1-125, further comprising a 5’ UTR sequence on a 5’ side of the cargo sequence.Attorney Docket No. 67098-712601127. The non-naturally occurring nucleic acid molecule of any one of claims 1-126, further comprising a stability element configured to protect the non-naturally occurring nucleic acid from degradation.

128. The nucleic acid construct of claim 127, wherein the stability element is selected from K3, K4, eK5, MALATl mm, MALATl hs, MALAT1 Jis-core, PAN KSHV, dENE os, HSL, U7, K9, K16, KEMCV, or KBDV1.

129. The non-naturally occurring nucleic acid molecule of any one of claims 1-126, further comprising a self-cleaving ribozyme motif.

130. The non-naturally occurring nucleic acid molecule of claim 129, wherein the selfcleaving ribozyme motif comprises a hepatitis delta virus fold.

131. The non-naturally occurring nucleic acid molecule of any one of claims 1-130, further comprising a 3’ tail region comprising an adenine-rich region.

132. The non-naturally occurring nucleic acid molecule of claim 131, wherein the adenine-rich region comprises a polyadenosine sequence.

133. The non-naturally occurring nucleic acid molecule of claim 132, wherein the polyadenosine sequence has a length of from 22 to 41 nucleotides.

134. The non-naturally occurring nucleic acid molecule of any one of claims 131-133, wherein the 3’ tail region comprises one or more non-adenine nucleobases.

135. The non-naturally occurring nucleic acid molecule of any one of claims 131-134, wherein the adenine-rich region has no more than 10%, 15%, 20%, or 25% non-adenine nucleobases.

136. The non-naturally occurring nucleic acid molecule of any one of claims 1-135, further comprising a complementarity region that is complementary to a region in a 28S rRNA gene.

137. The non-naturally occurring nucleic acid molecule of claim 136, wherein the complementarity region comprises a sequence of UAGC or TAGC.

138. A cell comprising the non-naturally occurring nucleic acid molecule of any one of claims 1-137.

139. The cell of claim 138, wherein the cell is a eukaryotic cell.

140. The cell of claim 139, wherein the cell is an epithelial cell.

141. A system for modifying a target nucleic acid in a cell of a subject, the system comprising:(a) the non-naturally occurring nucleic acid molecule of any one of claims 1-137; andAttorney Docket No. 67098-712601(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide, wherein the R2 retroelement polypeptide binds to the 3’ UTR sequence of the non-naturally occurring nucleic acid molecule.

142. The system of claim 141, wherein the R2 retroelement polypeptide is selected from the group consisting of a Geospiza fortis (GeFo) R2 retroelement polypeptide, a Zonotrichia albicollis (ZoAl) R2 retroelement polypeptide, a Taeniopygia guttata (TaGu) R2 retroelement polypeptide, and a Tinamus guttatus (TiGu) R2 retroelement polypeptide.

143. The system of claim 141 or 142, wherein the R2 retroelement polypeptide and the 3’ UTR sequence are derived from different species.

144. The system of claim 141 or 142, wherein the R2 retroelement polypeptide and the 3’ UTR sequence are derived from a same species.

145. The system of any one of claims 141-144, wherein the R2 retroelement polypeptide comprises a Geospiza fortis (GeFo) R2 retroelement polypeptide.

146. The system of any one of claims 141-144, wherein the R2 retroelement polypeptide comprises a Taeniopygia guttata (TaGu) R2 retroelement polypeptide.

147. The system of any one of claims 141-145, further comprising a target nucleic acid molecule.

148. The system of claim 146, wherein the target nucleic acid molecule comprises DNA.

149. The system of claim 146 or 148, wherein the target nucleic acid molecule comprises genomic DNA.

150. The system of any one of claims 146-149, wherein the target nucleic acid molecule comprises a rDNA sequence.

151. The system of any one of claims 146-150, wherein the target nucleic acid molecule comprises a 28S rRNA sequence.

152. A cell comprising the system of any one of claims 141-151.

153. The cell of claim 152, wherein the cell is a eukaryotic cell.

154. The cell of claim 153, wherein the cell is an epithelial cell.

155. A method of modifying a target nucleic acid in a cell of a subject, the method comprising providing the cell with:(a) the non-naturally occurring nucleic acid molecule of any one of any one of claims 1-137; and(b) an R2 retroelement polypeptide or a second nucleic acid encoding the R2 retroelement polypeptide;Attorney Docket No. 67098-712601 wherein the R2 retroelement polypeptide inserts a cargo nucleic acid sequence of the non- naturally occurring nucleic acid molecule, or a reverse complement thereof, into the target nucleic acid.

156. The method of claim 155, wherein the R2 retroelement polypeptide is selected from the group consisting of a Geospiza fortis (GeFo) R2 retroelement polypeptide, a Zonotrichia albicollis (ZoAl) R2 retroelement polypeptide, a Taeniopygia guttata (TaGu) R2 retroelement polypeptide, and a Tinamus guttatus (TiGu) R2 retroelement polypeptide.

157. The method of claim 155 or 156, wherein the R2 retroelement polypeptide binds to the 3’ UTR sequence of the non-naturally occurring nucleic acid molecule.

158. The method of any one of claims 155-157, wherein the target nucleic acid is in a cell.

159. The method of claim 158, wherein the cell is a eukaryotic cell.

160. The method of claim 159, wherein the cell is an epithelial cell.

161. The method of any one of claims 155-160, wherein the R2 retroelement polypeptide inserts the cargo nucleic acid sequence or the reverse complement thereof into a rDNA sequence of the target nucleic acid.

162. The method of any one of claims 155-161, wherein the R2 retroelement polypeptide inserts the cargo nucleic acid sequence or the reverse complement thereof into a 28S rDNA sequence of the target nucleic acid molecule.