RT editing compositions and methods
Patent Information
- Application Number
- PCT/IB2025/052079
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-11
- Filing Date
- 2025-02-26
- Publication Date
- 2025-10-09
AI Technical Summary
Existing gene editing approaches suffer from insufficient expression levels, inadequate efficacy, and lack of specificity, making them ineffective for therapeutic applications and in vivo use.
A reverse transcriptase (RT)-based editing system comprising a fusion protein with a Cas9 nickase and a template armed guide RNA (tagRNA) for precise nucleotide editing, including a DNA binding domain, endonuclease domain, and polymerase domain, along with an enhancer guide RNA for enhanced editing efficiency and specificity.
The RT-based editing system achieves greater than 30-80% nucleotide edit incorporation rates, providing improved strength, mechanistic efficacy, and duration of gene editing component expression for therapeutic applications.
Smart Images

Figure IB2025052079_09102025_PF_FP_ABST
Abstract
Description
80EM-341785-WO / CT229-PCT1 PATENT RT EDITING COMPOSITIONS AND METHODS RELATED APPLICATIONS
[0001] The present application claims priority to U.S. Provisional Application No. 63 / 558,338, filed February 27, 2024; U.S. Provisional Application No.63 / 649,882, filed May 20, 2024; and U.S. Provisional Application No. 63 / 718,994, filed November 11, 2024. The entire contents of these applications are hereby expressly incorporated by reference in their entireties. REFERENCE TO SEQUENCE LISTING
[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 80EM-341785- WO_SeqListing, created February 25, 2025, which is 10,377 kilobytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety. BACKGROUND Field
[0003] The present disclosure relates generally to the field of gene editing. Description of the Related Art
[0004] Many diseases and disorders have a genetic component, including those that involve pathogenic single nucleotide mutations. There is a need for nucleic acid editing compositions and methods with greater editing efficiency and specificity. Existing gene editing approaches wherein components are engineered to be expressed within target cells suffer from insufficient levels of expression, or inadequate efficacy in their action, to be therapeutically effective; or lack specificity in directing a precise editing outcome. There is a need for compositions and methods with improved strength, mechanistic efficacy, and duration of gene editing component expression and employment for in vivo applications. SUMMARY
[0005] Disclosed herein are reverse transcriptase (RT)-based editing systems. In some embodiments, the RT editing system comprises: a fusion protein comprising a Cas9 nickase and a reverse transcriptase, and a template armed guide RNA (tagRNA) comprising from 5’ to 3’ a spacer sequence, a scaffold sequence, an editing template and a flap binding sequence. In some embodiments, the RT editing system further comprises an enhancer guide RNA (egRNA).
[0006] Disclosed herein are reverse transcriptase (RT) editors. In some embodiments, the RT editor comprises: a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, wherein the DNA polymerase domain and the DNA endonuclease domain are fused or linked to form a fusion protein, wherein the DNA polymerase domain comprises areverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, or at least 95%, or 100% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, and 315-319, 899- 909, and 1015-1018, and wherein the DNA binding domain, DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80% identical, at least 85% identical, at least 90% identical, or at least 95%, or 100% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314.
[0007] Disclosed herein include template armed guide RNAs (tagRNAs). In some embodiments, the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a double stranded target DNA; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the double stranded target DNA; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, and wherein the first strand and the second strand of the double-stranded target DNA are complementary to each other.
[0008] In some embodiments, the tagRNA comprises a flap binding sequence at least partially complementary to the spacer. In some embodiments, the scaffold sequence is between the spacer and the editing template. In some embodiments, the tagRNA comprises from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the flap binding sequence. In some embodiments, the spacer, the scaffold sequence, the editing template, and the flap binding sequence form a contiguous sequence in a single molecule. In some embodiments, the editing template comprises an intended nucleotide edit compared to the double stranded target DNA. In some embodiments, the editing template comprises several intended nucleotide edits compared to the double stranded target DNA. In some embodiments, the tagRNA guides the RT editor to incorporate the intended nucleotide edit or edits into the double stranded target DNA when the tagRNA is contacted with the double stranded target DNA. In some embodiments, the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA. In some embodiments, the search target sequence is complementary to a protospacer sequence in the double stranded target DNA, and wherein the protospacer sequence is adjacent to a protospaceradjacent motif (PAM) in the double stranded target DNA. In some embodiments, the tagRNA results in incorporation of a nucleotide edit or nucleotide edits in the PAM when contacted with the double stranded target DNA. In some embodiments, the tagRNA results in incorporation of a nucleotide edit or nucleotide edits outside the PAM when contacted with the double-stranded target DNA. In some embodiments, the spacer of the tagRNA is from 16 to 25 nucleotides in length, optionally 20 nucleotides in length or 21-23 nucleotides in length. In some embodiments, the flap binding sequence is about 2 to 20 nucleotides in length, optionally about 8 to 16 nucleotides in length or 6 nucleotides in length. In some embodiments, the editing template is about 4 to 30 nucleotides in length, optionally about 10 to 30 nucleotides in length, further optionally 6 to 9 nucleotides in length.
[0009] In some embodiments, the tagRNA results in incorporation of the intended nucleotide edit about 0 to 30 base pairs downstream of the nickase cleavage site. In some embodiments, the intended nucleotide edit comprises a single nucleotide substitution compared to the region corresponding to the editing target sequence in the double stranded target DNA, optionally, the single nucleotide substitution is a T>G substitution or a T>C substitution. In some embodiments, the intended nucleotide edit comprises more than one nucleotide substitution compared to the region corresponding to the editing target sequence. In some embodiments, the intended nucleotide edit or edits comprises an insertion compared to the region corresponding to the editing target in the double stranded target DNA (e.g., an insertion of a nucleotide sequence at least 50, at least 45, at least 40, at least 35, at least 30, at least 25, at least 20, at least 15, at least 10, or at least 5, nucleotides in length). In some embodiments, the intended nucleotide edit or edits comprises a deletion compared to the region corresponding to the editing target in the double stranded target DNA. In some embodiments, the editing template comprises one or more silent nucleotide edits compared to the region corresponding to the editing target in the double stranded target DNA. In some embodiments, the editing template comprises a wild type DNA sequence. In some embodiments, the tagRNA results in correction of a mutation when contacted with the double stranded target DNA. In some embodiments, the tagRNA comprises any one of the sequences of SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of the sequences of SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203. In some embodiments, the tagRNA further comprises: 3’ mN*mN*mN*N and 5’ mN*mN*mN* modifications, where m indicates that the nucleotide contains a 2’-O-Me modification and a * indicates the presence of a phosphorothioate bond; and / or a structural motif at the 3’ terminus selected from the group consisting of: a prequeosine1-1 riboswitch aptamer (evopreQ1) and variants thereof, aframeshifting pseudoknot from Moloney murine leukemia virus (MMLV) (mpknot), G- quadruplexes, hairpin structures, xrRNA, and a P4-P6 domain of the group I intron; optionally, the structural motif is evopreQ1 or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 84-90. In some embodiments, a chemically modified 3’-terminus comprises inverted-dT. In some embodiments, the tagRNA further comprises: 3’ mN*mN*mN*N, 5’ mN*mN*mN*, and / or 3’ inverted-dT modifications.
[0010] Disclosed herein are reverse transcriptase (RT) editing systems. In some embodiments, the system comprises: a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; or a nucleic acid encoding the RT editor, wherein the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80%, at least about 85%, at least 90%, at least 95%, or 100% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, and wherein the DNA DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80%, at least about 85%, at least 90%, at least 95%, or 100% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314.
[0011] Disclosed herein include reverse transcriptase (RT) editing systems. In some embodiments, the system comprises: any of the tagRNAs, or a nucleic acid encoding the tagRNA, described herein; and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor.
[0012] In some embodiments, RT editing system further comprises: an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises a egRNA spacer that is complementary to a second search target sequence in the double stranded target DNA. In some embodiments, the second search target sequence is on the second strand of the double stranded target DNA. In some embodiments, the egRNA spacer is from 16 to 25 nucleotides in length, optionally 21-23 nucleotides in length, optionally 20 nucleotides in length.In some embodiments, the egRNA comprises a scaffold sequence. In some embodiments, the intended nucleotide edit incorporation rate of the RT editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80%.
[0013] Disclosed herein include reverse transcriptase (RT) editing complexes. In some embodiments, the RT editing complex comprises: template armed guide RNA (tagRNA) comprising: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; or a nucleic acid encoding the RT editor, wherein the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80%, at least about 85%, at least 90%, at least 95%, or 100% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015- 1018, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250- 269, 315-319, 899-909, and 1015-1018, and wherein the DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80%, at least about 85%, at least 90%, at least 95%, or 100% identical to any one of SEQ ID NOs: 142- 163 and 270-314, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314.
[0014] In some embodiments, the RT editing complex comprises: (i) any of the tagRNAs disclosed herein and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; or (ii) any of the RT editing systems disclosed herein. In some embodiments, the intended nucleotide edit incorporation rate of the RT editing complex is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80%.
[0015] In some embodiments, the DNA binding domain, a DNA endonuclease domain is a CRISPR associated (Cas) protein domain. In some embodiments, the Cas protein domain has nickase activity. In some embodiments, the Cas protein domain is a Cas9. In some embodiments, the Cas9 comprises a mutation in an HNH domain. In some embodiments, the Cas9 comprises an H840A mutation in the HNH domain. In some embodiments, the Cas protein domain is a Cas12b. In some embodiments, the Cas protein domain is a Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, Casφ, Ssch1Cas9, Sro1Cas9, Sha4Cas9, SsuCas9, iSpyMacCas9, Ssi5Cas9, Ssi8Cas9, Ssci4Cas9, Shy1Cas9,Sag3Cas9, Slutr1Cas9, Ssch3Cas9, SpRYCas9, SpRYcCas9, Sma2Cas9, SsaCas9, EvoCjCas9, or iSpyMac. In some embodiments, the DNA polymerase domain is a reverse transcriptase, optionally the editing template is a reverse transcription template. In some embodiments, the reverse transcriptase is a retrovirus reverse transcriptase. In some embodiments, the reverse transcriptase is a Moloney murine leukemia virus (MMLV) reverse transcriptase. In some embodiments, the DNA polymerase domain and the DNA binding domain, a DNA endonuclease domain are fused or linked to form a fusion protein. In some embodiments, the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80%, at least about 85%, at least 90%, at least 95%, or 100% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018; and / or the DNA binding domain, a DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80%, at least about 85%, at least 90%, at least 95%, or 100% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 142- 163 and 270-314.
[0016] In some embodiments, the editing template comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ Fluoro RNA base(s). In some embodiments, the FBS comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, or 82’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, or 82’ Fluoro RNA base(s).
[0017] In some embodiments, the scaffold sequence comprises a nucleotide sequence that is at least 80% identical (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to any one of SEQ ID NOs: 683-718, 894-898, and 964-1014 (e.g., a nucleotide sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 683-718, 894-898, and 964-1014). In some embodiments, the scaffold sequence comprises one or more nucleotide substitutions, insertions, and / or deletions at one or more nucleotide positions relative to the parental scaffold sequence of SEQ ID NO: 683, such as, for example: (i) dinucleotide substitutions at nucleotide positions 5-6 and 34-35, optionally configured such that the nucleotide sequence at positions 5-6 is complementary to the nucleotide sequence at positions 34-35; (ii) a deletion at nucleotide positions 55-87, 57-87, 59-87, 61-87, or 63-87; (iii) a substitution at one or more of nucleotide positions 17, 18, 19, and 20, optionally aUUCG substitution at nucleotide positions 17-20; (iv) a deletion at nucleotide positions 14-16 and 21-24; (v) a replacement of nucleotides at nucleotide positions 11-28 with GUUCGC; and / or (vi) a deletion at nucleotide positions 12-16 and 21-26.
[0018] In some embodiments, the scaffold sequence comprises one or more nucleotide substitutions, insertions, and / or deletions at one or more nucleotide positions relative to the parental scaffold sequence of SEQ ID NO: 700, such as, for example, a substitution at one or more of nucleotide positions 13, 18, 21, 40, 50, and 52-53 (e.g., a U-to-C substitution at nucleotide position 13; an A-to-G substitution at nucleotide position 18; an A-to-C or an A-to-U substitution at nucleotide position 21; a C-to-U or a C-to-A or a C-to-G substitution at nucleotide position 40; a U-to-C or a U-to-A or a U-to-G substitution at nucleotide position 50; an A-to-C or an A-to-U or an A-to-G substitution at nucleotide position 52; and / or an A-to-C or an A-to-U or an A-to-G substitution at nucleotide position 53). In some embodiments, the scaffold sequence comprises one or more modifications (e.g., nucleoside modification(s), sugar modification(s), modified internucleoside linkage(s), and / or backbone modification(s)) relative to the parental scaffold sequence of SEQ ID NO: 700, such as, for example, a 2’ O-methyl modification at one or more of nucleotide positions 5-9, 12-21, 29, 37-38, 40, 50, and 52-53. In some embodiments, the scaffold sequence comprises any one of the sequences of SEQ ID NOs: 894-898 and 964-1014, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 894-898 and 964-1014. In some embodiments, the scaffold sequence of the tagRNA, the scaffold sequence of the egRNA, or both, comprises the sequence of any one of SEQ ID NOs: 1210-1266 or a sequence that exhibits at least about 85% identity to any one of SEQ ID NOs: 1210-1266.
[0019] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase. In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising one or more mutations. In some embodiments, at least one of the one or more mutations is at an amino acid position functionally equivalent to V101, N200, A208, G248, P330, L435, K445, and / or A623 relative to SEQ ID NO: 8.
[0020] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 255. In some embodiments, the reverse transcriptase comprises one or more mutations. In some embodiments, at least one of the one or more mutations is at amino acid position V98, N197, A205, S245, P327, V432, R442, and / or A623 relative to SEQ ID NO: 255. In some embodiments, the reverse transcriptase comprises one or more transition mutations selected from the group consisting of V98R, N197C, N197D, N197G, A205T, S245C, P327E, P327Q, V432K, R442T, and A623F relative to SEQ ID NO: 255.
[0021] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 260. The reverse transcriptase can comprise one or more mutations. In some embodiments, at least one of the one or more mutations is at amino acid position V106, N204, E212, G252, P334, L440, E450, and / or A629 relative to SEQ ID NO: 260. In some embodiments, the reverse transcriptase comprises one or more transition mutations selected from the group consisting of V106R, N204C, N204D, N204G, E212T, G252C, P334E, P334Q, L440K, E450T, and / or A629F relative to SEQ ID NO: 260.
[0022] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 255. In some embodiments, the reverse transcriptase comprises an N197C or V98R transition mutation relative to SEQ ID NO: 255.
[0023] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 260. In some embodiments, the reverse transcriptase comprises an N204C or V106R transition mutation relative to SEQ ID NO: 260.
[0024] In some embodiments, the RT editor further comprises an accessory domain. In some embodiments, the accessory domain comprises a single-strand binding (SSB) protein domain or a stabilon. In some embodiments, the SSB protein domain is derived from RecA protein, Sso7d protein, or Sto7d protein. In some embodiments, the accessory domain is situated at the N-terminus, the C-terminus, or at an internal location of the RT editor. In some embodiments, the RT editor comprising an accessory domain comprises an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1276-1298. In some embodiments, the SSB protein domain is derived from Sso7d and comprises one or more mutations at an amino acid position functionally equivalent to K12 and / or E35 of a wild type Sso7d amino acid sequence. In some embodiments, the one or more mutations comprise K12L and / or E35L relative to a wild type Sso7d sequence. In some embodiments, the SSB protein domain comprises an amino acid sequence that is at least 85% identical to SEQ ID NO: 1205. In some embodiments, the RT editor comprising the SSB protein domain comprises from N-terminus to C-terminus: [N-terminus of nCas9]-[linker]-[SSB protein domain]-[linker]-[RT]-[linker]-[C-terminus of nCas9]. In some embodiments, the N-terminus of nCas9 comprises an amino acid sequence functionally equivalent to amino acids 1-1247 of SEQ ID NO: 6 and / or the C-terminus of nCas9 comprises an amino acid sequence functionally equivalent to amino acids 1248-1367 of SEQ ID NO: 6. In some embodiments, the RT editor comprising the SSB protein domain comprises an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1207-1209.
[0025] Disclosed herein include ribonucleoprotein (RNP) complexes. In some embodiments, the RNP complex comprises any of the RT editing complexes described herein, or a component thereof.
[0026] Disclosed herein include lipid nanoparticles (LNPs). In some embodiments, the LNP comprises any of the RT editing systems described herein, or a component thereof. In some embodiments, the LNP comprises the tagRNA and the nucleic acid encoding the RT editor. In some embodiments, the nucleic acid encoding the RT editor is mRNA. In some embodiments, the LNP comprises a egRNA.
[0027] Disclosed herein include polynucleotides. In some embodiments, the polynucleotide encodes any of the RT editors, the tagRNAs, the RT editing systems, or the RT editing complexes described herein. In some embodiments, the polynucleotide is an mRNA. In some embodiments, the polynucleotide is operably linked to a regulatory element, optionally the regulatory element is an inducible regulatory element.
[0028] Disclosed herein include vectors. In some embodiments, the vector comprises any of the polynucleotides described herein. In some embodiments, the vector is an AAV vector.
[0029] Also provided herein are isolated cells. In some embodiments, the isolated cell comprises an RT editor, a tagRNA, an RT editing system, an RT editing complex, an RNP, an LNP, a polynucleotide, or a vector of the disclosure. In some embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is from a subject having a disease or disorder, optionally Wilson’s disease, further optionally the subject is a human. In some embodiments, the cell is from a subject having a disease or disorder, optionally alpha1 antitrypsin deficiency disease, further optionally the subject is a human. The disease or disorder can be selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof. In some embodiments, the cell is from a subject having a disease or disorder, optionally phenylketonuria or hyperphenylalaninemia, further optionally the subject is a human.
[0030] Disclosed herein include pharmaceutical compositions. In some embodiments, the pharmaceutical composition comprises: (i) a RT editor, a tagRNA, a RT editing system, a RT editing complex, a RNP, a LNP, a polynucleotide, a vector, or a cell of the disclosure; and (ii) a pharmaceutically acceptable carrier.
[0031] Disclosed herein include methods for editing a double stranded target DNA. In some embodiments, the method comprises contacting the double stranded target DNA with (i) any of the tagRNAs of the disclosure and a reverse transcriptase (RT) editor comprising a DNAbinding domain, a DNA endonuclease domain and a DNA polymerase domain or (ii) any of the RT editing systems described herein, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, thereby editing the double stranded target DNA.
[0032] Disclosed herein include methods for editing a double stranded target DNA. In some embodiments, the method comprises contacting the double stranded target DNA with any of the RT editing complexes of the disclosure, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, thereby editing the double stranded target DNA.
[0033] In some embodiments, the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit or edits into a region corresponding to the editing target in the double stranded target DNA. In some embodiments, the double stranded target DNA is in a cell. In some embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is a stem cell; optionally, an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell. In some embodiments, the cell is in a subject, optionally the subject is a human. In some embodiments, the cell is from a subject having a disease or disorder, optionally Wilson’s disease. In some embodiments, the cell is from a subject having a disease or disorder, optionally alpha1 antitrypsin deficiency disease. In some embodiments, the cell is from a subject having a disease or disorder, optionally phenylketonuria or hyperphenylalaninemia. The disease or disorder can selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof. In some embodiments, the method further comprises administering the cell to the subject after incorporation of the intended nucleotide edit.
[0034] Disclosed herein include cells generated by any of the methods described herein. Disclosed herein include populations of cells generated by any of the methods disclosed herein.
[0035] Disclosed herein include methods for treating or preventing a disease or disorder in a subject in need thereof. In some embodiments, the method comprises administering to the subject any of the RT editing systems of the disclosure, wherein the editing template comprises an intended nucleotide edit compared to a double stranded target DNA of the subject, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the doublestranded target DNA, and wherein incorporation of the intended nucleotide edit corrects a mutation in the double stranded target DNA associated with the disease or disorder, thereby treating or preventing the disease or disorder in the subject. The disease or disorder can selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.
[0036] Disclosed herein include methods for treating or preventing a disease or disorder in a subject in need thereof. In some embodiments, the method comprises administering to the subject any of the RT editing complexes, the RNPs, the LNPs, or the pharmaceutical compositions of the disclosure, wherein the editing template comprises an intended nucleotide edit compared to a double stranded target DNA of the subject, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, and wherein incorporation of the intended nucleotide edit corrects a mutation in the double stranded target DNA associated with the disease or disorder, thereby treating or preventing the disease or disorder in the subject. The disease or disorder can selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.
[0037] Disclosed herein include messenger RNAs (mRNAs). In some embodiments, the mRNA encodes a reverse transcriptase (RT) editor. In some embodiments, the mRNA encodes any of the RT editors of the disclosure. In some embodiments, the mRNA comprises one or more of a 5'-cap structure, a 5’-UTR, a 3’-UTR, and a nuclear localization sequence (NLS). In some embodiments, the mRNA comprises (i) a 5′-cap, (ii) a 5’-untranslated region (UTR); (iii) an open reading frame (ORF) comprising a nucleotide sequence that encodes the RT editor; and (iv) a 3′ untranslated region (UTR). In some embodiments, the mRNA has a structure comprising or consisting of 5’ - [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]- [polyA sequence] - 3’. In some embodiments, one or more nucleosides of the mRNA sequence are chemically modified. In some embodiments, the mRNA comprises one or more additional segmented polyA sequences configured to increase mRNA stability and / or half-life. In some embodiments, the mRNA comprises one or more viral element sequences, optionally 3’ of the 3’- UTR. In some embodiments, the viral element comprises a woodchuck hepatitis virus post- transcriptional regulatory element (WPRE) sequence, optionally comprising the sequence of SEQ ID NO: 189, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to SEQ ID NO: 189. In some embodiments, the viral element comprises a eK5 sequence, optionally located 3’ of the 3’-UTR, optionally comprising thesequence of SEQ ID NO: 190, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to SEQ ID NO: 190. In some embodiments, the mRNA comprises one or more secondary structure motifs. In some embodiments, a secondary structure motif comprises a triple helix sequence, optionally a synthetic triple helix (STH) sequence. In some embodiments, the STH sequence comprises the sequence derived from a sequence element of a long non-coding RNA, optionally MALAT1. In some embodiments, the STH sequence comprises any one of the sequences of SEQ ID NOs: 191-192 and 245-249, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of the sequences of SEQ ID NOs: 191-192 and 245-249.
[0038] In some embodiments, the mRNA comprises: a polyA sequence comprising any one of the sequences of SEQ ID NOs: 182-187, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of the sequences of SEQ ID NOs: 182-187; one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 188-190, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of the sequences of SEQ ID NOs: 188-190; one or more 3’ UTR structural motifs selected from the group comprising any one of the sequences of SEQ ID NOs: 191-192, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of the sequences of SEQ ID NOs: 191-192; a 5’UTR sequence comprising the sequence of any one of SEQ ID NOS: 193 and 910-933, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of SEQ ID NOS: 193 and 910-933; and / or an RT editor sequence comprising the sequence of SEQ ID NO: 194, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to SEQ ID NO: 194. In some embodiments, the mRNA comprises a 3’ UTR comprising any one of the sequences of SEQ ID NOs: 195-219, or a sequence that exhibits at least about 80%, at least about 85%, at least 90%, at least 95%, or 100% identity to any one of the sequences of SEQ ID NOs: 195-219. The mRNA can comprise one or more 5’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 940-950 and 1075, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 940-950 and 1075. The mRNA can comprise one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 951-961 and 1076-1080, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 951-961 and 1076-1080. The mRNA can comprise one or more additional elements selected from the group comprising any one of the sequences of SEQ ID NOs: 962-963, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 962-963, and in someembodiments said one or more additional elements are situated downstream of the 3’UTR and / or the polyA tail. In some embodiments, the mRNA comprises one or more 5’-UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 1081- 1082, wherein the second codon of the mRNA is a “gcc”. In some embodiments, the mRNA is codon optimized for expression in human cells. In some embodiments, the mRNA (e.g., codon- optimized) comprises the sequence of any one of SEQ ID NOs: 1086-1088. Disclosed herein include ribonucleoprotein (RNP) complexes, lipid nanoparticles (LNPs), vectors, isolated cells, and pharmaceutical compositions comprising any of the mRNA of the disclosure. Disclosed herein include polynucleotides encoding the mRNAs provided herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] FIG. 1 displays ATP7B genomic and protein sequences near the location of amino acid position 1069. Also shown are exemplary locations of spacer sequences of the disclosure.
[0040] FIG. 2 displays non-limiting exemplary efficiencies for correction of the H1069Q mutation as assayed in HEK293T cells that contain the H1069Q mutation (H1069Q HEK293T). In this experiment, editing components were delivered by plasmid transfection.
[0041] FIG.3A-FIG.3C display editing efficiencies in H1069Q HEK293T cells using mRNA (encoding the RT editor) and synthetic gRNA.
[0042] FIG. 4A-FIG. 4C display editing efficiencies in H1069Q Huh7 cells using mRNA (encoding the RT editor) and synthetic gRNA.
[0043] FIG. 5 displays non-limiting exemplary results of a spacer screen in the indicated cell types containing the H1069Q mutation. Cells were transfected with Cas9 mRNA and a synthetic gRNA.
[0044] FIG.6 depicts an exemplary 3D structure of an RNA comprising SEQ ID NO: 245 (Comp14 transcript). FIG. 6 is original Figure 3A from Wilusz, Jeremy E., et al. “A triple helix stabilizes the 3’ ends of long noncoding RNAs that lack poly(A) tails.” GenesDev 26 (2012): 2392-2407.
[0045] FIG. 7 depicts exemplary 3D structure of an RNA comprising SEQ ID NO: 247, 248, or 249. FIG. 7 is original Fig. 1A from Brown, Jessica A., et al. “Formation of triple- helical structures by the 3’-end sequences of MALAT1 and MENβ noncoding RNAs.” PNAS 109 (2012): 19202-19207.
[0046] FIG.8 depicts an exemplary structure of an RNA comprising SEQ ID NO: 245 (predicted secondary structure of the 3’ end of the mature Comp.14 transcript).
[0047] FIG.9 depicts an exemplary structure of an RNA comprising SEQ ID NO: 246 (full-length segment from human MALAT1 (HTH)).
[0048] FIG. 10 depicts an exemplary structure of an RNA comprising SEQ ID NO: 191 (synthetic triple helix, trimmed, adaptation, based on synthetic mix MALAT1 (STH)).
[0049] FIG.11 depicts exemplary modified nucleotides that can be used in the RNAs (e.g., tagRNAs) of the disclosure.
[0050] FIG. 12 depicts a non-limiting exemplary structure of a scaffold RNA of the disclosure. Shown are various regions of the scaffold that were modified (See, e.g., Table 14A- Table 14C and Table 23).
[0051] FIG.13 depicts a non-limiting exemplary schematic of a tagRNA according to the methods and compositions of the disclosure.
[0052] FIG.14 depicts a non-limiting exemplary flow diagram for improvement of RT editing methods of the disclosure.
[0053] FIG. 15 depicts a non-limiting exemplary secondary structure of a tagRNA scaffold, with regions for optimization outlined by a box.
[0054] FIG.16 depicts non-limiting exemplary data related to sequence optimization of the tagRNA scaffold. Shown are editing efficiencies of tested tagRNA scaffold sequence variants.
[0055] FIG.17A-FIG.17B depict non-limiting exemplary data related to an ET-FBS screen with a SsaCas9-based RT editor in Huh7 cells. Shown are the % T>C correction data (FIG. 17A) and percent indel formation data (FIG. 17B) for flap binding sequence (FBS) and editing template (ET) length variants tested.
[0056] FIG.18A-FIG.18B depict non-limiting exemplary data related to an ET-FBS screen with a truncated scaffold and shorter lengths. Shown are editing data for the variants tested with a dose of 125 ng mRNA and either 0.25 pmol tagRNA (FIG. 18A) or 0.07 pmol tagRNA (FIG.18B).
[0057] FIG. 19 depicts non-limiting exemplary data related to optimization of the tagRNA spacer length.
[0058] FIG.20 depicts non-limiting exemplary data related to editing results derived chemically modified scaffolds.
[0059] FIG.21 depicts a non-limiting exemplary phylogenetic tree of MMLV, bat, and avian reverse transcriptases.
[0060] FIG.22 depicts non-limiting exemplary data of editing efficiencies of RTs of the disclosure in correction of ATP7B mutation in Huh7 cells.
[0061] FIG.23 depicts non-limiting exemplary data of editing efficiencies of RTs of the disclosure in ATP7B surrogate mutation correction in primary human hepatocytes.
[0062] FIG. 24A-FIG. 24B depict non-limiting exemplary data related to editing efficiencies of M. brandtii variants. In FIG. 24B, lower concentrations of all RNA components were used as compared to FIG. 24A. The experimental data of FIG. 24B employed improved chemical modifications in the tagRNA as compared to FIG.24A. P1_30 is also referred to herein as ChemMod_Set2_30.
[0063] FIG. 25A-FIG. 25B depict non-limiting exemplary data related to editing efficiencies of M. georgiana variants. In FIG.25B, lower concentrations of all RNA components were used as compared to FIG. 25A. The tagRNA of FIG. 25B used improved chemical modifications in the tagRNA as compared to FIG.25A.
[0064] FIG.26A-FIG.26B depict non-limiting exemplary data related to M. brandtii variants (FIG.26A) and M. georgiana variants (FIG.26B) in primary human hepatocytes.
[0065] FIG. 27 depicts non-limiting exemplary data related to improved editing exhibited by different RT variants.
[0066] FIG. 28 depicts non-limiting exemplary data related to editing efficiencies using different 5’ UTRs in the mRNA encoding the RT.
[0067] FIG. 29A-FIG. 29B depicts non-limiting exemplary data related to editing results in Groups 1, 2, and 4 (FIG.29A) and Group 3 (FIG.29B) of a first in vivo proof of concept study performed in a humanized AATD mouse model.
[0068] FIG. 30 depicts non-limiting exemplary data related to editing results in a second in vivo proof of concept study performed in a humanized AATD mouse model.
[0069] FIG.31 depicts a non-limiting exemplary schematic of mouse exon 12 being replaced with the human exon 12 in the humanized PKU mouse model generated.
[0070] FIG.32 depicts non-limiting exemplary data related to editing results in an in vivo proof of concept study performed in humanized PKU mice.
[0071] FIG. 33 depicts editing frequencies of codon-optimized RT editors for correction of ATP7B surrogate mutation in primary human hepatocytes.
[0072] FIG. 34 depicts editing frequencies of codon-optimized RT editors for correction of PAH surrogate mutation in primary human hepatocytes.
[0073] FIG.35 displays an exemplary schematic of SSB insertion sites within an RT editor.
[0074] FIG.36 depicts an exemplary schematic and 3D model for insertion of an SSB as a chimeric protein with nSsaCas9.
[0075] FIG.37 displays a bar graph of exemplary editing data using the indicated RT editor variants.
[0076] FIG.38 shows editing results from a stabilion variant screen.
[0077] FIG.39 displays results of in vivo screening of indicated mRNAs encoding an RT editor.
[0078] FIG.40 displays an exemplary schematic of an SSB (e.g., Sso7d) insertion sites within an RT editor.
[0079] FIG.41 displays exemplary editing data using the indicated RT editors.
[0080] FIG. 42 displays editing data for correction of a Wilson’s disease mutation from primary human hepatocytes using the indicated mRNA, for the purpose of testing different linkers, NLSs, UTRs, RTs, stabilizing elements, and polyA tails.
[0081] FIG.43 displays editing data from tested in Huh7 cells containing the ATP7B H1069Q mutation using the indicated chemically modified tagRNA.
[0082] FIG.44 displays editing data from tested in Huh7 cells containing the ATP7B H1069Q mutation using the indicated chemically modified tagRNA.
[0083] FIG.45 displays editing data from tested in Huh7 cells containing the ATP7B H1069Q mutation using the indicated chemically modified tagRNA.
[0084] FIG.46 displays exemplary in vivo data from mouse for correction of human ATP7B (H1069Q) using the indicated RT editor mRNA.
[0085] FIG.47 displays exemplary in vivo data from mouse for correction of human ATP7B mutation using the indicated mRNA encoding RT editor and tagRNAs.
[0086] FIG. 48 displays in vivo data from mouse for correction of PAH-R408W mutation using the indicated mRNA encoding RT editor, tagRNA, and egRNA. Shown from left to right on the graph are data using: pTPRT-U-493, pTPRT-U-540, pTPRT-U-543. pTPRT-U- 554, pTPRT-U-558, pVC313, pVC268, and pAM198.
[0087] FIG.49 displays exemplary editing data using the indicated tagRNA scaffold.
[0088] FIG.50 displays exemplary editing data using the indicated tagRNA scaffold.
[0089] FIG.51 displays exemplary editing data using the indicated egRNA scaffold. DETAILED DESCRIPTION
[0090] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.
[0091] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.
[0092] Disclosed herein include compositions for reverse transcriptase (RT) editing. Compositions disclosed herein comprise an RT editing system, comprising an RT editor, template armed guide RNA (tagRNAs). In some embodiments, the RT editing system comprises an additional gRNA designated herein as ‘egRNA’.
[0093] Also provided herein are ribonucleoprotein (RNP) complexes comprising any of the RT editing complexes disclosed herein, or a component thereof. Disclosed herein include lipid nanoparticles (LNPs). In some embodiments, the LNP comprises any of the RT editing systems disclosed herein, or a component thereof. Provided herein include polynucleotides. In some embodiments the polynucleotide encodes any of the tagRNAs, the RT editor, the egRNA, or the RT editing complexes of the disclosure. Disclosed herein include isolated cells. In some embodiments, the isolated cell comprises any of the tagRNAs, the RT editor, the egRNA, the RT editing complexes, the RNPs, the LNPs, or the vectors provided herein.
[0094] Disclosed herein include compositions. In some embodiments, the composition is a pharmaceutical composition. In some embodiments, the pharmaceutical composition comprises: any of the tagRNAs, the RT editor, the egRNA, the RT editing systems, the RT editing RT editing complexes, the RNPs, the LNPs, the polynucleotides, the vectors, or the cells disclosed herein; and (ii) a pharmaceutically acceptable carrier.
[0095] Disclosed herein include methods for editing a mutated gene. In some embodiments, the method comprises contacting the gene at the site of mutation with (i) any of the tagRNAs of the disclosure and a RT editor. In some embodiments, the tagRNA directs the RT editor to incorporate the intended nucleotide edit or edits in the mutated gene, thereby editing the mutated gene and correcting the mutation.
[0096] Disclosed herein include methods for treating a disease or disorder in a subject in need thereof. In some embodiments, the method comprises administering to the subject (i) any of the tagRNAs of the disclosure and a RT editor. In some embodiments, the tagRNA directs the RT editor to incorporate the intended nucleotide edit or edits in the mutated gene in the subject, thereby treating the disease or disorder in the subject.
[0097] Disclosed herein include methods for editing an ATP7B gene. In some embodiments, the method comprises contacting the ATP7B gene with (i) any of the tagRNAs of the disclosure and a RT editor. In some embodiments, the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the ATP7B gene, thereby editing the ATP7B gene.
[0098] Disclosed herein include methods for treating Wilson’s disease in a subject in need thereof. In some embodiments, the method comprises administering to the subject (i) any of the tagRNAs of the disclosure and a RT editor. In some embodiments, the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the ATP7B gene in the subject, thereby treating Wilson’s disease in the subject.
[0099] Provided herein are methods and compositions of editing of the SERPINA1 gene that encodes A1AT serine protease inhibitor for treating alpha1 antitrypsin deficiency (AATD) disease. Editing of the S allele mutation or the Z allele mutation in SERPINA1 gene are contemplated herein. The Z allele mutation on exon 5 of SERPINA1 is an E342K mutation. E342K mutation leads to misfolding of the AAT protein leading to polymers and liver damage. Contemplated are methods for a single nucleotide (T -> C) correction leading to a K342E correction at the SERPINA1 gene.
[0100] Provided herein are methods and compositions of editing of the PAH gene that encodes Phenylalanine hydroxylase enzyme for treating phenylketonuria or hyperphenylalaninemia. Editing of the R408W mutation is contemplated herein. The R408W mutation is located in exon 12 of the PAH gene, within the protein's catalytic domain, and causes the PAH protein to misfold. This structural alteration results in a significant reduction in enzyme activity. Contemplated are methods for a single nucleotide (T -> C) correction leading to a W408R correction at the PAH gene. Definitions
[0101] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g. Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.
[0102] As used herein, the term “about” can mean plus or minus 5% of the provided value.
[0103] As used herein, the terms “DNA binding domain” and “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which a Cas protein is an example, may be used interchangeably herein, and can refer to proteins that use RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid“programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence. As used herein, the term “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which Cas9 is an example, refer to proteins that use RNA:DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence.
[0104] As used herein, the term “guide RNA” or “gRNA” can refer to a site-specific targeting RNA that can bind an RNA-guided endonuclease to form a complex, and direct the activities of the bound RNA-guided endonuclease (such as a Cas endonuclease) to a specific sequence within a target nucleic acid (e.g., a specific gene or region within a gene). The guide RNA can include one or more RNA molecules. In some embodiments, the gRNA is a template armed gRNA (tagRNA). In some embodiments, the gRNA is an enhancer gRNA (egRNA).
[0105] As used herein, the term “target DNA” can refer to the specific region on a double-stranded DNA within a subject’s genome intended for editing by a gene editing system. In certain embodiments, the gene editing system is an RT editing system.
[0106] As used herein, the term “search target sequence” can refer to the sequence on target strand that is complementary or substantially complementary to the spacer sequence of the tagRNA.
[0107] As used herein, the term “editing target sequence” can refer to the sequence being edited on the non-target strand.
[0108] As used herein, the term “silent mutation” can refer to a nucleotide change or nucleotide changes (for e.g., substitution or substitutions) in a DNA sequence that does not result in a change to the amino acid sequence of the protein the DNA sequence encodes.
[0109] As used herein, the term “protospacer” refers to the sequence in DNA adjacent to the PAM (protospacer adjacent motif) sequence. In some embodiments, the protospacer is 20 nucleotides long. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand thereof, i.e., the “target strand” versus the “non-target strand” of the target DNA sequence). In order for Cas9 to function it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The skilled person will appreciate that the literature in the state of the art sometimes refers to the “protospacer” as the target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer.”Thus, in some cases, the term “protospacer” as used herein may be used interchangeably with the term “spacer.” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is in reference to the gRNA or the DNA target.
[0110] As used herein, the terms “upstream” and “downstream” define relevant positions at least two regions or sequences in a nucleic acid molecule orientated in a 5'-to-3' direction. For example, a first sequence is upstream of a second sequence in a DNA molecule where the first sequence is positioned 5’ to the second sequence. Accordingly, the second sequence is downstream of the first sequence.
[0111] As used herein, the term “protospacer adjacent sequence” or “PAM” refers to a DNA sequence that is an important targeting component of a Cas9 nuclease. The PAM sequence can be on either strand, and is downstream in the 5ʹ to 3ʹ direction of the Cas9 cut site. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins from different organisms.
[0112] As used herein, the term “spacer sequence” in connection with a guide RNA or a tagRNA refers to the portion of the guide RNA or tagRNA which contains a nucleotide sequence that shares the same sequence as the protospacer sequence in the target DNA sequence. The spacer sequence anneals to the complement of the protospacer sequence to form a ssRNA / ssDNA hybrid structure at the target site and a corresponding R loop ssDNA structure of the endogenous DNA strand.
[0113] As used herein, a “secondary structure” of a nucleic acid molecule (e.g., an RNA fragment, or a gRNA) refers to the base pairing interactions within the nucleic acid molecule.
[0114] As used herein, the term “Cas endonuclease” or “Cas nuclease” refers to an RNA-guided DNA endonuclease associated with and / or derived from the CRISPR adaptive immunity system. The term “nickase” refers to a Cas9 or other endonuclease with one of the two nuclease domains inactivated. This enzyme is capable of cleaving only one strand of a target DNA.
[0115] Unless otherwise indicated “nuclease” and “endonuclease” are used interchangeably herein to refer to an enzyme which possesses endonucleolytic catalytic activity for polynucleotide cleavage.
[0116] The terms “polynucleotide” and “nucleic acid” are used interchangeably herein and refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. A polynucleotide can be single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids / triple helices, or a polymer including purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. Any of the RNA sequences disclosed herein may also be DNA(either single-stranded or double-stranded), e.g., wherein “U” is converted to “T.” Any of the DNA sequences disclosed herein (e.g., FIG. 8-FIG. 10) may also be RNA, e.g., wherein “T” is converted to “U.”
[0117] A “functional variant” or “functional mutant”, as used herein, refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that encompasses one or more alterations to the amino acid sequence of the reference protein while retaining one or more of the functions, e.g., catalytic or binding functions. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions, insertions or deletions, or any combination thereof. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions. For example, a functional variant of a reverse transcriptase may comprise one or more amino acid substitutions compared to the amino acid sequence of a wild-type reverse transcriptase but retains the ability under at least one set of conditions to catalyze the polymerization of a polynucleotide. When the reference protein is a fusion of multiple functional domains, a functional variant thereof may retain one or more of the functions of at least one of the functional domains. For example, in some embodiments, a functional fragment of a Cas9 may comprise one or more amino acid substitutions in a nuclease domain, e.g., an H840A amino acid substitution, compared to the amino acid sequence of a wild type Cas9, but retains the DNA binding ability and lacks the nuclease activity partially or completely.
[0118] As used herein, the term “binding” refers to a non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” (e.g., when a molecule X is said to interact with a molecule Y, it means that the molecule X binds to molecule Y in a non-covalent manner). Binding interactions can be characterized by a dissociation constant (Kd), for example a Kdof, or a Kdless than, 10-6M, 10-7M, 10-8M, 10-9M, 10-10M, 10-11M, 10-12M, 10-13M, 10-14M,10-15M, or a number or a range between any two of these values. Kdcan be dependent on environmental conditions, e.g., pH and temperature. “Affinity” refers to the strength of binding, and increased binding affinity is correlated with a lower Kd.
[0119] As used herein, the term “hybridizing” or “hybridize” refers to the pairing of substantially complementary or complementary nucleic acid sequences within two different molecules. Pairing can be achieved by any process in which a nucleic acid sequence joins with a substantially or fully complementary sequence through base pairing to form a hybridization complex. “Hybridizing” or “hybridize” can comprise denaturing the molecules to disrupt the intramolecular structure(s) (e.g., secondary structure(s)) in the molecule. In some embodiments, denaturing the molecules comprises heating a solution comprising the molecules to a temperaturesufficient to disrupt the intramolecular structures of the molecules. In some instances, denaturing the molecules comprises adjusting the pH of a solution comprising the molecules to a pH sufficient to disrupt the intramolecular structures of the molecules. For purposes of hybridization, two nucleic acid sequences or segments of sequences are “substantially complementary” if at least 80% of their individual bases are complementary to one another. The complementary portion of each sequence can be referred to herein as a “segment”, and the segments are substantially complementary if they have 80% or greater identity.
[0120] The terms “complementarity” and “complementary” mean that a nucleic acid can form hydrogen bond(s) with another nucleic acid based on traditional Watson-Crick base paring rule, that is, adenine (A) pairs with thymine (T, or uracil (U) in RNA) and guanine (G) pairs with cytosine (C). Complementarity can be perfect (e.g., complete complementarity) or imperfect (e.g. partial complementarity). Perfect or complete complementarity indicates that each and every nucleic acid base of one strand is capable of forming hydrogen bonds according to Watson-Crick canonical base pairing with a corresponding base in another, antiparallel nucleic acid sequence. Partial complementarity indicates that only a percentage of the contiguous residues of a nucleic acid sequence can form Watson-Crick base pairing with the same number of contiguous residues in another, antiparallel nucleic acid sequence. In some embodiments, the complementarity can be at least 70%, 80%, 90%, 100% or a number or a range between any two of these values. In some embodiments, the complementarity is perfect, i.e., 100%. For example, the complementary candidate sequence segment is perfectly complementary to the candidate sequence segment, whose sequence can be deduced from the candidate sequence segment using the Watson-Crick base pairing rules.
[0121] As used herein, the terms “nucleic acid" and “polynucleotide” are interchangeable and refer to any nucleic acid, whether composed of phosphodiester linkages or modified linkages such as phosphotriester, phosphoramidate, siloxane, carbonate, carboxymethylester, acetamidate, carbamate, thioether, bridged phosphoramidate, bridged methylene phosphonate, bridged phosphoramidate, bridged phosphoramidate, bridged methylene phosphonate, phosphorothioate, methylphosphonate, phosphorodithioate, bridged phosphorothioate or sultone linkages, and combinations of such linkages. The terms “nucleic acid” and “polynucleotide” also specifically include nucleic acids composed of bases other than the five biologically occurring bases (adenine, guanine, thymine, cytosine and uracil).
[0122] The terms “DNA editing efficiency,” or “editing efficiency” may be used interchangeably herein and can refer to the number or proportion of intended target sequences that are edited. In some embodiments, the efficiency can be reported as % indel, e.g., the proportion of insertions and / or deletions detected in the target sequence. Indels (e.g., insertion-deletions) canresult from repair of double-stranded DNA breaks caused by Cas9 cleavage by processes including, but not limited to, non-homologous end joining (NHEJ) repair. The terms “correction efficiency”, “precise correction efficiency”, or “precise edit efficiency” may be used interchangeably herein and can refer to the number or proportion of target sequences that contain the desired or intended edit. In some embodiments, the efficiency can be reported as % correction or % edit, e.g. the proportion of intended edits detected in the target sequence. In some embodiments, the intended edit or edits revert a pathogenic mutation to wildtype sequence, thereby correcting a disease-associated mutation.
[0123] The term “off-target editing frequency,” as used herein, refers to the number or proportion of unintended DNA sequences that are edited. On-target and off- target editing frequencies may be measured by the methods and assays described herein, further in view of techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) with complementarity to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence or off-target sequence of interest. Since many of the Cas9-dependent off-target sites have high sequence identity to the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9- dependent off-target site may be designed using techniques and kits known in the art. These kits make use of polymerase chain reaction (PCR) amplification, which produces amplicons as intermediate products. The target and off-target sequences may comprise genomic loci that further comprise protospacers and PAMs. Accordingly, the term “amplicons,” as used herein, may refer to nucleic acid molecules that constitute the aggregates of genomic loci, protospacers and PAMs. High-throughput sequencing techniques used herein may further include Sanger sequencing and / or whole genome sequencing (WGS).
[0124] As used herein, the terms “transfection” or “infection” refer to the introduction of a nucleic acid into a host cell, such as by contacting the cell with liposomes or nanoparticles (e.g., lipid nanoparticles) as described herein.
[0125] As used herein, “treatment” refers to a clinical intervention made in response to a disease, disorder or physiological condition manifested by a patient or to which a patient may be susceptible. The aim of treatment includes, but is not limited to, the alleviation or prevention of symptoms, slowing or stopping the progression or worsening of a disease, disorder, or condition and / or the remission of the disease, disorder or condition. “Treatments” refer to one or both of therapeutic treatment and prophylactic or preventative measures. Subjects in need of treatment include those already affected by a disease or disorder or undesired physiological condition as well as those in which the disease or disorder or undesired physiological condition is to beprevented.
[0126] As used herein, the terms “effective amount” or “pharmaceutically effective amount” or “therapeutically effective amount” refer to an amount sufficient to effect beneficial or desirable biological and / or clinical results.
[0127] The term “pharmaceutically acceptable excipient” as used herein refers to any suitable substance that provides a pharmaceutically acceptable carrier, additive or diluent for administration of a compound(s) of interest to a subject. Pharmaceutically acceptable excipients can encompass substances referred to as pharmaceutically acceptable diluents, pharmaceutically acceptable additives, and pharmaceutically acceptable carriers.
[0128] As used herein, a “subject” refers to an animal for whom a diagnosis, treatment, or therapy is desired. In some embodiments, the subject is a mammal. “Mammal,” as used herein, refers to an individual belonging to the class Mammalia and includes, but not limited to, humans,. In some embodiments, the mammal is a primate. In some embodiments, the mammal is a human. In some embodiments, the mammal is not a human. In some embodiments, the subject has or is suspected of having a disease that can be corrected via gene editing. In some embodiments, the gene editing system is an RT editing system. In some embodiments, the subject has or is suspected of having Wilson’s disease. In some embodiments, the subject has or is suspected of having alpha1 antitrypsin deficiency disease. In some embodiments, the subject has or is suspected of having phenylketonuria or hyperphenylalaninemia. Reverse Transcriptase (RT) editing
[0129] RT editing methods and compositions disclosed herein are directed to correction of a disease mutation using reverse transcriptase (RT) editing. Compositions disclosed herein comprise and RT editor or an mRNA encoding a reverse transcriptase (RT) editor, a long guide RNA encoding the edit designated as ‘tagRNA’, and optionally, a second guide RNA designated as ‘egRNA’ (opposite strand gRNA). In some embodiments, the composition induces programmable editing of a target DNA using an RT editor complexed with a tagRNA to incorporate an intended nucleotide edit (also referred to herein as a nucleotide change) into the target DNA. In some embodiments, the composition comprises a second guide RNA (egRNA).
[0130] A target gene of the RT editing may comprise a double stranded DNA molecule having two complementary strands: a first strand that may be referred to as a “target strand” or a “non-edit strand”, and a second strand that may be referred to as a “non-target strand,” or an “edit strand” or an “opposite strand”. The RT editors provided herein can be employed to introduce more than one nucleotide change. In some embodiments, two (or more) RT editors provided herein can be employed to edit two (or more) different genes or two (or more) different sites of a gene. RT editors provided herein can correct more than one mutation. Two disclosed RT editors can beused to target two different genes or two different segments of a gene. In some embodiments, one RT editor (e.g., as an mRNA) can be delivered with two or more different tagRNAs (and optionally egRNAs) - either in the same or different LNP formulations. The mRNA of the RT editor can be the same for any two (or more) different sites. In some embodiments, one or more different editors can be employed. In some embodiments, the target is a non-coding element, e.g., a promoter, an enhancer, and the like. RT editor
[0131] The term “RT editor” refers to the polypeptide or polypeptide components involved in RT editing, or any polynucleotide(s) encoding the polypeptide or polypeptide components. In various embodiments, a RT editor includes a polypeptide domain having DNA endonuclease activity and a polypeptide domain having DNA polymerase activity. In some embodiments, the RT editor further comprises a polypeptide domain having nuclease activity. In some embodiments, the polypeptide domain having DNA binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain having nuclease activity comprises a nickase, or a fully active nuclease. As used herein, the term “nickase” refers to a nuclease capable of cleaving only one strand of a double-stranded DNA target. In some embodiments, the RT editor comprises a polypeptide domain that is an inactive nuclease, in some embodiments, the polypeptide domain having comprises a nucleic acid guided DNA endonuclease domain, for example, a CRISPR-Cas protein, for example, a Cas9 nickase, a Cpfl nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity comprises a template-dependent DNA polymerase, for example, a DNA- dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase, in some embodiments, the RT editor comprises additional polypeptides involved in RT editing, for example, a polypeptide domain having a 5’ endonuclease activity, e.g., a 5’ endogenous DNA flap endonucleases (e.g., FEN1), for helping to drive the RT editing process towards the edited product formation. In some embodiments, the RT editor further comprises an RNA-protein recruitment polypeptide, for example, a MS2 coat protein.
[0132] A RT editor may be engineered. In some embodiments, the polynucleotide or polypeptide components of a RT editor do not naturally occur in the same organism or cellular environment. In some embodiments, the polynucleotide or polypeptide components of a RT editor may be of different origins or from different organisms. In some embodiments, a RT editor comprises a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain that are derived from different species. In some embodiments, a RT editor comprises a Caspolypeptide (DNA endonuclease domain) and a reverse transcriptase polypeptide (DNA polymerase) that are derived from different species.
[0133] In some embodiments, polypeptide domains of a RT editor may be fused or linked by a peptide linker to form a fusion protein. In other embodiments, a RT editor comprises one or more polypeptide domains provided in trans as separate proteins, which are capable of being associated to each other through non-peptide linkages or through aptamers or recruitment sequences. For example, a RT editor may comprise a DNA binding domain and a reverse transcriptase domain associated with each other by an RNA-protein recruitment aptamer, e.g., an MS2 aptamer, which may be linked to a tagRNA. RT editor polypeptide components may be encoded by one or more polynucleotides in whole or in part, in some embodiments, a single polynucleotide, construct, or vector encodes the RT editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of a RT editor, or a portion of a RT editor fusion protein. For example, a RT editor fusion protein may comprise an N-terminal portion fused to an intein-N and a C-terminal portion fused to an intein-C, each of which is individually encoded by an AAV vector
[0134] In some embodiments, the RT Editor is transcribed from an mRNA. In some embodiments, the mRNA comprises a 5'-cap structure. In some embodiments, the mRNA comprises a nuclear localization sequence (NLS). In some embodiments, the mRNA sequence comprises a dead Cas9 sequence. In some embodiments, the Cas9 is a Cas nickase sequence. In some embodiments, the RT Editor mRNA sequence comprises a 5’-UTR. In some embodiments, the RT Editor mRNA sequence comprises a 3’-UTR. In some embodiments, the RT Editor comprises a sequence of a viral element. In some embodiments, the viral element is a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) sequence. In some embodiments, the RT editor comprises secondary structure motifs. In some embodiments, the viral element sequence comprises or consists of any one of SEQ ID Nos: 189-190. In some embodiments, the RT editor comprises an ek5 sequence. The ek5 sequence can comprise the sequence of SEQ ID NO: 190, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 190. The ek5 sequence can have one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 190. The viral element sequence can be derived from Aichi virus 1 (AiV-1), such as, for example, the 3’ UTR K5 element (GenBank: NC_001918.1, 8,122–8,251). In some embodiments, the viral element sequence comprises the extended form of K5 (‘‘ek5,’’ 8,067–8,251, 185 nt). Viral element sequences are described in Seo, Jenny J., et al. ("Functional viromic screens uncover regulatory RNA elements." Cell 186.15 (2023): 3291-3306), the content of which is incorporated herein by reference in its entirety. In some embodiments, the secondary structure motif comprises a triple helix sequence. In some embodiments, the RT editor mRNA sequence has a structurecomprising or consisting of a sequence comprising from 5’ to 3’ as [5’UTR]-[NLS]-[nCas9]- [linker]-[RT]-[NLS]-[3’UTR_and / or_viral_element]-[polyA_sequence]. In some embodiments, any nucleoside of the RT editor mRNA sequence may be chemically modified.
[0135] Compositions disclosed herein comprise a long guide RNA designated herein as “template armed guide RNA” or “tagRNA”. The tagRNA comprises a spacer sequence, a scaffold sequence, an editing template and a flap binding sequence. In some embodiments, the tagRNA comprises in 5’ to 3’ order: spacer sequence, scaffold sequence, editing template, and a flap binding sequence. RT Editor Nucleotide Polymerase Domain and Endonuclease Domain
[0136] In some embodiments, a RT editor comprises a nucleotide polymerase domain, e.g., a DNA polymerase domain. The DNA polymerase domain may be a wild-type DNA polymerase domain, a full-length DNA polymerase protein domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. In some embodiments, the polymerase domain is a template dependent polymerase domain. For example, the DNA polymerase may rely on a template polynucleotide strand, e.g.,the editing template sequence, for new strand DNA synthesis. In some embodiments, the RT editor comprises a DNA-dependent DNA polymerase. For example, a RT editor having a DNA-dependent DNA polymerase can synthesize a new single stranded DNA using a tagRNA editing template that comprises a DNA sequence as a template. In such cases, the tagRNA is a chimeric or hybrid tagRNA, and comprising an extension arm comprising a DNA strand. The chimeric or hybrid tagRNA may comprise an RNA portion (including the spacer and the gRNA core) and a DNA portion (the extension arm comprising the editing template that includes a strand of DNA).
[0137] In some embodiments, RT editor comprises a DNA endonuclease domain (e.g., a Cas9 nickase). In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild-type Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. A Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild-type Cas protein. A Cas protein, e.g., Cas9, can comprise an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof relative to a corresponding wild-type version of the Cas protein. In some embodiments, a Cas protein can be a polypeptide with at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild type exemplary Cas protein.
[0138] A Cas protein, e.g., Cas9, may comprise one or more domains. Non-limiting examples of Cas domains include, guide nucleic acid recognition and / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA binding domain, a DNA endonuclease domain, RNA binding domain, helicase domains, protein-protein interaction domains, and dimerization domains. In various embodiments, a Cas protein comprises a guide nucleic acid recognition and / or binding domain that can interact with a guide nucleic acid, and one or more nuclease domains that comprise catalytic activity for nucleic acid cleavage.
[0139] In some embodiments, a Cas protein, e.g., Cas9, comprises one or more nuclease domains. A Cas protein can comprise an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein. In some embodiments, a Cas protein comprises a single nuclease domain. For example, a Cpfl may comprise a RuvC domain but lacks HNH domain. In some embodiments, a Cas protein comprises two nuclease domains, e.g., a Cas9 protein can comprise an HNH nuclease domain and a RuvC nuclease domain.
[0140] In some embodiments, a RT editor comprises a Cas protein, e.g., Cas9, wherein all nuclease domains of the Cas protein are active. In some embodiments, a RT editor comprises a Cas protein having one or more inactive nuclease domains. One or a plurality of the nuclease domains (e.g., RuvC, HNH) of a Cas protein can be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. In some embodiments, a Cas protein, e.g., Cas9, comprising mutations in a nuclease domain has reduced (e.g., nickase) or abolished nuclease activity while maintaining its ability to target a nucleic acid locus at a search target sequence when complexed with a guide nucleic acid, e.g., a tagRNA.
[0141] In some embodiments, a RT editor comprises a Cas nickase that can bind to the target gene in a sequence-specific manner and generate a single-strand break at a protospacer within double-stranded DNA in the target gene, but not a double-strand break. For example, the Cas nickase can cleave the edit strand or the non-edit strand of the target gene, but may not cleave both. In some embodiments, a RT editor comprises a Cas nickase comprising two nuclease domains (e.g., Cas9), with one of the two nuclease domains modified to lack catalytic activity or deleted. In some embodiments, the Cas nickase of a RT editor comprises a nuclease inactive RuvC domain and a nuclease active HNH domain. In some embodiments, the Cas nickase of a RT editor comprises a nuclease inactive HNH domain and a nuclease active RuvC domain. In some embodiments, a RT editor comprises a Cas9 nickase having an amino acid substitution in the RuvC domain e.g., an amino acid substitution that reduces or abolishes nuclease activity of the RuvC domain. In some embodiments, the Cas9 nickase comprises a D10X amino acid substitutioncompared to a wild type S. pyogenes Cas9, wherein X is any amino acid other than D. In some embodiments, a RT editor comprises a Cas9 nickase having an amino acid substitution in the HNH domain, e.g., an amino acid substitution that reduces or abolishes nuclease activity of the HNH domain. In some embodiments, the Cas9 nickase comprises a H840X amino acid substitution compared to a wild type S. pyogenes Cas9, wherein X is any ammo acid other than H.
[0142] In some embodiments, a RT editor comprises a Cas protein that can bind to the target gene in a sequence-specific manner but lacks or has abolished nuclease activity and may not cleave either strand of a double stranded DNA in a target gene. Abolished activity or lacking activity can refer to an enzymatic activity less than 1%, less than 2%, less than 3%, less than 4%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, or less than 10% activity compared to a wild-type exemplary activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, a Cas protein of a RT editor completely lacks nuclease activity. A nuclease, e.g., Cas9, that lacks nuclease activity may be referred to as nuclease inactive or “nuclease dead” (abbreviated by “d”). A nuclease dead Cas protein (e.g., dCas, dCas9) can bind to a target polynucleotide but may not cleave the target polynucleotide. In some aspects, a dead Cas protein is a dead Cas9 protein. In some embodiments, a RT editor comprises a nuclease dead Cas protein wherein all of the nuclease domains (e.g., both RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are mutated to lack catalytic activity, or are deleted.
[0143] A Cas protein can be modified. A Cas protein, e.g., Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein.
[0144] A Cas protein can be a fusion protein. For example, a Cas protein can be fused to a cleavage domain, an epigenetic modification domain, a transcriptional regulation domain, or a polymerase domain. A Cas protein can also be fused to a heterologous polypeptide providing increased or decreased stability. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.
[0145] In some embodiments, the Cas protein of a RT editor is a Class 2 Cas protein. In some embodiments, the Cas protein is a type II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, a Cas9 protein homolog, mutant, variant, or a functional fragment thereof. As used herein, a Cas9, Cas9 protein, Cas9 polypeptideor a Cas9 nuclease refers to an RNA guided nuclease comprising one or more Cas9 nuclease domains and a Cas9 gRNA binding domain having the ability to bind a guide polynucleotide, e.g., a tagRNA. A Cas9 protein may refer to a wild-type Cas9 protein from any organism or a homolog, ortholog, or paralog from any organisms; any functional mutants or functional variants thereof; or any functional fragments or domains thereof. In some embodiments, a RT editor comprises a full- length Cas9 protein. In some embodiments, the Cas9 protein can generally comprises at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity to a wild-type reference Cas9 protein (e.g., Cas9 from S. pyogenes). In some embodiments, the Cas9 comprises an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof as compared to a wild-type reference Cas9 protein.
[0146] In some embodiments the DNA endonuclease domain of the RT editor comprises a Cas9 protein (e.g., a Cas9 nickase). The Cas9 protein can comprise an amino acid sequence that is at least 80% identical to any one of SEQ ID NOS: 142-163 and 270-314. The Cas9 protein can comprise an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOS: 142-163 and 270-314. Reverse transcriptases
[0147] In some embodiments, a RT editor comprises an RNA-dependent DNA polymerase domain, for example, a reverse transcriptase (RT). An RT or an RT domain may be a wild-type RT domain, a full-length RT domain, or may be a functional mutant, a functional variant, or a functional fragment thereof. An RT or an RT domain of a RT editor may comprise a wild-type RT, or may be engineered or evolved to contain specific amino acid substitutions, truncations, or variants. An engineered RT may comprise sequences or amino acid changes different from a naturally occurring RT. In some embodiments, the engineered RT may have improved reverse transcription activity over a naturally occurring RT or RT domain. In some embodiments, the engineered RT may have improved features over a naturally occurring RT, for example, improved thermostability, reverse transcription efficiency, or target fidelity. In some embodiments, a RT editor comprising the engineered RT has improved RT editing efficiency over a RT editor having a reference naturally occurring RT.
[0148] In some embodiments, a RT editor comprises a virus RT, for example, a retrovirus RT. Nonlimiting examples of virus RT include Moloney murine leukemia virus (M- MLV or MLVRT or M-MLV RT); human T-cell leukemia virus type 1 (HTLV-l) RT; bovine leukemia virus (BLV) RT; Rous Sarcoma Virus (RSV) RT; human immunodeficiency virus (HIV) RT, M-MFV RT, Avian Sarcoma-Leukosis Virus (ASLV) RT, Rous Sarcoma Virus (RSV) RT, Avian Myeloblastosis Virus (AMV) RT, Avian Erythroblastosis Virus (AEV) Helper Virus MCAV RT, Avian Myelocytomatosis Virus MC29 Helper Virus MCAV RT, AvianReticuloendotheliosis Virus (REV-T) Helper Virus REV-A RT, Avian Sarcoma Virus UR2 Helper Virus (LJR2AV) RT, Avian Sarcoma Virus ¥73 Helper Virus YAV RT, Rous Associated Virus (RAV) RT, and Myeloblastosis Associated Virus (MAV) RT, all of which may be suitably used in the methods and composition described herein. In some embodiments, an MMLV RT, e.g., reference MMLV RT, comprises a sequence as disclosed in SEQ ID NO: 15.
[0149] In some embodiments, the RT editor comprises a wild-type M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof. In some embodiments, the RT editor comprises a reference M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof.
[0150] In some embodiments, the RT is an RT (e.g., a retro-transposon RT) from an avian genome. The avian may be of the order Galliformes, Aseriformes, Passeriformes, Gruiformes, Struthioniformes, Rheiformes, Casuariformes, Apyerygiformes, Otidiformes, Columbiformes, Sphenisciformes, Cathartiformes, Accipitriformes, Strigiformes, Psittaciformes, Charadriiformes, or Falconiformes. Exemplary avians include those of Galliformes (e.g. chicken, quails, and turkey), Anseriformes (e.g. duck, and goose), Charadriiformes (e.g. gull, barred button quail, and plover), Columbiformes (e.g. pigeon), Struthioniformes (e.g. ostrich), Passeriformes (e.g. crow, finch, sparrow, starling, and swallow), Psittaciformes (e.g. parrot), Falconiformes (e.g. eagle, and falcon), Strigiformes (e.g. owl), Sphenisciformes (e.g. penguin), and Psittaciformes (e.g. parakeet, and parrot). The avian can belong to the order Passeriformes. The avian can belong to the family Passerellidae. The avian can belong to the genus Melospiza. In some embodiments, the RT is from the genome of M. georgiana. In some embodiments, the M. georgiana RT is an engineered RT. In some embodiments, the mutation sites on the M. georgiana RT may be one or more of G206N, M312K, W319F, E336P, and L611W. The DNA polymerase domain (e.g., RT) can comprise an amino acid sequence that is at least about 80% identical to SEQ ID NO: 141 (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to SEQ ID NO: 141). The DNA polymerase domain (e.g., RT) can comprise an amino acid sequence having one, two, three, four, five, six, seven, eight, nine, or ten mismatches relative to the sequence of SEQ ID NO: 141. There are provided, in some embodiments, RT editors. The RT editor can comprise a DNA endonuclease domain, a DNA binding domain and a DNA polymerase domain. The DNA polymerase domain can comprise an amino acid sequence that is at least 80% identical to SEQ ID NO: 141. The DNA polymerase domain can comprise an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 141. In some embodiments, the RT editor comprises a reverse transcriptase has a sequence of any one of SEQID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018. In certain embodiments, the reverse transcriptase comprises an amino acid sequence that is at least 80% identical to SEQ ID NO: 141.
[0151] In some embodiments, the RT is an RT from a mammalian genome. In some embodiments, the mammal may be of the order Chiroptera (bat). That bat can be of the family Phyllostomidae (e.g., leaf-nosed bats), Noctilionidae (bulldog bats), Cistugidae, Thyropteridae (e.g., disk-winged bats), Molossidae (e.g., free-tailed bats), Miniopteridae (e.g., long winged bat), Mormoopidae, Mystacinidae (e.g., New Zealand short-tailed bats), Myzopodidae (e.g., sucker- footed bats), Natalidae (e.g., funnel-eared bats), Emballonuridae (e.g., sheath-tailed bats), Nycteridae (e.g., slit-faced bats), Furipteridae (e.g., smoky bats), Vespertilionidae (e.g., vesper bats), Craseonycteridae, Megadermatidae (e.g., false vampire bats), Rhinolophidae (e.g., horseshoe bats), Pteropodidae (e.g., Old World fruit bats), Hipposideridae (e.g., Old World leaf- nosed bats), or Rhinopomatidae. In some embodiments, the bat is of the genus Myotis, Pipistrelles, or Eptesicus. In some embodiments, the bat species is M. brandtii, M. daubentonii, P. kuhlii, or E. nilssonii. In some embodiments, the RT editor comprises a reverse transcriptase having a sequence of any one of SEQ ID NOs: 250-255, 258-259, and 266-267. In some embodiments, the reverse transcriptase comprises an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 250-255, 258-259, and 266-267.
[0152] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase. In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising one or more mutations. In some embodiments, at least one of the one or more mutations is at an amino acid position functionally equivalent to V101, N200, A208, G248, P330, L435, K445, and / or A623 relative to SEQ ID NO: 8.
[0153] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 255 (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to SEQ ID NO: 255). In some embodiments, the reverse transcriptase comprises one or more mutations. In some embodiments, at least one of the one or more mutations is at amino acid position V98, N197, A205, S245, P327, V432, R442, and / or A623 relative to SEQ ID NO: 255. In some embodiments, the reverse transcriptase comprises one or more transition mutations selected from the group consisting of V98R, N197C, N197D, N197G, A205T, S245C, P327E, P327Q, V432K, R442T, and A623F relative to SEQ ID NO: 255.
[0154] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 260 (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%,96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to SEQ ID NO: 260). The reverse transcriptase can comprise one or more mutations. In some embodiments, at least one of the one or more mutations is at amino acid position V106, N204, E212, G252, P334, L440, E450, and / or A629 relative to SEQ ID NO: 260. In some embodiments, the reverse transcriptase comprises one or more transition mutations selected from the group consisting of V106R, N204C, N204D, N204G, E212T, G252C, P334E, P334Q, L440K, E450T, and / or A629F relative to SEQ ID NO: 260.
[0155] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 255. In some embodiments, the reverse transcriptase comprises an N197C or V98R transition mutation relative to SEQ ID NO: 255 (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to SEQ ID NO: 255).
[0156] In some embodiments, the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 260 (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to SEQ ID NO: 260). In some embodiments, the reverse transcriptase comprises an N204C and / or V106R transition mutation relative to SEQ ID NO: 260. Accessory domains
[0157] In some embodiments, the RT editor further comprises an accessory domain. In some embodiments, the accessory domain comprises a single-strand binding (SSB) protein domain or a stabilon. In some embodiments, the SSB protein domain is derived from RecA protein, Sso7d protein, or Sto7d protein. In some embodiments the accessory domain is situated at the N-terminus, the C-terminus, or at an internal location of the RT editor. In some embodiments, the RT editor comprising an accessory domain comprises an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1276-1298 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to any one of SEQ ID NOs: 1276-1298). In some embodiments, the SSB protein domain is derived from Sso7d and comprises one or more mutations at an amino acid position functionally equivalent to K12 and / or E35 of a wild type Sso7d amino acid sequence. In some embodiments, the one or more mutations comprise K12L and / or E35L relative to a wild type Sso7d sequence. In some embodiments, the SSB protein domain comprises an amino acid sequence that is at least 85% identical to SEQ ID NO: 1205 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, ora number or a range between any two of these values, identical to SEQ ID NO: 1205). In some embodiments, the RT editor comprising the SSB protein domain comprises from N-terminus to C-terminus: [N-terminus of nCas9]-[linker]-[SSB protein domain]-[linker]-[RT]-[linker]-[C- terminus of nCas9]. In some embodiments, the N-terminus of nCas9 comprises an amino acid sequence functionally equivalent to amino acids 1-1247 of SEQ ID NO: 6 and / or the C-terminus of nCas9 comprises an amino acid sequence functionally equivalent to amino acids 1248-1367 of SEQ ID NO: 6. In some embodiments, the RT editor comprising the SSB protein domain comprises an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1207- 1209 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to any one of SEQ ID NOs: 1207-1209). Additional sequence elements of the RT editor mRNA
[0158] The present disclosure provides optimized mRNAs encoding a RT editor, that provide effective genome editing of a target cell population when administered with one or more gRNAs (e.g., tagRNA and egRNA). In some embodiments, additional segmented polyA sequences increase mRNA stability and half-life of the RT editor. In some embodiments, the disclosure provides an mRNA comprising (i) a 5′-cap, (ii) a 5’-untranslated region (UTR); (ii) an open reading frame (ORF) comprising a nucleotide sequence that encodes a RT editor; and (iv) a 3′ untranslated region (UTR). In some embodiments, the mRNA further comprises viral element sequences 3’ of the 3’-UTR. In some embodiments, the viral element sequence comprises the sequence of SEQ ID NO: 189. In some embodiments, the mRNA further comprises an eK5 sequence that is located 3’ of the 3’-UTR. eK5 is a regulatory RNA element derived from a virus that when placed in the 3‘UTR of an mRNA can lead to increased protein expression (for instance, as in Seo et al Cell 2023, 186: 3291). In some embodiments, the ek5 sequence comprises the sequence of SEQ ID NO: 190. In some embodiments, the synthetic triple helix secondary structure sequence (STH) comprises the sequence derived from a sequence element of the long non-coding RNA, such as MALAT1. In some embodiments, the STH sequence comprises the sequence of SEQ ID NO: 246. FIG.6-FIG.10 provide exemplary secondary structure motifs employed in the methods and compositions provided herein. In some embodiments their presence in the 3’ UTR of mRNAs disclosed herein provides 3’ end stabilization and / or protection. Some embodiments provided herein employ the triple helix motif from MALAT1, or an engineered (minimal) version. In some embodiments, MALAT1 (Gene ID: 378938) based triple helix motif is situated at the 3’ end. MALAT1 (metastasis associated lung adenocarcinoma transcript 1) also known as NEAT2 (noncoding nuclear-enriched abundant transcript 2) is a large, infrequently spliced non-coding RNA, which is highly conserved amongst mammals and highly expressed in the nucleus.tagRNAs
[0159] Disclosed herein include target priming RNAs (tagRNAs). The term “target priming RNA”, or “tagRNA”, refers to a guide polynucleotide that comprises one or more intended nucleotide edits for incorporation into the target DNA. In some embodiments, the tagRNA associates with and directs a RT editor to incorporate the one or more intended nucleotide edits into the target gene via RT editing. “Nucleotide edit” or “intended nucleotide edit” refers to a specified deletion of one or more nucleotides at one specific position, insertion of one or more nucleotides at one specific position, substitution of a single nucleotide, or other alterations at one specific position to be incorporated into the sequence of the target gene. Intended nucleotide edit may refer to the edit on the editing template as compared to the sequence on the target strand of the target gene or may refer to the edit encoded by the editing template on the newly synthesized single stranded DNA that replaces the editing target sequence, as compared to the editing target sequence. In some embodiments, a tagRNA comprises a spacer sequence that is complementary or substantially complementary to a search target sequence on a target strand of the target gene, in some embodiments, the tagRNA comprises a gRNA core that associates with a DNA endonuclease domain, e.g., a CRISPR-Cas protein domain, of a RT editor. In some embodiments, the tagRNA further comprises an extended nucleotide sequence comprising one or more intended nucleotide edits compared to the endogenous sequence of the target gene, wherein the extended nucleotide sequence may be referred to as an extension arm.
[0160] In some embodiments, in a template armed gRNA (tagRNA), a spacer sequence is complementary or substantially complementary to a specific sequence on the target strand, which may be referred to as a “search target sequence.” In some embodiments, the spacer sequence anneals with the target strand at the search target sequence. The target strand may also be referred to as the “non-Protospacer Adjacent Motif (non-PAM strand).” In some embodiments, the non-target strand may also be referred to as the “PAM strand.” In some embodiments, the PAM strand comprises a protospacer sequence and optionally a protospacer adjacent motif (PAM) sequence. In RT editing using a Cas-protein-based RT editor, a PAM sequence refers to a short DNA sequence immediately adjacent to the protospacer sequence on the PAM strand of the target gene. A PAM sequence may be specifically recognized by a programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease, in some embodiments, a specific PAM is characteristic of a specific programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease. A protospacer sequence refers to a specific sequence in the PAM strand of the target gene that is complementary to the search target sequence. In a tagRNA, a spacer sequence may have a substantially identical sequence as the protospacer sequence on the edit strand of a target gene, except that the spacer sequence may comprise Uracil (U) and the protospacer sequence may comprise Thymine (T).
[0161] In some embodiments, the double stranded target DNA comprises a nick site on the PAM strand (or non-target strand). As used herein, a “nick site" refers to a specific position in between two nucleotides or two base pairs of the double stranded target DNA. In some embodiments, the position of a nick site is determined relative to the position of a specific PAM sequence. In some embodiments, the nick site is the particular position where a nick will occur when the double stranded target DNA is contacted with a nickase, for example, a Cas nickase, that recognizes a specific PAM sequence. In some embodiments, the nick site is upstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is downstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is upstream of a PAM sequence recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Streptococcus pyogenes Cas9 nickase. In some embodiments, the Cas nickase comprises any one of the sequences recited in SEQ ID NO: 142-163.
[0162] In some embodiments, the nick site is 3 base pairs upstream of the PAM sequence, and the PAM sequence is recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active HNH domain and a nuclease inactive RuvC domain. In some embodiments, the nick site is 2 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a S. thermophilus Cas9 nickase that comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the Cas nick site (e.g., a Cas9 nickase comprising an amino acid sequence that is at least 80% identical (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to any one of SEQ ID NOs: 142- 163and 270-314, or an amino acid sequence having at least one, at least two, at least three, at least four, or at least five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314), is 4, 5, 6 or more than 6 nucleotides base pairs upstream of the PAM sequence and the PAM sequence is recognized by a Cas9 nickase wherein the Cas9 nickase comprises a nuclease active HNH domain and a nuclease inactive RuvC domain.
[0163] An “editing template” of a tagRNA is a single-stranded portion of the tagRNA that is 5' of the FB sequence and comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand), and comprises one or more intended nucleotide edits compared to the endogenous sequence of the double stranded target DNA. In some embodiments, the editing template and the FB sequence are immediately adjacent to each other. Accordingly, in some embodiments, a tagRNA in RT editing comprises a single-stranded portion that comprisesthe editing template sequence and the FB sequence immediately adjacent to each other. In some embodiments, the single stranded portion of the tagRNA comprising both the editing template sequence and the flap binding sequence is complementary or substantially complementary to an endogenous sequence on the PAM strand (i.e., the non-target strand or the edit strand) of the double stranded target DNA except for one or more non-complementary nucleotides at the intended nucleotide edit positions. As used herein, regardless of relative 5 -3' positioning in other contexts, the relative positions as between the FB sequence and the editing template, and the relative positions as among elements of a tagRNA, are determined by the 5' to 3' order of the tagRNA as a single molecule regardless of the position of sequences in the double stranded target DNA that may have complementarity or identity to elements of the tagRNA. In some embodiments, the editing template is complementary or substantially complementary to a sequence on the PAM strand that is immediately downstream of the nick site, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. The endogenous, e.g., genomic, sequence that is complementary or substantially complementary to the editing template, except for the one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit, may be referred to as an “editing target sequence." In some embodiments, the editing template has identity or substantial identity to a sequence on the target strand that is complementary to, or having the same position in the genome as, the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions. In some embodiments, the editing template encodes a single stranded DNA, wherein the single stranded DNA has identity or substantial identity to the editing target sequence except for one or more insertions, deletions, or substitutions at the positions of the one or more intended nucleotide edits.
[0164] In some embodiments, the editing template comprises a nucleotide sequence selected from SEQ ID NOs: 95-98. In some embodiments, the editing template comprises a nucleotide sequence selected from SEQ ID NOs: 95-98 or a sequence having one, two, or three mismatches relative to a nucleotide sequence selected from SEQ ID NOs: 95-98. In some embodiments, the editing template comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 95-98 or a sequence at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a nucleotide sequence selected from SEQ ID NOs: 95-98.
[0165] A “flap binding (FB) sequence” is a single-stranded portion of the tagRNA that comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand). The FB sequence is complementary or substantially complementary to a sequence on the PAM strand of the double stranded target DNA that is immediately upstream of the nick site. Insome embodiments, in the process of RT editing, the tagRNA complexes with and directs a RT editor to bind the search target sequence on the target strand of the double stranded target DNA and the RT editor generates a nick at the nick site on the non-target strand (e.g., the PAM strand) of the double stranded target DNA. In some embodiments, the FB sequence is complementary to or substantially complementary to, and can anneal to, a free 3' end on the non-target strand of the double stranded target DNA at the nick site. In some embodiments, the FB sequence annealed to the free 3' end on the non-target strand can initiate target-primed DNA synthesis. In some embodiments, the FB sequence is about 2 to 20 nucleotides in length. In some embodiments, the FB sequence is about 8 to 16 nucleotides in length. In some embodiments, the FBS is 6 nucleotides in length. In some embodiments, the FB site comprises a nucleotide sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG. In some embodiments, the FB sequence comprises a nucleotide sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG. In some embodiments, the FB sequence comprises a nucleotide sequence having one, two, or three mismatches relative to a sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG. In some embodiments, the FB sequence comprises or consists of a sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG or a sequence at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG. Nucleic acid modifications
[0166] In some embodiments, any of the nucleic acids of the disclosure can comprise one or more modifications (e.g., gRNA, tagRNA, egRNA, or any nucleic acid encoding any component of the RT editor systems disclosed herein). In some embodiments, the gRNA (e.g., egRNA) or tagRNA is a chemically modified gRNA or tagRNA. Various types of RNA modifications can be introduced to the gRNAs or tagRNAs to enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes as described inthe art. The gRNAs or tagRNAs described herein can comprise one or more modifications including internucleoside linkages, purine or pyrimidine bases, or sugar. In some embodiments, a modification is introduced at the terminal of a gRNA or tagRNA with chemical synthesis or with a polymerase enzyme. Examples of modified nucleic acids and their synthesis are disclosed in WO2013 / 052523. Synthesis of modified polynucleotides is also described in Verma and Eckstein, Annual Review of Biochemistry, vol.76, 99-134 (1998).
[0167] In some embodiments, programmable editing of a target DNA comprises a template armed gRNA (tagRNA). In some embodiments, one or more nucleoside of the tagRNA is chemically modified. In some embodiments, the chemical modifications can be any one of an LNA, 2’-fluoro, DNA, 2’-OMe, and 2’ MethoxyEthoxy (2’-MOE)- chemical modification. In some embodiments, the tagRNA comprises one or more deoxyribonucleotides (e.g., DNA). In some embodiments, each nucleotide of the editing template comprises an LNA modification. In some embodiments, each nucleotide of the flap binding sequence comprises an LNA modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, every second nucleotide of the editing template comprises an LNA modification. In some embodiments, every second nucleotide of the flap binding sequence comprises an LNA modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, every third nucleotide of the editing template comprises an LNA modification. In some embodiments, every third nucleotide of the flap binding sequence comprises an LNA modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, each nucleotide of the editing template comprises a 2’-fluoro modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-fluoro modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-fluoro modification. In some embodiments, every second nucleotide of the editing template comprises a 2’-fluoro modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’-fluoro modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’-fluoro modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-fluoro modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’-fluoro modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’-fluoro modification. In some embodiments, each nucleotide of the editing template comprises a 2’-OMe modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-OMe modification. Insome embodiments, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-OMe modification. In some embodiments, every second nucleotide of the editing template comprises a 2’-OMe modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’-OMe modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’-OMe modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-OMe modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’-OMe modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’-OMe modification. In some embodiments, each nucleotide of the editing template comprises a 2’-MOE modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-MOE modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-MOE modification. In some embodiments, every second nucleotide of the editing template comprises a 2’-MOE modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’-MOE modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’-MOE modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-MOE modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’-MOE modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’-MOE modification. In some embodiments, the total number of tagRNA nucleotides comprising a chemical modification provided herein (e.g., an LNA, 2’-fluoro, DNA, 2’-OMe, and / or 2’-MOE chemical modification) can be at least, or can be at most, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100, nucleotides. As understood by the skilled artisan, any modification pattern disclosed herein (e.g., according to any of SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791- 792, 1046-1053, and 1089-1203) can be applied to a tagRNA comprising a desired spacer and editing template.
[0168] In some embodiments, a tagRNA complexes with and directs a RT editor to bind to the search target sequence of the target gene. In some embodiments, the bound RT editor generates a nick on the edit strand (PAM strand) of the target gene at the nick site. In some embodiments, a flap binding site (FB sequence) of the tagRNA anneals with a free 3’ end formed at the nick site, and the RT editor initiates DNA synthesis from the nick site, using the free 3’ endas a primer. Subsequently, a single-stranded DNA encoded by the editing template of the tagRNA is synthesized. In some embodiments, the newly synthesized single-stranded DNA comprises one or more intended nucleotide edits compared to an endogenous target gene sequence. Accordingly, in some embodiments, the editing template of a tagRNA is complementary to a sequence in the edit strand except for one or more mismatches at the intended nucleotide edit positions in the editing template. The endogenous, e.g., genomic, sequence that is partially complementary to the editing template may be referred to as an “editing target sequence.” Accordingly, in some embodiments, the newly synthesized single stranded DNA has identity or substantial identity to a sequence in the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions.
[0169] In some embodiments, the newly synthesized single-stranded DNA equilibrates with the editing target on the edit strand of the target gene for pairing with a target strand of a target gene. In some embodiments, an editing target sequence of a target gene is excised by a flap endonuclease (FEN), for example, FEN1. In some embodiments, the FEN is an endogenous FEN, for example, in a cell comprising a target gene. In some embodiments, the FEN is provided as part of the RT editor, either linked to other components of the RT editor or provided in trans. In some embodiments, the newly synthesized single stranded DNA, which comprises the intended nucleotide edit, replaces the endogenous single stranded editing target sequence on the edit strand of the target gene. In some embodiments, the newly synthesized single stranded DNA and the endogenous DNA on the target strand form a heteroduplex DNA structure at the region corresponding to the editing target sequence of the target gene. In some embodiments, the newly synthesized single-stranded DNA comprising the nucleotide edit is paired in the heteroduplex with the target strand of the target DNA that does not comprise the nucleotide edit, thereby creating a mismatch between the two otherwise complementary strands. In some embodiments, the mismatch is recognized by DNA repair machinery, e.g., an endogenous DNA repair machinery. In some embodiments, through DNA repair, the intended nucleotide edit is incorporated into the target gene.
[0170] In some embodiments, a RT editor comprises a Cas9 functional variant that is of smaller molecular weight than a wild-type SPYCas9 protein. In some embodiments, a smaller- sized Cas9 functional variant may facilitate delivery to cells, e.g., by an expression vector, nanoparticle, or other means of delivery. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type II Cas protein. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type V Cas protein. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type VI Cas protein.Nuclear Localization Sequences and linkers
[0171] In some embodiments, a RT editor further comprises one or more nuclear localization sequence (NLS). In some embodiments, the NLS helps promote translocation of a protein into the cell nucleus. In some embodiments, a RT editor comprises a fusion protein, e.g., a fusion protein comprising a DNA endonuclease domain and a DNA polymerase, that comprises one or more NLSs. In some embodiments, one or more polypeptides of the RT editor are fused to or linked to one or more NLSs. In some embodiments, the RT editor comprises a DNA endonuclease domain and a DNA polymerase domain that are provided in trans, wherein the DNA endonuclease domain and / or the DNA polymerase domain is fused or linked to one or more NLSs. In some embodiments the RT editor mRNA comprises or consists of the structure: [5’UTR]- [NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]-[polyA sequence]. The RT editor mRNA can comprises or consist of the structure [5’-cap]-[5’UTR]-[NLS]-[nCas9]-[linker and / or NLS]-[RT]-[NLS]-[3’UTR_and / or_viral_element]-[secondary structure motif and / or polyA sequence]. In some cases, the positions of nCas9 and RT are swapped to e.g. improve editing efficiency.
[0172] In some embodiments, a RT editor or RT editing complex comprises at least one NLS. In some embodiments, a RT editor or RT editing complex comprises at least two NLSs. In embodiments with at least two NLSs, the NLSs can be the same NLS, or they can be different NLSs.
[0173] In addition, the NLSs can be expressed as part of a RT editor complex. The location of the NLS fusion can be at the N-terminus, the C-terminus, or positioned anywhere within a sequence of a RT editor or a component thereof (e.g., inserted between the DNA endonuclease domain and the DNA polymerase domain of a RT editor fusion protein, between the DNA endonuclease domain and a linker sequence, between a DNA polymerase and a linker sequence, between two linker sequences of a RT editor fusion protein or a component thereof, in either N-terminus to C-terminus or C-terminus to N-terminus order).
[0174] Any NLSs that are known in the art are also contemplated herein. The NLSs may be any naturally occurring NLS, or any non-naturally occurring NLS (e.g., an NLS with one or more mutations relative to a wild-type NLS). In some embodiments, the one or more NLSs of a RT editor comprise bipartite NLSs. In some embodiments, a nuclear localization signal (NLS) is predominantly basic. In some embodiments, the one or more NLSs of a RT editor are rich in lysine and arginine residues. In some embodiments, the one or more NLSs of a RT editor comprise proline residues.
[0175] In some embodiments, a nuclear localization signal (NLS) comprises the sequence of any one of: MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 66),KRTADGSEFESPKKKRKV (SEQ ID NO: 11), KRTADGSEFEPKKKRKV (SEQ ID NO: 59), NLSKRPAAIKKAGQAKKKK (SEQ ID NO: 60), RQRRNELKRSF (SEQ ID NO: 61), and NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 62).
[0176] In some embodiments, an NLS is a monopartite NLS. For example, in some embodiments, a NLS is a SV40 large T antigen NLS PKKKRKV (SEQ ID NO: 65). In some embodiments, an NLS is a bipartite NLS. In some embodiments, a bipartite NLS comprises two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, an NLS is a bipartite NLS. In some embodiments, a bipartite NLS consists of two basic domains separated by a spacer sequence comprising a variable number of amino acids. In some embodiments, the spacer amino acid sequence comprises the sequence KRXXXXXXXXXXKKKL (Xenopus nucleoplasmin) (SEQ ID NO: 63), wherein X is any amino acid. In some embodiments, the NLS comprises a nucleoplasmin NLS sequence KRPAATKKAGQAKKKK (SEQ ID NO: 64). In some embodiments, an NLS is a noncanonical sequences such as M9 of the hnRNP Al protein, the influenza virus nucleoprotein NLS, and the yeast Gal4 protein NLS. In some embodiments, a NLS is a noncanonical sequences such as M9 of the hnRNP Al protein, the influenza virus nucleoprotein NLS, and the yeast Gal4 protein NLS.
[0177] In some embodiments, a bipartite NLS consists of two basic domains separated by a spacer sequence comprising a variable number of amino acids, in some embodiments, an NLS comprises an amino acid sequence that is at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of any one of SEQ ID NOs: 5, 12, and 65-73. In some embodiments, an NLS comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 5, 12, and 65-73. in some embodiments, a RT editing composition comprises a polynucleotide that encodes an NLS that comprises an amino acid sequence that is at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the amino acid sequence of any one of SEQ ID NOs: 12, and 65-73. In some embodiments, a RT editing composition comprises a polynucleotide that encodes an NLS that comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 5, 2, and 65-73.
[0178] Non-limiting examples of NLS sequences are provided in Table 1 below. TABLE 1: EXEMPLARY NUCLEAR LOCALIZATION SIGNALS Name / Description Sequence SEQ ID NO: NLS of SV40 Large T-AG PKKKRKV 65 NLS MKRTADGSEFESPKKKRKV 5 NLS MDSLLMNRRKFLYQFKNVRWAKGRRETYLC 66 NLS of Nucleoplasmin AVKRPAATKKAGQAKKKKLD 67 NLS of EGL-13 MSRRRKANPTKLSENAKKLAKEVEN 68 NLS of C-Myc PAAKRVKLD 12 NLS of Tus-protein KLKIKRPVK 69NLS of polyoma large T-AG VSRKRPRP 70 NLS of Hepatitis D virus antigen EGAPPAKRAR 71 NLS of murine p53 PPQPKKKPLDGE 72 C-terminal linker and NLS of an SGGSKRTADGSEFEPKKKRKV 73 exemplary RT editor fusion protein Hybrid NLS PAAKKKKLD 1054 VirD NLS PKRPRDRHDGELGGRKRARG 1055
[0179] In some embodiments, components of a RT editor are directly fused to each other. In some embodiments, components of a RT editor are associated to each other via a linker. As used herein, a linker can be any chemical group or a molecule linking two molecules or moieties, e.g., a DNA binding domain, a DNA endonuclease domain and a polymerase domain of a RT editor. In some embodiments, a linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker comprises a non-peptide moiety. The linker may be a covalent bond (e.g., a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.), or it may be a polymeric linker, for example, a polynucleotide sequence.
[0180] In some embodiments, two or more components of a RT editor are linked to each other by a peptide linker. In some embodiments, a peptide linker is 5-100 amino acids in length, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. In some embodiments, the peptide linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140,150, 160, 175, 180, 190, or 200 amino acids in length. In some embodiments, the peptide linker is 5-100 amino acids in length. In some embodiments, the peptide linker is 10-80 amino acids in length. In some embodiments, the peptide linker is 15- 70 amino acids in length. In some embodiments, the peptide linker is 16 amino acids in length, 24 amino acids in length, 64 amino acids in length, or 96 amino acids in length, in some embodiments, the peptide linker is at least 50 amino acids in length, in some embodiments, the peptide linker is at least 40 amino acids in length, in some embodiments, the peptide linker is at least 30 amino acids in length. In some embodiments, the peptide linker is 46 amino acids in length. In some embodiments, the peptide linker is 92 amino acids in length. In some embodiments, the peptide linker is 16 amino acids in length, 24 amino acids in length, 64 amino acids in length, or 96 amino acids in length.
[0181] In some embodiments, the linker comprises the amino acid sequence (GGGGS)n (SEQ ID NO: 74), (G)n, (EAAAK)n (SEQ ID NO: 75), (GGS)n, (SGGS)n (SEQ ID NO: 76), (XP)n, or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n, wherein n is 1, 3, or 7. In some embodiments, the linker comprises the aminoacid sequence SGSETPGTSESATPES (SEQ ID NO: 77). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 78). In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS (SEQ ID NO: 79). In some embodiments, the linker comprises the amino acid sequence SGGS (SEQ ID NO: 10). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGG S (SEQ ID NO: 80). In some embodiments, the linker comprises the amino acid sequence GGSGGS (SEQ ID NO: 81), GGSGGSGGS (SEQ ID NO: 82), or SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 83).
[0182] Components of a RT editor may be connected to each other in any order. In some embodiments, the DNA binding domain, a DNA endonuclease domain and the DNA polymerase domain of a RT editor may be fused to form a fusion protein or may be joined by a peptide or protein linker, in any order from the N-terminus to the C-terminus. In some embodiments, a RT editor comprises a DNA binding domain, a DNA endonuclease domain fused or linked to the C-terminal end of a DNA polymerase domain. In some embodiments, a RT editor comprises a DNA binding domain, a DNA endonuclease domain fused or linked to the N-terminal end of a DNA polymerase domain. In some embodiments, the RT editor comprises a fusion protein comprising the structure NH2-[DNA binding domain, a DNA endonuclease domain]- [polymerase]-COOH; or NH2-[polymerase]-[DNA binding domain, a DNA endonuclease domain]-COOH, wherein each instanceindicates the presence of an optional linker sequence. In some embodiments, a RT editor comprises a fusion protein and a DNA polymerase domain provided in trans, wherein the fusion protein comprises the structure NH2-[DNA binding domain, a DNA endonuclease domain]-[RNA-protein recruitment polypeptide]-COOH. In some embodiments, a RT editor comprises a fusion protein and a DNA binding domain, a DNA endonuclease domain provided in trans, wherein the fusion protein comprises the structure NH2- [DNA polymerase domain]-[RNA-protein recruitment polypeptide]-COOH.
[0183] In some embodiments, a RT editor fusion protein, a polypeptide component of a RT editor, or a polynucleotide encoding the RT editor fusion protein or polypeptide component, may be split into an N-terminal half and a C-terminal half or polypeptides that encode the N- terminal half and the C-terminal half, and provided to a target DNA in a cell separately. For example, in some embodiments, a RT editor fusion protein may be split into a N-terminal and a C-terminal half for separate delivery in AAV vectors, and subsequently translated and colocalized in a target cell to reform the complete polypeptide or RT editor protein. In such cases, separate halves of a protein or a fusion protein may each comprise a split-intein to facilitate colocalization and reformation of the complete protein or fusion protein by the mechanism of intein facilitatedtrans splicing. In some embodiments, a RT editor comprises a N-terminal half fused to an intein- N, and a C-terminal half fused to an intein-C, or polynucleotides or vectors (e.g., AAV vectors) encoding each thereof. When delivered and / or expressed in a target cell, the intein-N and the intein-C can be excised via protein trans-splicing, resulting in a complete RT editor fusion protein in the target cell. tagRNAs
[0184] Disclosed herein include template armed gRNAs (tagRNAs). In some embodiments, the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a target gene; a scaffold sequence, an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the target gene; and a flap binding sequence, wherein the first strand and the second strand are complementary to each other. In some embodiments, the editing target sequence is within the mutation site of the target gene. In some embodiments, the editing target sequence is within exon 14 of the ATP7B gene. In some embodiments, the editing target sequence is within exon 5 of the SERPINA1 gene. In some embodiments, the editing target sequence is within exon 12 of the PAH gene.
[0185] The tagRNA can comprise from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the FB sequence. In some embodiments, spacer, the scaffold sequence, the editing template, and the FB sequence form a contiguous sequence in a single molecule. The editing template can comprise an intended nucleotide edit compared to the target gene. In some embodiments, the tagRNA guides the RT editor to incorporate the intended nucleotide edit into the target gene when the tagRNA is contacted with the target gene. In some embodiments, the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the target gene. The search target sequence can be complementary to a protospacer sequence in the target gene. The protospacer sequence can be adjacent to a protospacer adjacent motif (PAM) in the target gene. In some embodiments, the tagRNA results in incorporation of the intended nucleotide edit in the PAM when contacted with the target gene. In some embodiments, the target gene is the ATP7B gene. In some embodiments, the target gene is the SERPINA1 gene. In some embodiments, the target gene is the ATP7B gene. In some embodiments, the target gene is the PAH gene.
[0186] In some embodiments, the extension arm comprises a flap binding site sequence (FB sequence) that can initiate target-primed DNA synthesis. In some embodiments, the FB sequence is complementary or substantially complementary to a free 3’ end on the edit strand of the target gene at a nick site generated by the RT editor. In some embodiments, the extension arm further comprises an editing template that comprises one or more intended nucleotide edits tobe incorporated in the target gene by RT editing. In some embodiments, the editing template is a template for an RNA-dependent DNA polymerase domain or polypeptide of the RT editor, for example, a reverse transcriptase domain. The reverse transcriptase editing template may also be referred to herein as an editing template. In some embodiments, the editing template comprises partial complementarity to an editing target sequence in the target gene, e.g., an ATP7B gene, a PAH gene, or a SERPINA1 gene. In some embodiments, the editing template comprises substantial or partial complementarity to the editing target sequence except at the position of the intended nucleotide edits to be incorporated into the target gene.
[0187] In some embodiments, a tagRNA includes only RNA nucleotides and forms an RNA polynucleotide. In some embodiments, a tagRNA is a chimeric polynucleotide that includes both RNA and DNA nucleotides. For example, a tagRNA can include DNA in the spacer sequence, the scaffold sequence, or the extension arm. In some embodiments, a tagRNA comprises DNA in the spacer sequence. In some embodiments, the entire spacer sequence of a tagRNA is a DNA sequence. In some embodiments, the tagRNA comprises DNA in the scaffold sequence, for example, in a stem region of the scaffold sequence. In some embodiments, the tagRNA comprises DNA in the extension arm, for example, in the editing template. An editing template that comprises a DNA sequence may serve as a DNA synthesis template for a DNA polymerase in a RT editor, for example, a DNA-dependent DNA polymerase. Accordingly, the tagRNA may be a chimeric polynucleotide that comprises RNA in the spacer, scaffold sequence, and / or the FB sequence sequences and DNA in the editing template.
[0188] In some embodiments, a spacer sequence comprises a region that has substantial complementarity to a search target sequence on the target strand of a double stranded target DNA, e.g., an AT7B gene, a PAH gene, or a SERPINA1 gene. In some embodiments, the spacer sequence of a tagRNA is identical or substantially identical to a protospacer sequence on the edit strand of the target gene (except that the protospacer sequence comprises thymine and the spacer sequence may comprise uracil). In some embodiments, the spacer sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to a search target sequence in the target gene. In some embodiments, the spacer comprises is substantially complementary to the search target sequence.
[0189] The spacer of the tagRNA can be from 16 to 25 nucleotides in length. The spacer can be 20 nucleotides in length. The spacer can be 21-23 nucleotides in length. In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is from 15 nucleotides to 30 nucleotides in length, 15 to25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, or 20 to 30 nucleotides in length. In some embodiments, the spacer is 16 to 22 nucleotides in length. In some embodiments, the spacer is 16 to 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length.
[0190] As used herein in a gRNA (e.g., a tagRNA or a enhancer guide egRNA sequence), or fragments thereof such as a spacer, FB sequence, or editing template sequence, unless indicated otherwise, it should be appreciated that the letter “T” or “thymine” indicates a nucleobase in a DNA sequence that encodes the tagRNA or guide RNA sequence, and is intended to refer to a uracil (U) nucleobase of the tagRNA or guide RNA or any chemically modified uracil nucleobase known in the art, such as 5-methoxyuracil.
[0191] The extension arm of a tagRNA may comprise a flap binding sequence (FB sequence; FBS) and an editing template (e.g., an ET). The extension arm may be partially complementary to the spacer. In some embodiments, the editing template (e.g., ET) is partially complementary to the spacer. In some embodiments, the editing template (e.g., ET) and the flap binding sequence (FB sequence; FBS) are each partially complementary to the spacer. An extension arm of a tagRNA may comprise a flap binding sequence (FB sequence, or FBS) that comprises complementarity to and can hybridize with a free 3’ end of a single stranded DNA in the target gene (e.g., the ATP7B gene, PAH gene, or SERPINA1 gene) generated by nicking with a RT editor at the nick site on the PAM strand.
[0192] The length of the FB sequence may vary depending on, e.g., the RT editor components, the search target sequence and other components of the tagRNA. The FB sequence can be about 2 to 20 nucleotides in length. The FB sequence can be about 8 to 16 nucleotides in length. In some embodiments, the FB sequence is 6 nucleotides in length In some embodiments, the FB sequence is about 3 to 19 nucleotides in length, or about 3 to 17 nucleotides in length. In some embodiments, the FB sequence is about 4 to 16 nucleotides, about 6 to 16 nucleotides, about 6 to 18 nucleotides, about 6 to 20 nucleotides, about 8 to 20 nucleotides, about 10 to 20 nucleotides, about 12 to 20 nucleotides, about 14 to 20 nucleotides, about 16 to 20 nucleotides, or about 18 to 20 nucleotides in length. In some embodiments, the FB sequence is 8 to 17 nucleotides in length. In some embodiments, the FB sequence is 8 to 16 nucleotides in length. In some embodiments, the FB sequence is 8 to 15 nucleotides in length. In some embodiments, the FB sequence is 8 to 14 nucleotides in length. In some embodiments, the FB sequence is 8 to 13 nucleotides in length. In some embodiments, the FB sequence is 8 to 12 nucleotides in length. In some embodiments, the FB sequence is 8 to 11 nucleotides in length. In some embodiments, the FB sequence is 8 to 10 nucleotides in length. In some embodiments, the FB sequence is 8 or 9 nucleotides in length. In some embodiments, the FB sequence is 16 or 17 nucleotides in length, insome embodiments, the FB sequence is 15 to 17 nucleotides in length. In some embodiments, the FB sequence is 14 to 17 nucleotides in length. In some embodiments, the FB sequence is 13 to 17 nucleotides in length. In some embodiments, the FB sequence is 12 to 17 nucleotides in length. In some embodiments, the FB sequence is 11 to 17 nucleotides in length. In some embodiments, the FB sequence is 10 to 17 nucleotides in length. In some embodiments, the FB sequence is 9 to 17 nucleotides in length. In some embodiments, the FB sequence is about 7 to 15 nucleotides in length. In some embodiments, the FB sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 nucleotides in length. In some embodiments, the FB sequence is 8, 9, 10, 11, 12, 13, or 14 nucleotides in length.
[0193] The FB sequence may be complementary or substantially complementary to a DNA sequence in the edit strand of the target gene. By annealing with the edit strand at a free hydroxy group, e.g., a free 3’ end generated by RT editor nicking activity, the FB sequence may initiate synthesis of a new single stranded DNA encoded by the editing template at the nick site. In some embodiments, the FB sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to a region of the edit strand of the target gene (e.g., the ATP7B gene, PAH gene, or SERPINA1 gene). In some embodiments, the FB sequence is perfectly complementary, or 100% complementary, to a region of the edit strand of the target gene (e.g., the ATP7B gene, PAH gene, or SERPINA1 gene).
[0194] In some embodiments, the FB sequence comprises a nucleotide sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG. In some embodiments, the FB sequence comprises a nucleotide sequence having one, two, or three mismatches relative to a sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG. In some embodiments, the FB sequence comprises or consists of a sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG or a sequence at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a sequence selected from GGCGTGGCAGTCA (SEQ ID NO: 91), GGCGTGGCAGTC (SEQ ID NO: 92), GGCGTGGCAGT (SEQ ID NO: 93), GGCGTGGCAG (SEQ ID NO: 94), GGCGTGGCA, GGCGTGGC, GGCGTGG, and GGCGTG.
[0195] An extension arm of a tagRNA may comprise an editing template that serves as a DNA synthesis template for the DNA polymerase in a RT editor during RT editing. The length of an editing template may vary depending on, e.g., the RT editor components, the search target sequence and other components of the tagRNA. In some embodiments, the editing template serves as a DNA synthesis template for a reverse transcriptase. The editing template can be about 4 to 30 nucleotides in length. The editing template can be about 10 to 30 nucleotides in length. The editing template can be 6 to 9 nucleotides in length. In some embodiments, the editing template is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.
[0196] In some embodiments, the editing template sequence is about 70%, 75%, 80%, 85%, 90%, 95%, or 99% complementary to the editing target sequence on the edit strand of the target gene. In some embodiments, the editing template sequence is complementary to the editing target sequence except at positions of the intended nucleotide edits to be incorporated int the target gene. In some embodiments, the editing template comprises a nucleotide sequence comprising about 85% to about 95% complementarity to (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, complementarity to) an editing target sequence in the edit strand in the target gene. In some embodiments, the editing template comprises about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% complementarity to an editing target sequence in the edit strand of the target gene (e.g., the ATP7B gene, PAH gene, or SERPINA1 gene).
[0197] In some embodiments, the editing template comprises a nucleotide sequence selected from SEQ ID NOs: 95-98. In some embodiments, the editing template comprises a nucleotide sequence selected from SEQ ID NOs: 95-98 or a sequence having one, two, or three mismatches relative to a nucleotide sequence selected from SEQ ID NOs: 95-98. In some embodiments, the editing template comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 95-98 or a sequence at least 85% identical (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a nucleotide sequence selected from SEQ ID NOs: 95-98.
[0198] An intended nucleotide edit or intended nucleotide edits in an editing template of a tagRNA may comprise various types of alterations as compared to the target gene sequence, in some embodiments, the nucleotide edit or edits is nucleotide substitution(s) as compared to the target gene sequence. In some embodiments, the nucleotide edit comprise deletion(s) as compared to the target gene sequence. In some embodiments, the nucleotide edit comprise an insertion ascompared to the target gene sequence. In some embodiments, the editing template comprises one to ten intended nucleotide edits as compared to the target gene sequence, in some embodiments, the editing template comprises one or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises two or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises three or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises four or more, five or more, or six or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises two or more single nucleotide substitutions, insertions, deletions, or any combination thereof, as compared to the target gene sequence. In some embodiments, the editing template comprises three single nucleotide substitutions, insertions, deletions, or any combination thereof as compared to the target gene sequence. In some embodiments, the editing template comprises four, five, or six single nucleotide substitutions, insertions, deletions, or any combination thereof, as compared to the target gene sequence. In some embodiments, a nucleotide substitution comprises an adenine (A)-to-thymine (T) substitution. In some embodiments, a nucleotide substitution comprises an A-to-guanine (G) substitution. In some embodiments, a nucleotide substitution comprises an A-to-cytosine (C) substitution. In some embodiments, a nucleotide substitution comprises a T-A substitution. In some embodiments, a nucleotide substitution comprises a T-G substitution, in some embodiments, a nucleotide substitution comprises a T-C substitution. In some embodiments, a nucleotide substitution comprises a G-to- A substitution. In some embodiments, a nucleotide substitution comprises a G-to-T substitution. In some embodiments, a nucleotide substitution comprises a G-to-C substitution. In some embodiments, a nucleotide substitution comprises a C-to-A substitution, in some embodiments, a nucleotide substitution comprises a C-to-T substitution. In some embodiments, a nucleotide substitution comprises a C-to-G substitution.
[0199] In some embodiments, a nucleotide insertion is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 22 nucleotides, at least 24 nucleotides, at least 26 nucleotides, at least 28 nucleotides, at least 30 nucleotides, at least 32 nucleotides, at least 34 nucleotides, at least 36 nucleotides, at least 38 nucleotides, at least 40 nucleotides, at least 42 nucleotides, at least 44 nucleotides, at least 46 nucleotides, at least 48 nucleotides, or at least 50 nucleotides, in length. In some embodiments, a nucleotide insertion is from 1 to 2 nucleotides, from 1 to 3 nucleotides, from 1 to 4 nucleotides, from 1 to 5 nucleotides,form 2 to 5 nucleotides, from 3 to 5 nucleotides, from 3 to 6 nucleotides, from 3 to 8 nucleotides, from 4 to 9 nucleotides, from 5 to 10 nucleotides, from 6 to 11 nucleotides, from 7 to 12 nucleotides, from 8 to 13 nucleotides, from 9 to 14 nucleotides, from 10 to 15 nucleotides, from 11 to 16 nucleotides, from 12 to 17 nucleotides, from 13 to 18 nucleotides, from 14 to 19 nucleotides, from 15 to 20 nucleotides in length. In some embodiments, a nucleotide insertion is a single nucleotide insertion. In some embodiments, a nucleotide insertion comprises insertion of two nucleotides.
[0200] The editing template of a tagRNA may comprise one or more intended nucleotide edits, compared to the target gene to be edited. Position of the intended nucleotide edit(s) relevant to other components of the tagRNA, or to particular nucleotides (e.g., mutations) in the target gene may vary. In some embodiments, the nucleotide edit is in a region of the tagRNA corresponding to or homologous to the protospacer sequence. In some embodiments, the nucleotide edit is in a region of the tagRNA corresponding to a region of the target gene outside of the protospacer sequence.
[0201] In some embodiments, the position of a nucleotide edit incorporation in the target gene may be determined based on position of the protospacer adjacent motif (PAM). For instance, the intended nucleotide edit may be installed in a sequence corresponding to the protospacer adjacent motif (PAM) sequence. In some embodiments, a nucleotide edit in the editing template is at a position corresponding to the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit in the editing template is at a position corresponding to the 3’ most nucleotide of the PAM sequence, in some embodiments, position of an intended nucleotide edit in the editing template may be referred to by aligning the editing template with the partially complementary edit strand of the target gene, and referring to nucleotide positions on the editing strand where the intended nucleotide edit is incorporated, in some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream of the 5’ most nucleotide of the PAM sequence in the edit strand of the target gene. By 0 base pair upstream or downstream of a reference position, it is meant that the intended nucleotide is immediately upstream or downstream of the reference position. In some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to l6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 3 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in is incorporated at a position corresponding to 4 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 5 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in the editing template is at a position corresponding to 6 nucleotides upstream of the 5’ most nucleotide of the PAM sequence.
[0202] In some embodiments, an intended nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides downstream of the 5’ most nucleotide of the PAM sequence in the edit strand of the target gene. In some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to 16 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 3 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to4 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 5 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 6 nucleotides downstream of the 5’ most nucleotide of the PAM sequence.
[0203] In some embodiments, the position of a nucleotide edit incorporation in the target gene can be determined based on position of the nick site. In some embodiments, position of an intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides apart from the nick site. In some embodiments, position of an intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides downstream of the nick site on the PAM strand (or the non-target strand, or the edit strand) of the double stranded target DNA. In some embodiments, position of the intended nucleotide edit in the editing template can be referred to by aligning the editing template with the partially complementary editing target sequence on the edit strand and referring to nucleotide positions on the editing strand where the intended nucleotide edit is incorporated. Accordingly, in some embodiments, a nucleotide edit in an editing template is at a position corresponding to a position about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides apart from the nick site. In some embodiments, a nucleotide edit in an editing template is at a position corresponding to a position about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, 40 to 50 nucleotides,50 to 60 nucleotides, 60 to 70 nucleotides, 70 to 80 nucleotides, 80 to 90 nucleotides, 90 to 100 nucleotides, 100 to 110 nucleotides, 110 to 120 nucleotides, 120 to 130 nucleotides, 130 to 140 nucleotides, or 140 to 150 nucleotides apart from the nick site. In some embodiments, when referred to in the context of the PAM strand (or the non-target strand, or the edit strand), a nucleotide edit in an editing template is at a position corresponding to a position about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, 40 to 50 nucleotides, 50 to 60 nucleotides, 60 to 70 nucleotides, 70 to 80 nucleotides, 80 to 90 nucleotides, 90 to 100 nucleotides, 100 to 110 nucleotides, 110 to 120 nucleotides, 120 to 130 nucleotides, 130 to 140 nucleotides, or 140 to 150 nucleotides downstream from the nick site. The relative positions of the intended nucleotide edit(s) and nick site may be referred to by numbers. For example, in some embodiments, the nucleotide immediately downstream of the nick site on a PAM strand (or the non-target strand, or the edit strand) may be referred to as at position 0. The nucleotide immediately upstream of the nick site on the PAM strand (or the non-target strand, or the edit strand) may be referred to as at position -1. The nucleotides downstream of position 0 on the PAM strand can be referred to as at positions +1, +2, +3, +4, ... +n, and the nucleotides upstream of position -1 on the PAM strand may be referred to as at positions -2, -3, -4, .. -n. Accordingly, in some embodiments, the nucleotide in the editing template that corresponds to position 0 when the editing template is aligned with the partially complementary editing target sequence by complementarity can also be referred to as position 0 in the editing template, the nucleotides in the editing template corresponding to the nucleotides at positions +1, +2, +3, +4, ..., +n on the PAM strand of the double stranded target DNA can also be referred to as at positions +1, +2, +3, +4, ..., -in in the editing template, and the nucleotides in the editing template corresponding to the nucleotides at positions -1, -2, -3, -4, ..., -n on the PAM strand on the double stranded target DNA may also be referred to as at positions -1, -2, -3, -4 -n on the editing template, even though whenthe tagRNA is viewed as a standalone nucleic acid, positions +1, +2, +3, +4, ..., +n are 5' of position 0 and positions -1, -2, -3, -4, ...-n are 3' of position 0 in the editing template. In some embodiments, an intended nucleotide edit is at position +n of the editing template relative to position 0. Accordingly, the intended nucleotide edit may be incorporated at position +n of the PAM strand of the double stranded target DNA (and subsequently, the target strand of the double stranded target DNA) by RT editing. The number n may be referred to as the nick to edit distance.
[0204] When referred to within the tagRNA, positions of the one or more intended nucleotide edits may be referred to relevant to components of the tagRNA. For example, an intended nucleotide edit may be 5’ or 3’ to the FB sequence. In some embodiments, a tagRNA comprises the structure, from 5’ to 3’: a spacer, a gRNA core, an editing template, and a FB sequence. In some embodiments, the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream to the 5’ most nucleotide of the FB sequence. in some embodiments, the intended nucleotide edit is 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides upstream to the 5’ most nucleotide of the FB sequence.
[0205] The corresponding positions of the intended nucleotide edit incorporated in the target gene may also be referred to based on the nicking position (i.e., the nick site) generated by a RT editor based on sequence homology and complementarity. For example, in some embodiments, the distance between the intended nucleotide edit to be incorporated into the target gene and the nick site (also referred to as the “nick to edit distance”) may be determined by the position of the nick site and the position of the nucleotide(s) corresponding to the intended nucleotide edit(s), for example, by identifying sequence complementarity between the spacer and the search target sequence and sequence complementarity between the editing template and the editing target sequence. In some embodiments, the position of the nucleotide edit can be in anyposition downstream of the nick site on the edit strand (or the PAM strand) generated by the RT editor, such that the distance between the nick site and the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, the position of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the nick site on the edit strand. In some embodiments, the position of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream of the nick site on the edit strand. In some embodiments, the position of the nucleotide edit is 0 base pair from the nick site on the edit strand, that is, the editing position is at the same position as the nick site. As used herein, the distance between the nick site and the nucleotide edit, for example, where the nucleotide edit comprises an insertion or deletion, refers to the 5’ most position of the nucleotide edit for a nick that creates a 3’ free end on the edit strand (i.e., the “near position” of the nucleotide edit to the nick site). Similarly, as used herein, the distance between the nick site and a PAM position edit, for example, where the nucleotide edit comprises an insertion, deletion, or substitution of two or more contiguous nucleotides, refers to the 5’ most position of the nucleotide edit and the 5’ most position of the PAM sequence.
[0206] In some embodiments, the editing template extends beyond a nucleotide edit to be incorporated into the target gene sequence. For example, in some embodiments, the editing template comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides 3’ to the nucleotide edit to be incorporated into the target gene sequence. In some embodiments, the editing template comprises 1 to 80 nucleotides 3’ to the nucleotide edit to be incorporated into the target gene sequence.
[0207] In some embodiments, the editing template can comprise a second editing sequence comprising a second mutation relative to a target sequence. The second mutation can be designed to mutate or otherwise silence a PAM sequence such that a corresponding nucleic acid guided nuclease or CRISPR nuclease is no longer able to cleave the target sequence. In some embodiments, this mutation or silencing of a PAM can serve as a method for selecting transformants in which the first editing sequence has been incorporated. In some embodiments, the mutation is in at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleic acids in a PAM motif.
[0208] The editing template of a tagRNA may encode a new single stranded DNA (e.g., by reverse transcription) to replace a editing target sequence in the target gene. In some embodiments, the editing target sequence in the edit strand of the target gene is replaced by thenewly synthesized strand, and the nucleotide edit(s) are incorporated into the region of the target gene.
[0209] In some embodiments, the target gene is ATP7B gene. In some embodiments, the editing template of the tagRNA encodes a newly synthesized single stranded DNA that comprises a wild type ATP7B gene sequence. In some embodiments, the newly synthesized DNA strand replaces the editing target sequence in the target ATP7B gene, wherein the editing target sequence (or the endogenous sequence complementary to the editing target sequence on the target strand of the ATP7B gene) comprises a mutation compared to a wild-type ATP7B gene. In some embodiments, the mutation is associated with Wilson’s disease.
[0210] In some embodiments, the target gene is SERPINA1 gene. In some embodiments, the editing template of the tagRNA encodes a newly synthesized single stranded DNA that comprises a wild type SERPINA1 gene sequence. In some embodiments, the newly synthesized DNA strand replaces the editing target sequence in the target SERPINA1 gene, wherein the editing target sequence (or the endogenous sequence complementary to the editing target sequence on the target strand of the SERPINA1 gene) comprises a mutation compared to a wild-type SERPINA1 gene. In some embodiments, the mutation is associated with A1ATD disease or disorder.
[0211] In some embodiments, the target gene is the PAH gene. In some embodiments, the editing template of the tagRNA encodes a newly synthesized single stranded DNA that comprises a wild type PAH gene sequence. In some embodiments, the newly synthesized DNA strand replaces the editing target sequence in the target PAH gene, wherein the editing target sequence (or the endogenous sequence complementary to the editing target sequence on the target strand of the PAH gene) comprises a mutation compared to a wild-type PAH gene. In some embodiments, the mutation is associated with phenylketonuria or hyperphenylalaninemia.
[0212] In some embodiments, the editing target sequence comprises a mutation in exon 1, exon 2, exon 3, exon 4, exon 5, exon 6, exon 7, exon 8, exon 9, exon 10, exon 11, exon 12, exon 13, exon 14, exon 15, exon 16, exon 17, exon 18, exon 19, exon 20, or exon 21 of the target gene as compared to a wild type target gene, in some embodiments, the editing target sequence comprises a mutation in exon 8, exon 13, exon 14, exon 15, or exon 17 of the ATP7B gene as compared to a wild type ATP7B gene. In some embodiments, the editing target sequence comprises a mutation in exon 14 of the ATP7B gene as compared to a wild type ATP7B gene. In some embodiments, the editing target sequence comprises a mutation that encodes an amino acid substitution H1069Q relative to a wild type ATP7B polypeptide. In some embodiments, the ATP7B gene sequence comprises a C>A (e.g., C-A, C to A) mutation. In some embodiments, the mutation is the S allele mutation or the Z allele mutation in SERPINA1 gene associated withalpha-1 antitrypsin deficiency disease (A1ATD). In some embodiments, the mutation is the Z allele mutation on exon 5 of SERPINA1. In some embodiment, the mutation is an E342K mutation commonly associated with A1ATD. E342K mutation leads to misfolding of the alpha-1- antitrypsin (AAT) protein leading to polymers and liver damage. In some embodiments, the mutation is the R408W mutation in PAH gene associated with phenylketonuria or hyperphenylalaninemia.
[0213] The compositions and methods described herein can be used to correct a disease-causing mutation or otherwise edit the PAH gene at any position. Exemplary mutations in PAH that can be edited with the disclosed compositions and methods include, but are not limited to: p.M1L; p.M1V; p.M1R; p.M1T; p.M1I; p.V5Sfs*3; p.N8Ifs*30; p.G10G; p.R13Qfs*5; p.L11*; p.L15Qfs*24; p.S16P; p.D17Lfs*22; p.S16*; p.S16Y; p.D17*; p.Q20Kfs*17; p.Q20*; p.Q20L; p.Q20P; p.Q20Pfs*7; p.Q20H; p.T22K; p.E26Lfs*13; p.I25Mfs*13; p.N28=; p.C29*; p.I35M; p.S36*; p.L37P; p.I38Dfs*19; p.F39del; p.F39L; p.S40L; p.L41F; p.L41P; p.K42del; p.K42E; p.K42I; p.E43*; p.E44del; p.V45A; p.G46R; p.G46S; p.G46Vfs*15; p.A47E; p.A47V; p.L48=; p.L48V; p.L48S; p.L48W; pA49T; p.V51L; p.L52Cfs*9; p.L52S; p.L52F; p.R53C; p.R53H; p.L54S; p.F55del; p.F55S; p.F55L; p.F55Lfs*6; p.E56=; p.E56D; p.E57Rfs*3; p.E57*; p.E57Cfs*77; p.E57del; p.E57K; p.N58*; p.N58*; p.D59Y; p.D59G; p.D59V; p.N61D; p.N61S; p.N61K; p.L62*; p.L62Pfs*7; p.L62V; p.H64Dfs*4; p.L62P; p.L62Pfs*3; p.T63_H64delinsPN; p.T63P; p.H64Pfs*10; p.H64*; p.H64N; p.H64Tfs*9; p.I65V; p.I65Kfs*9; p.I65N; p.I65S; p.I65T; p.I65M; p.E66*; p.E66K; p.E66Afs*8; p.S67A; p.S67P; p.R68G; p.R68S; p.P69_S70dup; p.P69S; p.P69T; p.S70del; p.S70Ffs*7; p.S70del; p.S70P; p.S70F; p.R71C; p.R71H; p.R71P; p.D75N; p.N75H; p.D75G; p.D75V; p.E76*; p.E76K; p.E76A; p.E76G; p.E76D; p.Y77H; p.Y77*; p.E78K; p.E78Q; p.E78V; p.E78Ffs*13; p.F79Ifs*7; p.T81P; p.T81Vfs*6; p.T81N; p.L83Wfs*9; p.D84Y; p.D84G; p.K85*; p.R86P; p.S87R; p.P89S; p.P89T; p.A90Cfs*12; p.T92I; p.Q93Sfs*5; p.N93Kfs*5; p.I95F; p.I95del; p.I95T; p.I95I; p.I97L; p.L98V; p.L98S; p.H100R; p.D101N; p.I102T; p.G103C; p.G103S; p.G103D; p.A104_V106del; p.A104S; p.A104D; p.A104V; p.V106A; p.H107P; p.H107R; p.E108*; p.E108Dfs*4; p.S110L; p.R111*; p.R111Q; p.D112H; p.D112Efs*2; p.K113_K114fs*81; p.K115Rfs*81; p.K115Tfs*79; p.D116Hfs*29; p.T117I; p.T117Kfs*78; p.P119S; p.P119L; p.W120Gfs*75; p.W120*; p.F121L; p.F121V; p.F121S; p.P122S; p.P122Q; p.R123I; p.T124I; p.E127K; p.E127G; p.L128P; p.D129Y; p.D129G; p.D129V; p.A132V; p.N133Rfs*61; p.Q134*; p.S137S; p.L142P; p.D143G; p.D143V; p.D145N; p.D145V; p.H146Y; p.P147S; p.P147L; p.P147=; p.G148Lfs*105; p.G148R; p.G148S; p.G148Wfs*29; p.G148D; p.G148V; p.F149S; p.D151RfsTer13; p.D151H; p.D151G; p.D151E; p.P152Lfs*43; p.P152T; p.R155Lfs*43; p.Y154H; p.Y154N; p.Y154C; p.Y154F; p.Y154*; p.R155C; p.R155Vfs*40; p.R155H; p.R155P; p.A156P; p.R157I; p.R157K; p.R157N; p.R157T;p.R157S; p.R158G; p.R158W; p.R158P; p.R158Q; p.K159*; p.Q160*; p.Q160P; p.F161S; p.D163N; p.I164V; p.I164T; p.A165P; p.A165T; p.A165Vfs*35; p.A165D; p.Y166*; p.N167Y; p.N167I; p.N167S; p.Y168H; p.Y168N; p.Y168Sfs*27; p.Y168*; p.R169C; p.R169G; p.R169S; p.R169H; p.R169Pfs*26; p.H170D; p.H170P; p.H170R; p.H170Q; p.G171R; p.G171W; p.G171A; p.G171E; p.Q172*; p.Q172R; p.Q172H; p.P173T; p.I174V; p.I174N; p.I174T; p.P175A; p.P175S; p.P175R; p.R176*; p.R176L; p.R176P; p.R176Q; p.V177L; p.V177M; p.V177A; p.E178K; p.E178G; p.E178V; p.Y179H; p.Y179N; p.M180V; p.M180Ifs*18; p.E181Kfs*13; p.E183del; p.E182K; p.E182G; p.E183L; p.E183Q; p.E183G; p.K184Rfs*11; p.K184N; p.T186Hfs*9; p.T186I; p.W187Gfs*12; p.W187Gfs*8; p.W187R; p.W187*; p.W187C; p.G188Afs*7; p.G188D; p.G188V; p.V190M; p.V190A; p.V190G; p.K192*; p.T193P; p.T193I; p.L194Dfs*6; p.L194Efs*5; p.L194P; p.L194R; p.S196Vfs*4; p.S196Lfs*2; p.S196T; p.S196Y; p.L197*; p.L197W; p.Y198Sfs*136; p.L197F; p.Y198Vfs*9; p.Y198D; p.Y198Sfs*136; p.Y198*; p.T200Nfs*6; p.T200N; p.H201Y; p.H201P; p.H201R; p.H201Q; p.A202T; p.A202V; p.C203G; p.C203Lfs*3; p.C203S; p.C203Y; p.C203*; p.C203C; p.C203W; p.Y204*; p.Y204Y; p.E205K; p.E205A; p.E205D; p.Y206D; p.Y206C; p.Y206*; p.N207D; p.N207S; p.P211T; p.P211Hfs*130; p.P211L; p.L212P; p.E214Kfs*127; p.L213P; p.E214G; p.Y216*; p.C217G; p.C217R; p.C217F; p.C217Y; p.G218Afs*123; p.G218V; p.F219S; p.H220P; p.E221G; p.D222*; p.D222G; p.D222V; p.N223Y; p.Q226Tfs*118; p.N223I; p.N223Kfs*118; p.I224T; p.I224M; p.P225A; p.P225T; p.P225L; p.P225R; p.Q226*; p.Q226K; p.Q226H; p.L227M; p.L227V; p.L227Q; p.E228*; p.E228K; p.E228G; p.E228D; p.D229Efs*54; p.D229G; p.V230I; p.V230A; p.V230G; p.S231Vfs*52; p.S231P; p.S231F; p.Q232*; p.Q232E; p.Q232P; p.Q232R; p.Q232Q; p.F233I; p.F233L; p.Q235*; p.Q235P; p.T236Mfs*60; p.C237F; p.T238A; p.T238P; p.G239S; p.G239A; p.G239D; p.G239V; p.F240V; p.F240S; p.F240=; p.R241C; p.R241H; p.R241L; p.R241P; p.R241Pfs*100; p.L242F; p.L242=; p.R243*; p.R243L; p.R243Q; p.P244S; p.P244L; p.V245I; p.V245L; p.V245M; p.V245A; p.V245E; p.V245V; p.A246D; p.A246V; p.A246Vfs*95; p.G247R; p.G247S; p.G247Afs*94; p.G247D; p.G247V; p.L248P; p.L248R; p.L248Rfs*93; p.L249F; p.L249Ffs*92; p.L249H; p.L249P; p.S250F; p.R252Gfs*30; p.R252Gfs*89; p.R252G; p.R252W; p.R252L; p.R252P; p.R252Q; p.F254I; p.L255=; p.L255V; p.L255S; p.L255W; p.G257C; p.G257S; p.G257D; p.G257V; p.L258=; p.L258P; p.L258R; p.A259T; p.A259V; p.F260I; p.F260L; p.R261*; p.R261G; p.R261L; p.R261P; p.R261Q; p.V262G; p.F263S; p.F263L; p.H264L; p.C265G; p.C265R; p.C265Y; p.C265*; p.T266A; p.T266E; p.T266P; p.T266I; p.Q267*; p.Q267E; p.Q267L; p.Q267R; p.Q267H; p.Y268H; p.Y268C; p.Y268*; p.I269L; p.I269N; p.I269Tfs*72; p.R270G; p.R270I; p.R270K; p.H271Ifs*10; p.R270S; p.H271Y; p.H271L; p.H271R; p.H271Q; p.G272*; p.S273P; p.S273F; p.K274E; p.K274Nfs*5; p.P275A; p.P275S; p.P275L; p.P275R; p.P275=; p.M276V;p.M276K; p.M276R; p.M276T; p.M276I; p.Y277D; p.Y277C; p.T278A; p.E280Nfs*61; p.T278I; p.T278N; p.T278S; p.P279T; p.P279L; p.E280Nfs*61; p.P279P; p.E280_IVS7+3fs; p.E280*; p.E280K; p.E280Q; p.E280A; p.E280Dfs*3; p.E280G; p.P281A; p.P281fs; p.P281S; p.P281L; p.P281R; p.D282N; p.D282G; p.I283F; p.I283V; p.I283N; p.C284R; p.C284Y; p.C284*; p.H285Y; p.E286K; p.E286=; p.L287V; p.L288F; p.G289R; p.G289R; p.H290Y; p.H290L; p.H290R; p.H290Q; p.V291L; p.V291M; p.P292S; p.P292H; p.P292L; p.292=; p.L293M; p.L293S; p.S295*; p.D296H; p.D296G; p.R297C; p.R297H; p.R297L; p.F299del; p.F299C; p.F299S; p.A300S; p.A300D; p.A300V; p.Q301*; p.Q301P; p.Q301H; p.F302fs*39; p.F302V; p.S303A; p.S303P; p.S303Pfs*38; p.Q304*; p.Q304K; p.Q304P; p.Q304R; p.Q304Q; p.I306Lfs*35; p.I306V; p.G307Afs*34; p.G307D; p.L308F; p.L308V; p.A309T; p.A309D; p.A309V; p.S310_L311del; p.S310P; p.L311*; p.S310C; p.S310F; p.S310Y; p.L311*; p.L311Gfs*4; p.L311P; p.L311R; p.G312C; p.G312R; p.G312S; p.G312D; p.G312V; p.G312Vfs*29; p.A313T; p.A313V; p.P314A; p.P314S; p.P314T; p.P314H; p.P314Lfs*27; p.D315Y; p.Y317H; p.I318T; p.K320N; p.L321F; p.L321L; p.A322T; p.A322D; p.A322G; p.A322V; p.T323del; p.T323I; p.T323T; p.I324V; p.I324N; p.Y325C; p.Y325S; p.Y325*; p.W326Gfs*15; p.W326*; p.W326S; p.F327L; p.T328A; p.T328P; p.T328I; p.T328N; p.E330D; p.F331L; p.F331C; p.F331S; p.G332R; p.G332E; p.G332V; p.L333F; p.L333P; p.C334R; p.C334S; p.C334*; p.K335E; p.K335T; p.Q336*; p.Q336R; p.G337V; p.D338Y; p.S339F; p.S339Y; p.I340T; p.K341*; p.K341R; p.K341T; p.A342Hfs*59; p.K341K; p.K341N; p.A342Hfs*58; p.A342P; p.A342S; p.A342T; p.A342E; p.Y343D; p.Y343N; p.Y343C; p.Y343F; p.Y343*; p.G344R; p.G344S; p.G344D; p.G344V; p.A345S; p.A345T; p.G346R; p.G346E; p.L347Sfs*53; p.L347F; p.L347=; p.L348Cfs*52; p.L348V; p.L348P; p.L348Rfs*2; p.S350Vfs*5; p.S349A; p.S349P; p.S349*; p.S349L; p.S350T; p.S350Y; p.G352C; p.G352Cfs*48; p.G352R; p.G352Vfs*48; p.E353Nfs*47; p.E353*; p.E353Vfs*42; p.L354Ffs*40; p.Q355*; p.Y356D; p.Y356H; p.Y356*; p.C357G; p.C357R; p.C357Y; p.C357*; p.L358V; p.L358F; p.S359*; p.S359L; p.E360*; p.K361Q; p.P362S; p.P362T; p.P362L; p.K363Afs*30; p.K363N; p.K363Nfs*37; p.L364F; p.L365_L369del; p.L365del; p.P366A; p.P366H; p.L367Pfs*27; p.L367V; p.L367Wfs*33; p.L367P; p.L367R; p.L367L; p.E368K; p.E368G; p.L369V; p.E370Afs*25; p.L369L; p.E370G; p.E370D; p.K371R; p.T372S; p.T372R; p.A373Hfs*20; p.A373T; p.A373D; p.A373=; p.I374Yfs*20; p.Q375E; p.Q375R; p.Q375H; p.N376Ifs*24; p.N376T; p.Y377D; p.Y377Tfs*23; p.Y377C; p.T378S; p.V379A; p.T380Gfs*13; p.T380M; p.F382I; p.F382L; p.F382L; p.Q383*; p.Q383E; p.P384S; p.L385P; p.L385=; p.Y386D; p.Y386H; p.Y386C; p.Y386Ffs*14; p.Y387D; p.Y387H; p.Y387*; p.Y387=; p.V388Gfs*5; p.V388L; p.V388M; p.V388A; p.V388Gfs*5; p.A389E; p.A389Efs*11; p.A389G; p.E390_S391delin; sGG; p.E390Dfs*4; p.E390G; p.S391Ffs*2; p.S391G; p.S391I; p.S391T;p.F392I; p.F392S; p.N393Ifs*2; p.D394H; p.D394Y; p.D394A; p.A395P; p.A395S; p.A395D; p.A395G; p.K396R; p.K398E; p.K398K; p.K398N; p.V399A; p.V399Gfs*52; p.V399V; p.R400Gfs*52; p.R400R; p.R400K; p.R400T; p.N401_S439del; p.N401Tfs*51; p.R400S; p.F402I; p.F402L; p.F402V; p.F402C; p.A403V; p.I406Sfs*17; p.I406Sfs*15; p.I406V; p.I406T; p.I406M; p.P407S; p.P407T; p.P407L; p.P407Lfs*45; p.R408Gfs*44; p.R408W; p.R408L; p.R408Q; p.P409S; p.P409L; p.F410I; p.F410L; p.F410C; p.F410S; p.S411*; p.V412D; p.V412G; p.R413C; p.R413G; p.R413S; p.R413H; p.R413P; p.Y414H; p.Y414C; p.Y414*; p.Y414Y; p.D415N; p.D415Y; p.D415A; p.D415V; p.P416Hfs*36; p.P416T; p.P416Q; p.Y417D; p.Y417H; p.Y417N; p.Y417C; p.Y417*; p.T418P; p.T418N; p.Q419P; p.Q419R; p.R420M; p.I421S; p.I421T; p.E422K; p.L424Wfs*28; p.L424*; p.L424S; p.N426N; p.Q428Sfs*24; p.Q429K; p.Q429P; p.L430P; p.L430R; p.L433Ffs*3; p.A434D; p.A434V; p.D435N; p.D435V; p.S436Pfs*16; p.N438D; p.N438fs; p.G442Wfs*47; p.L444F; p.L444P; p.A447P; p.A447D; p.L448Rfs*3; p.*453Vext*35; p.*453Pext*33; 1378G>T; 1503A>G; 1546G>A; 4171_-405del; -4173_-407del; -626G>A; -480delACT; -224G>A; -147C>T; -81C>T; -71A>C; -52G>C; -1C>T; c.60+4A>T; c.60+5_60+6del; c.60+5G>A; c.60+5G>C; c.60+5G>T; c.60+40T>G; c.60+62C>T; c.61-13del9; c.61-3T>C; c.168+1G>A; c.168+1G>T; c.168+2T>C; c.168+5G>A; c.168+5G>C; c.168+5G>T; c.168+6T>G; c.168+19T>C; c.169-13T>G; c.169- 2A>G; c.169-1G>A; c.352+1G>A; c.352+7757C>G; c.353-22C>A; c.353-22C>T; c.353-6T>A; c.353-2A>G; c.353-2A>T; c.353-1G>A; c.353-1G>C; c.353-1G>T; c.441+1delGT; c.441+1G>A; c.441+1G>C; c.441+1G>T; c.441+2T>C; c.441+2_441+3del; c.441+2T>A; c.441+2T>G; c.441+3G>C; c.441+4A>G; c.441+5G>A; c.441+5G>T; c.441+6T>A; c.441+6T>C; c.441+47C>T; c.442-14C>T; c.442-5C>G; c.442-2A>C; c.442-1G>A; c.509+1delG; c.509+1G>A; c.509+5delG; c.509+54C>G; c.509+101A>C; c.510-21_665del177; c.510-54G>A; c.510-6T>A; c.510-6T>G; c.510-2A>G; c.510-1G>A; c.706+1G>T; c.706+4A>T; c.706+5G>A; c.706+17G>T; c.706+36T>G; c.706+44T>G; c.707-96A>G; c.707- 59C>G; c.707-7A>T; c.707-2A>G; c.707-1G>A; c.707-1G>C; c.842+1G>A; c.842+1G>T; c.842+2T>A; c.842+3G>C; c.842+4A>G; c.842+4A>T; c.842+5G>A; c.842+6T>A; c.842+25G>T; c.843-14_11del; c.843-6T>C; c.843-5T>C; c.843-4del; c.843-2A>T; c.912+1G>A; c.912+1G>C; c.912+1G>T; c.912+2T>C; c.912+3A>C; c.912+16T>A; c.913- 8A>G; c.913-7A>G; c.913-5T>C; c.913-5T>G; c.913-3C>G; c.913-2A>G; c.969+1G>A; c.969+2insT; c.969+4A>T; c.969+5G>A; c.969+5G>C; c.969+6T>A; c.969+6T>C; c.969+34G>A; c.969+43G>T; c.970-7A>G; c.970-6T>G; c.970-5T>A; c.970-3C>T; c.970- 2A>C; c.970-2A>G; c.970-1G>A; c.970-1G>C; c.970-1G>T; c.1065+1G>A; c.1065+1G>T; c.1065+3A>C; c.1065+3A>G; c.1065+7C>A; c.1065+32T>A; c.1065+39G>T; c.1065+97G>A; c.1066-193G>C; c.1066-31G>A; c.1066-15A>C; c.1066-14C>G; c.1066-13del; c.1066-13T>G;c.1066-12del; c.1066-11G>A; c.1066-10G>A; c.1066-7C>A; c.1066-3C>G; c.1066-3C>T; c.1066-2A>G; c.1066-2A>T; c.1066-1G>A; c.1066-1G>C; c.1066-1G>T; c.1199+1G>A; c.1199+1G>C; c.1199+1G>T; c.1199+2T>C; c.1199+2T>G; c.1199+4A>G; c.1199+5G>A; c.1199+5G>A; c.1199+17G>A; c.1199+19T>C; c.1199+20G>C; c.1199+50G>A; c.1199+88del; c.1200-35C>T; c.1200-12del; c.1200-8G>A; c.1200-3T>G; c.1200-2A>G; c.1200-2A>C; c.1200-1del; c.1200-1G>A; c.1200-1G>C; c.1200-1G>T; c.1315+1G>A; c.1315+1G>T; c.1315+2T>C; c.1315+4A>G; c.1315+5G>A; c.1315+5G>C; c.1315+5G>T; c.1315+6T>A; c.1316-35C>T; c.1316-16T>G; c.1316-15T>C; c.1316-5T>C; c.1316-2A>C; and c.1316-1G>A. Mutations in the PAH gene are also described in Chen, Ting, et al. ("Mutational and phenotypic spectrum of phenylalanine hydroxylase deficiency in Zhejiang Province, China." Scientific Reports 8.1 (2018): 17137) and Hillert, Alicia, et al. ("The genetic landscape and epidemiology of phenylketonuria." The American Journal of Human Genetics 107.2 (2020): 234-250), the contents of which are hereby incorporated by reference in their entireties.
[0214] In some embodiments, the tagRNA results in incorporation of the intended nucleotide edit about 0 to 27 base pairs downstream (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27 base pairs downstream) of the 5’ end of the PAM when contacted with the target gene. The intended nucleotide edit can comprise a single nucleotide substitution compared to the region corresponding to the editing target in the target gene. The single nucleotide substitution can be a T>G substitution. The intended nucleotide edit can comprise an insertion compared to the region corresponding to the editing target in the target gene. The intended nucleotide edit can comprise a deletion compared to the region corresponding to the editing target in the target gene. The editing target sequence can comprise a mutation associated with a disease or disorder. In some embodiments, the mutation encodes an amino acid substitution. In some embodiments, the amino acid substitution is H1069Q. The editing template can comprise a wild-type ATP7B gene sequence. In some embodiments, the tagRNA results in correction of the mutation when contacted with the ATP7B gene. In some embodiments, the mutation is the S allele mutation or the Z allele mutation in SERPINA1 gene associated with alpha-1 antitrypsin deficiency disease (A1ATD). The Z allele mutation on exon 5 of SERPINA1 is an E342K mutation commonly associated with A1ATD. E342K mutation leads to misfolding of the alpha-1-antitrypsin (AAT) protein leading to polymers and liver damage. The editing template can comprise a wild-type SERPINA1 gene sequence. In some embodiments, the tagRNA results in correction of the mutation when contacted with the SERPINA1 gene.
[0215] In some embodiments, the tagRNA comprises a nucleotide sequence selected from SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089- 1203. In some embodiments, the tagRNA comprises a nucleotide sequence selected from SEQ IDNOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203 or a sequence have one, two, or three mismatches relative to a nucleotide sequence selected from SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203. In some embodiments, the tagRNA comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203 or a sequence at least 85% identical to a nucleotide sequence selected from SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to a nucleotide sequence selected from SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203).
[0216] A scaffold sequence (also referred to herein as the gRNA core, gRNA scaffold, or gRNA backbone sequence) of a tagRNA may contain a polynucleotide sequence that binds to a DNA binding domain, a DNA endonuclease domain (e.g., Cas9) of a RT editor. The scaffold sequence may interact with a RT editor as described herein, for example, by association with a DNA binding domain, a DNA endonuclease domain, such as a DNA nickase of the RT editor.
[0217] One of skill in the art will recognize that different RT editors having different DNA binding domain and a DNA endonuclease domains from different DNA binding domain and a DNA endonuclease proteins may require different gRNA core sequences specific to the DNA binding domain and a DNA endonuclease protein. In some embodiments, the scaffold sequence is capable of binding to a Cas9-based RT editor. In some embodiments, the scaffold sequence is capable of binding to a Cpfl -based RT editor. In some embodiments, the scaffold sequence is capable of binding to a Casl2b-based RT editor. In some embodiments, the scaffold sequence is capable of binding to a Cas9 nickase of any one of the sequences selected from SEQ ID NOs: 6, 142-163 and 270-314.
[0218] In some embodiments, the scaffold sequence comprises regions and secondary structures involved in binding with specific CRISPR Cas proteins. For example, in a Cas9 based RT editing system, the scaffold sequence of a tagRNA may comprise one or more regions of a base paired “lower stem” adjacent to the spacer sequence and a base paired “upper stem” following the lower stem, where the lower stem and upper stem may be connected by a “bulge” comprising unpaired RNAs. The scaffold sequence may further comprise a “nexus” distal from the spacer sequence, followed by a hairpin structure, e.g., at the 3’ end. In some embodiments, the scaffold sequence comprises modified nucleotides as compared to a wild type gRNA core in the lower stem, upper stem, and / or the hairpin. For example, nucleotides in the lower stem, upper stem, an / or the hairpin regions may be modified, deleted, or replaced. In some embodiments, RNA nucleotides in the lower stem, upper stem, and / or the hairpin regions may be replaced with one ormore DNA sequences. In some embodiments, the scaffold sequence comprises unmodified or wild type RNA sequences in the nexus and / or the bulge regions. In some embodiments, the scaffold sequence does not include long stretches of A-T pairs, for example, a GUUUU-AAAAC (SEQ ID NO: 889) pairing element.
[0219] A tagRNA may also comprise optional modifiers, e.g., 3’ end modifier region and / or an 5' end modifier region. In some embodiments, a tagRNA comprises at least one nucleotide that is not part of a spacer, a scaffold sequence, or an extension arm. The optional sequence modifiers can be positioned within or between any of the other regions of the tagRNA, and not limited to being located at the 3' and 5' ends. In some embodiments, the tagRNA comprises secondary RNA structure, such as, but not limited to, aptamers, hairpins, stem / loops, toeloops, and / or RNA-binding protein recruitment domains (e.g., the MS2 aptamer which recruits and binds to the MS2cp protein). In some embodiments, a tagRNA comprises a short stretch of uracil at the 5’ end or the 3’ end. For example, in some embodiments, a tagRNA comprising a 3’ extension arm comprises a “UUU” sequence at the 3’ end of the extension arm. In some embodiments, a tagRNA comprises a toeloop sequence at the 3’ end. In some embodiments, the tagRNA comprises a 3’ extension arm and a toeloop sequence at the 3’ end of the extension arm. In some embodiments, the tagRNA comprises a 5’ extension arm and a toeloop sequence at the 5’ end of the extension arm. In some embodiments, the tagRNA comprises a toeloop element having the sequence S’-GAAANNNNN-3’, wherein N is any nucleobase. In some embodiments, the secondary RNA structure is positioned within the spacer. In some embodiments, the secondary structure is positioned within the extension arm. In some embodiments, the secondary structure is positioned within the gRNA core. In some embodiments, the secondary structure is positioned between the spacer and the gRNA core, between the gRNA core and the extension arm, or between the spacer and the extension arm. In some embodiments, the secondary structure is positioned between the FB sequence and the editing template. In some embodiments the secondary structure is positioned at the 3’ end or at the 5’ end of the tagRNA. In some embodiments, the tagRNA comprises a transcriptional lamination signal at the 3' end of the tagRNA. In addition to secondary RNA structures, the tagRNA may comprise a chemical linker or a poly(N) linker or tail, where “N” can be any nucleobase. In some embodiments, the chemical linker may function to prevent reverse transcription of the gRNA core.
[0220] The tagRNAs may be modified in one or more ways to improve their overall stability and / or performance in RT editing. The tagRNA can comprise 3’ mN*mN*mN*N and 5’ mN*mN*mN* modifications, where m indicates that the nucleotide contains a 2’-O-Me modification and a * indicates the presence of a phosphorothioate bond. The tagRNA can comprise a structural motif at the 3’ terminus selected from the group consisting of: an inverted-dT, aprequeosine1-1 riboswitch aptamer (evopreQ1) and variants thereof, a frameshifting pseudoknot from Moloney murine leukemia virus (MMLV) (mpknot), G-quadruplexes, hairpin structures, xrRNA, and a P4-P6 domain of the group I intron. In some embodiments, the structural motif is evopreQ1 or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 84- 90. In some embodiments, the structural motif is evopreQ1 or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 84-90 or a sequence having one, two or three mismatches relative a sequence selected from SEQ ID NOs: 84-90. In some embodiments, the structural motif is evopreQ1 or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 84-90 or a sequence at least 85% identical (85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a nucleotide sequence selected from SEQ ID NOs: 84-90. In some embodiments, the evopreQ1 is trimmed (TevopreQ1).
[0221] In some embodiments, appending one or more RNA structural motifs to a tagRNA can protect against degradation of the tagRNA. Such RNA structural motifs can include, but are not limited to (i) a prequeosine1-1 riboswitch aptamer (evopreQ1) and variants thereof, (ii) a frameshifting pseudoknot from Moloney murine leukemia virus (MMLV), hereafter referred to as “mpknot,” and variants thereof (iii) G-quadruplexes, (iv) hairpin structures (e.g., 15-bp hairpins), (v) xrRNA, and (vi) a P4-P6 domain of the group I intron.
[0222] The present disclosure provides modified tagRNAs with improved properties, including but not limited to, increased stability and cellular lifespan, and improved binding affinity for a RT editor. These modified tagRNAs result in improved genome editing as demonstrated by increase editing efficiency at a wide variety of genomic sites. In some embodiments, by appending certain nucleic acid structural motifs to terminus of the extension arm of a tagRNA, including but limited to, a prequeosin1-1 riboswitch aptamer (“evopreQ1-1”) or variant thereof, a pseudoknot from the MMLV viral genome (“evopreQ1-1”) or variant thereof, a modified tRNA used by MMLV RT as a primer for reverse transcription or variant thereof, and a G quadruplex or variant thereof, a consistent increase in editing activity can be achieved. In some embodiments, the 3’ terminus of the tagRNA comprises a evopreQ1 aptamer or variant thereof, comprising or consisting of a sequence of any one of: TTGACGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 84), CGCGAGTCTAGGGGATAACGCGTTAAACTTCCTAGAAGGCGGTT (SEQ ID NO: 85), CGCGGATCTAGATTGTAACGCGTTAAACCATCTAGAAGGCGGTT (SEQ ID NO: 86), CGCGTCGCTACCGCCCGGCGCGTTAAACACACTAGAAGGCGGTT (SEQ ID NO: 87), CGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAA (SEQ ID NO: 88), TTGACGCGCTTCTATCTAGTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 89), andTTGACGCGGTTCTATCTACTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 90). For RNA sequences, T can be U.
[0223] In one embodiment, the modified tagRNAs include a nucleic acid moiety at the 3′ end of the tagRNA. Optionally, the 3′ end of the tagRNA is fused to the nucleic acid moiety through a nucleotide linker. In various embodiments, it will be appreciated that a wide variety of nucleotide sequences will work reasonably well for each genomic target site. Linker length can also be variable. In some cases, linkers ranging in length from 3-18 nucleotides can be used. In other cases, the linker may be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, or at least 30 nucleotides.
[0224] In general, the nucleic acid moieties that may be used to modify a tagRNA, for example, by attaching it to the 3′ end of a tagRNA, may include any nucleic acid moiety, including, for instance, a nucleic acid molecule comprising or which forms a double-helix moiety, toeloop moiety, hairpin moiety, stem-loop moiety, pseudoknot moiety, aptamer moiety, G quadraplex moiety, tRNA moiety, or a ribozyme moiety. The nucleic acid moiety may be characterized as forming a secondary nucleic acid structure, a tertiary nucleic acid structure, or a quadruple nucleic acid structure. In other words, the nucleic acid moiety may form any two dimensional or three dimensional structure known to be formed by such structures. The nucleic acid moiety may be DNA or RNA.
[0225] In some embodiments, a tagRNA or a nick guide RNA (egRNA) can be chemically synthesized, or can be assembled or cloned and transcribed from a DNA sequence, e.g., a plasmid DNA sequence, or by any RNA oligonucleotide synthesis method known in the art. In some embodiments, DNA sequence that encodes a tagRNA (or egRNA) can be designed to append one or more nucleotides at the 5' end or the 3' end of the tagRNA (or nick guide RNA) encoding sequence to enhance tagRNA transcription. For example, in some embodiments, a DNA sequence that encodes a tagRNA (or an egRNA) can be designed to append a nucleotide G at the 5' end. Accordingly, in some embodiments, the tagRNA (or nick guide RNA) can comprise an appended nucleotide G at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or nick guide RNA) can be designed to append a sequence that enhances transcription, e.g., a Kozak sequence, at the 5' end. In some embodiments, a DNA sequence that encodes atagRNA (or nick guide RNA) can be designed to append the sequence CACC or CCACC at the 5' end. Accordingly, in some embodiments, the tagRNA (or nick guide RNA) can comprise an appended sequence CACC or CCACC at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or nick guide RNA) can be designed to append the sequence TTT, TTTT, TTTTT, TTTTTT, TTTTTTT at the 3' end. Accordingly, in some embodiments, the tagRNA (or nick guide RNA) can comprise an appended sequence UUU, UUUU, UUUUU, UUUUUU, or UUUUUUU at the 3' end.
[0226] Disclosed herein include RT editing systems. In some embodiments, the RT editing system comprises: any of the tagRNAs disclosed herein, or a nucleic acid encoding the tagRNA; and a RT editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor. In some embodiments, the intended nucleotide edit incorporation rate of the RT editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).
[0227] Disclosed herein include RT editing complexes. In some embodiments, the RT editing complex comprises: (i) any of the tagRNA disclosed herein and a RT editor comprising a DNA endonuclease domain and a DNA polymerase domain; or (ii) any of the RT editing systems of the disclosure. In some embodiments, the intended nucleotide edit incorporation rate of the RT editing complex is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values). Enhancer gRNA (egRNA)
[0228] Disclosed herein include RT editing systems. In some embodiments, the RT editing system comprises: any of the tagRNAs disclosed herein, or a nucleic acid encoding the tagRNA; and a RT editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor. The RT editing system can comprise: an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein theegRNA comprises a egRNA spacer that is complementary to a second search target sequence in the target gene.
[0229] In some embodiments, a RT editing system or composition further comprises an enhancer guide polynucleotide, such as an enhancer guide RNA (egRNA). Without wishing to be bound by any particular theory, the non-edit strand of a double stranded target DNA in the target gene may be nicked by a CRISPR-Cas nickase directed by an egRNA. In some embodiments, the nick on the non-edit strand directs endogenous DNA repair machinery to use the edit strand as a template for repair of the non-edit strand, which may increase efficiency of RT editing. In some embodiments, the non-edit strand is nicked by a RT editor localized to the non- edit strand by the egRNA. Accordingly, also provided herein are tagRNA systems comprising at least one tagRNA and at least one egRNA.
[0230] In some embodiments, the egRNA is a guide RNA which contains a variable spacer sequence and a guide RNA scaffold or core region that interacts with the DNA binding domain, a DNA endonuclease domain, e.g., Cas9 of the RT editor, in some embodiments, the egRNA comprises a spacer sequence (referred to herein as an ng spacer, or a second spacer) that is substantially complementary to a second search target sequence (or ng search target sequence), which is located on the edit strand, or the non-target strand. Thus, in some embodiments, the egRNA search target sequence recognized by the egRNA spacer and the search target sequence recognized by the spacer sequence of the tagRNA are on opposite strands of the double stranded target DNA of target gene, e.g., the ATP7B gene, PAH gene, or SERPINA1 gene. In some embodiments, an egRNA spacer sequence is complementary to, and may hybridize with the second search target sequence only after an intended nucleotide edit has been incorporated on the edit strand, by the editing template of a tagRNA.
[0231] In some embodiments, the egRNA search target sequence is located on the non- target strand, within 10 base pairs to 100 base pairs of an intended nucleotide edit incorporated by the tagRNA on the edit strand, in some embodiments, the egRNA target search target sequence is within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp of an intended nucleotide edit incorporated by the tagRNA on the edit strand. In some embodiments, the 5’ ends of the egRNA search target sequence and the tagRNA search target sequence are within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp apart from each other. In some embodiments, the 5’ ends of the egRNA search target sequence and the tagRNA search target sequence are within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp apart from each other.
[0232] In some embodiments, an egRNA spacer sequence is complementary to, and may hybridize with the second search target sequence only after an intended nucleotide edit has been incorporated on the edit strand, by the editing template of a tagRNA. In some embodiments, the egRNA comprises a spacer sequence that matches only the edit strand after incorporation of the nucleotide edits, but not the endogenous target gene sequence on the edit strand. Accordingly, in some embodiments, an intended nucleotide edit is incorporated within the egRNA search target sequence. In some embodiments, the intended nucleotide edit is incorporated within about 1-10 nucleotides of the position corresponding to the PAM of the egRNA search target sequence.
[0233] The RT editing system can comprise: a nick guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises an egRNA spacer that is complementary to a second search target sequence in the target gene. The second search target sequence can be on the second strand of the target gene. The egRNA spacer can be from 16 to 22 nucleotides in length. In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is from 15 nucleotides to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, or 20 to 30 nucleotides in length. In some embodiments, the spacer is 16 to 22 nucleotides in length. In some embodiments, the spacer is 16 to 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length. In some embodiments, the egRNA spacer is 20 nucleotides in length. In some embodiments, the egRNA comprises a scaffold sequence. The scaffold sequence can comprise a nucleotide sequence selected from SEQ ID NOs: 1210-1266.
[0234] In some embodiments, the egRNA comprises a nucleotide sequence selected from SEQ ID NOs: 48-49, 498-555, and 1048-1049. In some embodiments, the egRNA comprises a nucleotide sequence selected from SEQ ID NOs: 48-49, 498-555, and 1048-1049 or a sequence have one, two, or three mismatches relative to a nucleotide sequence selected from SEQ ID NOs: 48-49, 498-555, and 1048-1049. In some embodiments, the egRNA comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 48-49, 498-555, and 1048-1049 or a sequence at least 85% identical to a nucleotide sequence selected from SEQ ID NOs: 48-49, 498-555, and 1048-1049 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to a nucleotide sequence selected from SEQ ID NOs: 48-49, 498-555, and 1048-1049).
[0235] Different tagRNAs can be used with different egRNAs. In some embodiments, the tagRNA comprises or consists of any one of SEQ ID NOs: 16-47 and the egRNA comprisesor consists of SEQ ID NO: 48. In some embodiments, the tagRNA comprises or consists of any one of SEQ ID NOs: 16-47 and the egRNA comprises or consists of SEQ ID NO: 49.
[0236] In some embodiments, the egRNA spacer comprises a nucleotide sequence selected from SEQ ID NOs: 51-52. In some embodiments, the egRNA spacer comprises a nucleotide sequence selected from SEQ ID NOs: 51-52 or a sequence have one, two, or three mismatches relative to a nucleotide sequence selected from SEQ ID NOs: 51-52. In some embodiments, the egRNA spacer comprises or consists of a nucleotide sequence selected from SEQ ID NOs: 51-52 or a sequence at least 85% identical to a nucleotide sequence selected from SEQ ID NOs: 51-52 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to a nucleotide sequence selected from SEQ ID NOs: 51-52).
[0237] In some embodiments, the intended nucleotide edit incorporation rate of the RT editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).
[0238] Provided in Table 2A-Table 2B below are exemplary protein (Table 2A) and mRNA (Table 2B) sequences of RT Editors and components thereof. Table 2C provides exemplary RT Editor linker sequences. TABLE 2A: RT EDITING PROTEIN AND COMPONENTS NAME and SEQ ID NO RT editor fusion (SEQ ID NO: 4) bipartite SV40 NLS (SEQ ID NO: 5) SPYCas9 R221K N394K H840A (SEQ ID NO: 6) Linker (SEQ ID NO: 7) codon optimized MMLV RT pentamutant (SEQ ID NO: 8) Linker – SGGS* (SEQ ID NO: 9 or 10) *may be repeated one or more times bipartite SV40 NLS (SEQ ID NO: 11) c-Myc NLS (SEQ ID NO: 12) RT editor fusion cDNA (SEQ ID NO: 13) Reference / wild-type SPYCas9 (SEQ ID NO: 14) wild type MMLV RT (SEQ ID NO: 15) M. georgiana RT (SEQ ID NO: 141) TABLE 2B: EXEMPLARY mRNA SEQUENCES ENCODING RT EDITORS NAME and SEQ ID NO pPG275 mRNA (SEQ ID NO: 137) pPG276 mRNA (SEQ ID NO: 138)pPG277 mRNA (SEQ ID NO: 139) pPG278 mRNA (SEQ ID NO: 140) *For RNA, U are T in the sequence listing. TABLE 2C: EXEMPLARY RT EDITOR LINKER SEQUENCESNameShorthandSEQ Sequence AA ResiduesID NO GS3 (GGGGS)3 GGGGSGGGGSGGGGS 1026 GS4 (GGGGS)4 GGGGSGGGGSGGGGSGGGGS 1027 GS5 (GGGGS)5 GGGGSGGGGSGGGGSGGGGSGGGGS 1056 GS6 (GGGGS)6 GGGGSGGGGSGGGGSGGGGSGGGGSGGGGS 1057 GS7 (GGGGS)7 GGGGSGGGGSGGGGSGGGGSGGGGSGGGGSGGGGS 1028 EAK3 (EAAAK)3 EAAAKEAAAKEAAAK 1022 EAK4 (EAAAK)4 EAAAKEAAAKEAAAKEAAAK 1023 EAK5 (EAAAK)5 EAAAKEAAAKEAAAKEAAAKEAAAK 1058 EAK6 (EAAAK)6 EAAAKEAAAKEAAAKEAAAKEAAAKEAAAK 1024 EAK7 (EAAAK)7 EAAAKEAAAKEAAAKEAAAKEAAAKEAAAKEAAAK 1059 AEAK4- A(EAAAK)4ALEA AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEA ALEA (EAAAK)4A AAKEAAAKA 1020 GS- AEAK4- GSA(EAAAK)4AL GSAEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKE ALEA EA(EAAAK)4ASG AAAKEAAAKASG 1060 GGGGS(EAAAK)4 GGS- ALEA AEAK4- (EAAAAK)4SGGG GGGGSEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAA ALEA G AKEAAAKEAAAKSGGGG 1021 AEAAAKEAAAK AEAK A AEAAAKEAAAKA 1061 (AEAAAKEAAAK AEAK2 A)2 AEAAAKEAAAKAAEAAAKEAAAKA 1019 (AEAAAKEAAAK AEAK3 A)3 AEAAAKEAAAKAAEAAAKEAAAKAAEAAAKEAAAKA 1062 GGS- EAK3- GGGGS(EAAAK)3 SGG GGGGS GGGGSEAAAKEAAAKEAAAKGGGGS 1025 GGS- 1063 EAK4- GGGGS(EAAAK)4 SGG GGGGS GGGGSEAAAKEAAAKEAAAKEAAAKGGGGS GGS- 1064 EAK5- GGGGS(EAAAK)5 SGG GGGGS GGGGSEAAAKEAAAKEAAAKEAAAKEAAAKGGGGS GGS- GGGGS(EAAAK)3 1065 EAK3- GGSGGSG(EAAA GGGGSEAAAKEAAAKEAAAKGGSGGSGEAAAKEAAAKE GGS K)3GGGGS AAAKGGGGS SGAGSAAGSGEF GSA-F GS GSAGSAAGSGEF 1066 GSA-F- SGAGSAAGSGEF flip EGSGSAASGAGS GSAGSAAGSGEFEGSGSAASGASG 1067 KESGSVSSEQLA K19D QFRSLD KESGSVSSEQLAQFRSLD 1033 EGKSSGSGSESKS E14T T EGKSSGSGSESKST 1032 PAPAP PAPAP PAPAP 1068 PAPAP-2 (PAPAP)2 PAPAPPAPAP 1029 PAPAP-3 (PAPAP)3 PAPAPPAPAPPAPAP 1069 PAPAP-4 (PAPAP)4 PAPAPPAPAPPAPAPPAPAP 1030 PAPAP-5 (PAPAP)5 PAPAPPAPAPPAPAPPAPAPPAPAP 1070 PAPAP-6 (PAPAP)6 PAPAPPAPAPPAPAPPAPAPPAPAPPAPAP 1071 PAPAP-7 (PAPAP)7 PAPAPPAPAPPAPAPPAPAPPAPAPPAPAPPAPAP 1072PAPAP-8 (PAPAP)8 PAPAPPAPAPPAPAPPAPAPPAPAPPAPAPPAPAPPAPAP 1031 GS- XTEN16- GS SGGGGSSGSETPGTSESATPESSGGGGS 1034 SGGSx2- XTEN16- SGGSx2 SGGSSGGSSGSETPGTSESATPESSGGSSGGS 1035 GGS- SGGGGSAEAAAKEAAAKEAAAKEAAAKALEAEAAAKEA AEAK- AAKEAAAKEAAAKASGGGGS ALEA_m od1-gs 1073 GS-NLS- SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS GS_4-gs 1074
[0239] Provided in Table 3 below are exemplary spacer and edit target sequences for editing ATP7B gene. TABLE 3: EXEMPLARY SPACER AND PROTOSPACER SEQUENCES SEQ ID SPACER SEQUENCE1SEQ ID PROTOSPACER SEQUENCE NO: NO: 50 UUGGUGACUGCCACGCCCAA 53 TTGGTGACTGCCACGCCCAA 51 GGCCAGCAGUGAACACCCCU 54 GGCCAGCAGTGAACACCCCT 52 GCCAGCAGUGAACACCCCUU 55 GCCAGCAGTGAACACCCCTT 133 UUUGGUGACUGCCACGCCCA 135 TTTGGTGACTGCCACGCCCA 134 CAGUGAACACCCCUUGGGCG 136 CAGTGAACACCCCTTGGGCG 1For RNA, U are T in the sequence listing.
[0240] The tables below provide sequences related to the gRNAs of the disclosure (e.g., tagRNA or egRNA). For guide RNA sequences, the spacer sequences are underlined, and the extension arms are bold. Descriptions provide information regarding, e.g., the length of the FBS and the homology or editing template (ET) region of the extension arm. As would be understood by a person of ordinary skill in the art, the extension arm can comprise, from 5’ to 3’, an editing template (comprising a homology region) and flap binding sequence (FB sequence; FBS). As also understood by the skilled person, the nucleotide thymine (T) is used in DNA sequences and uracil (U) in RNA sequences. In some embodiments, a “T” can represent “U” when the molecule is an RNA molecule. Provided in Table 4 below are exemplary tagRNA (e.g., long gRNA) and nicking guide RNA (egRNA) sequences of the disclosure. TABLE 4: EXEMPLARY tagRNA AND egRNA SEQUENCES NAME DESCRIPTION SEQ ID SEQUENCE1NO: BK g050 Mut1.Hom14.FBS 16 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 13 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGCAGUCAUUUU BK g051 Mut1.Hom14.FBS 17 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 12 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGCAGUCUUUU BK g052 Mut1.Hom14.FBS 18 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 11 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGCAGUUUUU BK g053 Mut1.Hom14.FBS 19 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 10 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGCAGUUUU BK g054 Mut1.Hom14.FBS 20 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 9 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGCAUUUU BK g055 Mut1.Hom14.FBS 21 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 8 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGCUUUU BK g056 Mut1.Hom14.FBS 22 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 7 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGGUUUU BK g057 Mut1.Hom14.FBS 23 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 6 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCCAGCAGUGAACAcCCCUUGGG CGUGUUUU BK g058 Mut1.Hom10.FBS 24 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 13 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GCAGUCAUUUU BK g059 Mut1.Hom10.FBS 25 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 12 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GCAGUCUUUU BK g060 Mut1.Hom10.FBS 26 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 11 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GCAGUUUUU BK g061 Mut1.Hom10.FBS 27 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 10 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GCAGUUUU BK g062 Mut1.Hom10.FBS 28 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 9 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GCAUUUU BK g063 Mut1.Hom10.FBS 29 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 8 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GCUUUU BK g064 Mut1.Hom10.FBS 30 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 7 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG GUUUU BK g065 Mut1.Hom10.FBS 31 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 6 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGCAGUGAACAcCCCUUGGGCGUG UUUU BK g066 Mut1.Hom7.FBS1 32 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 3 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGCA GUCAUUUU BK g067 Mut1.Hom7.FBS1 33 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 2 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGCA GUCUUUUBK g068 Mut1.Hom7.FBS1 34 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 1 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGCA GUUUUU BK g069 Mut1.Hom7.FBS1 35 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 0 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGCA GUUUU BK g070 Mut1.Hom7.FBS9 36 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGCA UUUU BK g071 Mut1.Hom7.FBS8 37 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGCU UUU BK g072 Mut1.Hom7.FBS7 38 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGGUU UU BK g073 Mut1.Hom7.FBS6 39 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGUGAACAcCCCUUGGGCGUGUUU U BK g074 Mut1.Hom5.FBS1 40 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 3 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGCAGU CAUUUU BK g075 Mut1.Hom5.FBS1 41 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 2 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGCAGU CUUUU BK g076 Mut1.Hom5.FBS1 42 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 1 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGCAGU UUUU BK g077 Mut1.Hom5.FBS1 43 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC 0 AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGCAGU UUU BK g078 Mut1.Hom5.FBS9 44 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGCAUU UU BK g079 Mut1.Hom5.FBS8 45 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGCUUU U BK g080 Mut1.Hom5.FBS7 46 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGGUUUU BK g081 Mut1.Hom5.FBS6 47 UUGGUGACUGCCACGCCCAAgUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCGAACAcCCCUUGGGCGUGUUUU BK g040248 GGCCAGCAGUGAACACCCCUGUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCUUUU BK g041249 GCCAGCAGUGAACACCCCUUGUUUUAGAGCUAGAAAUAGC AAGUUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGU GGCACCGAGUCGGUGCUUUU1For RNA, U are T in the sequence listing. For each entry, the spacer sequence is underlined and the extension arm is bold. The description indicates the length of the FBS and ET or homology region (Hom) of the extensionarm. 2egRNA TABLE 5: EXEMPLARY FLAP BINDING SEQUENCE (FBS) AND EDITING TEMPLATE SEQUENCES NAME DESCRIPTION FBS1SEQ ID NO: EDITING TEMPLATE1SEQ ID NO: BK g050 Mut1.Hom14.FBS13 GGCGTGGCAGTCA 91 GCCAGCAGTGAACA 95 cCCCTTG BK g051 Mut1.Hom14.FBS12 GGCGTGGCAGTC 92 GCCAGCAGTGAACA 95 cCCCTTG BK g052 Mut1.Hom14.FBS11 GGCGTGGCAGT 93 GCCAGCAGTGAACA 95 cCCCTTG BK g053 Mut1.Hom14.FBS10 GGCGTGGCAG 94 GCCAGCAGTGAACA 95 cCCCTTG BK g054 Mut1.Hom14.FBS9 GGCGTGGCA GCCAGCAGTGAACA 95 cCCCTTG BK g055 Mut1.Hom14.FBS8 GGCGTGGC GCCAGCAGTGAACA 95 cCCCTTG BK g056 Mut1.Hom14.FBS7 GGCGTGG GCCAGCAGTGAACA 95 cCCCTTG BK g057 Mut1.Hom14.FBS6 GGCGTG GCCAGCAGTGAACA 95 cCCCTTG BK g058 Mut1.Hom10.FBS13 GGCGTGGCAGTCA 91 GCAGTGAACAcCCC 96 TTG BK g059 Mut1.Hom10.FBS12 GGCGTGGCAGTC 92 GCAGTGAACAcCCC 96 TTG BK g060 Mut1.Hom10.FBS11 GGCGTGGCAGT 93 GCAGTGAACAcCCC 96 TTG BK g061 Mut1.Hom10.FBS10 GGCGTGGCAG 94 GCAGTGAACAcCCC 96 TTG BK g062 Mut1.Hom10.FBS9 GGCGTGGCA GCAGTGAACAcCCC 96 TTG BK g063 Mut1.Hom10.FBS8 GGCGTGGC GCAGTGAACAcCCC 96 TTG BK g064 Mut1.Hom10.FBS7 GGCGTGG GCAGTGAACAcCCC 96 TTG BK g065 Mut1.Hom10.FBS6 GGCGTG GCAGTGAACAcCCC 96 TTG BK g066 Mut1.Hom7.FBS13 GGCGTGGCAGTCA 91 GTGAACAcCCCTTG 97 BK g067 Mut1.Hom7.FBS12 GGCGTGGCAGTC 92 GTGAACAcCCCTTG 97 BK g068 Mut1.Hom7.FBS11 GGCGTGGCAGT 93 GTGAACAcCCCTTG 97 BK g069 Mut1.Hom7.FBS10 GGCGTGGCAG 94 GTGAACAcCCCTTG 97 BK g070 Mut1.Hom7.FBS9 GGCGTGGCA GTGAACAcCCCTTG 97 BK g071 Mut1.Hom7.FBS8 GGCGTGGC GTGAACAcCCCTTG 97 BK g072 Mut1.Hom7.FBS7 GGCGTGG GTGAACAcCCCTTG 97 BK g073 Mut1.Hom7.FBS6 GGCGTG GTGAACAcCCCTTG 97 BK g074 Mut1.Hom5.FBS13 GGCGTGGCAGTCA 91 GAACAcCCCTTG 98 BK g075 Mut1.Hom5.FBS12 GGCGTGGCAGTC 92 GAACAcCCCTTG 98 BK g076 Mut1.Hom5.FBS11 GGCGTGGCAGT 93 GAACAcCCCTTG 98 BK g077 Mut1.Hom5.FBS10 GGCGTGGCAG 94 GAACAcCCCTTG 98 BK g078 Mut1.Hom5.FBS9 GGCGTGGCA GAACAcCCCTTG 98 BK g079 Mut1.Hom5.FBS8 GGCGTGGC GAACAcCCCTTG 98 BK g080 Mut1.Hom5.FBS7 GGCGTGG GAACAcCCCTTG 98 BK g081 Mut1.Hom5.FBS6 GGCGTG GAACAcCCCTTG 981For RNA, T is U; U are T in the sequence listing. The description indicates the length of the FBS and ET or homology region (Hom) of the extension arm.TABLE 6: EXEMPLARY gRNAs NAME SEQUENCE1, 2SEQ ID NO: DESCRIPTION BK g0503mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 99 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS13.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGrCrArGrU rCrAmU*mU*mU*mU BK g0513mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 100 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS12.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGrCrArGrU rCmU*mU*mU*mU BK g0523mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 101 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS11.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGrCrArGrU mU*mU*mU*mU BK g0533mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 102 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS10.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGrCrArGm U*mU*mU*mU BK g0543mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 103 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS9.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGrCrAmU* mU*mU*mU BK g0553mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 104 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS8.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGrCmU*mU *mU*mU BK g0563mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 105 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS7.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGrGmU*mU* mU*mU BK g0573mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 106 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m14.FBS6.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrCrArGrCrArGrUrGrArArCrArC rCrCrCrUrUrGrGrGrCrGrUrGmU*mU*mU *mU BK g0583mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 107 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS13.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGrCrArGrUrCrAmU* mU*mU*mU BK g0593mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 108 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS12.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGrCrArGrUrCmU*mU *mU*mU BK g0603mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 109 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS11.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGrCrArGrUmU*mU* mU*mU BK g0613mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 110 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS10.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGrCrArGmU*mU*mU *mU BK g0623mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 111 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS9.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGrCrAmU*mU*mU*m U BK g0633mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 112 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS8.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGrCmU*mU*mU*mU BK g0643mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 113 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS7.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGrGmU*mU*mU*mU BK g0653mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 114 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m10.FBS6.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrCrArGrUrGrArArCrArCrCrCrCrU rUrGrGrGrCrGrUrGmU*mU*mU*mU BK g0663mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 115 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS13.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGrCrArGrUrCrAmU*mU*mU *mU BK g0673mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 116 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS12.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGrCrArGrUrCmU*mU*mU*m U BK g0683mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 117 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS11.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGrCrArGrUmU*mU*mU*mU BK g0693mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 118 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS10.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGrCrArGmU*mU*mU*mU BK g0703mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 119 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS9.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGrCrAmU*mU*mU*mUBK g0713mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 120 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS8.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGrCmU*mU*mU*mU BK g0723mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 121 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS7.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGrGmU*mU*mU*mU BK g0733mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 122 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m7.FBS6.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrUrGrArArCrArCrCrCrCrUrUrGrG rGrCrGrUrGmU*mU*mU*mU BK g0743mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 123 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS13.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGrCrArGrUrCrAmU*mU*mU*mU BK g0753mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 124 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS12.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGrCrArGrUrCmU*mU*mU*mU BK g0763mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 125 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS11.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGrCrArGrUmU*mU*mU*mU BK g0773mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 126 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS10.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGrCrArGmU*mU*mU*mU BK g0783mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 127 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS9.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGrCrAmU*mU*mU*mU BK g0793mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 128 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS8.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGrCmU*mU*mU*mU BK g0803mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 129 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS7.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGrGmU*mU*mU*mU BK g0813mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCr 130 Human.ATP7B.NGG1.Mut1.Ho CrCrArArGrUrUrUrUrArGrAmGmCmUmAmG m5.FBS6.mU*mU*mU*mU mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCrGrArArCrArCrCrCrCrUrUrGrGrGrC rGrUrGmU*mU*mU*mU BK g0404mG*mG*mC*rCrArGrCrArGrUrGrArArCrArCr 131 Human.ATP7B.Mut1.3b.1.mU* CrCrCrUrGrUrUrUrUrArGrAmGmCmUmAmG mU*mU*mU.NGG2 mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCmU*mU*mU*mU BK g0414mG*mC*mC*rArGrCrArGrUrGrArArCrArCrCr 132 Human.ATP7B.Mut1.3b.2.mU* CrCrUrUrGrUrUrUrUrArGrAmGmCmUmAmG mU*mU*mU.NGG2 mAmAmAmUmAmGmCrArArGrUrUrArArArA rUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCm AmAmCmUmUmGmAmAmAmAmAmGmUmG mGmCmAmCmCmGmAmGmUmCmGmGmUm GmCmU*mU*mU*mU1For RNA, U are T in the sequence listing. For each entry, the spacer sequence is underlined and the extension arm is bold. The description indicates the length of the FBS and ET or homology region (Hom) of the extension arm.2For the sequences: *, phosphorothioate modification; mN, 2’O-methyl modification; rN, unmodified ribonucleotide. N=A,C,T / U,G.3tagRNA4egRNA TABLE 7: EXEMPLARY tagRNAs NAME DESCRIPTION SEQUENCE1, 2SEQ ID NO: PG g060 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 164 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET14OMe UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCmGmUmGmAmAmCmAmC mCmCmCmUmUmGrGrGrCrGrUrGrGrCmU*mU *mU*mU PG g061 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 165 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.FBS8OMe UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUrGrArArCrArCrCrCrC rUrUrGmGmGmCmGmUmGmGmCmU*mU*mU* mU PG g062 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 166 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET14OMe.FBS8 UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr OMe ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCmGmUmGmAmAmCmAmC mCmCmCmUmUmGmGmGmCmGmUmGmGmC mU*mU*mU*mU PG g063 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 167 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET7OMe UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGmUrGmArAmCrAmCrC mCrCmUrUmGrGrGrCrGrUrGrGrCmU*mU*mU *mU PG g064 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 168 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.FBS4OMe UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUrGrArArCrArCrCrCrC rUrUrGrGmGrCmGrUmGrGmCmU*mU*mU*mU PG g065 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 169 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET7OMe.FBS4O UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr Me ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGmUrGmArAmCrAmCrC mCrCmUrUmGrGmGrCmGrUmGrGmCmU*mU* mU*mU PG g066 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 170 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET4OMe UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUmGrArAmCrArCmCrC rCmUrUrGrGrGrCrGrUrGrGrCmU*mU*mU*mU PG g067 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 171 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.FBS3OMe UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUrGrArArCrArCrCrCrC rUrUrGmGrGrCmGrUrGmGrCmU*mU*mU*mU PG g068 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 172 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET4OMe.FBS3O UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr Me ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUmGrArAmCrArCmCrC rCmUrUrGmGrGrCmGrUrGmGrCmU*mU*mU* mU PG g069 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 173 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET14F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmC / i2FG / / i2FU / / i2FG / / i2FA / / i2F A / / i2FC / / i2FA / / i2FC / / i2FC / / i2FC / / i2FC / / i2FU / / i2FU / / i2FG / rGrGrCrGrUrGrGrCmU*mU*mU*mU PG g070 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 174 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.FBS8F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUrGrArArCrArCrCrCrC rUrUrG / i2FG / / i2FG / / i2FC / / i2FG / / i2FU / / i2FG / / i2FG / / i2FC / mU*mU*mU*mU PG g071 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 175 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET14F.FBS8F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmC / i2FG / / i2FU / / i2FG / / i2FA / / i2F A / / i2FC / / i2FA / / i2FC / / i2FC / / i2FC / / i2FC / / i2FU / / i2FU / / i2FG / / i2FG / / i2FG / / i2FC / / i2FG / / i2FU / / i2FG / / i2FG / / i 2FC / mU*mU*mU*mU PG g072 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 176 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET7F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrG / i2FU / rG / i2FA / rA / i2FC / rA / i2FC / rC / i2FC / rC / i2FU / rU / i2FG / rGrGrCrGrUrGr GrCmU*mU*mU*mU PG g073 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 177 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.FBS4F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUrGrArArCrArCrCrCrC rUrUrGrG / i2FG / rC / i2FG / rU / i2FG / rG / i2FC / mU*m U*mU*mU PG g074 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 178 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET7F.FBS4F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrG / i2FU / rG / i2FA / rA / i2FC / rA / i2FC / rC / i2FC / rC / i2FU / rU / i2FG / rG / i2FG / rC / i2FG / r U / i2FG / rG / i2FC / mU*mU*mU*mU PG g075 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 179 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET4F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrU / i2FG / rArA / i2FC / rArC / i2FC / rCrC / i2FU / rUrGrGrGrCrGrUrGrGrCmU*m U*mU*mU PG g076 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 180 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.FBS3F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrUrGrArArCrArCrCrCrC rUrUrG / i2FG / rGrC / i2FG / rUrG / i2FG / rCmU*mU*m U*mUPG g077 Human.ATP7B.NGG1.Mu mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrA 181 t1.Hom7.FBS8.mU*mU* rArGrUrUrUrUrArGrAmGmCmUmAmGmAmAmAm mU*mU.ET4F.FBS3F UmAmGmCrArArGrUrUrArArArArUrArArGrGrCrUr ArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAm AmAmAmAmGmUmGmGmCmAmCmCmGmAmGm UmCmGmGmUmGmCrGrU / i2FG / rArA / i2FC / rArC / i2FC / rCrC / i2FU / rUrG / i2FG / rGrC / i2FG / rUrG / i2FG / rCmU*mU*mU*mU 1For RNA, T is U; U are T in the sequence listing. For each entry, the spacer sequence is underlined and the extension arm is bold. The description indicates the length of the FBS and ET or homology region of the extension arm. 2 For the sequences: *, phosphorothioate modification; mN, 2’O-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A,C,T / U,G. TABLE 8A: EXEMPLARY CAS9 SEQUENCES DESCRIPTION SEQ ID NO: Ssch1Cas9_N583A 142 Sro1Cas9_N584A 143 Sha4Cas9_N583A 144 SsuCas9_N622A 145 evoCjCas9_H559A 146 iSpyMac_H840A 147 Ssi5Cas9_N580A 148 Ssi8Cas9_N580A 149 Ssci4Cas9_N580A 150 Shy1Cas9_N582A 151 Sag3Cas9_N582A 152 Slutr1Cas9_N583A 153 Ssch3Cas9_N583A 154 SpRYCas9_H840A 155 SpRYcCas9_H849A 156 Sma2Cas9_N622A 157 SsaCas9_N628A 158 EvoSluCas9_N582A 159 SluCas9_N582A 160 EvosRGN3.1_N585A 161 EvoNme2Cas9_H588A 162 EvoAnaCas9_H582A 163 TABLE 8B: EXEMPLARY CAS9 SEQUENCES NAME SPECIES (native or derived) Length SEQ ID NO: St1Cas9- Methanobacterium sp 1097 270 sim3_MetCas9_MBW4258562.1_Methanobacterium-sp- YSL_01 SsaCas9-sim1_ARC34060.1_Streptococcus-equinus Streptococcus equinus 1128 271 SsaCas9-sim1_SsaCas9_AR00_CCB95040.1_Streptococcus- Streptococcus salivarius 1128 272 salivarius-JIM8777_01 SsaCas9-sim1_SsaCas9_AR01_ancestral-seq- Streptococcus salivarius 1128 273 recon_Streptococcus-salivarius_01 SsaCas9-sim1_SsaCas9_AR02_ancestral-seq- Streptococcus salivarius 1128 274 recon_Streptococcus-salivarius_01 SsaCas9-sim1_WP_003097512.1_Streptococcus-vestibularis Streptococcus vestibularis 1128 275 SsaCas9-sim1_WP_014634195.1_Streptococcus-salivarius Streptococcus salivarius 1127 276 SsaCas9-sim1_WP_037611471.1_Streptococcus-sp-SR4 Streptococcus sp SR4 1128 277 SsaCas9-sim1_WP_048787825.1_Streptococcus Streptococcus 1122 278 SsaCas9-sim1_WP_048790568.1_Streptococcus-salivarius Streptococcus salivarius 1139 279 SsaCas9-sim1_WP_064519118.1_Streptococcus-vestibularis Streptococcus vestibularis 1128 280SsaCas9-sim1_WP_145516278.1_Streptococcus-salivarius Streptococcus salivarius 1121 281 SsaCas9-sim1_WP_227278337.1_Streptococcus-vestibularis Streptococcus vestibularis 1122 282 SsaCas9-sim2_clustered_ARC34060.1_Streptococcus- Streptococcus equinus 1128 283 equinus SsaCas9-sim2_clustered_HEM3702072.1_Streptococcus-suis Streptococcus suis 1122 284 SsaCas9-sim2_clustered_HEM4744366.1_Streptococcus-suis Streptococcus suis 1122 285 SsaCas9-sim2_clustered_MDU6911040.1_Streptococcus- Streptococcus salivarius 1127 286 salivarius SsaCas9-sim2_clustered_WP_002299171.1_Streptococcus- Streptococcus mutans 1125 287 mutans SsaCas9-sim2_clustered_WP_021002704.1_Streptococcus- Streptococcus intermedius 1125 288 intermedius SsaCas9-sim2_clustered_WP_084910940.1_Streptococcus- Streptococcus oralis 1129 289 oralis SsaCas9-sim2_clustered_WP_115265868.1_Streptococcus- Streptococcus macedonicus 1130 290 macedonicus SsaCas9-sim2_clustered_WP_126440388.1_Streptococcus- Streptococcus equinus 1130 291 equinus SsaCas9-sim2_clustered_WP_155974233.1_Streptococcus- Streptococcus ruminantium 1125 292 ruminantium SsaCas9-sim2_clustered_WP_198459777.1_Streptococcus- Streptococcus oralis 1122 293 oralis SsaCas9-sim2_clustered_WP_219967735.1_Streptococcus- Streptococcus gordonii 1136 294 gordonii SsaCas9-sim2_clustered_WP_223322622.1_Streptococcus- Streptococcus sanguinis 1121 295 sanguinis SsaCas9-sim2_clustered_WP_257737268.1_Streptococcus- Streptococcus uberis 1122 296 uberis SsaCas9-sim2_clustered_WP_268734931.1_Streptococcus- Streptococcus cristatus 1127 297 cristatus SsaCas9-sim2_clustered_WP_270615522.1_Streptococcus- Streptococcus koreensis 1123 298 koreensis SsaCas9-sim2_clustered_WP_281734300.1_Streptococcus- Streptococcus lutetiensis 1123 299 lutetiensis SsaCas9-sim2_clustered_WP_311154762.1_uncultured- Streptococcus sp 1123 300 Streptococcus-sp SsaCas9-sim3_Structure-Ssa-Streptococcus- Streptococcus cristatus 1130 301 cristatus_A0A3R9LSK4_STRCR_Streptococcus-cristatus_01 SsaCas9-sim3_structure-Ssa-Streptococcus- Streptococcus massiliensis 1126 302 massiliensis_A0A380KY87_9STRE_Streptococcus- massiliensis_01 SsaCas9-sim3_structure-Ssa-Streptococcus- Streptococcus vestibularis 1128 303 vestibularis_A0A3S4QG93_Streptococcus-vestibularis_01 SsaCas9-sim3_structure-St1_A0A7X2QGI6_Streptococcus- Streptococcus uberis 1122 304 uberis_01 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1123 305 asini_A0A413AJC9_9ENTE_Enterococcus-asini_01 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1123 306 asini_MCD5030292.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1123 307 asini_RGW10367.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1130 308 asini_WP_231453593.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1130 309 asini_WP_259305198.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1131 310 asini_WP_303220258.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1130 311 asini_WP_311832475.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1130 312 asini_WP_311835806.1SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1130 313 asini_WP_311877876.1 SsaCas9-sim3_structure-St1-Enterococcus- Enterococcus asini 1130 314 asini_WP_311896923.1 TABLE 9A: EXEMPLARY mRNA SEQUENCES RELATED TO RT EDITORS DESCRIPTION SEQ ID NO:1SegPolyA_1 182 SegPolyA_2 183 SegPolyA_3 184 SegPolyA_4 185 120_PolyA 186 130_PolyA 187 3'UTR 188 WPRE 189 ek5 190 STH 191 STH-2 192 5' UTR 193 RT editor mRNA sequence 194 3'UTR-WPRE-SegPolyA_1 195 3'UTR-WPRE-SegPolyA_2 196 3'UTR-WPRE-SegPolyA_3 197 3'UTR-WPRE-SegPolyA_4 198 3'UTR-WPRE-120_PolyA 199 3'UTR-WPRE-130_PolyA 200 3'UTR-STH 201 3'UTR-eK5-120_PolyA 202 3'UTR-eK5-130_PolyA 203 3'UTR-eK5-SegPolyA_1 204 3'UTR-eK5-SegPolyA_2 205 3'UTR-eK5-SegPolyA_3 206 3'UTR-eK5-SegPolyA_4 207 eK5-120_PolyA 208 eK5-130_PolyA 209 eK5-SegPolyA_1 210 eK5-SegPolyA_2 211 eK5-SegPolyA_3 212 eK5-SegPolyA_4 213 3'UTR-STH-120_PolyA 214 3'UTR-STH-130_PolyA 215 3'UTR-STH-SegPolyA_1 216 3'UTR-STH-SegPolyA_2 217 3'UTR-STH-SegPolyA_3 218 3'UTR-STH-SegPolyA_4 219 RT_editor-3'UTR-WPRE-SegPolyA_1 220 RT_editor-3'UTR-WPRE-SegPolyA_2 221 RT_editor-3'UTR-WPRE-SegPolyA_3 222 RT_editor-3'UTR-WPRE-SegPolyA_4 223 RT_editor-3'UTR-WPRE-120_PolyA 224 RT_editor-3'UTR-WPRE-130_PolyA 225 RT_editor-3'UTR-STH 226 RT_editor-3'UTR-eK5-120_PolyA 227 RT_editor-3'UTR-eK5-130_PolyA 228 RT_editor-3'UTR-eK5-SegPolyA_1 229 RT_editor-3'UTR-eK5-SegPolyA_2 230 RT_editor-3'UTR-eK5-SegPolyA_3 231 RT_editor-3'UTR-eK5-SegPolyA_4 232 RT_editor-eK5-120_PolyA 233RT_editor-eK5-130_PolyA 234 RT_editor-eK5-SegPolyA_1 235 RT_editor-eK5-SegPolyA_2 236 RT_editor-eK5-SegPolyA_3 237 RT_editor-eK5-SegPolyA_4 238 RT_editor-3'UTR-STH-120_PolyA 239 RT_editor-3'UTR-STH-130_PolyA 240 RT_editor-3'UTR-STH-SegPolyA_1 241 RT_editor-3'UTR-STH-SegPolyA_2 242 RT_editor-3'UTR-STH-SegPolyA_3 243 RT_editor-3'UTR-STH-SegPolyA_4 244 Comp14 transcript 245 Full-length segment from human 246 MALAT1 (HTH) 1For RNA, T is U; U are T in the sequence listing. TABLE 9B: EXEMPLARY mRNA Elements Name Element Sequence SEQ Position ID NO: SAA1 5UTR AGGGACCCGCAGCTCAGCTACAGCACAGATCAGCACC 940 APOC3 5UTR CTGCTCAGTTCATCCCTAGAGGCAGCTGCTCCAGGAACAGAGGT 941 GCC ALB 5UTR CTAGCTTTTCTCTTCTGTCAACCCCACACGCCTTTGGCACA 942 HP 5UTR ACTGGAAAAGATAGTGACCTTACCAGGGCCAAAGTTTGTAGAC 943 ACAGGAATTACGAAATGGAGAAGGGGGAGAAGTGAGCTAGTGG CAGCATAAAAAGACCAGCAGATGCCCCACAGCACTGCTCTTCCA GAGGCAAGACCAACCAAG ORM1 5UTR AGCACTGCCTGGCTCCACGTGCCTCCTGGTCTCAGT 944 APOA2 5UTR AGGCACAGACACCAAGGACAGAGACGCTGGCTAGGCCGCCCTC 945 CCCACTGTTACCAAC SAA2 5UTR ACTATAAATAGCAGCCACCTCTCCCTGGCAGACAGGGACCCGCA 946 GCTCAGCTACAGCACAGATCAGCACC SERPINA1 5UTR CTCCTCAGCTTCAGGCACCACCACTGACCTGGGACAGTGAATCG 947 ACA APOC1 5UTR AGGCGGTCAGGGGAAGGCTCAGGAGGAGGGAGATCAACATCAA 948 CCTGCCCCGCCCCCTCCCCAGCCTGATAAAGGTCCTGCGGGCAG GACAGGACCTCCCAACCAAGCCCTCCAGCAAGGATTCAGAGTG CCCCTCCGGCCTCGCC Apt17_custom 5UTR ACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCG 949 custom 5UTR AGGATAATATACTTACATACTTACTAATTAATACTAAACTCAAC 950 GCCACC Syn_5'UTR 5UTR AGGAAATAAGAGAGAAAAGAAGAGTAAGAAGAAATATAAGAG 1075 CCACC SAA1 3UTR GCTTCCTCTTCACTCTGCTCTCAGGAGATCTGGCTGTGAGGCCCT 951 CAGGGCAGGGATACAAAGCGGGGAGAGGGTACACAATGGGTAT CTAATAAATACTTAAGAGGTGGAA APOC3 3UTR GACCTCAATACCCCAAGTCCACCTGCCTATCCATCCTGCGAGCT 952 CCTTGGGTCCTGCAATCTCCAGGGCTGCCCCTGTAGGTTGCTTA AAAGGGACAGTATTCTCAGTGCTCTCCTACCCCACCTCATGCCT GGCCCCCCTCCAGGCATGCTGGCCTCCCAATAAAGCTGGACAAG AAGCTGCTATGA ALB 3UTR CATCACATTTAAAAGCATCTCAGCCTACCATGAGAATAAGAGAA 953 AGAAAATGAAGATCAAAAGCTTATTCATCTGTTTTTCTTTTTCGT TGGTGTAAAGCCAACACCCTGTCTAAAAAACATAAATTTCTTTA ATCATTTTGCCTCTTTTCTCTGTGCTTCAATTAATAAAAAATGGA AAGAATCTAATAGAGTGGTACAGCACTGTTATTTTTCAAAGATG TGTTGCTATCCTGAAAATTCTGTAGGTTCTGTGGAAGTTCCAGTGTTCTCTCTTATTCCACTTCGGTAGAGGATTTCTAGTTTCTTGTGG GCTAATTAAATAAATCATTAATACTCTTCTAAGTTATGGATTATA AACATTCAAAATAATATTTTGACATTATGATAATTCTGAATAAA AGAACAAAAACCA HP 3UTR TGCAAGGCTGGCCGGAAGCCCTTGCCTGAAAGCAAGATTTCAGC 954 CTGGAAGAGGGCAAAGTGGACGGGAGTGGACAGGAGTGGATGC GATAAGATGTGGTTTGAAGCTGATGGGTGCCAGCCCTGCATTGC TGAGTCAATCAATAAAGAGCTTTCTTTTGACCCA ORM1 3UTR CAGGACACAGCCTTGGATCAGGACAGAGACTTGGGGGCCATCC 955 TGCCCCTCCAACCCGACATGTGTACCTCAGCTTTTTCCCTCACTT GCATCAATAAAGCTTCTGTGTTTGGAACAGCTAA APOA2 3UTR AGTGTCCAGACCATTGTCTTCCAACCCCAGCTGGCCTCTAGAAC 956 ACCCACTGGCCAGTCCTAGAGCTCCTGTCCCTACCCACTCTTTGC TACAATAAATGCTGAATGAATCCA SAA2 3UTR GCTTCCTCTTCACTCTGCTCTCAGGAGACCTGGCTATGAGGCCCT 957 CGGGGCAGGGATACAAAGTTAGTGAGGTCTATGTCCAGAGAAG CTGAGATATGGCATATAATAGGCATCTAATAAATGCTTAAGAGG TGGAA SERPINA1 3UTR CTGCCTCTCGCTCCTCAACCCCTCCCCTCCATCCCTGGCCCCCTC 958 CCTGGATGACATTAAAGAAGGGTTGAGCTGGTCCCTGCCTGCAT GTGACTGTAAATCCCTCCCATGTTTTCTCTGAGTCTCCCTTTGCC TGCTGAGGCTGTATGTGGGCTCCAGGTAACAGTGCTGTCTTCGG GCCCCCTGAACTGTGTTCATGGAGCATCTGGCTGGGTAGGCACA TGCTGGGCTTGAATCCAGGGGGGACTGAATCCTCAGCTTACGGA CCTGGGCCCATCTGTTTCTGGAGGGCTCCAGTCTTCCTTGTCCTG TCTTGGAGTCCCCAAGAAGGAATCACAGGGGAGGAACCAGATA CCAGCCATGACCCCAGGCTCCACCAAGCATCTTCATGTCCCCCT GCTCATCCCCCACTCCCCCCCACCCAGAGTTGCTCATCCTGCCA GGGCTGGCTGTGCCCACCCCAAGGCTGCCCTCCTGGGGGCCCCA GAACTGCCTGATCGTGCCGTGGCCCAGTTTTGTGGCATCTGCAG CAACACAAGAGAGAGGACAATGTCCTCCTCTTGACCCGCTGTCA CCTAACCAGACTCGGGCCCTGCACCTCTCAGGCACTTCTGGAAA ATGACTGAGGCAGATTCTTCCTGAAGCCCATTCTCCATGGGGCA ACAAGGACACCTATTCTGTCCTTGTCCTTCCATCGCTGCCCCAGA AAGCCTCACATATCTCCGTTTAGAATCAGGTCCCTTCTCCCCAG ATGAAGAGGAGGGTCTCTGCTTTGTTTTCTCTATCTCCTCCTCAG ACTTGACCAGGCCCAGCAGGCCCCAGAAGACCATTACCCTATAT CCCTTCTCCTCCCTAGTCACATGGCCATAGGCCTGCTGATGGCTC AGGAAGGCCATTGCAAGGACTCCTCAGCTATGGGAGAGGAAGC ACATCACCCATTGACCCCCGCAACCCCTCCCTTTCCTCCTCTGAG TCCCGACTGGGGCCACATGCAGCCTGACTTCTTTGTGCCTGTTGC TGTCCCTGCAGTCTTCAGAGGGCCACCGCAGCTCCAGTGCCACG GCAGGAGGCTGTTCCTGAATAGCCCCTGTGGTAAGGGCCAGGA GAGTCCTTCCATCCTCCAAGGCCCTGCTAAAGGACACAGCAGCC AGGAAGTCCCCTGGGCCCCTAGCTGAAGGACAGCCTGCTCCCTC CGTCTCTACCAGGAATGGCCTTGTCCTATGGAAGGCACTGCCCC ATCCCAAACTAATCTAGGAATCACTGTCTAACCACTCACTGTCA TGAATGTGTACTTAAAGGATGAGGTTGAGTCATACCAAATAGTG ATTTCGATAGTTCAAAATGGTGAAATTAGCAATTCTACATGATT CAGTCTAATCAATGGATACCGACTGTTTCCCACACAAGTCTCCT GTTCTCTTAAGCTTACTCACTGACAGCCTTTCACTCTCCACAAAT ACATTAAAGATATGGCCATCACCAAGCCCCCTAGGATGACACCA GACCTGAGAGTCTGAAGACCTGGATCCAAGTTCTGACTTTTCCC CCTGACAGCTGTGTGACCTTCGTGAAGTCGCCAAACCTCTCTGA GCCCCAGTCATTGCTAGTAAGACCTGCCTTTGAGTTGGTATGAT GTTCAAGTTAGATAACAAAATGTTTATACCCATTAGAACAGAGA ATAAATAGAACTACATTTCTTGCA APOC1 3UTR GGACCTGAAGGGTGACATCCCAGGAGGGGCCTCTGAAATTTCCC 959 ACACCCCAGCGCCTGTGCTGAGGACTCCCTCCATGTGGCCCCAG GTGCCACCAATAAAAATCCTACAGAAAA 3WJ1_custom 3UTR TTGCCATGTGTATGTGGGTTTTTTTTTTCCCACATACTCTGATGA 960 TCCTTTTTTTTTTGGATCATTCATGGCAA3WJ2_custom 3UTR TTGCCATGTGTATGTGGGAAAAAAAAAACCCACATACTCTGATG 961 ATCCAAAAAAAAAAGGATCATTCATGGCAA gtgctggtctgtgtgctggcccatcactttggcaaagaattcaccccaccagtgcaggctgcctatcagaaa 1076 gtggtggctggtgtggctaatgccctggcccacaagtatcactaagctcgctttcttgctgtccaatttctatta HBB aaggttcctttgttccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctg Extended 3UTR cctaataaaaaacatttattttcattgca GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCC 1077 CAGCCCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGA hHBA 3UTR ATAAAGTCTGAGTGGGCGGC GCGGCCGCTTAATTAAGCTGCCTTCTGCGGGGCTTGCCTTCTGG 1078 CCATGCCCTTCTTCTCTCCCTTGCACCTGTACCTCTTGGTCTTTGA mHBA 3UTR ATAAAGCCTGAGTAGGAAGTCTAG HBB_extende 3UTR GTGCTGGTCTGTGTGCTGGCCCATCACTTTGGCAAAGAATTCAC 1079 d (HBBext) CCCACCAGTGCAGGCTGCCTATCAGAAAGTGGTGGCTGGTGTGG CTAATGCCCTGGCCCACAAGTATCACTAAGCTCGCTTTCTTGCTG TCCAATTTCTATTAAAGGTTCCTTTGTTCCCTAAGTCCAACTACT AAACTGGGGGATATTATGAAGGGCCTTGAGCATCTGGATTCTGC CTAATAAAAAACATTTATTTTCATTGCA AES-mtRNR1 3UTR CTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTG 1080 GGTACCCCGAGTCTCCCCCGACCTCGGGTCCCAGGTATGCTCCC ACCTCCACCTGCCCCACTCACCACCTCTGCTAGTTCCAGACACCT CCCAAGCACGCAGCAATGCAGCTCAAAACGCTTAGCCTAGCCA CACCCCCACGGGAAACAGCAGTGATTAACCTTTAGCAATAAAC GAAAGTTTAACTAAGCTATACTAACCCCAGGGTTGGTCAATTTC GTGCCAGCCACACC EK5 AE1cactcctccatggtgatataaagaccacccacttccttcgggtgagcccctaagccatggttgtactgcacta 962 tcatcctaagacggtccttcttcggatcgcaatctcaccctggtgccgcgcttccttcgggaactgcacccg cggaccagggccgtctttgaacttttctaactgttcttac STH AE1GTCATGAAGGTTTTTCTTTTCCTGAGAAAACAAgaaaTTGTTTTCT 963 CAGGTTTTGCTTTTTGAAAAAAAAGCAAAA 1AE (additional elements) can be attached either downstream of the 3’-UTR or the polyA tail
[0241] Table 9C below additionally discloses 5’-UTR sequence variants wherein the difference in sequences are not within the UTR itself but in the second codon, as shown in the fourth column. For CESvar1, CESvar2, and CESvar3, ‘gcc’ was inserted in the second codon position. Column 4 shows the sequence of the first three codons. TABLE 9C: EXEMPLARY mRNA Elements Name ElementSEQ ID PositionSequence of UTR First 3 codonsNO CES15UTRagACAGAGACCTCGCAGGCCCCGAGAACTGTCGCCCT 1081 var1 TCCACGATGgccAAACES15UTR agacagagacctcgcaggccccgag1082 var2aactgtcgcccttccacc ATGgccAAACES15UTR agacagagacctcgcaggccccgagaactgtc1083 var3gccctgccacc ATGgccAAACES15UTRAGACAGAGACCTCGCAGGCCCCGAGAACTGTCGCCCAT1084 var4 TTCCACGG---AAA
[0242] Presented in Table 10 below are mRNA sequences of the components of exemplary RT editors comprising variant 3’UTR sequences. The components are listed from left- to-right in the 5’ to 3’ direction, using SEQ ID NOs to denote the sequence.TABLE 10: EXEMPLARY 3’UTR VARIANTS DESCRIPTION 5’ UTR RT EDITOR 3'UTR 3'UTR Structure PolyA SEQ ID SEQ ID NO: element#1 element#2 motif SEQ ID NO: SEQ ID NO: SEQ ID NO: SEQ ID NO: NO: RT_editor-3'UTR- 193 194 188 189 182 WPRE- SegPolyA_1 (SEQ ID NO: 220) RT_editor-3'UTR- 193 194 188 189 183 WPRE- SegPolyA_2 (SEQ ID NO: 221) RT_editor-3'UTR- 193 194 188 189 184 WPRE- SegPolyA_3 (SEQ ID NO: 222) RT_editor-3'UTR- 193 194 188 189 185 WPRE- SegPolyA_4 (SEQ ID NO: 223) RT_editor-3'UTR- 193 194 188 189 186 WPRE- 120_PolyA (SEQ ID NO: 224) RT_editor-3'UTR- 193 194 188 189 187 WPRE- 130_PolyA (SEQ ID NO: 225) RT_editor-3'UTR- 193 194 188 191 STH (SEQ ID NO: 226) RT_editor-3'UTR- 193 194 188 190 186 eK5-120_PolyA (SEQ ID NO: 227) RT_editor-3'UTR- 193 194 188 190 187 eK5-130_PolyA (SEQ ID NO: 228) RT_editor-3'UTR- 193 194 188 190 182 eK5-SegPolyA_1 (SEQ ID NO: 229) RT_editor-3'UTR- 193 194 188 190 183 eK5-SegPolyA_2 (SEQ ID NO: 230) RT_editor-3'UTR- 193 194 188 190 184 eK5-SegPolyA_3 (SEQ ID NO: 231) RT_editor-3'UTR- 193 194 188 190 185 eK5-SegPolyA_4 (SEQ ID NO: 232) RT_editor-eK5- 193 194 190 186 120_PolyA (SEQ ID NO: 233) RT_editor-eK5- 193 194 190 187 130_PolyA (SEQ ID NO: 234) RT_editor-eK5- 193 194 190 182 SegPolyA_1 (SEQ ID NO: 235) RT_editor-eK5- 193 194 190 183 SegPolyA_2 (SEQ ID NO: 236)RT_editor-eK5- 193 194 190 184 SegPolyA_3 (SEQ ID NO: 237) RT_editor-eK5- 193 194 190 185 SegPolyA_4 (SEQ ID NO: 238) RT_editor-3'UTR- 193 194 188 192 186 STH-120_PolyA (SEQ ID NO: 239) RT_editor-3'UTR- 193 194 188 192 187 STH-130_PolyA (SEQ ID NO: 240) RT_editor-3'UTR- 193 194 188 192 182 STH-SegPolyA_1 (SEQ ID NO: 241) RT_editor-3'UTR- 193 194 188 192 183 STH-SegPolyA_2 (SEQ ID NO: 242) RT_editor-3'UTR- 193 194 188 192 184 STH-SegPolyA_3 (SEQ ID NO: 243) RT_editor-3'UTR- 193 194 188 192 185 STH-SegPolyA_4 (SEQ ID NO: 244) TABLE 11: EXEMPLARY REVERSE TRANSCRIPTASE VARIANTS NAME SPECIES (native or derived) Length SEQ ID NO: RT_RTBat-A_ARS01_RTEni-V-derived Eptesicus nilssonii, bat 672 250 RT_RTBat-A_ARS01_With-RT5m_Prot- Eptesicus nilssonii, bat 668 251 neg2 RT_RTEni_With-RT5m_Eptesicus- Eptesicus nilssonii, bat 679 252 nilssonii RT_RTEni_With-RT5m_Eptesicus- Eptesicus nilssonii, bat 672 253 nilssonii_within-prot RT_RTEni_With-RT5m_Eptesicus- Eptesicus nilssonii, bat 668 315 nilssonii_Prot-neg2 RT_RTMbr-A_Myotis-brandtii Myotis brandtii, bat 673 254 RT_RTMbr-A_Myotis-brandtii_With- Myotis brandtii, bat 669 255 RT5m_Prot-neg2 RT_RTMcr-A Melozone crissalis, bird 741 256 RT_RTMcr-A_Prot-neg2 Melozone crissalis, bird 675 257 RT_RTMcr-A_Melozone-crissalis_Prot- Melozone crissalis, bird 675 316 neg2_With-RT5m RT_RTMda-A Myotis daubertonii, bat 673 258 RT_RTMda-A_Myotis-daubentonii_With- Myotis daubentonii, bat 669 259 RT5m_Prot-neg2 RT_RTMge_With-RT5m Melospiza georgiana, bird 685 141 RT_RTMge_Melospiza-georgiana_With- Melospiza georgiana, bird 675 260 RT5m_Prot-neg2 RT_RTMge-A Melospiza georgiana, bird 713 261 RT_RTMge-A_Prot-neg2 Melospiza georgiana, bird 671 262 RT_RTMge-A_Prot-neg2_With-RT5m Melospiza georgiana, bird 671 317 RT_RTMge-B Melospiza georgiana, bird 741 263 RT_RTMge-C_var-of-B_Prot-neg2 Melospiza georgiana, bird 675 264 RT_RTMge-C_var-of-B_With- Melospiza georgiana, bird 675 318 RT5m_Prot-neg2 RT_RTMge-C_var-of-B_Within-prot- Melospiza georgiana, bird 679 265 cleavageRT_RTPku-A_Pipistrellus-kuhlii Pipistrellus kuhlii, bat 673 266 RT_RTPku-A_Pipistrellus-kuhlii_With- Pipistrellus kuhlii, bat 669 267 RT5m_Prot-neg2 RT_RTZal-A Zonotrichia albicollis, bird 741 268 RT_RTZal-A_Prot-neg2 Zonotrichia albicollis, bird 675 269 RT_RTZal-A_Zonotrichia-albicollis_Prot- Zonotrichia albicollis, bird 675 319 neg2_With-RT5m Klebsiella pneumoniae mut1Klebsiella pneumonia (added 425 899 mutations E48K,S339K, Y395E) Burkholderia cepacia mut1Burkholderia cepacia (added 423 900 mutations T47K, S87E, S338K, Y394E) Caldibacillus thermoamylovorans Caldibacillus 442 901 thermoamylovorans Caldibacillus thermoamylovorans mut1Caldibacillus 442 902 thermoamylovorans (added mutations T53K, S357K, Y408E) Lysinibacillus sphaericus Lysinibacillus sphaericus 459 903 Lysinibacillus sphaericus mut1Lysinibacillus sphaericus 459 904 (added mutations E63K, S374K, Y425E) Bacillus sp. Bacillus sp. 438 905 Bacillus sp. mut1Bacillus sp. (added mutations 438 906 S339I, S351K, Y402E) Truncated RT from Myotis brandtii Myotis brandtii 492 907 Truncated RT from Pipistrellus kuhlii Pipistrellus kuhlii 492 908 Truncated RT from MMLV MMLV 492 909 pPG273_MMLV_protneg_RT 667 1015 PG270_Woolly_monkey_RT 672 1016 PG273_Melospiza_consensus_protneg_R 675 1017 T PG273_Avian_protneg_RT 694 1018 TABLE 12A: EXEMPLARY tagRNA SEQUENCES FOR EDITING PAH NAME SEQUENCE1SEQ ID NO: hPAH_NGG1_F15E6 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 320 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAGUUCGC hPAH_NGG1_F14E6 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 321 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAGUUCG hPAH_NGG1_F13E6 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 322 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAGUUC hPAH_NGG1_F12E6 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 323 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAGUU hPAH_NGG1_F11E6 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 324 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAGU hPAH_NGG1_F10E6 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 325 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAG hPAH_NGG1_F9E6_ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 326 TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUCAhPAH_NGG1_F8E6_ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 327 TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCUC hPAH_NGG1_F7E6_ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 328 TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCUCGGCCCUUCU hPAH_NGG1_F15E8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 329 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCAGUUCGC hPAH_NGG1_F14E8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 330 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCAGUUCG hPAH_NGG1_F13E8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 331 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCAGUUC hPAH_NGG1_F12E8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 332 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCAGUU hPAH_NGG1_F11E8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 333 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCAGU hPAH_NGG1_F10E8 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 334 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCAG hPAH_NGG1_F9E8_ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 335 TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUCA hPAH_NGG1_F8E8_ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 336 TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCUC hPAH_NGG1_F7E8_ UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 337 TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCUACCUCGGCCCUUCU hPAH_NGG1_F15E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 338 0_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCAGUUCGC hPAH_NGG1_F14E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 339 0_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCAGUUCG hPAH_NGG1_F13E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 340 0_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCAGUUC hPAH_NGG1_F12E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 341 0_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCAGUU hPAH_NGG1_F11E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 342 0_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCAGU hPAH_NGG1_F10E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 343 0_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCAG hPAH_NGG1_F9E10 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 344 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUCA hPAH_NGG1_F8E10 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 345 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUC hPAH_NGG1_F7E10 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 346 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCAAUACCUCGGCCCUUCUhPAH_NGG1_F14E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 347 2_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUCAGUUCG hPAH_NGG1_F13E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 348 2_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUCAGUUC hPAH_NGG1_F12E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 349 2_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUCAGUU hPAH_NGG1_F11E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 350 2_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUCAGU hPAH_NGG1_F10E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 351 2_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUCAG hPAH_NGG1_F9E12 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 352 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUCA hPAH_NGG1_F8E12 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 353 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCUC hPAH_NGG1_F7E12 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 354 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCACAAUACCUCGGCCCUUCU hPAH_NGG1_F15E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 355 4_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCAGUUCGC hPAH_NGG1_F14E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 356 4_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCAGUUCG hPAH_NGG1_F13E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 357 4_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCAGUUC hPAH_NGG1_F12E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 358 4_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCAGUU hPAH_NGG1_F11E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 359 4_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCAGU hPAH_NGG1_F10E1 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 360 4_TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCAG hPAH_NGG1_F9E14 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 361 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUCA hPAH_NGG1_F8E14 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 362 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCUC hPAH_NGG1_F7E14 UAGCGAACUGAGAAGGGCCAGUUUUAGAGCUAGAAAUAGCAAG 363 _TtoC UUAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCG AGUCGGUGCCCACAAUACCUCGGCCCUUCU hPAH_NGG2_F15E7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 364 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGCAGCAAA hPAH_NGG2_F14E7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 365 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGCAGCAA hPAH_NGG2_F13E7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 366 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGCAGCAhPAH_NGG2_F12E7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 367 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGCAGC hPAH_NGG2_F11E7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 368 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGCAG hPAH_NGG2_F10E7 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 369 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGCA hPAH_NGG2_F9E7_ ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 370 AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGGC hPAH_NGG2_F8E7_ ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 371 AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUGG hPAH_NGG2_F7E7_ ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 372 AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGCCGAGGUAUUGUG hPAH_NGG2_F15E9 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 373 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGCAGCAAA hPAH_NGG2_F14E9 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 374 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGCAGCAA hPAH_NGG2_F13E9 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 375 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGCAGCA hPAH_NGG2_F12E9 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 376 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGCAGC hPAH_NGG2_F11E9 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 377 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGCAG hPAH_NGG2_F10E9 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 378 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGCA hPAH_NGG2_F9E9_ ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 379 AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGGC hPAH_NGG2_F8E9_ ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 380 AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUGG hPAH_NGG2_F7E9_ ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 381 AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCGGGCCGAGGUAUUGUG hPAH_NGG2_F15E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 382 1_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGCAGCAAA hPAH_NGG2_F14E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 383 1_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGCAGCAA hPAH_NGG2_F13E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 384 1_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGCAGCA hPAH_NGG2_F12E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 385 1_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGCAGC hPAH_NGG2_F11E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 386 1_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGCAGhPAH_NGG2_F10E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 387 1_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGCA hPAH_NGG2_F9E11 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 388 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGGC hPAH_NGG2_F8E11 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 389 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUGG hPAH_NGG2_F7E11 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 390 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAAGGGCCGAGGUAUUGUG hPAH_NGG2_F15E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 391 3_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGCAGCAAA hPAH_NGG2_F14E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 392 3_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGCAGCAA hPAH_NGG2_F13E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 393 3_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGCAGCA hPAH_NGG2_F12E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 394 3_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGCAGC hPAH_NGG2_F11E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 395 3_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGCAG hPAH_NGG2_F10E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 396 3_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGCA hPAH_NGG2_F9E13 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 397 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGGC hPAH_NGG2_F8E13 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 398 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUGG hPAH_NGG2_F7E13 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 399 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCAGAAGGGCCGAGGUAUUGUG hPAH_NGG2_F15E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 400 5_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGCAGCAAA hPAH_NGG2_F14E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 401 5_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGCAGCAA hPAH_NGG2_F13E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 402 5_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGCAGCA hPAH_NGG2_F12E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 403 5_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGCAGC hPAH_NGG2_F11E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 404 5_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGCAG hPAH_NGG2_F10E1 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 405 5_AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGCA hPAH_NGG2_F9E15 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 406 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGGChPAH_NGG2_F8E15 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 407 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUGG hPAH_NGG2_F7E15 ACUUUGCUGCCACAAUACCUGUUUUAGAGCUAGAAAUAGCAAGU 408 _AtoG UAAAAUAAGGCUAGUCCGUUAUCAACUUGAAAAAGUGGCACCGA GUCGGUGCUGAGAAGGGCCGAGGUAUUGUG1For RNA, T is U; U are T in the sequence listing. For each entry, the spacer sequence is underlined and the extension arm is bold. The description indicates the length of the FBS (F) and ET (E) or homology region of the extension arm. TABLE 12B: EXEMPLARY MODIFIED tagRNA SEQUENCES FOR EDITING PAH NAME SEQUENCE1,2SEQ ID NO: hPAH_NGG1_F15E6 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 409 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrCrUrCrArGrUrU*mC*mG*mC hPAH_NGG1_F14E6 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 410 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrCrUrCrArGrU*mU*mC*mG hPAH_NGG1_F13E6 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 411 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrCrUrCrArG*mU*mU*mC hPAH_NGG1_F12E6 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 412 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrCrUrCrA*mG*mU*mU hPAH_NGG1_F11E6 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 413 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrCrUrC*mA*mG*mU hPAH_NGG1_F10R6 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 414 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrCrU*mC*mA*mG hPAH_NGG1_F9E6_ mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 415 TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrUrC*mU*mC*mA hPAH_NGG1_F8E6_ mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 416 TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrUrU*mC*mU*mC hPAH_NGG1_F7R6_ mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 417 TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrUrCrGrGrCrCrCrU*mU*mC*mU hPAH_NGG1_F15E8 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 418 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrUrU*mC*mG*mC hPAH_NGG1_F14E8 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 419 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrU*mU*mC*mG hPAH_NGG1_F13E8 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 420 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArG*mU*mU*mC hPAH_NGG1_F12E8 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 421 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrCrUrCrA*mG*mU*mU hPAH_NGG1_F11E8 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 422 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrCrUrC*mA*mG*mU hPAH_NGG1_F10E8 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 423 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrCrU*mC*mA*mG hPAH_NGG1_F9E8_ mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 424 TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrUrC*mU*mC*mA hPAH_NGG1_F8E8_ mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 425 TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrUrU*mC*mU*mC hPAH_NGG1_F7E8_ mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 426 TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrU rArCrCrUrCrGrGrCrCrCrU*mU*mC*mU hPAH_NGG1_F15E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 427 0_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrUrU*mC*mG*mC hPAH_NGG1_F14E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 428 0_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrU*mU*mC*mG hPAH_NGG1_F13E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 429 0_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArG*mU*mU*mC hPAH_NGG1_F12E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 430 0_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrA*mG*mU*mUhPAH_NGG1_F11E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 431 0_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrC*mA*mG*mU hPAH_NGG1_F10E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 432 0_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrCrU*mC*mA*mG hPAH_NGG1_F9E10 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 433 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrUrC*mU*mC*mA hPAH_NGG1_F8E10 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 434 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrUrU*mC*mU*mC hPAH_NGG1_F7E10 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 435 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rArUrArCrCrUrCrGrGrCrCrCrU*mU*mC*mU hPAH_NGG1_F14E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 436 2_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrU*mU*mC* mG hPAH_NGG1_F13E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 437 2_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArG*mU*mU*mC hPAH_NGG1_F12E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 438 2_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrA*mG*mU*mU hPAH_NGG1_F11E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 439 2_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrC*mA*mG*mU hPAH_NGG1_F10E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 440 2_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrU*mC*mA*mG hPAH_NGG1_F9R12 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 441 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrUrC*mU*mC*mA hPAH_NGG1_F8E12 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 442 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrUrU*mC*mU*mChPAH_NGG1_F7E12 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 443 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrA rCrArArUrArCrCrUrCrGrGrCrCrCrU*mU*mC*mU hPAH_NGG1_F15E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 444 4_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrUrU*m C*mG*mC hPAH_NGG1_F14E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 445 4_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArGrU*mU* mC*mG hPAH_NGG1_F13R1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 446 4_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrArG*mU*m U*mC hPAH_NGG1_F12E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 447 4_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrCrA*mG*mU* mU hPAH_NGG1_F11E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 448 4_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrUrC*mA*mG*mU hPAH_NGG1_F10E1 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 449 4_TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrCrU*mC*mA*mG hPAH_NGG1_F9E14 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 450 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrUrC*mU*mC*mA hPAH_NGG1_F8R14 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 451 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrUrU*mC*mU*mC hPAH_NGG1_F7E14 mU*mA*mG*rCrGrArArCrUrGrArGrArArGrGrGrCrCrArGrUrUrUrUrAr 452 _TtoC GrAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUr ArArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmA mAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrC rCrArCrArArUrArCrCrUrCrGrGrCrCrCrU*mU*mC*mU hPAH_NGG2_F15E7 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 453 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrUrGrGrCrArGrC*mA*mA*mA hPAH_NGG2_F14E7 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 454 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrUrGrGrCrArG*mC*mA*mA hPAH_NGG2_F13E7 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 455 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrUrGrGrCrA*mG*mC*mA hPAH_NGG2_F12E7 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 456 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrUrGrGrC*mA*mG*mC hPAH_NGG2_F11E7 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 457 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrUrGrG*mC*mA*mG hPAH_NGG2_F10E7 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 458 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrUrG*mG*mC*mA hPAH_NGG2_F9E7_ mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 459 AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrGrU*mG*mG*mC hPAH_NGG2_F8E7_ mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 460 AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrUrG*mU*mG*mG hPAH_NGG2_F7E7_ mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 461 AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr CrCrGrArGrGrUrArUrU*mG*mU*mG hPAH_NGG2_F15E9 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 462 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArGrC*mA*mA*mA hPAH_NGG2_F14E9 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 463 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArG*mC*mA*mA hPAH_NGG2_F13E9 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 464 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrA*mG*mC*mA hPAH_NGG2_F12E9 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 465 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrUrGrGrC*mA*mG*mC hPAH_NGG2_F11E9 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 466 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrUrGrG*mC*mA*mGhPAH_NGG2_F10E9 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 467 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrUrG*mG*mC*mA hPAH_NGG2_F9E9_ mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 468 AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrGrU*mG*mG*mC hPAH_NGG2_F8E9_ mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 469 AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrUrG*mU*mG*mG hPAH_NGG2_F7E9_ mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 470 AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrGr GrGrCrCrGrArGrGrUrArUrU*mG*mU*mG hPAH_NGG2_F15E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 471 1_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArGrC*mA*mA* mA hPAH_NGG2_F14E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 472 1_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArG*mC*mA*mA hPAH_NGG2_F13E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 473 1_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrA*mG*mC*mA hPAH_NGG2_F12E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 474 1_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrC*mA*mG*mC hPAH_NGG2_F11E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 475 1_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrG*mC*mA*mG hPAH_NGG2_F10E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 476 1_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrUrG*mG*mC*mA hPAH_NGG2_F9E11 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 477 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrGrU*mG*mG*mC hPAH_NGG2_F8E11 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 478 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrUrG*mU*mG*mGhPAH_NGG2_F7E11 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 479 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr ArGrGrGrCrCrGrArGrGrUrArUrU*mG*mU*mG hPAH_NGG2_F15E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 480 3_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArGrC*mA* mA*mA hPAH_NGG2_F14E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 481 3_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArG*mC*m A*mA hPAH_NGG2_F13E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 482 3_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrA*mG*mC* mA hPAH_NGG2_F12E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 483 3_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrC*mA*mG*mC hPAH_NGG2_F11E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 484 3_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrG*mC*mA*mG hPAH_NGG2_F10E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 485 3_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrG*mG*mC*mA hPAH_NGG2_F9E13 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 486 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrGrU*mG*mG*mC hPAH_NGG2_F8E13 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 487 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrUrG*mU*mG*mG hPAH_NGG2_F7E13 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 488 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrAr GrArArGrGrGrCrCrGrArGrGrUrArUrU*mG*mU*mG hPAH_NGG2_F15E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 489 5_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArGrC *mA*mA*mA hPAH_NGG2_F14E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 490 5_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAmAmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrArG*m C*mA*mA hPAH_NGG2_F13E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 491 5_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrCrA*mG* mC*mA hPAH_NGG2_F12E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 492 5_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrGrC*mA*m G*mC hPAH_NGG2_F11E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 493 5_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrGrG*mC*mA* mG hPAH_NGG2_F10E1 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 494 5_AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrUrG*mG*mC*mA hPAH_NGG2_F9E15 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 495 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrGrU*mG*mG*mC hPAH_NGG2_F8E15 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 496 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrUrG*mU*mG*mG hPAH_NGG2_F7E15 mA*mC*mU*rUrUrGrCrUrGrCrCrArCrArArUrArCrCrUrGrUrUrUrUrArG 497 _AtoG rAmGmCmUmAmGmAmAmAmUmAmGmCrArArGrUrUrArArArArUrAr ArGrGrCrUrArGrUrCrCrGrUrUrArUrCmAmAmCmUmUmGmAmAmAm AmAmGmUmGmGmCmAmCmCmGmAmGmUmCmGmGmUmGmCrUr GrArGrArArGrGrGrCrCrGrArGrGrUrArUrU*mG*mU*mG1For RNA, T is U; U are T in the sequence listing. For each entry, the spacer sequence is underlined and the extension arm is bold. The description indicates the length of the FBS (F) and ET (E) or homology region of the extension arm.2For the sequences: *, phosphorothioate modification; mN, 2’O-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A,C,T / U,G. TABLE 12C: EXEMPLARY SPACER AND PROTOSPACER SEQUENCES FOR EDITING PAH SEQ ID SPACER SEQUENCE1SEQ ID PROTOSPACER SEQUENCE NO: NO: 793 UAGCGAACUGAGAAGGGCCA 795 TAGCGAACTGAGAAGGGCCA 794 ACUUUGCUGCCACAAUACCU 796 ACTTTGCTGCCACAATACCT 1For RNA, U are T in the sequence listing. TABLE 12D: EXEMPLARY FLAP BINDING SEQUENCES AND TEMPLATES FOREDITING PAH DESCRIPTION FBS1SEQ ID NO: EDITING TEMPLATE1SEQ ID NO: hPAH_NGG1_F15E6_TtoC CCCUUCUCAGUUCGC 803 CCUCGG hPAH_NGG1_F14E6_TtoC CCCUUCUCAGUUCG 804 CCUCGG hPAH_NGG1_F13E6_TtoC CCCUUCUCAGUUC 805 CCUCGG hPAH_NGG1_F12E6_TtoC CCCUUCUCAGUU 806 CCUCGG hPAH_NGG1_F11E6_TtoC CCCUUCUCAGU 807 CCUCGG hPAH_NGG1_F10E6_TtoC CCCUUCUCAG 808 CCUCGG hPAH_NGG1_F9E6_TtoC CCCUUCUCA CCUCGG hPAH_NGG1_F8E6_TtoC CCCUUCUC CCUCGG hPAH_NGG1_F7E6_TtoC CCCUUCU CCUCGG hPAH_NGG1_F15E8_TtoC CCCUUCUCAGUUCGC 803 UACCUCGG hPAH_NGG1_F14E8_TtoC CCCUUCUCAGUUCG 804 UACCUCGG hPAH_NGG1_F13E8_TtoC CCCUUCUCAGUUC 805 UACCUCGG hPAH_NGG1_F12E8_TtoC CCCUUCUCAGUU 806 UACCUCGG hPAH_NGG1_F11E8_TtoC CCCUUCUCAGU 807 UACCUCGG hPAH_NGG1_F10E8_TtoC CCCUUCUCAG 808 UACCUCGG hPAH_NGG1_F9E8_TtoC CCCUUCUCA UACCUCGG hPAH_NGG1_F8E8_TtoC CCCUUCUC UACCUCGG hPAH_NGG1_F7E8_TtoC CCCUUCU UACCUCGG hPAH_NGG1_F15E10_TtoC CCCUUCUCAGUUCGC 803 AAUACCUCGG 797 hPAH_NGG1_F14E10_TtoC CCCUUCUCAGUUCG 804 AAUACCUCGG 797 hPAH_NGG1_F13E10_TtoC CCCUUCUCAGUUC 805 AAUACCUCGG 797 hPAH_NGG1_F12E10_TtoC CCCUUCUCAGUU 806 AAUACCUCGG 797 hPAH_NGG1_F11E10_TtoC CCCUUCUCAGU...
Claims
WHAT IS CLAIMED IS:
1. A reverse transcriptase (RT) editing system comprising: a fusion protein comprising a Cas9 nickase and a reverse transcriptase, and a template armed guide RNA (tagRNA) comprising from 5’ to 3’ a spacer sequence, a scaffold sequence, an editing template and a flap binding sequence.
2. The RT editing system of claim 1, wherein the RT editing system further comprises an enhancer guide RNA (egRNA).
3. A reverse transcriptase (RT) editor comprising: a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein, wherein the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, and wherein the DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314.
4. A template armed guide RNA (tagRNA) comprising: a spacer that is complementary to a search target sequence on a first strand of a double stranded target DNA; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the double stranded target DNA; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, and wherein the first strand and the second strand are complementary to each other.
5. The tagRNA of claim 4, wherein the tagRNA comprises a flap binding sequence at least partially complementary to the spacer.
6. The tagRNA of any one of claims 4-5, wherein the scaffold sequence is between the spacer and the editing template.
7. The tagRNA of any one of claims 5-6, comprising from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the flap binding sequence.
8. The tagRNA of any one of claims 5-7, wherein the spacer, the scaffold sequence, the editing template, and the flap binding sequence form a contiguous sequence in a single molecule.
9. The tagRNA of any one of claims 4-8, wherein the editing template comprises an intended nucleotide edit compared to the double stranded target DNA.
10. The tagRNA of claim 9, wherein the tagRNA guides the RT editor to incorporate the intended nucleotide edit into the double stranded target DNA when the tagRNA is contacted with the double stranded target DNA.
11. The tagRNA of any one of claims 9-10, wherein the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA.
12. The tagRNA of any one of claims 4-11, wherein the search target sequence is complementary to a protospacer sequence in the double stranded target DNA, and wherein the protospacer sequence is adjacent to a protospacer adjacent motif (PAM) in the double stranded target DNA.
13. The tagRNA of claim 12, wherein the tagRNA results in incorporation of a nucleotide edit in the PAM when contacted with the double stranded target DNA.
14. The tagRNA of any one of claims 4-13, wherein the spacer of the tagRNA is from 16 to 25 nucleotides in length, optionally 20 nucleotides in length or 21-23 nucleotides in length.
15. The tagRNA of any one of claims 5-14, wherein the flap binding sequence is about 2 to 20 nucleotides in length, optionally about 8 to 16 nucleotides in length or 6 nucleotides in length.
16. The tagRNA of any one of claims 4-15, wherein the editing template is about 4 to 30 nucleotides in length, optionally about 10 to 30 nucleotides in length, further optionally 6 to 9 nucleotides in length.
17. The tagRNA of any one of claims 9-16, wherein the tagRNA results in incorporation of the intended nucleotide edit about 0 to 30 base pairs downstream of the nickase cleavage site.
18. The tagRNA of any one of claims 9-17, wherein the intended nucleotide edit comprises a single nucleotide substitution compared to the region corresponding to the editingtarget in the double stranded target DNA, optionally, the single nucleotide substitution is a T>G or T>C substitution.
19. The tagRNA of any one of claims 9-18, wherein the intended nucleotide edit comprises an insertion compared to the region corresponding to the editing target in the double stranded target DNA, optionally an insertion of a nucleotide sequence at least 50, at least 45, at least 40, at least 35, at least 30, at least 25, at least 20, at least 15, at least 10, or at least 5, nucleotides in length.
20. The tagRNA of any one of claims 9-19, wherein the intended nucleotide edit comprises a deletion compared to the region corresponding to the editing target in the double stranded target DNA.
21. The tagRNA of any one of claims 4-20, wherein the editing template comprises one or more silent nucleotide edits compared to the region corresponding to the editing target in the double stranded target DNA, optionally said silent nucleotide edits do not alter the amino acid sequence of the protein encoded by the double stranded target DNA, further optionally said one or more silent nucleotide edits comprise a substitution of 2 to 5 contiguous nucleotides.
22. The tagRNA of any one of claims 4-21, wherein the editing template comprises a wild type double stranded target DNA sequence.
23. The tagRNA of claim 22, wherein the tagRNA results in correction of a mutation when contacted with the double stranded target DNA.
24. The tagRNA of any one of claims 4-23, wherein the tagRNA comprises any one of the sequences of SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203 or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 16-47, 99-130, 164-181, 320-497, 563-682, 791-792, 1046-1053, and 1089-1203.
25. The tagRNA of any one of claims 4-24, further comprising: 3’ mN*mN*mN*N and 5’ mN*mN*mN* modifications, where m indicates that the nucleotide contains a 2’-O-Me modification and a * indicates the presence of a phosphorothioate bond; and / or a structural motif at the 3’ terminus selected from the group consisting of: an inverted-dT, a prequeosine1-1 riboswitch aptamer (evopreQ1) and variants thereof, a frameshifting pseudoknot from Moloney murine leukemia virus (MMLV) (mpknot), G- quadruplexes, hairpin structures, xrRNA, and a P4-P6 domain of the group I intron; optionally, the structural motif is evopreQ1 or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 84-90.
26. A reverse transcriptase (RT) editing system comprising:a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor, wherein the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, and wherein the DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314.
27. A reverse transcriptase (RT) editing system comprising: the tagRNA according to any one of claims 4-25, or a nucleic acid encoding the tagRNA; and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor.
28. The RT editing system of any one of claims 26-27, further comprising: an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises a egRNA spacer that is complementary to a second search target sequence in the double stranded target DNA, optionally the egRNA comprises a scaffold sequence.
29. The RT editing system of claim 28, wherein the second search target sequence is on the second strand of the double stranded target DNA.
30. The RT editing system of any one of claims 28-29, wherein the egRNA spacer is from 16 to 22 nucleotides in length, optionally 20 nucleotides in length.
31. The RT editing system of any one of claims 26-30, wherein the intended nucleotide edit incorporation rate of the RT editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80%.
32. A reverse transcriptase (RT) editing complex comprising: template armed guide RNA (tagRNA) comprising: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the RT editor, wherein the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, and wherein the DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 142-163 and 270-314.
33. A reverse transcriptase (RT) editing complex comprising: (i) the tagRNA of any one of claims 4-25 and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; or (ii) the RT editing system of any one of claims 26-31.
34. The RT editing complex of any one of claims 32-33, wherein the intended nucleotide edit incorporation rate of the RT editing complex is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80%.
35. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-34, wherein the DNA endonuclease domain is a CRISPR associated (Cas) protein domain.
36. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of claim 35, wherein the Cas protein domain has nickase activity.
37. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 35-36, wherein the Cas protein domain is a Cas9.
38. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of claim 37, wherein the Cas9 comprises a mutation in an HNH domain.
39. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 37-38, wherein the Cas9 comprises an H840A mutation in the HNH domain.
40. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 35-36, wherein the Cas protein domain is a Cas12b.
41. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 35-36, wherein the Cas protein domain is a Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas14a, Cas14b, Cas14c, Cas14d, Cas14e, Cas14f, Cas14g, Cas14h, Cas14u, Casφ, Ssch1Cas9, Sro1Cas9, Sha4Cas9, SsuCas9, iSpyMacCas9, Ssi5Cas9, Ssi8Cas9, Ssci4Cas9, Shy1Cas9, Sag3Cas9, Slutr1Cas9, Ssch3Cas9, SpRYCas9, SpRYcCas9, Sma2Cas9, SsaCas9, EvoCjCas9, or iSpyMac.
42. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-41, wherein the DNA polymerase domain is a reverse transcriptase, optionally the editing template is a reverse transcription template.
43. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of claim 42, wherein the reverse transcriptase is a retrovirus reverse transcriptase.
44. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 42-43, wherein the reverse transcriptase is a Moloney murine leukemia virus (MMLV) reverse transcriptase.
45. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-44, wherein the DNA polymerase domain and the DNA binding domain, a DNA endonuclease domain are fused or linked to form a fusion protein.
46. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of claim 45, wherein:the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 141, 250-269, 315-319, 899-909, and 1015-1018; and / or the DNA endonuclease domain comprises a Cas9 nickase, optionally a Cas9 nickase comprising an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 142-163 and 270-314, optionally an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 142- 163 and 270-314.
47. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-46, wherein the editing template comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ Fluoro RNA base(s).
48. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-47, wherein the FBS comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, or 82’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, or 82’ Fluoro RNA base(s).
49. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-48, wherein the scaffold sequence comprises a nucleotide sequence that is at least 80% identical to any one of SEQ ID NOs: 683-718, 894-898, and 964-1014, optionally a nucleotide sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 683-718, 894-898, and 964-1014.
50. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-49, wherein the scaffold sequence comprises one or more nucleotide substitutions, insertions, and / or deletions at one or more nucleotide positions relative to the parental scaffold sequence of SEQ ID NO: 683, optionally: dinucleotide substitutions at nucleotide positions 5-6 and 34-35, optionally configured such that the nucleotide sequence at positions 5-6 is complementary to the nucleotide sequence at positions 34-35; a deletion at nucleotide positions 55-87, 57-87, 59-87, 61-87, or 63-87; a substitution at one or more of nucleotide positions 17, 18, 19, and 20, optionally a UUCG substitution at nucleotide positions 17-20; a deletion at nucleotide positions 14-16 and 21-24; a replacement of nucleotides at nucleotide positions 11-28 with GUUCGC; and / ora deletion at nucleotide positions 12-16 and 21-26.
51. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-50, wherein the scaffold sequence comprises one or more modifications and / or one or more nucleotide substitutions, insertions, and / or deletions at one or more nucleotide positions relative to the parental scaffold sequence of SEQ ID NO: 700, optionally: a substitution at one or more of nucleotide positions 13, 18, 21, 40, 50, and 52-53, further optionally: a U-to-C substitution at nucleotide position 13; an A-to-G substitution at nucleotide position 18; an A-to-C or an A-to-U substitution at nucleotide position 21; a C-to-U or a C-to-A or a C-to-G substitution at nucleotide position 40; a U-to-C or a U-to-A or a U-to-G substitution at nucleotide position 50; an A-to-C or an A-to-U or an A-to-G substitution at nucleotide position 52; and / or an A-to-C or an A-to-U or an A-to-G substitution at nucleotide position 53; the one or more modifications comprise nucleoside modification(s), sugar modification(s), modified internucleoside linkage(s), and / or backbone modification(s), optionally the scaffold sequence comprises (i) a 2’ O-methyl modification at one or more of nucleotide positions 1, 3-9, 12-27, 29-32, 36-38, 40, 44-45, 49-50, and 52-53, and / or (ii) a 2’-fluoro modification at one or more of nucleotide positions 1-4, 6, 8, 10-11, 20, 22- 32, 35-38, 40-46, and 49-51; and / or the scaffold sequence comprises any one of the sequences of SEQ ID NOs: 894- 898 and 964-1014, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 894-898 and 964-1014.
52. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 1-51, wherein the scaffold sequence of the tagRNA, the scaffold sequence of the egRNA, or both, comprises the sequence of any one of SEQ ID NOs: 1210-1266 or a sequence that exhibits at least about 85% identity to any one of SEQ ID NOs: 1210-1266.
53. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 3-52, wherein: the DNA polymerase domain comprises a reverse transcriptase, optionally a reverse transcriptase comprising one or more mutations, wherein at least one of the one or more mutations is at an amino acid position functionally equivalent to V101, N200, A208, G248, P330, L435, K445, and / or A623 relative to SEQ ID NO: 8.
54. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 3-52, wherein: the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 255, optionally the reverse transcriptase comprises one or more mutations, optionally at least one of the one or more mutations is at amino acid position V98, N197, A205, S245, P327, V432, R442, and / or A623 relative to SEQ ID NO: 255, further optionally the reverse transcriptase comprises one or more transition mutations selected from the group consisting of V98R, N197C, N197D, N197G, A205T, S245C, P327E, P327Q, V432K, R442T, and A623F relative to SEQ ID NO:
255.
55. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 3-52, wherein: the DNA polymerase domain comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 260, optionally the reverse transcriptase comprises one or more mutations, optionally at least one of the one or more mutations is at amino acid position V106, N204, E212, G252, P334, L440, E450, and / or A629 relative to SEQ ID NO: 260, further optionally the reverse transcriptase comprises one or more transition mutations selected from the group consisting of V106R, N204C, N204D, N204G, E212T, G252C, P334E, P334Q, L440K, E450T, and A629F relative to SEQ ID NO:
260.
56. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 3-52, wherein the DNA polymerase domain: comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 255, and wherein the reverse transcriptase comprises an N197C and / or V98R transition mutation relative to SEQ ID NO: 255; or comprises a reverse transcriptase comprising an amino acid sequence that is at least 80% identical to SEQ ID NO: 260, and wherein the reverse transcriptase comprises an N204C and / or V106R transition mutation relative to SEQ ID NO:
260.
57. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 3-56, wherein the RT editor further comprises an accessory domain, optionally the accessory domain comprises a single-strand binding (SSB) protein domain or a stabilon, further optionally the SSB protein domain is derived from RecA protein, Sso7d protein, or Sto7d protein.
58. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of claim 57, wherein the accessory domain is situated at the N-terminus, the C-terminus, or at an internal location of the RT editor.
59. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 57-58, wherein the RT editor comprising an accessory domain comprises an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1276-1298.
60. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 57-58, wherein the SSB protein domain is derived from Sso7d and comprises one or more mutations at an amino acid position functionally equivalent to K12 and / or E35 of a wild type Sso7d amino acid sequence, optionally the one or more mutations comprise K12L and / or E35L relative to a wild type Sso7d sequence, further optionally, the SSB protein domain comprises an amino acid sequence that is at least 85% identical to SEQ ID NO: 1205.
61. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of any one of claims 57-60, wherein the RT editor comprising the SSB protein domain comprises from N-terminus to C-terminus: [N-terminus of nCas9]-[linker]-[SSB protein domain]-[linker]- [RT]-[linker]-[C-terminus of nCas9], optionally the N-terminus of nCas9 comprises an amino acid sequence functionally equivalent to amino acids 1-1247 of SEQ ID NO: 6 and / or the C-terminus of nCas9 comprises an amino acid sequence functionally equivalent to amino acids 1248-1367 of SEQ ID NO:
6.
62. The RT editor, the tagRNA, the RT editing system, or the RT editing complex of claim 61, wherein the RT editor comprising the SSB protein domain comprises an amino acid sequence that is at least 85% identical to any one of SEQ ID NOs: 1207-1209.
63. A ribonucleoprotein (RNP) complex comprising the RT editing complex of any one of claims 32-62, or a component thereof.
64. A lipid nanoparticle (LNP) comprising the RT editing system of any one of claims 26-31 and 35-62, or a component thereof.
65. The LNP of claim 64, comprising the tagRNA and the nucleic acid encoding the RT editor.
66. The LNP of claim 65, wherein the nucleic acid encoding the RT editor is mRNA.
67. The LNP of any one of claims 64-66, further comprising the egRNA.
68. A polynucleotide encoding the RT editor of claim 1, the tagRNA of any one of claims 4-25 and 35-62, the RT editing system of any one of claims 26-31 and 35-56, or the RT editing complex of any one of claims 32-62.
69. The polynucleotide of claim 68, wherein the polynucleotide is an mRNA.
70. The polynucleotide of any one of claims 68-69, wherein the polynucleotide is operably linked to a regulatory element, optionally the regulatory element is an inducible regulatory element.
71. A vector comprising the polynucleotide of any one of claims 68-70.
72. The vector of claim 71, wherein the vector is an AAV vector.
73. An isolated cell comprising the RT editor of claim 1, the tagRNA of any one of claims 4-25 and 35-62, the RT editing system of any one of claims 26-31 and 35-62, the RT editing complex of any one of claims 32-62, the RNP of claim 63, the LNP of any one of claims 64-67, the polynucleotide of any one of claims 68-70, or the vector of any one of claims 71-72.
74. The cell of claim 73, wherein the cell is a mammalian cell, optionally a human cell.
75. The cell of any one of claims 73-74, wherein the cell is a primary cell.
76. The cell of any one of claims 73-75, wherein the cell is a hepatocyte.
77. The cell of any one of claims 73-76, wherein the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally Wilson’s disease, alpha1 antitrypsin deficiency disease, phenylketonuria, or hyperphenylalaninemia, further optionally the subject is a human.
78. A pharmaceutical composition comprising: (i) the RT editor of claim 1, the tagRNA of any one of claims 4-25 and 35-62, the RT editing system of any one of claims 26-31 and 35-62, the RT editing complex of any one of claims 32-62, the RNP of claim 63, the LNP of any one of claims 64-67, the polynucleotide of any one of claims 68-70, the vector of any one of claims 71-72, or the cell of any one of claims 73-77; and (ii) a pharmaceutically acceptable carrier.
79. A method for editing a double stranded target DNA, the method comprising contacting the double stranded target DNA with (i) the tagRNA of any one of claims 4-25 and 35- 62 and a reverse transcriptase (RT) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain or (ii) the RT editing system of any one of claims 26-31 and 35-62, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, thereby editing the double stranded target DNA.
80. A method for editing a double stranded target DNA, comprising contacting the double stranded target DNA with the RT editing complex of any one of claims 32-62, wherein thetagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, thereby editing the double stranded target DNA.
81. The method of any one of claims 79-80, wherein the RT editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA.
82. The method of any one of claims 79-81, wherein the double stranded target DNA is in a cell.
83. The method of claim 82, wherein the cell is a mammalian cell, optionally a human cell.
84. The method of any one of claims 82-83, wherein the cell is a primary cell.
85. The method of any one of claims 82-84, wherein the cell is a hepatocyte.
86. The method of any one of claims 82-83, wherein the cell is a stem cell; optionally, an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell.
87. The method of any one of claims 82-86, wherein the cell is in a subject, optionally the subject is a human.
88. The method of any one of claims 82-86, wherein the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally Wilson’s disease, alpha1 antitrypsin deficiency disease, phenylketonuria, or hyperphenylalaninemia.
89. The method of claim 88, further comprising administering the cell to the subject after incorporation of the intended nucleotide edit.
90. A cell generated by the method of any one of claims 79-89.
91. A population of cells generated by the method of any one of claims 79-89.
92. A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject the RT editing system of any one of claims 26, 28-31 and 35-62, wherein the editing template comprises an intended nucleotide edit compared to a double stranded target DNA of the subject, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, and wherein incorporation of the intended nucleotide edit corrects a mutation in the double stranded target DNA associated with the disease or disorder, thereby treating or preventing the disease or disorder in the subject, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, acardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.
93. A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject the RT editing complex of any one of claims 33-62, the RNP of claim 63, the LNP of any one of claims 64-67, or the pharmaceutical composition of claim 78, wherein the editing template comprises an intended nucleotide edit compared to a double stranded target DNA of the subject, wherein the tagRNA directs the RT editor to incorporate the intended nucleotide edit in the double stranded target DNA, and wherein incorporation of the intended nucleotide edit corrects a mutation in the double stranded target DNA associated with the disease or disorder, thereby treating or preventing the disease or disorder in the subject, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.
94. A messenger RNA (mRNA) encoding a reverse transcriptase (RT) editor, optionally the RT editor of any one of claims 3-62.
95. The mRNA of claim 94, wherein the mRNA comprises one or more of a 5'-cap structure, a 5’-UTR, a 3’-UTR, and a nuclear localization sequence (NLS).
96. The mRNA of any one of claims 94-95, wherein the mRNA comprises (i) a 5′-cap, (ii) a 5’-untranslated region (UTR); (iii) an open reading frame (ORF) comprising a nucleotide sequence that encodes the RT editor; and (iv) a 3′ untranslated region (UTR).
97. The mRNA of any one of claims 94-96, wherein mRNA has a structure comprising or consisting of 5’ - [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]- [polyA sequence] - 3’.
98. The mRNA of any one of claims 94-97, wherein one or more nucleosides of the mRNA sequence are chemically modified.
99. The mRNA of any one of claims 94-98, wherein the mRNA comprises one or more additional segmented polyA sequences configured to increase mRNA stability and / or half-life.
100. The mRNA of any one of claims 94-99, wherein the mRNA comprises one or more viral element sequences, optionally 3’ of the 3’-UTR.
101. The mRNA of claim 100, wherein the viral element comprises a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) sequence, optionally comprising the sequence of SEQ ID NO: 189, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 189.
102. The mRNA of any one of claims 100-101, wherein the viral element comprises a eK5 sequence, optionally located 3’ of the 3’-UTR, optionally comprising the sequence of SEQ ID NO: 190, or a sequence that exhibits at least about 85% identity to SEQ ID NO:
190.
103. The mRNA of any one of claims 94-102, wherein the mRNA comprises one or more secondary structure motifs.
104. The mRNA of claim 103, wherein a secondary structure motif comprises a triple helix sequence, optionally a synthetic triple helix (STH) sequence.
105. The mRNA of claim 104, wherein the STH sequence comprises the sequence derived from a sequence element of a long non-coding RNA, optionally MALAT1.
106. The mRNA of any one of claims 104-105, wherein the STH sequence comprises any one of the sequences of SEQ ID NOs: 191-192 and 245-249, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 191-192 and 245-249.
107. The mRNA of any one of claims 94-106, wherein the mRNA comprises: a polyA sequence comprising any one of the sequences of SEQ ID NOs: 182-187, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 182-187; one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 188-190, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 188-190; one or more 3’ UTR structural motifs selected from the group comprising any one of the sequences of SEQ ID NOs: 191-192, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 191-192; a 5’UTR sequence comprising the sequence of any one of SEQ ID NOs: 193 and 910-933, or a sequence that exhibits at least about 85% identity to any one of SEQ ID NOs: 193 and 910-933; and / or an RT editor sequence comprising the sequence of SEQ ID NO: 194, or a sequence that exhibits at least about 85% identity to SEQ ID NO:
194.
108. The mRNA of any one of claims 94-107, wherein the mRNA comprises a 3’ UTR comprising any one of the sequences of SEQ ID NOs: 195-219, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 195-219.
109. The mRNA of any one of claims 94-108, wherein the mRNA comprises: one or more 5’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 940-950 and 1075, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 940-950 and 1075;one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 951-961 and 1076-1080, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 951-961 and 1076- 1080; and / or one or more additional elements selected from the group comprising any one of the sequences of SEQ ID NOs: 962-963, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 962-963, optionally the one or more additional elements are situated downstream of the 3’UTR and / or the polyA tail.
110. The mRNA of any one of claims 94-109, wherein the mRNA comprises one or more 5’-UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 1081-1082, and wherein the second codon of the mRNA is a “gcc”.
111. The mRNA of any one of claims 94-110, wherein the mRNA is codon optimized for expression in human cells, optionally, the mRNA comprises the sequence of any one of SEQ ID NOs: 1086-1088.
112. A ribonucleoprotein (RNP) complex, a lipid nanoparticle (LNP), a vector, an isolated cell, or a pharmaceutical composition comprising the mRNA of any one of claims 94- 111.
113. A polynucleotide encoding the mRNA of any one of claims 94-111.
Citation Information
Patent Citations
Serpina-modulating compositions and methods
WO2023039447A2
CFTR-modulating compositions and methods
WO2023108153A2
Methods and compositions for editing nucleotide sequences
WO2023192655A2