Synthetic nucleotide template polymerase (syntase) editing
Patent Information
- Application Number
- PCT/IB2026/051851
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-10-08
- Filing Date
- 2026-02-26
- Publication Date
- 2026-09-03
Smart Images

Figure IB2026051851_03092026_PF_FP_ABST
Abstract
Description
80EM-800004-WO / CT244-PCT1 PATENT SYNTHETIC NUCLEOTIDE TEMPLATE POLYMERASE (SYNTASE) EDITING RELATED APPLICATIONS
[0001] The present application claims priority to U. S. Provisional Application No.63 / 763,742, filed February 26, 2025; U. S. Provisional Application No. 63 / 846,811, filed July 18, 2025; U. S. Provisional Application No. 63 / 891,211, filed September 30, 2025; and U. S. Provisional Application No. 63 / 895,913, filed October 8, 2025. The entire contents of these applications are hereby expressly incorporated by reference in their entireties.REFERENCE TO SEQUENCE LISTING
[0002] The present application is being filed along with a Sequence Listing in electronic format. The Sequence Listing is provided as a file entitled 80EM-800004-WO_SequenceListing, created February 21, 2026, which is 3,557,335 bytes in size. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.BACKGROUNDField
[0003] The present disclosure relates generally to the field of gene editing.Description of the Related Art
[0004] Many diseases and disorders have a genetic component, including those that involve pathogenic single nucleotide mutations. There is a need for nucleic acid editing compositions and methods with greater editing efficiency and specificity. Existing gene editing approaches wherein components are engineered to be employed and / or expressed within target cells suffer from insufficient durability and / or levels of expression, or inadequate efficacy in their action, to be therapeutically effective; or lack specificity in directing a precise editing outcome. There is a need for compositions and methods with improved durability, strength, mechanistic efficacy, and duration of gene editing component expression and employment for in vivo applications.SUMMARY
[0005] Disclosed herein include engineered polymerases. In some embodiments, the engineered polymerase comprises one or more amino acid substitutions as compared to a parent polymerase comprising a sequence selected from SEQ ID NOs: 232-234 and 579-619. The engineered polymerase can exhibit improved activity, reduced stalling, reduced premature termination, and / or improved fidelity when employing a chemically modified template (e.g., anediting template of a tagRNA). The engineered polymerase can be capable of reading through partially or fully chemically modified templates. The chemically modified template can comprise O-methyl RNAbase(s) (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ O-methyl RNA base(s)), 2’ Fluoro RNA base(s) (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ Fluoro RNA base(s)), and / or phosphorothioate linkage(s). In some embodiments, the parent polymerase comprises the sequence of SEQ ID NO: 234 or SEQ ID NOs: 604-605. In some embodiments, each of the one or more amino acid substitutions is at an amino acid position functionally equivalent to S67, E69, Q84, V101, K103, D114, I125, T128, C157, K193, N200, V223, R278, P297, E302, K306, E372, K397, T420, L435, or G490 of SEQ ID NO: 234. In some embodiments, each of the one or more amino acid substitutions from the group comprising S67T, E69V, Q84G, V101R, K103R, D114N, I125V, T128N, C157Q, K193R, D / N200C, V223Y, V223M, R278L, P297Q, E302T, K306Q, E372V, K397R, T420E, L435K, or G490N of SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, and a V223 Y amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, a V223Y amino acid substitution, and a L435K amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, a V223 Y amino acid substitution, and an E302T amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, an N200C amino acid substitution, and an E302T amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, and a V223M amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C154Q amino acid substitution, an 1222 V amino acid substitution, a K100R substitution, or any combination thereof, relative to any one of SEQ ID NOs: 604-605.
[0006] Disclosed herein include polynucleotides encoding an engineered polymerase provided herein. Disclosed herein include Synthetic Nucleotide Template Polymerase (SyNTase) editors. In some embodiments, the SyNTase editor comprises: a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein, wherein the DNA polymerase domain comprises an engineered polymerase disclosedherein. The SyNTase editor can exhibit improved activity, reduced stalling, reduced premature termination, and / or improved fidelity when employing a chemically modified template (e.g., an editing template of a tagRNA). The SyNTase editor can be capable of reading through partially or fully chemically modified templates. The chemically modified template can comprise O-methyl RNA base(s) (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ O-methyl RNA base(s)), 2’ Fluoro RNA base(s) (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ Fluoro RNA base(s)), and / or phosphorothioate linkage(s). Disclosed herein include a SynNTase editor system, wherein the editing system can read synthetic editing templates or chemically modified templates. In some embodiments, the SyNTase editor is a reverse transcriptase (RT) editor.
[0007] Disclosed herein include Synthetic Nucleotide Template Polymerase (SyNTase) editing systems. In some embodiments, the SyNTase editing system comprises: a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor, wherein the DNA polymerase domain comprises an engineered polymerase disclosed herein. In some embodiments, the editing template is chemically modified. In some embodiments, the editing template comprises RNA or DNA nucleotides.
[0008] Conventional RT Editors such as Prime Editors and DNA-Polymerase Editors (DPE), are only capable of utilizing natural nucleic acid templates, such as, RNA and DNA as editing templates. Disclosed herein are SyNTase editors that contain engineered polymerases capable of reading through synthetic editing templates in addition to RNA and DNA. The synthetic editing templates consist of chemical modifications such as 2’-fluoro modified RNAs, 2’ O methyl modified RNAs, and combinations of those modifications.
[0009] In some embodiments, the SyNTase editing system is an RT editing system. In some embodiments, the tagRNA comprises a flap binding sequence at least partially complementary to the spacer. In some embodiments, the scaffold sequence is between the spacer and the editing template. In some embodiments, the tagRNA comprises from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the flap binding sequence. In some embodiments, the spacer, the scaffold sequence, the editing template, and the flap binding sequence form a contiguous sequence in a single molecule. In some embodiments, the editing template comprises an intended nucleotide edit compared to the double stranded target DNA. In some embodiments,the tagRNA guides the SyNTase editor to incorporate the intended nucleotide edit into the double stranded target DNA when the tagRNA is contacted with the double stranded target DNA. In some embodiments, the SyNTase editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA. In some embodiments, the search target sequence is complementary to a protospacer sequence in the double stranded target DNA, and wherein the protospacer sequence is adjacent to a protospacer adjacent motif (PAM) in the double stranded target DNA. In some embodiments, the tagRNA results in incorporation of a nucleotide edit in the PAM when contacted with the double stranded target DNA. In some embodiments, the spacer of the tagRNA is from 16 to 25 nucleotides in length, optionally 20 nucleotides in length or 21-23 nucleotides in length. In some embodiments, the flap binding sequence is about 2 to 20 nucleotides in length, optionally about 8 to 16 nucleotides in length or 6 nucleotides in length. In some embodiments, the editing template is about 4 to 30 nucleotides in length, optionally about 10 to 30 nucleotides in length, further optionally 6 to 9 nucleotides in length.
[0010] In some embodiments, the tagRNA results in incorporation of the intended nucleotide edit about 0 to 30 base pairs downstream of the nickase cleavage site. In some embodiments, the intended nucleotide edit comprises a single nucleotide substitution compared to the region corresponding to the editing target in the double stranded target DNA. In some embodiments, the intended nucleotide edit comprises an insertion compared to the region corresponding to the editing target in the double stranded target DNA, optionally an insertion of a nucleotide sequence at least 50, at least 45, at least 40, at least 35, at least 30, at least 25, at least 20, at least 15, at least 10, or at least 5, or at least 1 nucleotides in length. In some embodiments, the intended nucleotide edit comprises a deletion compared to the region corresponding to the editing target in the double stranded target DNA. In some embodiments, the editing template comprises one or more silent nucleotide edits compared to the region corresponding to the editing target in the double stranded target DNA, optionally said silent nucleotide edits do not alter the amino acid sequence of the protein encoded by the double stranded target DNA, further optionally said one or more silent nucleotide edits comprise a substitution of 2 to 5 contiguous nucleotides. In some embodiments, the editing template is chemically modified. In some embodiments, the editing template comprises a wild type double stranded target DNA sequence. In some embodiments, the tagRNA results in correction of a mutation when contacted with the double stranded target DNA. The SyNTase editing system of can comprise an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises a egRNA spacer that is complementary to a second search target sequence in the double stranded target DNA,optionally the egRNA comprises a scaffold sequence.
[0011] In some embodiments, the DNA endonuclease domain is a CRISPR associated (Cas) protein domain, optionally the Cas protein domain is Cas9. In some embodiments, the Cas protein domain has nickase activity. In some embodiments, the Cas9 comprises a mutation in an HNH domain. In some embodiments, the Cas9 comprises an H840A mutation in the HNH domain. In some embodiments, the Cas protein domain is a Casl2b. In some embodiments, the Cas protein domain is a Casl2a, Casl2b, Casl2c, Casl2d, Casl2e, Casl4a, Casl4b, Casl4c, Casl4d, Casl4e, Casl4f, Casl4g, Casl4h, Casl4u, Cascp, SschlCas9, SrolCas9, Sha4Cas9, SsuCas9, iSpyMacCas9, Ssi5Cas9, Ssi8Cas9, Ssci4Cas9, ShylCas9, Sag3Cas9, SlutrlCas9, Ssch3Cas9, SpRYCas9, SpRYcCas9, Sma2Cas9, SsaCas9, EvoCjCas9, iSpyMac, SsaCas9_AR12, or SveCas9. In some embodiments, the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein.
[0012] In some embodiments, the scaffold sequence of the tagRNA comprises the sequence of any one of SEQ ID NOs: 640-685, optionally the scaffold sequence of the tagRNA comprises the sequence of any one of SEQ ID NOs: 640, 651, and 667. In some embodiments, the editing template comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ Fluoro RNA base(s). In some embodiments, all of the nucleotides of the editing template are chemically modified. In some embodiments, select nucleotides of the editing template are chemically modified. In some embodiments, none of the nucleotides of the editing template are chemically modified. In some embodiments, the chemical modification is a 2’ O-methyl RNA or a 2’ Fluoro RNA base(s). In some embodiments, about 25%, about 50%, about 75%, or about 100% of the nucleotides of the editing template are chemically modified. In some embodiments, all of the nucleotides of the flap binding sequence are chemically modified. In some embodiments, select nucleotides of the flap binding sequence are chemically modified. In some embodiments, none of the nucleotides of the flap binding sequence are chemically modified. In some embodiments, the chemical modification is a 2’ O-methyl RNA or a 2’ Fluoro RNA base(s). In some embodiments, about 25%, about 50%, about 75%, or about 100% of the nucleotides of the flap binding sequence are chemically modified. In some embodiments, the first five nucleotides of the editing template each comprise a 2’ Fluoro RNA base, and the last 3 nucleotides of the flap binding sequence each comprise a 2’ O-methyl RNA and a phosphorothioate linkage, from 5’ to 3’. In some embodiments, the tagRNA comprises the sequence of any one of SEQ ID NOs: 686-720, optionally the tagRNA comprises the sequence of SEQ ID NO: 712. In some embodiments, the SyNTase editor comprises an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 238-249, 266-273, 282-315, 350-371, 423-471, and 628-633. In some embodiments, the tagRNA comprises a sequenceat least 80% identical to the sequence of SEQ ID NO: 712; and the SyNTase editor comprises an amino acid sequence at least 80% identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723. The tagRNA can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of any one of SEQ ID NOs: 686-720. The SyNTase editor can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to any one of SEQ ID NOs: 238-249, 266-273, 282-315, 350-371, 423-471, and 628-633.
[0013] In some embodiments, the SyNTase editing system comprises: a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor, wherein the tagRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 712; and the SyNTase editor comprises an amino acid sequence at least 80% identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723. The tagRNA can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 712. The SyNTase editor can comprise an amino acid sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 726. The nucleic acid encoding the SyNTase editor can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 723. The SyNTase editing system can comprise: a tagRNA comprising a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 712; and a SyNTase editor comprising an amino acid sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 726. The SyNTase editing system can comprise: a tagRNA comprising a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 712; and a nucleic acid encoding a SyNTase editor comprising a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 723.
[0014] Disclosed herein include ribonucleoprotein (RNP) complexes comprising aSyNTase editing system provided herein, or a component thereof. Disclosed herein include lipid nanoparticles (LNPs) comprising a SyNTase editing system provided herein, or a component thereof. The LNP can comprise the tagRNA and the nucleic acid encoding the SyNTase editor. In some embodiments, the nucleic acid encoding the SyNTase editor is mRNA. In some embodiments, the LNP further comprises the egRNA.
[0015] Disclosed herein include polynucleotides encoding the SyNTase editor or SyNTase editing system provided herein. In some embodiments, the polynucleotide is an mRNA. In some embodiments, the mRNA comprises one or more of a 5 ’-cap structure, a 5 ’-UTR, a 3’-UTR, and a nuclear localization sequence (NLS). In some embodiments, the mRNA comprises (i) a 5’-cap, (ii) a 5 ’-untranslated region (UTR); (iii) an open reading frame (ORF) comprising a nucleotide sequence that encodes the SyNTase editor; and (iv) a 3’ untranslated region (UTR). In some embodiments, the mRNA has a structure comprising or consisting of 5’ - [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]-[polyA sequence] - 3’. In some embodiments, one or more nucleotides of the mRNA sequence are chemically modified. In some embodiments, the mRNA comprises one or more additional segmented polyA sequences configured to increase mRNA stability and / or half-life. In some embodiments, the mRNA comprises one or more viral element sequences, optionally 3’ of the 3 ’-UTR. In some embodiments, the viral element comprises a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) sequence, optionally comprising the sequence of SEQ ID NO: 141, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 141. In some embodiments, the viral element comprises a eK5 sequence, optionally located 3’ of the 3 ’-UTR, optionally comprising the sequence of SEQ ID NO: 142, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 142. In some embodiments, the mRNA comprises one or more secondary structure motifs.
[0016] In some embodiments, a secondary structure motif comprises a triple helix sequence, optionally a synthetic triple helix (STH) sequence. In some embodiments, the STH sequence comprises the sequence derived from a sequence element of a long non-coding RNA, optionally MALAT1. In some embodiments, the STH sequence comprises any one of the sequences of SEQ ID NOs: 143-144 and 172-173, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 143-144 and 172-173. In some embodiments, the mRNA comprises: a polyA sequence comprising any one of the sequences of SEQ ID NOs: 134-139, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 134-139; one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 140-142, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 140-142; one or more 3’ UTRstructural motifs selected from the group comprising any one of the sequences of SEQ ID NOs: 143-144, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 143-144; and / or a 5’UTR sequence comprising the sequence of any one of SEQ ID NOs: 145 and 207-230, or a sequence that exhibits at least about 85% identity to any one of SEQ ID NOs: 145 and 207-230. In some embodiments, the mRNA comprises a 3’ UTR comprising any one of the sequences of SEQ ID NOs: 147-171, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 147-171. In some embodiments, the mRNA comprises: one or more 5’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 174-185, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 174-185; one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 186-200, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 186-200; and / or one or more additional elements selected from the group comprising any one of the sequences of SEQ ID NOs: 201-202, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 201-202, optionally the one or more additional elements are situated downstream of the 3 ’UTR and / or the polyA tail. In some embodiments, the mRNA comprises one or more 5 ’-UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 203-206, and wherein the second codon of the mRNA is a “gcc”. In some embodiments, the mRNA is codon optimized for expression in human cells. In some embodiments, the mRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723. In some embodiments, the polynucleotide is operably linked to a regulatory element, optionally the regulatory element is an inducible regulatory element.
[0017] Disclosed herein include vectors comprising a polynucleotide provided herein. In some embodiments, the vector is an AAV vector. Disclosed herein include isolated cells. In some embodiments, the isolated cell comprises a SyNTase editor or a SyNTase editing system disclosed herein, an RNP disclosed herein, a LNP disclosed herein, a polynucleotide disclosed herein, or a vector disclosed herein. In some embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is from a subject having a disease or disorder, optionally P -thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, Wilson’s disease, phenylketonuria, or hyperphenylalaninemia, further optionally the subject is a human.
[0018] Disclosed herein include pharmaceutical compositions. In some embodiments, the pharmaceutical composition comprises: (i) a SyNTase editor or a SyNTase editing system disclosed herein, an RNP disclosed herein, a LNP disclosed herein, a polynucleotide disclosed herein, a vector disclosed herein, or a cell disclosed herein; and (ii) a pharmaceutically acceptablecarrier. In some embodiments, the pharmaceutical composition comprises: a SyNTase editor or a SyNTase editing system disclosed herein, an RNP disclosed herein, a LNP disclosed herein, a polynucleotide disclosed herein, a vector disclosed herein, or a cell disclosed herein. Disclosed herein include compositions (e.g., pharmaceutical composition) comprising means provided herein for gene editing and replacement. Disclosed herein include methods for editing a double stranded target DNA. In some embodiments, the method comprises contacting the double stranded target DNA with a SyNTase editor or a SyNTase editing system disclosed herein, thereby editing the double stranded target DNA. In some embodiments the method is performed in vitro or ex vivo. In some embodiments the method is performed in vivo. In some embodiments, the double stranded target DNA is in a cell. In some embodiments, the cell is a mammalian cell, optionally a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the cell is a hepatocyte. In some embodiments, the cell is a stem cell; optionally, an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell. In some embodiments, the cell is in a subject, optionally the subject is a human. In some embodiments, the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally P-thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, Wilson’s disease, phenylketonuria, or hyperphenylalaninemia. The method can comprise administering the cell to the subject after incorporation of the intended nucleotide edit. Disclosed herein include cell generated by a method provided herein. Disclosed herein include populations of cells generated by a method provided herein. Disclosed herein include methods for treating or preventing a disease or disorder in a subject in need thereof. In some embodiments, the method comprises administering to the subject a SyNTase editor or a SyNTase editing system disclosed herein, an RNP disclosed herein, a LNP disclosed herein, a pharmaceutical composition disclosed herein, a cell of disclosed herein, or a population of cells disclosed herein, thereby treating or preventing the disease or disorder in the subject. In some embodiments, the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof. In some embodiments, the disease or disorder comprises alpha-1 antitrypsin deficiency (AATD).
[0019] Provided herein are methods for treating or preventing or reversing AATD in a subject in need thereof, including correcting underlying genetic mutations of AATD to wildtype in cells of a patient. In some embodiments, the method comprises administering to the subjecta SyNTase editor or SyNTase editing system of the disclosure, an RNP of the disclosure, an LNP of the disclosure, a pharmaceutical composition of the disclosure, a cell of the disclosure, or a population of cells of the disclosure, thereby treating or preventing or reversing the AATD in the subject.
[0020] In some embodiments, the tagRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 712. In some embodiments, the SyNTase editor comprises an amino acid sequence at least 80% identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723. In some embodiments, the method is performed in vivo, in vitro, or ex vivo. The method can comprise administering to the subject a single dose of about 0.05 to 0.5 mg / kg of total nucleic acids comprising (a) the tagRNA and (b) the nucleic acid encoding the SynTase editor. In some embodiments, the method provides a SERPINA1 gene editing efficiency of at least 20%. In some embodiments, the method provides a SERPINA1 mRNA editing efficiency of at least 80%. In some embodiments, the method provides a SERPINA1 mRNA editing efficiency of at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%. In some embodiments, a clinically relevant increase in M-AAT protein levels is detected in the serum of the subject following the administration. In some embodiments, an at least five-fold increase in total serum AAT protein levels is detected in the subject following the administration. In some embodiments, the ratio of M-AAT: Z-AAT in the serum of the subject is at least 99% following the administration. In some embodiments, an increase in total serum AAT protein levels is detected in the subject at least seven days after the administration, optionally at least 7 weeks, further optionally at least 9 weeks. In some embodiments, a linear correlation is observed between SERPINA1 gene and / or mRNA editing efficiency and the increase in AAT protein levels in the subject. In some embodiments, the M-AAT is capable of functional rescue in a human neutrophil elastase inhibition assay. The tagRNA can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 712. The SyNTase editor can comprise an amino acid sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 726. The nucleic acid encoding the SyNTase editor can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 723. The tagRNA can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 712, and the SyNTase editor can comprise an amino acid sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 726. The tagRNA can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence ofSEQ ID NO: 712, and the nucleic acid encoding a SyNTase editor can comprise a sequence at least 85%, more preferably 90%, and most preferably 95%, or even 99%, identical to the sequence of SEQ ID NO: 723.BRIEF DESCRIPTION OF THE DRAWINGS
[0021] FIG. 1 displays exemplary data showing the MMLV-RT Editor fails to read through fully chemically modified template RNA.
[0022] FIG. 2 displays exemplary data showing low editing efficiency of Evo Tfl Editor with fully chemically modified template RNA.
[0023] FIG. 3 A-FIG. 3B show exemplary schematics related to sequence and structure guided alignments to identify beneficial mutations for MMLV RT. SEQ ID NOs shown are the full-length sequences.
[0024] FIG. 4A-FIG. 4B display exemplary data and schematics showing only the C157Q mutation in the MMLV RT editor pEY037, but not KI 03R or II 25 V, enables read-through of the fully chemically modified template.
[0025] FIG. 5 A-FIG. 5B display exemplary schematics of how C157Q enables MMLV RT to accommodate modified RNA templates.
[0026] FIG. 6 displays an exemplary schematic related to structure-guided engineering, e.g., E302T and K397R form H-bond with DNA / RNA substrate.
[0027] FIG. 7 displays exemplary data related to additional beneficial mutations that slightly improve read-through of fully modified template.
[0028] FIG. 8 displays exemplary data related to a screen of MMLV RT editors with partially or fully modified tagRNA templates in Huh7 cells. Shown from left to right for each variant is: Ome-F1 full mod, Min1 full mod, Min4 Ome-1F-2, Alt2 Ome-1F-2.
[0029] FIG. 9 displays a graph of 2ndround SyNTase rational engineering results.
[0030] FIG. 10 shows a graph of results from screening SyNTase variants with unmodified RNA template in PHH.
[0031] FIG. 11 shows a graph of results from a screen of SyNTase variants with fully modified RNA template in PHH.
[0032] FIG. 12 depicts exemplary editing frequencies in Huh7 cells with fully modified tagRNA vs PHH with unmodified tagRNA.
[0033] FIG. 13 displays exemplary data showing further engineered Evo Tfl editors still exhibit low editing efficiency with fully chemically modified template RNA.
[0034] FIG. 14 shows exemplary results from a screen of templates partially or fully chemically modified tagRNA variants in Huh7 cells using EvoTfl RT editor and MMLV RTeditor. For each template, results with EvoTfl RT editor and MMLV RT editor are shown from left to right, respectively.
[0035] FIG. 15 depicts overall SyNTase (e.g., SPE) alignment with different RTs.
[0036] FIG. 16 shows alignments of SyNTases (e.g., SPEs) with different RTs at C157Q. SEQ ID NOs shown are the full-length sequences.
[0037] FIG. 17 shows alignments of SyNTases (e.g., SPEs) with different RTs at E302T. SEQ ID NOs shown are the full-length sequences.
[0038] FIG. 18 shows alignments of SyNTase (e.g., SPEs) with different RTs at K193R. SEQ ID NOs shown are the full-length sequences.
[0039] FIG. 19 shows alignments of SyNTase (e.g., SPEs) with different RTs at K397R. SEQ ID NOs shown are the full-length sequences.
[0040] FIG. 20 shows alignments of SyNTase (e.g., SPEs) with different RTs at L435K. SEQ ID NOs shown are the full-length sequences.
[0041] FIG. 21 shows alignments of SyNTase (e.g., SPEs) with different RTs at N200C. SEQ ID NOs shown are the full-length sequences.
[0042] FIG. 22 shows alignments of SyNTase (e.g., SPEs) with different RTs at T128N. SEQ ID NOs shown are the full-length sequences.
[0043] FIG. 23 shows alignments of SyNTase (e.g., SPEs) with different RTs at V233M. SEQ ID NOs shown are the full-length sequences.
[0044] FIG. 24 shows alignments of SyNTase (e.g., SPEs) with different RTs at V233Y. SEQ ID NOs shown are the full-length sequences.
[0045] FIG. 25 displays exemplary results from screening top 47 AATD mRNA with fully modified tagRNA in Huh7.
[0046] FIG. 26 shows graphs of data according to lead mRNA components from the screen in Huh7 with fully modified template. Nucleic acid dose is 0.5 ng of mRNA encoding an editor and 1 pmol tagRNA.
[0047] FIG. 27 displays exemplary results from screening top 48 AATD mRNA with standard tagRNA in Huh7. Nucleic acid dose is 0.1 ng of mRNA encoding an editor and 1 pmol tagRNA.
[0048] FIG. 28 displays exemplary results from screening top 48 AATD mRNA with standard tagRNA in Huh7. Nucleic acid dose is 0.05 ng of mRNA encoding an editor and 1 pmol tagRNA.
[0049] FIG. 29 shows exemplary results from screening top 48 AATD mRNA with standard tagRNA in PHH.
[0050] FIG. 30 displays exemplary data related to in vivo screening of mRNAs.
[0051] FIG. 31 displays exemplary results related to in vivo editing from a dose response curve.
[0052] FIG. 32 displays as a heatmap editing results using the indicated RT variant, editing template (tagRNA) and egRNA.
[0053] FIG. 33 shows the modifications tested for the editing template in FIG. 32 described above.
[0054] FIG. 34A-FIG. 34C display the chemical modifications (FIG. 34A) and nonlimiting exemplary data related to the chemically modified tagRNA screen with SynRT81 in Huh7 (FIG. 34B-C).
[0055] FIG. 35 displays non-limiting exemplary data showing in vivo comparison of chemical modifications identified an ET5F chemical modification pattern that improved editing.
[0056] FIG. 36A-FIG. 36E display non-limiting exemplary data related to scaffold optimization (FIG. 36 A) and non-limiting exemplary designs of scaffold secondary structures (FIG. 36B-E).
[0057] FIG. 37 displays non-limiting exemplary data related to an in vivo scaffold screen.
[0058] FIG. 38 displays non-limiting exemplary data related to a final tagRNA nomination in vivo screen.
[0059] FIG. 39 displays non-limiting exemplary data related to in vivo DNA and mRNA editing. For each tagRNA, DNA Editing and RNA editing are displayed from left to right.
[0060] FIGS. 40A-40C depict non-limiting exemplary editing component combinations and data. FIG. 40A-FIG. 40B display non-limiting exemplary combinations (FIG.40A) and related data of editor mRNA components in an in vitro screen of editing in Huh7 with fully modified tagRNA (FIG. 40B). FIG. 40C displays non-limiting exemplary data of editor mRNA in vitro screen of editing in PHH with standard tagRNA.
[0061] FIG. 41 displays non-limiting exemplary in vivo results of editor mRNA screening.
[0062] FIG. 42 displays non-limiting exemplary data from codon optimization of the editor mRNA in in vitro screens.
[0063] FIG. 43 displays non-limiting exemplary data showing codon optimization of the editor mRNA increases editing efficiency in vivo.
[0064] FIG. 44 shows non-limiting exemplary results of in vivo DNA editing by the codon optimized editor mRNA. For each plasmid, precise editing and then indel formation is shown from left to right.
[0065] FIG. 45 shows non-limiting exemplary results of in vivo RNA editing by the codon optimized editor mRNA. For each plasmid, precise editing and then indel formation is shown from left to right.
[0066] FIG. 46A-FIG. 46C show in vivo data of a SyNTase editor mRNA in a rat disease model. FIG. 46A shows generation of a custom rat model with the PiZ mutation that results in development of AATD. FIG. 46B displays data related to in vivo DNA editing of SERPINA1 gene in a rat disease model for AATD using the methods and compositions of the disclosure. FIG.46C displays data related to in vivo editing of SERPINA1 gene in rat using the methods and compositions of the disclosure as determined by RNA analysis.
[0067] FIG. 47A-FIG. 47B display non-limiting exemplary data related to in vivo DNA and mRNA editing of SERPINA1 gene in a rat disease model for AATD using the methods and compositions of the disclosure. Experiments were performed with SEQ ID NO: 528 tagRNA and pEY192 and pEY200.
[0068] FIG. 48A-FIG. 48D display non-limiting exemplary data related to AAT levels (ELISA) following in vivo editing (FIG. 47A-FIG. 47B). AAT levels are shown in FIG. 48A-FIG.48B. Correlation between edits and AAT levels is shown in FIG. 48C. Correlation between body weight (BW) and liver weight is shown in FIG. 48D.
[0069] FIG. 49A-FIG. 49C display non-limiting exemplary data related to in vivo DNA and mRNA editing of SERPINA1 gene in a rat disease model for AATD using the methods and compositions of the disclosure. Editing was performed with SEQ ID NO: 699 tagRNA and pEY190 mRNA. Homozygous and heterozygous rats were observed.
[0070] FIG. 50A-FIG. 50C display non-limiting exemplary data related to AAT levels (ELISA) following in vivo editing (FIG. 49A-FIG. 49C).
[0071] FIG. 51A-FIG. 5 IB display non-limiting exemplary data related to in vivo DNA and mRNA editing of SERPINA1 gene in a rat disease model for AATD using the methods and compositions of the disclosure. Experiments were performed with SEQ ID NO: 712 tagRNA and pAM320 mRNA. Homozygous rats are observed.
[0072] FIG. 52A-FIG. 52B display non-limiting exemplary data related to AAT levels (ELISA) and correlation following in vivo editing (FIG. 51A-FIG. 5 IB).
[0073] FIG. 53A-FIG. 53B display non-limiting exemplary data showing potent in-vivo editing of AATD-E342K in NSG-PiZ mice.
[0074] FIG. 54A-FIG. 54B display non-limiting exemplary data showing AATD correction with disclosed mRNA and tagRNA in PiZ+ / +rats.DETAILED DESCRIPTION
[0075] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.
[0076] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.Definitions
[0077] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure belongs. See, e.g. Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of the present disclosure, the following terms are defined below.
[0078] As used herein, the term “about” can mean plus or minus 5% of the provided value.
[0079] As used herein, the term “gene editing” (including genomic editing) is a type of genetic engineering in which nucleotide(s) / nucleic acid(s) is / are inserted, deleted, and / or substituted in a DNA sequence, such as in the genome of a targeted cell. Targeted gene editing enables insertion, deletion, and / or substitution at pre-selected sites in the genome of a targeted cell (e.g., in a targeted gene or targeted DNA sequence). When a sequence of an endogenous gene is edited, for example by deletion, insertion or substitution of nucleotide(s) / nucleic acid(s), the endogenous gene comprising the affected sequence can be knocked-out or knocked-down due to the sequence alteration. Therefore, targeted editing can be used to disrupt endogenous gene expression. In some embodiments, gene editing is precise editing, and comprises an intended nucleotide edit at an intended location (e.g., a single nucleotide substitution, such as an E342K correction). In some embodiments, precise gene editing involves the generation of a single nick or the generation of two or more nicks. Said nicks can be simultaneous or sequential, can occur on the same strand on or opposite strands, and can occur near each other (leading to a double-strand break) or separated from each such that a double-strand break does not occur. Precise editing does not comprise a double-strand break in some embodiments.
[0080] As used herein, the term “RNA-guided endonuclease” refers to a polypeptide capable of binding a RNA (e.g., a gRNA) to form a complex targeted to a specific DNA sequence (e.g., in a target DNA). A non-limiting example of RNA-guided endonuclease is a Cas polypeptide (e.g., a Cas endonuclease, such as a Cas9 endonuclease). In some embodiments, the RNA-guided endonuclease as described herein is targeted to a specific DNA sequence in a target DNA by an RNA molecule to which it is bound. The RNA molecule can include a sequence that is complementary to and capable of hybridizing with a target sequence within the target DNA, thus allowing for targeting of the bound polypeptide to a specific location within the target DNA.
[0081] As used herein, the term “invariable region” of a gRNA refers to the nucleotide sequence of the gRNA that associates with the RNA-guided endonuclease. In some embodiments, the gRNA comprises a crRNA and a transactivating crRNA (tracrRNA), wherein the crRNA and tracrRNA hybridize to each other to form a duplex. In some embodiments, the crRNA comprises 5’ to 3’: a spacer sequence and minimum CRISPR repeat sequence (also referred to as a “crRNA repeat sequence” herein); and the tracrRNA comprises a minimum tracrRNA sequence complementary to the minimum CRISPR repeat sequence (also referred to as a “tracrRNA antirepeat sequence” herein) and a 3’ tracrRNA sequence. In some embodiments, the invariable region of the gRNA refers to the portion of the crRNA that is the minimum CRISPR repeat sequence and the tracrRNA.
[0082] As used herein, the term “fusion protein” refers to a hybrid polypeptide which comprises protein domains from at least two different proteins. One protein may be located at the amino-terminal (N-terminal) portion of the fusion protein or at the carboxy -terminal (C-terminal) protein thus forming an “amino-terminal fusion protein” or a “carboxy-terminal fusion protein,” respectively. A protein may comprise different domains, for example, a nucleic acid binding domain (e.g., the gRNA binding domain of Cas9 that directs the binding of the protein to a target site) and a nucleic acid cleavage domain or a catalytic domain of a recombinase. In some embodiments, a protein comprises a proteinaceous part, e.g., an amino acid sequence constituting a nucleic acid binding domain, and an organic compound, e.g., a compound that can act as a nucleic acid cleavage agent. In some embodiments, a protein is in a complex with, or is in association with, a nucleic acid, e.g., RNA. Any of the proteins provided herein can be produced by any method known in the art. For example, the proteins provided herein may be produced via recombinant protein expression and purification, which is especially suited for fusion proteins comprising a peptide linker. Methods for recombinant protein expression and purification are well known, and include those described by Green and Sambrook, Molecular Cloning: A LaboratoryManual (4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N. Y. (2012)), the entire contents of which are incorporated herein by reference.
[0083] As used herein, the terms “DNA binding domain” and “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which a Cas protein is an example, may be used interchangeably herein, and can refer to proteins that use RNA: DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence. As used herein, the term “nucleic acid programmable DNA binding protein” or “napDNAbp,” of which Cas9 is an example, refer to proteins that use RNA: DNA hybridization to target and bind to specific sequences in a DNA molecule. Each napDNAbp is associated with at least one guide nucleic acid (e.g., guide RNA), which localizes the napDNAbp to a DNA sequence that comprises a DNA strand (i.e., a target strand) that is complementary to the guide nucleic acid, or a portion thereof (e.g., the protospacer of a guide RNA). In other words, the guide nucleic-acid “programs” the napDNAbp (e.g., Cas9 or equivalent) to localize and bind to a complementary sequence. DNA binding domains can refer to proteins that bind DNA in a sequence-specific manner, such as, for example, TALENs and Zinc finger proteins. In some embodiments, DNA binding domains can operate without a guide nucleic acid. In some embodiments, Cas9 comprises a napDNAbp and a DNA endonuclease domain. In some embodiments, a Cas9 polypeptide possess both napDNAbp and DNA endonuclease activities. The term “endonuclease” can refer to an enzyme that can cleave the phosphodiester bond within a polynucleotide chain. The polynucleotide may be double-stranded DNA (dsDNA), single-stranded DNA (ssDNA), RNA, double-stranded hybrids of DNA and RNA, and synthetic DNA (for example, containing bases other than A, C, G, and T). An endonuclease may cut a polynucleotide symmetrically, leaving “blunt” ends, or in positions that are not directly opposing, creating overhangs, which may be referred to as “sticky ends.” Examples of endonucleases include Cas, Argonaute protein (AGO), TAL Effector Nuclease” (TALEN), and meganucleases such as MegaTAL, or a fusion protein comprising a domain of an endonuclease, for example, Cas9, Ago, TALEN, or MegaTAL, or one or more portions thereof. An endonuclease can be an RNA-guided endonuclease.
[0084] As used herein, the term “guide RNA” or “gRNA” can refer to a site-specific targeting RNA that can bind an RNA-guided endonuclease to form a complex, and direct the activities of the bound RNA-guided endonuclease (such as a Cas endonuclease) to a specific sequence within a target nucleic acid (e.g., a specific gene or region within a gene). The guideRNA can include one or more RNA molecules. In some embodiments, the gRNA is a template armed gRNA (tagRNA). In some embodiments, the gRNA is an enhancer gRNA (egRNA).
[0085] As used herein, the term “target DNA” can refer to the specific region on a double-stranded DNA within a subject’s genome intended for editing by a gene editing system. In certain embodiments, the gene editing system is a SyNTase editing system.
[0086] As used herein, the term “search target sequence” can refer to the sequence on target strand that is complementary or substantially complementary to the spacer sequence of the tagRNA.
[0087] As used herein, the term “editing target sequence” can refer to the sequence being edited on the non-target strand.
[0088] As used herein, the term “silent mutation” can refer to a nucleotide change or nucleotide changes (for e.g., substitution or substitutions) in a DNA sequence that does not result in a change to the amino acid sequence of the protein the DNA sequence encodes.
[0089] As used herein, the term “protospacer” refers to the sequence in DNA adjacent to the PAM (protospacer adjacent motif) sequence. In some embodiments, the protospacer is 20 nucleotides long. The protospacer shares the same sequence as the spacer sequence of the guide RNA. The guide RNA anneals to the complement of the protospacer sequence on the target DNA (specifically, one strand thereof, i.e., the “target strand” versus the “non-target strand” of the target DNA sequence). In order for Cas9 to function, it also requires a specific protospacer adjacent motif (PAM) that varies depending on the bacterial species of the Cas9 gene. The skilled person will appreciate that the literature in the state of the art sometimes refers to the “protospacer” as the target-specific guide sequence on the guide RNA itself, rather than referring to it as a “spacer.” Thus, in some cases, the term “protospacer” as used herein may be used interchangeably with the term “spacer.” The context of the description surrounding the appearance of either “protospacer” or “spacer” will help inform the reader as to whether the term is in reference to the gRNA or the DNA target.
[0090] As used herein, the terms “upstream” and “downstream” define relevant positions at least two regions or sequences in a nucleic acid molecule orientated in a 5'-to-3' direction. For example, a first sequence is upstream of a second sequence in a DNA molecule where the first sequence is positioned 5’ to the second sequence. Accordingly, the second sequence is downstream of the first sequence.
[0091] As used herein, the term “protospacer adjacent sequence” or “PAM” refers to a DNA sequence that is an important targeting component of a Cas9 nuclease. The PAM sequence can be on either strand, and is downstream in the 5' to 3' direction of the Cas9 cut site. Different PAM sequences can be associated with different Cas9 nucleases or equivalent proteins fromdiff erent organisms.
[0092] As used herein, the term “spacer sequence” in connection with a guide RNA or a tagRNA refers to the portion of the guide RNA or tagRNA which contains a nucleotide sequence that shares the same sequence as the protospacer sequence in the target DNA sequence. The spacer sequence anneals to the complement of the protospacer sequence to form an ssRNA / ssDNA hybrid structure at the target site and a corresponding R loop ssDNA structure of the endogenous DNA strand.
[0093] As used herein, a “secondary structure” of a nucleic acid molecule (e.g., an RNA fragment, or a gRNA) refers to the base pairing interactions within the nucleic acid molecule.
[0094] As used herein, the term “Cas endonuclease” or “Cas nuclease” refers to an RNA-guided DNA endonuclease associated with and / or derived from the CRISPR adaptive immunity system. The term “nickase” refers to a Cas9 or other endonuclease with one of the two nuclease domains inactivated. This enzyme is capable of cleaving only one strand of a target DNA.
[0095] Unless otherwise indicated “nuclease” and “endonuclease” are used interchangeably herein to refer to an enzyme which possesses endonucleolytic catalytic activity for polynucleotide cleavage.
[0096] The terms “polynucleotide” and “nucleic acid” are used interchangeably herein and refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. A polynucleotide can be single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids / triple helices, or a polymer including purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. Any of the RNA sequences disclosed herein may also be DNA (either single-stranded or double-stranded), e.g., wherein “U” is converted to “T.” Any of the DNA sequences disclosed herein may also be RNA, e.g., wherein “T” is converted to “U ”
[0097] A “functional variant” or “functional mutant”, as used herein, refers to any variant or mutant of a reference protein (e.g., a wild-type protein) that encompasses one or more alterations to the amino acid sequence of the reference protein while retaining one or more of the functions, e.g., catalytic or binding functions. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions, insertions or deletions, or any combination thereof. In some embodiments, the one or more alterations to the amino acid sequence comprises amino acid substitutions. For example, a functional variant of a reverse transcriptase may comprise one or more amino acid substitutions compared to the amino acid sequence of a wild-type reverse transcriptase but retains the ability under at least one set of conditions to catalyze the polymerization of a polynucleotide. When the reference protein is a fusion of multiple functional domains, a functional variant thereof may retain one or more of thefunctions of at least one of the functional domains. For example, in some embodiments, a functional fragment of a Cas9 may comprise one or more amino acid substitutions in a nuclease domain, e.g., an H840A amino acid substitution, compared to the amino acid sequence of a wild type Cas9, but retains the DNA binding ability and lacks the nuclease activity partially or completely.
[0098] As used herein, the term “binding” refers to a non-covalent interaction between macromolecules (e.g., between a protein and a nucleic acid). While in a state of non-covalent interaction, the macromolecules are said to be “associated” or “interacting” or “binding” (e.g., when a molecule X is said to interact with a molecule Y, it means that the molecule X binds to molecule Y in a non-covalent manner). Binding interactions can be characterized by a dissociation constant (Kd), for example a Kd of, or a Kd less than, 10'6M, 10'7M, 10'8M, 10'9M, IO'10M, 10"11M, 10'12M, 10'13M, 10'14M, 10'15M, or a number or a range between any two of these values. Kd can be dependent on environmental conditions, e.g., pH and temperature. “Affinity” refers to the strength of binding, and increased binding affinity is correlated with a lower Kd.
[0099] As used herein, the term “hybridizing” or “hybridize” refers to the pairing of substantially complementary or complementary nucleic acid sequences within two different molecules. Pairing can be achieved by any process in which a nucleic acid sequence joins with a substantially or fully complementary sequence through base pairing to form a hybridization complex. “Hybridizing” or “hybridize” can comprise denaturing the molecules to disrupt the intramolecular structure(s) e.g., secondary structure(s)) in the molecule. In some embodiments, denaturing the molecules comprises heating a solution comprising the molecules to a temperature sufficient to disrupt the intramolecular structures of the molecules. In some instances, denaturing the molecules comprises adjusting the pH of a solution comprising the molecules to a pH sufficient to disrupt the intramolecular structures of the molecules. For purposes of hybridization, two nucleic acid sequences or segments of sequences are “substantially complementary” if at least 80% of their individual bases are complementary to one another. The complementary portion of each sequence can be referred to herein as a “segment”, and the segments are substantially complementary if they have 80% or greater identity.
[0100] The terms “complementarity” and “complementary” mean that a nucleic acid can form hydrogen bond(s) with another nucleic acid based on traditional Watson-Crick base paring rule, that is, adenine (A) pairs with thymine (T, or uracil (U) in RNA) and guanine (G) pairs with cytosine (C). Complementarity can be perfect (e.g., complete complementarity) or imperfect (e.g. partial complementarity). Perfect or complete complementarity indicates that each and every nucleic acid base of one strand is capable of forming hydrogen bonds according to Watson-Crick canonical base pairing with a corresponding base in another, antiparallel nucleicacid sequence. Partial complementarity indicates that only a percentage of the contiguous residues of a nucleic acid sequence can form Watson-Crick base pairing with the same number of contiguous residues in another, antiparallel nucleic acid sequence. In some embodiments, the complementarity can be at least 70%, 80%, 90%, 100% or a number or a range between any two of these values. In some embodiments, the complementarity is perfect, z.e., 100%. For example, the complementary candidate sequence segment is perfectly complementary to the candidate sequence segment, whose sequence can be deduced from the candidate sequence segment using the Watson-Crick base pairing rules.
[0101] As used herein, the terms “nucleic acid" and “polynucleotide” are interchangeable and refer to any nucleic acid, whether composed of phosphodiester linkages or modified linkages such as phosphotriester, phosphoramidate, siloxane, carbonate, carboxymethylester, acetamidate, carbamate, thioether, bridged phosphoramidate, bridged methylene phosphonate, bridged phosphoramidate, bridged phosphoramidate, bridged methylene phosphonate, phosphorothioate, methylphosphonate, phosphorodithioate, bridged phosphorothioate or sultone linkages, and combinations of such linkages. The terms “nucleic acid” and “polynucleotide” also specifically include nucleic acids composed of bases other than the five biologically occurring bases (adenine, guanine, thymine, cytosine and uracil).
[0102] The terms “DNA editing efficiency,” or “editing efficiency” may be used interchangeably herein and can refer to the number or proportion of intended target sequences that are edited. In some embodiments, the efficiency can be reported as % indel, e.g., the proportion of insertions and / or deletions detected in the target sequence. Indels (e.g., insertion-deletions) can result from repair of double-stranded DNA breaks caused by Cas9 cleavage by processes including, but not limited to, non-homologous end joining (NHEJ) repair.
[0103] The term “off-target editing frequency,” as used herein, refers to the number or proportion of unintended DNA sequences that are edited. On-target and off-target editing frequencies may be measured by the methods and assays described herein, further in view of techniques known in the art, including high-throughput sequencing reads. As used herein, high-throughput sequencing involves the hybridization of nucleic acid primers (e.g., DNA primers) with complementarity to nucleic acid (e.g., DNA) regions just upstream or downstream of the target sequence or off-target sequence of interest. Since many of the Cas9-dependent off-target sites have high sequence identity to the target site of interest, nucleic acid primers with sufficient complementarity to regions upstream or downstream of the Cas9-dependent off-target site may be designed using techniques and kits known in the art. These kits make use of polymerase chain reaction (PCR) amplification, which produces amplicons as intermediate products. The target and off-target sequences may comprise genomicloci that further comprise protospacers and PAMs. Accordingly, the term “amplicons,” as used herein, may refer to nucleic acid molecules that constitute the aggregates of genomic loci, protospacers and PAMs. High-throughput sequencing techniques used herein may further include Sanger sequencing and / or whole genome sequencing (WGS).
[0104] As used herein, the terms “transfection” or “infection” refer to the introduction of a nucleic acid into a host cell, such as by contacting the cell with liposomes or nanoparticles (e.g., lipid nanoparticles) as described herein.
[0105] As used herein, “treatment” refers to a clinical intervention made in response to a disease, disorder or physiological condition manifested by a patient or to which a patient may be susceptible. The aim of treatment includes, but is not limited to, the alleviation or prevention of symptoms, slowing or stopping the progression or worsening of a disease, disorder, or condition and / or the remission of the disease, disorder or condition. “Treatments” refer to one or both of therapeutic treatment and prophylactic or preventative measures. Subjects in need of treatment include those already affected by a disease or disorder or undesired physiological condition as well as those in which the disease or disorder or undesired physiological condition is to be prevented.
[0106] As used herein, the terms “effective amount” or “pharmaceutically effective amount” or “therapeutically effective amount” refer to an amount sufficient to effect beneficial or desirable biological and / or clinical results.
[0107] The term “pharmaceutically acceptable excipient” as used herein refers to any suitable substance that provides a pharmaceutically acceptable carrier, additive or diluent for administration of a compound(s) of interest to a subject. Pharmaceutically acceptable excipients can encompass substances referred to as pharmaceutically acceptable diluents, pharmaceutically acceptable additives, and pharmaceutically acceptable carriers.
[0108] As used herein, a “subject” refers to an animal for whom a diagnosis, treatment, or therapy is desired. In some embodiments, the subject is a mammal. “Mammal,” as used herein, refers to an individual belonging to the class Mammalia and includes, but is not limited to, humans. In some embodiments, the mammal is a primate. In some embodiments, the mammal is a human. In some embodiments, the mammal is not a human. In some embodiments, the subject has or is suspected of having a disease that can be corrected via gene editing. In some embodiments, the gene editing system is a polymerase-based editing (e.g., SyNTase editing, RT editing) system. In some embodiments, the subject has or is suspected of having Wilson’s disease. In some embodiments, the subject has or is suspected of having alphal antitrypsin deficiency disease. In some embodiments, the subject has or is suspected of having phenylketonuria or hyperphenylal aninemia.Engineered polymerases
[0109] Provided herein are engineered polymerases. An engineered polymerase can be comprised within an editor, e.g., a polymerase-based editor (e.g., SyNTase editor, RT editor). In some embodiments, the engineered polymerase advantageously provides for read-through of partially or fully chemically modified template polynucleotides. Provided herein are methods of gene editing using the described SyNTases, e.g., Synthetic Nucleotide Template Polymerase (SyNTase) Editing = Syntase Editing = SE.
[0110] Conventional polymerases used in genome editing systems, including commonly used reverse transcriptases, are generally evolved to copy natural nucleic acid templates and can exhibit reduced activity, stalling, premature termination, or complete loss of function when encountering chemically modified template nucleotides and / or modified intemucleoside linkages within the templating region. Chemically modified template polynucleotides are known in the art to inhibit polymerase read-through; for example, templates that contain sugar modifications such as 2’-0Me modifications have been shown to completely block activity (See, e.g., Chen Z, Kelly K, Cheng H, Dong X, Hedger AK, Li L, Sontheimer EJ, Watts JK. In Vivo Prime Editing by Lipid Nanoparticle Co-delivery of Chemically Modified pegRNA and Prime Editor mRNA. GEN Biotechnol. 2023 Dec;2(6):490-502, which is hereby incorporated by reference in its entirety). Specifically, Chen et al. reports that pegRNA designs designated “mod-5” (including an RTT that is fully modified with 2’-O-methyl ribose) and “mod-6” (including an RTT that is fully 2’-O-methyl modified and additionally includes phosphorothioate backbone linkages across the RTT) are inactive, reflecting a failure of the employed prime editor reverse transcriptase to read the modified RTT. In contrast, the present disclosure provides engineered polymerases (SyNTases) that are engineered to efficiently process and extend through chemically modified template regions, including templates comprising sugar modifications such as 2’ -fluoro and 2’-0Me ribose substitutions (and, in some embodiments, in combination with backbone modifications), thereby enabling robust templated synthesis and installation of an intended edit from chemically stabilized templates that would otherwise be incompatible with conventional reverse transcriptases. The engineered polymerases provided herein fulfill an unmet need for read-through of partially or fully chemically modified template polynucleotides.
[0111] Described herein are engineered polymerases. In some embodiments, the engineered polymerases comprise one or more amino acid substitutions as compared to a parent polymerase comprising a sequence selected from SEQ ID NOs: 232-234 and 579-619. In some embodiments, the parent polymerase comprises the sequence of SEQ ID NO: 234 or SEQ ID NOs: 604-605.
[0112] In some embodiments, each of the one or more amino acid substitutions is at an amino acid position functionally equivalent to S67, E69, Q84, V101, K103, D114, I125, T128, C157, K193, N200, V223, R278, P297, E302, K306, E372, K397, T420, L435, or G490 of SEQ ID NO: 234. In some embodiments, each of the one or more amino acid substitutions from the group comprising S67T, E69V, Q84G, V101R, K103R, D114N, I125V, T128N, C157Q, K193R, D / N200C, V223Y, V223M, R278L, P297Q, E302T, K306Q, E372V, K397R, T420E, L435K, or G490N of SEQ ID NO: 234.
[0113] In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, and a V223Y amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, anN200C amino acid substitution, a V223Y amino acid substitution, and a L435K amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, a V223Y amino acid substitution, and an E302T amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, an N200C amino acid substitution, and an E302T amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, and a V223M amino acid substitution relative to SEQ ID NO: 234. In some embodiments, the engineered polymerase comprises a C154Q amino acid substitution, an I222V amino acid substitution, a K100R substitution, or any combination thereof, relative to a parent sequence derived from Pipistrellus kuhlii. e.g., any of SEQ ID NOs: 604-605.
[0114] In some embodiments, the engineered polymerase comprises an amino acid sequence that is at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to any one of SEQ ID NOs: 529-534. Also provided herein are polynucleotides encoding any engineered polymerase of the disclosure. In some embodiments, an editor (e.g., SyNTase editor, RT editor) comprising an engineered polymerase comprises the amino acid sequence of any one of SEQ ID NOs: 238-249, 266-273, 282-315, 350-371, 423-471, and 628-633 or a sequence at least 80 % (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a rangebetween any two of these values) identical to any one of SEQ ID NOs: 238-249, 266-273, 282-315, 350-371, 423-471, and 628-633.
[0115] In some embodiments, an engineered polymerase is derived from a DNA polymerase or an RNA-dependent DNA polymerase (e.g., an RT). The term “derived from” as used throughout the present specification can have its ordinary meaning, and for example, in the context of amino acid sequences (e.g. an engineered polymerase) the term “derived from” can mean that the amino acid sequence, which is derived from (another) amino acid sequence, shares e.g. at least 60%, 70%, 75%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% sequence identity with the amino acid sequence from which it is derived (e.g., the parent sequence). In some embodiments, the derived sequence (e.g., an engineered polymerase, e.g., SyNTase) can have one, two, three, four, five, six, seven, eight, nine, ten, or more, amino acid substitutions relative to the sequences from which it is derived (e.g., a parent sequence).
[0116] The RT (e.g., parent sequence) can be, e.g., a virus RT, for example, a retrovirus RT. Nonlimiting examples of virus RT include Moloney murine leukemia virus (M-MLV or MLVRT or M-MLV RT); human T-cell leukemia virus type 1 (HTLV-1) RT; bovine leukemia virus (BLV) RT; Rous Sarcoma Virus (RSV) RT; human immunodeficiency virus (HIV) RT, M-MLV RT, Avian Sarcoma-Leukosis Virus (ASLV) RT, Rous Sarcoma Virus (RSV) RT, Avian Myeloblastosis Virus (AMV) RT, Avian Erythroblastosis Virus (AEV) Helper Virus MCAV RT, Avian Myelocytomatosis Virus MC29 Helper Virus MCAV RT, Avian Reticuloendotheliosis Virus (REV-T) Helper Virus REV-A RT, Avian Sarcoma Virus UR2 Helper Virus (LJR2AV) RT, Avian Sarcoma Virus ¥73 Helper Virus YAV RT, Rous Associated Virus (RAV) RT, and Myeloblastosis Associated Virus (MAV) RT, all of which may be suitably used in the methods and compositions described herein.
[0117] In some embodiments, the engineered polymerase is derived from a wild-type M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof. In some embodiments, the engineered polymerase is derived from a reference M-MLV RT, a functional mutant, a functional variant, or a functional fragment thereof.
[0118] In some embodiments, the RT is an RT (e.g., a retro-transposon RT) from an avian genome. The avian may be of the order Galliformes, Aseriformes. Passeriformes, Gruiformes, Slriilhioniformes, Pheiformes, Casuariformes, Apyerygiformes, Olidiformes, Columbiformes, Sphenisciformes, Cathartiformes, Accipilriformes, Slrigiformes, PsiUaciformes, Charadriiformes, or Falconiformes. Exemplary avians include those of Galliformes (e.g. chicken, quails, and turkey), Anseriformes (e.g. duck, and goose), Charadriiformes (e.g. gull, barred button quail, and plover), Columbiformes (e.g. pigeon), Struthioniformes (e.g. ostrich), Passeriformes(e.g. crow, finch, sparrow, starling, and swallow), Psittaciformes (e.g. parrot), Falconiformes (e.g. eagle, and falcon), Strigiformes (e.g. owl), Sphenisciformes (e.g. penguin), and Psittaciformes (e.g. parakeet, and parrot). The avian can belong to the order Passeriformes. The avian can belong to the family Passerellidae. The avian can belong to the genus Melospiza. In some embodiments, the RT is from the genome of M. georgiana. In some embodiments, the M. georgiana RT is an engineered RT. In some embodiments, the mutation sites on the M. georgiana RT may be one or more of G206N, M312K, W319F, E336P, and L611W.
[0119] In some embodiments, the RT is an RT from a mammalian genome. In some embodiments, the mammal may be of the order Chiroptera (bat). That bat can be of the family Phyllostomidae (e.g., leaf-nosed bats), Noctilionidae (bulldog bats), Cistugidae, Thyropteridae (e.g., disk-winged bats), Molossidae (e.g., free-tailed bats), Miniopteridae (e.g., long winged bat), Mormoopidae, Mystacinidae (e.g., New Zealand short-tailed bats), Myzopodidae (e.g., suckerfooted bats), Natalidae (e.g., funnel-eared bats), Emballonuridae (e.g., sheath-tailed bats), Nycteridae (e.g., slit-faced bats), Furipteridae (e.g., smoky bats), Vespertilionidae (e.g., vesper bats), Craseonycteridae, Megadermatidae (e.g., false vampire bats), Rhinolophidae (e.g., horseshoe bats), Pteropodidae (e.g., Old World fruit bats), Hipposideridae (e.g., Old World leaf-nosed bats), or Rhinopomatidae. In some embodiments, the bat is of the genus Myotis, Pipistrelles, or Eptesicus. In some embodiments, the bat species is M. brandtii, M. daubentonii, P. kuhlii, or E. nilssonii.
[0120] An engineered polymerase of the disclosure can be a component of a polymerase-based editor (e.g., SyNTase editor, RT editor). In some embodiments, the editor comprises: a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are optionally fused or linked to form a fusion protein, wherein the DNA polymerase domain comprises an engineered polymerase as provided herein. Exemplary editing systems include, e.g., RT editing, Prime Editing, DNA Polymerase Editing (DPE), and Split systems (e.g., split prime editing, DPE editing).
[0121] Also provided herein are Synthetic Nucleotide Template Polymerase (SyNTase) editing systems comprising: a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and aDNA polymerase domain, or a nucleic acid encoding the SyNTase editor, wherein the DNA polymerase domain comprises the engineered polymerase described herein.
[0122] The engineered polymerase comprised within an editor (e.g., SyNTase editor, RT editor) as described herein can advantageously confer efficiency of gene editing from partially or fully modified templates, e.g., as compared to a parent polymerase. In some embodiments, the scaffold sequence is engineered to improve, e.g., stability of the tagRNA. In some embodiments, the scaffold sequence of the tagRNA comprises the sequence of any one of SEQ ID NOs: 640-685, optionally the scaffold sequence of the tagRNA comprises the sequence of any one of SEQ ID NOs: 640, 651, and 667. In some embodiments, the editing template comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 2’ Fluoro RNA base(s). In some embodiments, all of the nucleotides of the editing template are chemically modified. In some embodiments, none of the nucleotides of the editing template are chemically modified. In some embodiments, the chemical modification is a 2’ O-methyl RNA or a 2’ Fluoro RNA base(s). In some embodiments, about 25%, about 50%, about 75%, or about 100% of the nucleotides of the editing template are chemically modified. In some embodiments, the editing template is at least or at most ten nucleotides in length, and one, two, three, four, five, six, seven, eight, nine, or ten of the nucleotides of the editing template are chemically modified.
[0123] In some embodiments, all of the nucleotides of the flap binding sequence are chemically modified. In some embodiments, none of the nucleotides of the flap binding sequence are chemically modified. In some embodiments, the chemical modification is a 2’ O-methyl RNA or a 2’ Fluoro RNA base(s). In some embodiments, about 25%, about 50%, about 75%, or about 100% of the nucleotides of the flap binding sequence are chemically modified. In some embodiments, the flap binding sequence is at least or at most six nucleotides in length, and one, two, three, four, five, or six of the nucleotides of the flap binding sequence are chemically modified.
[0124] In some embodiments, the first five nucleotides of the editing template each comprise a 2’ Fluoro RNA base, and the last 3 nucleotides of the flap binding sequence each comprise a 2’ O-methyl RNA and a phosphorothioate linkage, from 5’ to 3’. In some embodiments, the tagRNA comprises the sequence of any one of SEQ ID NOs: 686-720, optionally the tagRNA comprises the sequence of SEQ ID NO: 712. In some embodiments, the SyNTase editor comprises an amino acid sequence that is at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to any one of SEQ ID NOs: 238-249, 266-273, 282-315, 350-371, 423-471, and 628-633.
[0125] In some embodiments, the editing system is capable of editing the SERPINA1 gene. In some embodiments, the tagRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 712; and the editor (e.g., SyNTase editor) comprises an amino acid sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723.Editing methods
[0126] Editing methods (e.g., genome editing) and compositions disclosed herein can, e.g., be directed to correction of a disease mutation using reverse transcriptase (RT) editing. Compositions disclosed herein can comprise a SyNTase editor or an mRNA encoding a SyNTase editor, a long guide RNA encoding the edit designated as ‘tagRNA’, and optionally, a second guide RNA designated as ‘egRNA’ (opposite strand gRNA). In some embodiments, the composition induces programmable editing of a target DNA using a SyNTase editor complexed with a tagRNA to incorporate an intended nucleotide edit (also referred to herein as a nucleotide change) into the target DNA. In some embodiments, the composition comprises a second guide RNA (egRNA).
[0127] A target gene of the editing may comprise a double stranded DNA molecule having two complementary strands: a first strand that may be referred to as a “target strand” or a “non-edit strand”, and a second strand that may be referred to as a “non-target strand,” or an “edit strand” or an “opposite strand”.
[0128] The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for RT editing described in PCT Application No. PCT / IB2025 / 052079, entitled, “RT EDITING COMPOSITIONS AND METHODS,” filed February 26, 2025, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the Cas9 variants described in the U. S. Provisional Patent Application No. 63 / 763,733, entitled “PAM DIVERSIFICATION OF CAS9 VARIANTS,” filed February 26, 2025, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for prime editing described in W02023015309, W02022150790, W02022067130, WO2020191233, WO2020191234, WO2020191239, W02020191241, WO2020191242, WO2020191243, WO2020191245, WO2020191246, WO2020191248,WO2020191249, W02020191153, and W02020191171, the contents of which are incorporated herein by reference in their entireties. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for click editing described in WO2024211883A1, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for DNA polymerase editing (DPE) described in WO2023235501A1, the content of which is incorporated herein by reference in its entirety. The systems, methods, compositions, and kits provided herein can, in some embodiments, be employed in concert with the systems, methods, compositions, and kits for writing / rewriting described in WO2021188840A1, the content of which is incorporated herein by reference in its entirety.Editors
[0129] Provided herein are nucleotide sequence editors (e.g., SyNTase editors, RT editors). There are provided, for example, SyNTase editors. The term “SyNTase editor” refers to the polypeptide or polypeptide components involved in gene editing, or any polynucleotide(s) encoding the polypeptide or polypeptide components. In various embodiments, an editor includes a polypeptide domain having DNA endonuclease activity and a polypeptide domain having DNA polymerase activity. In some embodiments, the editor further comprises a polypeptide domain having nuclease activity. In some embodiments, the polypeptide domain having DNA binding activity comprises a nuclease domain or nuclease activity. In some embodiments, the polypeptide domain having nuclease activity comprises a nickase, or a fully active nuclease. As used herein, the term “nickase” refers to a nuclease capable of cleaving only one strand of a double-stranded DNA target. In some embodiments, the editor comprises a polypeptide domain that is an inactive nuclease, in some embodiments, the polypeptide domain having comprises a nucleic acid guided DNA endonuclease domain, for example, a CRISPR-Cas protein, for example, a Cas9 nickase, a Cpfl nickase, or another CRISPR-Cas nuclease. In some embodiments, the polypeptide domain having DNA polymerase activity comprises a template-dependent DNA polymerase, for example, a DNA-dependent DNA polymerase or an RNA-dependent DNA polymerase. In some embodiments, the DNA polymerase is a reverse transcriptase, in some embodiments, the editor comprises additional polypeptides involved in editing, for example, a polypeptide domain having a 5’ endonuclease activity, e.g., a 5’ endogenous DNA flap endonucleases (e.g., FEN1), for helping to drive the editing process towards the edited product formation. In some embodiments, the editor further comprises an RNA-protein recruitment polypeptide, for example, an MS2 coat protein.
[0130] An editor (e.g., SyNTase editor, RT editor) may be engineered. In some embodiments, the polynucleotide or polypeptide components of an editor do not naturally occur in the same organism or cellular environment. In some embodiments, the polynucleotide or polypeptide components of an editor may be of different origins or from different organisms. In some embodiments, an editor comprises a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain that are derived from different species. In some embodiments, an editor comprises a Cas polypeptide (DNA endonuclease domain) and a reverse transcriptase polypeptide (DNA polymerase) that are derived from different species. As described herein, a SyNTase editor comprising an engineered polymerase of the disclosure can be referred to as an RT editor.
[0131] In some embodiments, polypeptide domains of an editor (e.g., SyNTase editor, RT editor) may be fused or linked by a peptide linker to form a fusion protein. In other embodiments, an editor comprises one or more polypeptide domains provided in trans as separate proteins, which are capable of being associated to each other through non-peptide linkages or through aptamers or recruitment sequences. For example, an editor may comprise a DNA binding domain and a reverse transcriptase domain associated with each other by an RNA-protein recruitment aptamer, e.g., an MS2 aptamer, which may be linked to a tagRNA. Editor polypeptide components may be encoded by one or more polynucleotides in whole or in part, in some embodiments, a single polynucleotide, construct, or vector encodes the editor fusion protein. In some embodiments, multiple polynucleotides, constructs, or vectors each encode a polypeptide domain or portion of a domain of an editor, or a portion of an editor fusion protein. For example, an editor fusion protein may comprise an N-terminal portion fused to an intein-N and a C-terminal portion fused to an intein-C, each of which is individually encoded by an AAV vector
[0132] In some embodiments, the editor (e.g., SyNTase editor, RT editor) is transcribed from an mRNA. In some embodiments, the mRNA comprises a 5'-cap structure. In some embodiments, the mRNA comprises a nuclear localization sequence (NLS). In some embodiments, the mRNA sequence comprises a dead Cas9 sequence. In some embodiments, the Cas9 is a Cas nickase sequence. In some embodiments, the editor mRNA sequence comprises a 5’-UTR. In some embodiments, the editor mRNA sequence comprises a 3’-UTR. In some embodiments, the editor comprises a sequence of a viral element. In some embodiments, the viral element is a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) sequence. In some embodiments, the editor comprises secondary structure motifs. In some embodiments, the viral element sequence comprises or consists of any one of SEQ ID NOs: 141-142. In some embodiments, the editor comprises an ek5 sequence. The ek5 sequence can comprise the sequence of SEQ ID NO: 142, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 142.The ek5 sequence can have one, two, three, four, or five mismatches relative to the sequence of SEQ ID NO: 142. The viral element sequence can be derived from Aichi virus 1 (AiV-1), such as, for example, the 3’ UTR K5 element (GenBank: NC 001918.1, 8,122-8,251). In some embodiments, the viral element sequence comprises the extended form of K5 (“ek5,” 8,067-8,251, 185 nt). Viral element sequences are described in Seo, Jenny J., et al. (" Functional viromic screens uncover regulatory RNA elements." Cell 186.15 (2023): 3291-3306), the content of which is incorporated herein by reference in its entirety. In some embodiments, the secondary structure motif comprises a triple helix sequence. In some embodiments, the editor mRNA sequence has a structure comprising or consisting of a sequence comprising from 5’ to 3’ as [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR_and / or_viral_element]-[polyA_sequence]. In some embodiments, any nucleotide of the editor mRNA sequence may be chemically modified.
[0133] Compositions disclosed herein comprise a long guide RNA designated herein as “template armed guide RNA” or “tagRNA”. The tagRNA comprises a spacer sequence, a scaffold sequence, an editing template and a flap binding sequence. In some embodiments, the tagRNA comprises in 5’ to 3’ order: spacer sequence, scaffold sequence, editing template, and a flap binding sequence.Endonuclease domains
[0134] In some embodiments, an editor (e.g., SyNTase editor, RT editor) comprises a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain. In some embodiments, the DNA endonuclease domain is a CRISPR associated (Cas) protein domain, optionally the Cas protein domain is Cas9.
[0135] In some embodiments, a Cas protein (e.g., a Cas protein domain), e.g., Cas9, can be a wild type or a modified form of a Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild-type Cas protein. In some embodiments, a Cas protein, e.g., Cas9, can be a wild type or a modified form of a Cas protein. A Cas protein, e.g., Cas9, can be a nuclease active variant, nuclease inactive variant, a nickase, or a functional variant or functional fragment of a wild-type Cas protein. A Cas protein, e.g., Cas9, can comprise an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof relative to a corresponding wild-type version of the Cas protein. In some embodiments, a Cas protein can be a polypeptide with at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild type exemplary Cas protein.
[0136] A Cas protein (e.g., a Cas protein domain), e.g., Cas9, may comprise one or more domains. Non-limiting examples of Cas domains include, guide nucleic acid recognitionand / or binding domain, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA binding domain, a DNA endonuclease domain, RNA binding domain, helicase domains, proteinprotein interaction domains, and dimerization domains. In various embodiments, a Cas protein comprises a guide nucleic acid recognition and / or binding domain that can interact with a guide nucleic acid, and one or more nuclease domains that comprise catalytic activity for nucleic acid cleavage.
[0137] In some embodiments, a Cas protein (e.g., a Cas protein domain), e.g., Cas9, comprises one or more nuclease domains. A Cas protein can comprise an amino acid sequence having at least about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity to a nuclease domain (e.g., RuvC domain, HNH domain) of a wild-type Cas protein. In some embodiments, a Cas protein comprises a single nuclease domain. For example, a Cpfl may comprise a RuvC domain but lacks HNH domain. In some embodiments, a Cas protein comprises two nuclease domains, e.g., a Cas9 protein can comprise an HNH nuclease domain and a RuvC nuclease domain.
[0138] In some embodiments, an editor (e.g., SyNTase editor, RT editor) comprises a Cas protein (e.g., a Cas protein domain), e.g., Cas9, wherein all nuclease domains of the Cas protein are active. In some embodiments, an editor comprises a Cas protein having one or more inactive nuclease domains. One or a plurality of the nuclease domains (e.g., RuvC, HNH) of a Cas protein can be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. In some embodiments, a Cas protein, e.g., Cas9, comprising mutations in a nuclease domain has reduced (e.g., nickase) or abolished nuclease activity while maintaining its ability to target a nucleic acid locus at a search target sequence when complexed with a guide nucleic acid, e.g., a tagRNA.
[0139] In some embodiments, an editor (e.g., SyNTase editor, RT editor) comprises a Cas nickase that can bind to the target gene in a sequence-specific manner and generate a singlestrand break at a protospacer within double-stranded DNA in the target gene, but not a doublestrand break. For example, the Cas nickase can cleave the edit strand or the non-edit strand of the target gene, but may not cleave both. In some embodiments, an editor comprises a Cas nickase comprising two nuclease domains (e.g., Cas9), with one of the two nuclease domains modified to lack catalytic activity or deleted. In some embodiments, the Cas nickase of an editor comprises a nuclease inactive RuvC domain and a nuclease active HNH domain. In some embodiments, the Cas nickase of an editor comprises a nuclease inactive HNH domain and a nuclease active RuvC domain. In some embodiments, an editor comprises a Cas9 nickase having an amino acid substitution in the RuvC domain e.g., an amino acid substitution that reduces or abolishes nuclease activity of the RuvC domain. In some embodiments, the Cas9 nickase comprises a D10X aminoacid substitution compared to a wild type S. pyogenes Cas9, wherein X is any amino acid other than D. In some embodiments, an editor comprises a Cas9 nickase having an amino acid substitution in the HNH domain, e.g., an amino acid substitution that reduces or abolishes nuclease activity of the HNH domain. In some embodiments, the Cas9 nickase comprises a H840X amino acid substitution compared to a wild type S. pyogenes Cas9, wherein X is any amino acid other than H.
[0140] In some embodiments, an editor (e.g., SyNTase editor, RT editor) comprises a Cas protein (e.g., a Cas protein domain) that can bind to the target gene in a sequence-specific manner but lacks or has abolished nuclease activity and may not cleave either strand of a double stranded DNA in a target gene. Abolished activity or lacking activity can refer to an enzymatic activity less than 1%, less than 2%, less than 3%, less than 4%, less than 5%, less than 6%, less than 7%, less than 8%, less than 9%, or less than 10% activity compared to a wild-type exemplary activity (e.g., wild-type Cas9 nuclease activity). In some embodiments, a Cas protein of an editor completely lacks nuclease activity. A nuclease, e.g., Cas9, that lacks nuclease activity may be referred to as nuclease inactive or “nuclease dead” (abbreviated by “d”). A nuclease dead Cas protein (e.g., dCas, dCas9) can bind to a target polynucleotide but may not cleave the target polynucleotide. In some aspects, a dead Cas protein is a dead Cas9 protein. In some embodiments, an editor comprises a nuclease dead Cas protein wherein all of the nuclease domains (e.g., both RuvC and HNH nuclease domains in a Cas9 protein; RuvC nuclease domain in a Cpfl protein) are mutated to lack catalytic activity, or are deleted.
[0141] A Cas protein can be modified. A Cas protein (e.g., a Cas protein domain), e.g., Cas9, can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein.
[0142] A Cas protein (e.g., a Cas protein domain) can be a fusion protein. For example, a Cas protein can be fused to a cleavage domain, an epigenetic modification domain, a transcriptional regulation domain, or a polymerase domain. A Cas protein can also be fused to a heterologous polypeptide providing increased or decreased stability. The fused domain or heterologous polypeptide can be located at the N-terminus, the C-terminus, or internally within the Cas protein.
[0143] In some embodiments, the Cas protein (e.g., a Cas protein domain) of an editor (e.g., SyNTase editor, RT editor) is a Class 2 Cas protein. In some embodiments, the Cas proteinis atype II Cas protein. In some embodiments, the Cas protein is a Cas9 protein, a modified version of a Cas9 protein, a Cas9 protein homolog, mutant, variant, or a functional fragment thereof. As used herein, a Cas9, Cas9 protein, Cas9 polypeptide or a Cas9 nuclease refers to an RNA guided nuclease comprising one or more Cas9 nuclease domains and a Cas9 gRNA binding domain having the ability to bind a guide polynucleotide, e.g., a tagRNA. A Cas9 protein may refer to a wild-type Cas9 protein from any organism or a homolog, ortholog, or paralog from any organisms; any functional mutants or functional variants thereof; or any functional fragments or domains thereof. In some embodiments, an editor comprises a full-length Cas9 protein. In some embodiments, the Cas9 protein can generally comprises at least about 50%, 60%, 70%, 80%, 90%, 100% sequence identity to a wild-type reference Cas9 protein (e.g., Cas9 from S. pyogenes). In some embodiments, the Cas9 comprises an amino acid change such as a deletion, insertion, substitution, fusion, chimera, or any combination thereof as compared to a wild-type reference Cas9 protein.
[0144] In some embodiments the DNA endonuclease domain of the editor (e.g., SyNTase editor, RT editor) comprises a Cas9 protein (e.g., a Cas9 nickase). The Cas9 protein can comprise an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 66-133. The Cas9 protein can comprise an amino acid sequence having one, two, three, four, or five mismatches relative to the sequence of any one of SEQ ID NOs: 66-133.Accessory Domains
[0145] In some embodiments, the editor (e.g., SyNTase editor, RT editor) further comprises an accessory domain. The accessory domain can be a single-strand binding protein domain or a stabilon. The SSB protein domain can be derived from RecA protein, Sso7d protein, or Sto7d protein. In some embodiments, the SSB protein domain is at the N-terminus, the C-terminus, or an internal location of the editor. In some embodiments, the SSB protein domain is derived from Sso7d and comprises one or more mutations at an amino acid position functionally equivalent to K12 and / or E35 of a wild type Sso7d amino acid sequence. In some embodiments, the one or more mutations comprise K12L and / or E35L relative to a wild type Sso7d sequence. In some embodiments, the SSB protein domain comprises an amino acid sequence that is at least 85% identical to SEQ ID NO: 231 (e.g., 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, identical to SEQ ID NO: 231). In some embodiments, the editor comprising the SSB protein domain comprises from N-terminus to C-terminus: [N-terminus of nCas9]-[linker]-[SSB protein domain]-[linker]-[RT]-[linker]-[C-terminus of nCas9],Additional sequence elements o f the editor mRNA
[0146] The present disclosure provides optimized mRNAs encoding an editor (e.g., SyNTase editor, RT editor), that provide effective genome editing of a target cell population when administered with one or more gRNAs (e.g., tagRNA and an optional egRNA). In some embodiments, additional segmented polyA sequences increase mRNA stability and half-life of the editor. In some embodiments, the disclosure provides an mRNA comprising (i) a 5 '-cap, (ii) a 5’-untranslated region (UTR); (ii) an open reading frame (ORF) comprising a nucleotide sequence that encodes an editor; and (iv) a 3' untranslated region (UTR). In some embodiments, the mRNA further comprises viral element sequences 3’ of the 3 ’-UTR. In some embodiments, the viral element sequence comprises the sequence of SEQ ID NO: 141. In some embodiments, the mRNA further comprises an eK5 sequence that is located 3’ of the 3 ’-UTR. eK5 is a regulatory RNA element derived from a virus that when placed in the 3 ‘UTR of an mRNA can lead to increased protein expression (for instance, as in Seo et al Cell 2023, 186: 3291). In some embodiments, the ek5 sequence comprises the sequence of SEQ ID NO: 142. In some embodiments, the synthetic triple helix secondary structure sequence (STH) comprises the sequence derived from a sequence element of the long non-coding RNA, such as MALAT1 In some embodiments, the STH sequence comprises the sequence of SEQ ID NO: 173. Some embodiments provided herein employ the triple helix motif from MALAT1, or an engineered (minimal) version. In some embodiments, MALAT1 (Gene ID: 378938) based triple helix motif is situated at the 3’ end. MALAT1 (metastasis associated lung adenocarcinoma transcript 1) also known as NEAT2 (noncoding nucl ear-enriched abundant transcript 2) is a large, infrequently spliced non-coding RNA, which is highly conserved amongst mammals and highly expressed in the nucleus. In some embodiments, the mRNA comprises sequences that are codon-optimized for expression in human cells. In some embodiments, the mRNA (e.g., codon-optimized mRNA) comprises a sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 723.Nucleic acid modi fications
[0147] In some embodiments, any of the nucleic acids of the disclosure can comprise one or more modifications (e.g., gRNA, tagRNA, egRNA, or any nucleic acid encoding any component of the editor systems disclosed herein). In some embodiments, the gRNA (e.g., egRNA) or tagRNA is a chemically modified gRNA or tagRNA. Various types of RNA modifications can be introduced to the gRNAs or tagRNAs to enhance stability, reduce the likelihood or degree of innate immune response, and / or enhance other attributes as described in the art. The gRNAs or tagRNAs described herein can comprise one or more modificationsincluding internucleoside linkages, purine or pyrimidine bases, or sugar. In some embodiments, a modification is introduced at the terminal of a gRNA or tagRNA with chemical synthesis or with a polymerase enzyme. Examples of modified nucleic acids and their synthesis are disclosed in WO2013 / 052523. Synthesis of modified polynucleotides is also described in Verma and Eckstein, Annual Review of Biochemistry, vol. 76, 99-134 (1998).
[0148] In some embodiments, programmable editing of a target DNA comprises a template armed gRNA (tagRNA). As provided herein, the engineered polymerases (e.g., SyNTases) of the disclosure can improve read-through of partially or fully modified editing templates, e.g., comprised within a tagRNA. In some embodiments, one or more positions (e.g., nucleotides) of the tagRNA is chemically modified. In some embodiments, the chemical modifications can be any one of an LNA, 2’-fluoro, DNA, 2’-OMe, and 2’ MethoxyEthoxy (2’-MOE)-chemical modification. In some embodiments, the tagRNA comprises one or more deoxyribonucleotides (e.g., DNA). In some embodiments, each nucleotide of the editing template comprises an LNA modification. In some embodiments, each nucleotide of the flap binding sequence comprises an LNA modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, every second nucleotide of the editing template comprises an LNA modification. In some embodiments, every second nucleotide of the flap binding sequence comprises an LNA modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, every third nucleotide of the editing template comprises an LNA modification. In some embodiments, every third nucleotide of the flap binding sequence comprises an LNA modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises an LNA modification. In some embodiments, each nucleotide of the editing template comprises a 2’ -fluoro modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-fluoro modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises a 2’ -fluoro modification. In some embodiments, every second nucleotide of the editing template comprises a 2’ -fluoro modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’ -fluoro modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’ -fluoro modification. In some embodiments, every third nucleotide of the editing template comprises a 2’ -fluoro modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’ -fluoro modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’ -fluoro modification. In some embodiments, each nucleotide of the editing templatecomprises a 2’-0Me modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-0Me modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-0Me modification. In some embodiments, every second nucleotide of the editing template comprises a 2’-0Me modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’-0Me modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’-0Me modification. In some embodiments, every third nucleotide of the editing template comprises a 2’-0Me modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’-0Me modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’-0Me modification. In some embodiments, each nucleotide of the editing template comprises a 2’-M0E modification. In some embodiments, each nucleotide of the flap binding sequence comprises a 2’-M0E modification. In some embodiments, each nucleotide of both the editing template and the flap binding sequences comprises a 2’-M0E modification. In some embodiments, every second nucleotide of the editing template comprises a 2’ -MOE modification. In some embodiments, every second nucleotide of the flap binding sequence comprises a 2’ -MOE modification. In some embodiments, every second nucleotide of both the editing template and the flap binding sequences comprises a 2’ -MOE modification. In some embodiments, every third nucleotide of the editing template comprises a 2’ -MOE modification. In some embodiments, every third nucleotide of the flap binding sequence comprises a 2’-M0E modification. In some embodiments, every third nucleotide of both the editing template and the flap binding sequences comprises a 2’ -MOE modification. In some embodiments, the total number of tagRNA nucleotides comprising a chemical modification provided herein (e.g., an LNA, 2’-fluoro, DNA, 2’-0Me, and / or 2’-M0E chemical modification) can be at least, or can be at most, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100, nucleotides.
[0149] In some embodiments, a tagRNA complexes with and directs an editor (e.g., SyNTase editor, RT editor) to bind to the search target sequence of the target gene. In some embodiments, the bound editor generates a nick on the edit strand (PAM strand) of the target gene at the nick site. In some embodiments, a flap binding site (FB sequence) of the tagRNA anneals with a free 3’ end formed at the nick site, and the editor initiates DNA synthesis from the nick site, using the free 3 ’ end as a primer. Subsequently, a single-stranded DNA encoded by the editing template of the tagRNA is synthesized. In some embodiments, the newly synthesized single-stranded DNA comprises one or more intended nucleotide edits compared to an endogenous target gene sequence. Accordingly, in some embodiments, the editing template of a tagRNA is complementary to a sequence in the edit strand except for one or more mismatches at the intended nucleotide edit positions in the editing template. The endogenous, e.g., genomic, sequence that is partially complementary to the editing template may be referred to as an “editing target sequence.” Accordingly, in some embodiments, the newly synthesized single stranded DNA has identity or substantial identity to a sequence in the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions.
[0150] In some embodiments, the newly synthesized single-stranded DNA equilibrates with the editing target on the edit strand of the target gene for pairing with a target strand of a target gene. In some embodiments, an editing target sequence of a target gene is excised by a flap endonuclease (FEN), for example, FEN1. In some embodiments, the FEN is an endogenous FEN, for example, in a cell comprising a target gene. In some embodiments, the FEN is provided as part of the editor (e.g., SyNTase editor, RT editor), either linked to other components of the editor or provided in trans. In some embodiments, the newly synthesized single stranded DNA, which comprises the intended nucleotide edit, replaces the endogenous single stranded editing target sequence on the edit strand of the target gene. In some embodiments, the newly synthesized single stranded DNA and the endogenous DNA on the target strand form a heteroduplex DNA structure at the region corresponding to the editing target sequence of the target gene. In some embodiments, the newly synthesized single-stranded DNA comprising the nucleotide edit is paired in the heteroduplex with the target strand of the target DNA that does not comprise the nucleotide edit, thereby creating a mismatch between the two otherwise complementary strands. In some embodiments, the mismatch is recognized by DNA repair machinery, e.g., an endogenous DNA repair machinery. In some embodiments, through DNA repair, the intended nucleotide edit is incorporated into the target gene.
[0151] In some embodiments, an editor (e.g., SyNTase editor, RT editor) comprises a Cas9 functional variant that is of smaller molecular weight than a wild-type SPYCas9 protein. In some embodiments, a smaller-sized Cas9 functional variant may facilitate delivery to cells, e.g., by an expression vector, nanoparticle, or other means of delivery. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type II Cas protein. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type V Cas protein. In some embodiments, a smaller-sized Cas9 functional variant is a Class 2 Type VI Cas protein.Nuclear Localization Sequences and linkers
[0152] In some embodiments, an editor (e.g., SyNTase editor, RT editor) further comprises one or more nuclear localization sequence (NLS). In some embodiments, the NLS helpspromote translocation of a protein into the cell nucleus. In some embodiments, an editor comprises a fusion protein, e.g., a fusion protein comprising a DNA endonuclease domain and a DNA polymerase, that comprises one or more NLSs. In some embodiments, one or more polypeptides of the editor are fused to or linked to one or more NLSs. In some embodiments, the editor comprises a DNA endonuclease domain and a DNA polymerase domain that are provided in trans, wherein the DNA endonuclease domain and / or the DNA polymerase domain is fused or linked to one or more NLSs. In some embodiments the editor mRNA comprises or consists of the structure:[5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]-[polyA sequence]. The editor mRNA can comprises or consist of the structure [5’-cap]-[5’UTR]-[NLS]-[nCas9]-[linker and / or NLS]-[RT]-[NLS]-[3’UTR_and / or_viral_element]-[secondary structure motif and / or polyA sequence]. In some cases, the positions of nCas9 and RT are swapped to e.g., improve editing efficiency.
[0153] In some embodiments, an editor (e.g., SyNTase editor, RT editor) or editing complex comprises at least one NLS. In some embodiments, an editor or editing complex comprises at least two NLSs. In embodiments with at least two NLSs, the NLSs can be the same NLS, or they can be different NLSs.
[0154] In addition, the NLSs can be expressed as part of an editor complex. The location of the NLS fusion can be at the N-terminus, the C-terminus, or positioned anywhere within a sequence of an editor (e.g., SyNTase editor, RT editor) or a component thereof (e.g., inserted between the DNA endonuclease domain and the DNA polymerase domain of an editor fusion protein, between the DNA endonuclease domain and a linker sequence, between a DNA polymerase and a linker sequence, between two linker sequences of an editor fusion protein or a component thereof, in either N-terminus to C-terminus or C-terminus to N-terminus order).
[0155] Any NLSs that are known in the art are also contemplated herein. The NLSs may be any naturally occurring NLS, or any non-naturally occurring NLS (e.g., an NLS with one or more mutations relative to a wild-type NLS). In some embodiments, the one or more NLSs of an editor (e.g., SyNTase editor, RT editor) comprise bipartite NLSs. In some embodiments, a nuclear localization signal (NLS) is predominantly basic. In some embodiments, the one or more NLSs of an editor are rich in lysine and arginine residues. In some embodiments, the one or more NLSs of an editor comprise proline residues.
[0156] Non-limiting examples of NLS sequences are provided in Table 1 below.Table 1: Exemplary Nuclear Localization SignalsName / Description Sequence SEQ ID NO NLS of SV40 Large T-AG PKKKRKV 1NLS KRTADGSEFESPKKKRKV 2NLS MDSLLMNRRKFLYQFKNVRWAKGRRETYLC 3NLS of Nucleoplasmin AVKRPAATKKAGQAKKKKLD 4 NLS ofEGL-13 MSRRRKANPTKLSENAKKLAKEVEN 5NLS ofC-Myc PAAKRVKLD 6NLS of Tus-protein KLKIKRPVK 7NLS of polyoma large T-AG VSRKRPRP 8NLS of Hepatitis D virus antigen EGAPPAKRAR 9NLS of murine p53 PPQPKKKPLDGE 10C-terminal linker and NLS of an SGGSKRTADGSEFESPKKKRKV 11 exemplary editor fusion proteinHybrid NLS PAAKKKKLD 12VirD NLS PKRPRDRHDGELGGRKRARG 13
[0157] In some embodiments, components of an editor (e.g., SyNTase editor, RT editor) are directly fused to each other. In some embodiments, components of an editor are associated to each other via a linker. As used herein, a linker can be any chemical group or a molecule linking two molecules or moi eties, e.g., a DNA binding domain, a DNA endonuclease domain and a polymerase domain of an editor. In some embodiments, a linker is an organic molecule, group, polymer, or chemical moiety. In some embodiments, the linker comprises a nonpeptide moiety. The linker may be a covalent bond (e.g., a carbon-carbon bond, disulfide bond, carbon-heteroatom bond, etc.), or it may be a polymeric linker, for example, a polynucleotide sequence.
[0158] In some embodiments, two or more components of an editor (e.g., SyNTase editor, RT editor) are linked to each other by a peptide linker. In some embodiments, a peptide linker is 5-100 amino acids in length, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 30-35, 35-40, 40-45, 45-50, 50-60, 60-70, 70-80, 80-90, 90-100, 100-150, or 150-200 amino acids in length. In some embodiments, the peptide linker is 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 35, 45, 50, 55, 60, 60, 65, 70, 70, 75, 80, 85, 90, 90, 95, 100, 101, 102, 103, 104, 105, 110, 120, 130, 140,150, 160, 175, 180, 190, or 200 amino acids in length. In some embodiments, the peptide linker is 5-100 amino acids in length. In some embodiments, the peptide linker is 10-80 amino acids in length. In some embodiments, the peptide linker is 15-70 amino acids in length. In some embodiments, the peptide linker is 16 amino acids in length, 24 amino acids in length, 64 amino acids in length, or 96 amino acids in length, in some embodiments, the peptide linker is at least 50 amino acids in length, in some embodiments, the peptide linker is at least 40 amino acids in length, in some embodiments, the peptide linker is at least 30 amino acids in length. In some embodiments, the peptide linker is 46 amino acids in length. In some embodiments, the peptide linker is 92 amino acids in length. In some embodiments, the peptide linker is 16 amino acids in length, 24 amino acids in length, 64 amino acids in length, or 96 amino acids in length.
[0159] In some embodiments, the linker comprises the amino acid sequence (GGGGS)n (SEQ ID NO: 14), (G)n, (EAAAK)n (SEQ ID NO: 15), (GGS)n, (SGGS)n (SEQ ID NO: 16), (XP)n, or any combination thereof, wherein n is independently an integer between 1 and 30, and wherein X is any amino acid. In some embodiments, the linker comprises the amino acid sequence (GGS)n, wherein n is 1, 3, or 7. In some embodiments, the linker comprises the amino acid sequence SGSETPGTSESATPES (SEQ ID NO: 17). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESSGGSSGGS (SEQ ID NO: 18). In some embodiments, the linker comprises the amino acid sequence SGGSGGSGGS (SEQ ID NO: 19). In some embodiments, the linker comprises the amino acid sequence SGGSSGGSSGSETPGTSESATPESAGSYPYDVPDYAGSAAPAAKKKKLDGSGSGGSSGGS (SEQ ID NO: 20). In some embodiments, the linker comprises the amino acid sequence GGSGGS (SEQ ID NO: 21), GGSGGSGGS (SEQ ID NO: 22), or SGGSSGGSSGSETPGTSESATPESSGGSSGGSS (SEQ ID NO: 23).
[0160] Components of an editor (e.g., SyNTase editor, RT editor) may be connected to each other in any order. In some embodiments, the DNA binding domain, a DNA endonuclease domain and the DNA polymerase domain of an editor may be fused to form a fusion protein or may be joined by a peptide or protein linker, in any order from the N-terminus to the C-terminus. In some embodiments, an editor comprises a DNA binding domain, a DNA endonuclease domain fused or linked to the C-terminal end of a DNA polymerase domain. In some embodiments, an editor comprises a DNA binding domain, a DNA endonuclease domain fused or linked to the N-terminal end of a DNA polymerase domain. In some embodiments, the editor comprises a fusion protein comprising the structure NH2-[DNA binding domain, a DNA endonuclease domain]-[polymerase]-COOH; or NH2- [polymerase] -[DNA binding domain, a DNA endonuclease domain]-COOH, wherein each instance of “]-[“ indicates the presence of an optional linker sequence. In some embodiments, an editor comprises a fusion protein and a DNA polymerase domain provided in trans, wherein the fusion protein comprises the structure NH2-[DNA binding domain, a DNA endonuclease domain] -[RNA-protein recruitment polypeptide]-COOH. In some embodiments, an editor comprises a fusion protein and a DNA binding domain, a DNA endonuclease domain provided in trans, wherein the fusion protein comprises the structure NH2-[DNA polymerase domain]-[RNA-protein recruitment polypeptide]-COOH.
[0161] In some embodiments, an editor (e.g., SyNTase editor, RT editor) fusion protein, a polypeptide component of an editor, or a polynucleotide encoding the editor fusion protein or polypeptide component, may be split into an N-terminal half and a C-terminal half or polypeptides that encode the N-terminal half and the C-terminal half, and provided to a target DNA in a cell separately. For example, in some embodiments, an editor fusion protein may besplit into an N-terminal and a C-terminal half for separate delivery in AAV vectors, and subsequently translated and co-localized in a target cell to reform the complete polypeptide or editor protein. In such cases, separate halves of a protein or a fusion protein may each comprise a split-intein to facilitate co-localization and reformation of the complete protein or fusion protein by the mechanism of intein facilitated trans splicing. In some embodiments, an editor comprises a N-terminal half fused to an intein-N, and a C-terminal half fused to an intein-C, or polynucleotides or vectors (e.g., AAV vectors) encoding each thereof. When delivered and / or expressed in a target cell, the intein-N and the intein-C can be excised via protein trans-splicing, resulting in a complete editor fusion protein in the target cell.tagRNAs
[0162] Disclosed herein include target priming RNAs (tagRNAs). The term “target priming RNA”, or “tagRNA”, refers to a guide polynucleotide that comprises one or more intended nucleotide edits for incorporation into the target DNA. In some embodiments, the tagRNA associates with and directs an editor (e.g., SyNTase editor, RT editor) to incorporate the one or more intended nucleotide edits into the target gene via, e.g., SyNTase editing. “Nucleotide edit” or “intended nucleotide edit” refers to a specified deletion of one or more nucleotides at one specific position, insertion of one or more nucleotides at one specific position, substitution of a single nucleotide, or other alterations at one specific position to be incorporated into the sequence of the target gene. Intended nucleotide edit may refer to the edit on the editing template as compared to the sequence on the target strand of the target gene or may refer to the edit encoded by the editing template on the newly synthesized single stranded DNA that replaces the editing target sequence, as compared to the editing target sequence. In some embodiments, a tagRNA comprises a spacer sequence that is complementary or substantially complementary to a search target sequence on a target strand of the target gene, in some embodiments, the tagRNA comprises a gRNA core that associates with a DNA endonuclease domain, e.g., a CRISPR-Cas protein domain, of a SyNTase editor. In some embodiments, the tagRNA further comprises an extended nucleotide sequence comprising one or more intended nucleotide edits compared to the endogenous sequence of the target gene, wherein the extended nucleotide sequence may be referred to as an extension arm.
[0163] In some embodiments, in a template armed gRNA (tagRNA), a spacer sequence is complementary or substantially complementary to a specific sequence on the target strand, which may be referred to as a “search target sequence.” In some embodiments, the spacer sequence anneals with the target strand at the search target sequence. The target strand may also be referred to as the “non-Protospacer Adjacent Motif (non-PAM strand).” In some embodiments, the non-target strand may also be referred to as the “PAM strand.” In some embodiments, thePAM strand comprises a protospacer sequence and optionally a protospacer adjacent motif (PAM) sequence. In SyNTase editing using a Cas-protein-based SyNTase editor, a PAM sequence refers to a short DNA sequence immediately adjacent to the protospacer sequence on the PAM strand of the target gene. A PAM sequence may be specifically recognized by a programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease, in some embodiments, a specific PAM is characteristic of a specific programmable DNA binding protein, e.g., a Cas nickase or a Cas nuclease. A protospacer sequence refers to a specific sequence in the PAM strand of the target gene that is complementary to the search target sequence. In a tagRNA, a spacer sequence may have a substantially identical sequence as the protospacer sequence on the edit strand of a target gene, except that the spacer sequence may comprise Uracil (U) and the protospacer sequence may comprise Thymine (T).
[0164] In some embodiments, the double stranded target DNA comprises a nick site on the PAM strand (or non-target strand). As used herein, a “nick site" refers to a specific position in between two nucleotides or two base pairs of the double stranded target DNA. In some embodiments, the position of a nick site is determined relative to the position of a specific PAM sequence. In some embodiments, the nick site is the particular position where a nick will occur when the double stranded target DNA is contacted with a nickase, for example, a Cas nickase, that recognizes a specific PAM sequence. In some embodiments, the nick site is upstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is downstream of a specific PAM sequence on the PAM strand of the double stranded target DNA. In some embodiments, the nick site is upstream of a PAM sequence recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the nick site is 3 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a Streptococcus pyogenes Cas9 nickase. In some embodiments, the Cas nickase comprises any one of the sequences recited in SEQ ID NOs: 66-87.
[0165] In some embodiments, the nick site is 3 base pairs upstream of the PAM sequence, and the PAM sequence is recognized by a Cas9 nickase, wherein the Cas9 nickase comprises a nuclease active HNH domain and a nuclease inactive RuvC domain. In some embodiments, the nick site is 2 nucleotides upstream of the PAM sequence, and the PAM sequence is recognized by a S. thermophilus Cas9 nickase that comprises a nuclease active RuvC domain and a nuclease inactive HNH domain. In some embodiments, the Cas nick is 4, 5, 6 or more than 6 nucleotides base pairs upstream of the PAM sequence and the PAM sequence is recognized by a Cas9 nickase wherein the Cas9 nickase comprises a nuclease active HNH domain and a nuclease inactive RuvC domain.
[0166] An “editing template” of a tagRNA is a single-stranded portion of the tagRNA that is 5' of the FB sequence and comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand), and comprises one or more intended nucleotide edits compared to the endogenous sequence of the double stranded target DNA. In some embodiments, the editing template and the FB sequence are immediately adjacent to each other. Accordingly, in some embodiments, a tagRNA in SyNTase editing comprises a single-stranded portion that comprises the editing template sequence and the FB sequence immediately adjacent to each other. In some embodiments, the single stranded portion of the tagRNA comprising both the editing template sequence and the flap binding sequence is complementary or substantially complementary to an endogenous sequence on the PAM strand (i.e., the non-target strand or the edit strand) of the double stranded target DNA except for one or more non-complementary nucleotides at the intended nucleotide edit positions. As used herein, regardless of relative 5 -3' positioning in other contexts, the relative positions as between the FB sequence and the editing template, and the relative positions as among elements of a tagRNA, are determined by the 5' to 3' order of the tagRNA as a single molecule regardless of the position of sequences in the double stranded target DNA that may have complementarity or identity to elements of the tagRNA. In some embodiments, the editing template is complementary or substantially complementary to a sequence on the PAM strand that is immediately downstream of the nick site, except for one or more non-complementary nucleotides at the intended nucleotide edit positions. The endogenous, e.g., genomic, sequence that is complementary or substantially complementary to the editing template, except for the one or more non-complementary nucleotides at the position corresponding to the intended nucleotide edit, may be referred to as an “editing target sequence." In some embodiments, the editing template has identity or substantial identity to a sequence on the target strand that is complementary to, or having the same position in the genome as, the editing target sequence, except for one or more insertions, deletions, or substitutions at the intended nucleotide edit positions. In some embodiments, the editing template encodes a single stranded DNA, wherein the single stranded DNA has identity or substantial identity to the editing target sequence except for one or more insertions, deletions, or substitutions at the positions of the one or more intended nucleotide edits.
[0167] A “flap binding (FB) sequence” is a single-stranded portion of the tagRNA that comprises a region of complementarity to the PAM strand (i.e., the non-target strand or the edit strand). The FB sequence is complementary or substantially complementary to a sequence on the PAM strand of the double stranded target DNA that is immediately upstream of the nick site. In some embodiments, in the process of SyNTase editing, the tagRNA complexes with and directs a SyNTase editor to bind the search target sequence on the target strand of the double stranded targetDNA and the SyNTase editor generates a nick at the nick site on the non-target strand (e.g., the PAM strand) of the double stranded target DNA. In some embodiments, the FB sequence is complementary to or substantially complementary to, and can anneal to, a free 3' end on the nontarget strand of the double stranded target DNA at the nick site. In some embodiments, the FB sequence annealed to the free 3' end on the non-target strand can initiate target-primed DNA synthesis. In some embodiments, the FB sequence is about 2 to 20 nucleotides in length. In some embodiments, the FB sequence is about 8 to 16 nucleotides in length. In some embodiments, the FBS is 6 nucleotides in length.
[0168] The spacer of the tagRNA can be from 16 to 25 nucleotides in length. The spacer can be 20 nucleotides in length. The spacer can be 21-23 nucleotides in length. In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is from 15 nucleotides to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, or 20 to 30 nucleotides in length. In some embodiments, the spacer is 16 to 22 nucleotides in length. In some embodiments, the spacer is 16 to 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length.
[0169] As used herein in a gRNA (e.g., a tagRNA or a enhancer guide egRNA sequence), or fragments thereof such as a spacer, FB sequence, or editing template sequence, unless indicated otherwise, it should be appreciated that the letter “T” or “thymine” indicates a nucleobase in a DNA sequence that encodes the tagRNA or guide RNA sequence, and is intended to refer to a uracil (U) nucleobase of the tagRNA or guide RNA or any chemically modified uracil nucleobase known in the art, such as 5 -methoxyuracil.
[0170] The extension arm of a tagRNA may comprise a flap binding sequence (FB sequence; FBS) and an editing template (e.g., an ET). The extension arm may be partially complementary to the spacer. In some embodiments, the editing template (e.g., ET) is partially complementary to the spacer. In some embodiments, the editing template (e.g., ET) and the flap binding sequence (FB sequence; FBS) are each partially complementary to the spacer. An extension arm of a tagRNA may comprise a flap binding sequence (FB sequence, or FBS) that comprises complementarity to and can hybridize with a free 3’ end of a single stranded DNA in the target gene generated by nicking with an editor (e.g., a SyNTase editor) at the nick site on the PAM strand.
[0171] The length of the FB sequence may vary depending on, e.g., the SyNTase editor components, the search target sequence and other components of the tagRNA. The FB sequencecan be about 2 to 20 nucleotides in length. The FB sequence can be about 8 to 16 nucleotides in length. In some embodiments, the FB sequence is 6 nucleotides in length. In some embodiments, the FB sequence is about 3 to 19 nucleotides in length, or about 3 to 17 nucleotides in length. In some embodiments, the FB sequence is about 4 to 16 nucleotides, about 6 to 16 nucleotides, about 6 to 18 nucleotides, about 6 to 20 nucleotides, about 8 to 20 nucleotides, about 10 to 20 nucleotides, about 12 to 20 nucleotides, about 14 to 20 nucleotides, about 16 to 20 nucleotides, or about 18 to 20 nucleotides in length. In some embodiments, the FB sequence is 8 to 17 nucleotides in length. In some embodiments, the FB sequence is 8 to 16 nucleotides in length. In some embodiments, the FB sequence is 8 to 15 nucleotides in length. In some embodiments, the FB sequence is 8 to 14 nucleotides in length. In some embodiments, the FB sequence is 8 to 13 nucleotides in length. In some embodiments, the FB sequence is 8 to 12 nucleotides in length. In some embodiments, the FB sequence is 8 to 11 nucleotides in length. In some embodiments, the FB sequence is 8 to 10 nucleotides in length. In some embodiments, the FB sequence is 8 or 9 nucleotides in length. In some embodiments, the FB sequence is 16 or 17 nucleotides in length, in some embodiments, the FB sequence is 15 to 17 nucleotides in length. In some embodiments, the FB sequence is 14 to 17 nucleotides in length. In some embodiments, the FB sequence is 13 to 17 nucleotides in length. In some embodiments, the FB sequence is 12 to 17 nucleotides in length. In some embodiments, the FB sequence is 11 to 17 nucleotides in length. In some embodiments, the FB sequence is 10 to 17 nucleotides in length. In some embodiments, the FB sequence is 9 to 17 nucleotides in length. In some embodiments, the FB sequence is about 7 to 15 nucleotides in length. In some embodiments, the FB sequence is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 nucleotides in length. In some embodiments, the FB sequence is 8, 9, 10, 11, 12, 13, or 14 nucleotides in length.
[0172] The FB sequence may be complementary or substantially complementary to a DNA sequence in the edit strand of the target gene. By annealing with the edit strand at a free hydroxy group, e.g., a free 3’ end generated by SyNTase editor nicking activity, the FB sequence may initiate synthesis of a new single stranded DNA encoded by the editing template at the nick site. In some embodiments, the FB sequence is at least about 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementary to a region of the edit strand of the target gene. In some embodiments, the FB sequence is perfectly complementary, or 100% complementary, to a region of the edit strand of the target gene.
[0173] An extension arm of a tagRNA may comprise an editing template that serves as a DNA synthesis template for the DNA polymerase in a SyNTase editor during SyNTase editing. The length of an editing template may vary depending on, e.g., the SyNTase editor components, the search target sequence and other components of the tagRNA. In someembodiments, the editing template serves as a DNA synthesis template for a reverse transcriptase. The editing template can be about 4 to 30 nucleotides in length. The editing template can be about 10 to 30 nucleotides in length. The editing template can be 6 to 9 nucleotides in length. In some embodiments, the editing template is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.
[0174] In some embodiments, the editing template sequence is about 70%, 75%, 80%, 85%, 90%, 95%, or 99% complementary to the editing target sequence on the edit strand of the target gene. In some embodiments, the editing template sequence is complementary to the editing target sequence except at positions of the intended nucleotide edits to be incorporated into the target gene. In some embodiments, the editing template comprises a nucleotide sequence comprising about 85% to about 95% complementarity to (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values, complementarity to) an editing target sequence in the edit strand in the target gene. In some embodiments, the editing template comprises about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, or about 99% complementarity to an editing target sequence in the edit strand of the target gene.
[0175] An intended nucleotide edit or intended nucleotide edits in an editing template of a tagRNA may comprise various types of alterations as compared to the target gene sequence, in some embodiments, the nucleotide edit or edits is nucleotide substitution(s) as compared to the target gene sequence. In some embodiments, the nucleotide edit comprise deletion(s) as compared to the target gene sequence. In some embodiments, the nucleotide edit comprise an insertion as compared to the target gene sequence. In some embodiments, the editing template comprises one to ten intended nucleotide edits as compared to the target gene sequence, in some embodiments, the editing template comprises one or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises two or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises three or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises four or more, five or more, or six or more intended nucleotide edits as compared to the target gene sequence. In some embodiments, the editing template comprises two or more single nucleotide substitutions, insertions, deletions, or any combination thereof, as compared to the target gene sequence. In some embodiments, the editing template comprises three single nucleotide substitutions, insertions, deletions, or any combination thereof as compared to the target gene sequence. In some embodiments, the editingtemplate comprises four, five, or six single nucleotide substitutions, insertions, deletions, or any combination thereof, as compared to the target gene sequence. In some embodiments, a nucleotide substitution comprises an adenine (A)-to-thymine (T) substitution. In some embodiments, a nucleotide substitution comprises an A-to-guanine (G) substitution. In some embodiments, a nucleotide substitution comprises an A-to-cytosine (C) substitution. In some embodiments, a nucleotide substitution comprises a T-A substitution. In some embodiments, a nucleotide substitution comprises a T-G substitution, in some embodiments, a nucleotide substitution comprises a T-C substitution. In some embodiments, a nucleotide substitution comprises a G-to-A substitution. In some embodiments, a nucleotide substitution comprises a G-to-T substitution. In some embodiments, a nucleotide substitution comprises a G-to-C substitution. In some embodiments, a nucleotide substitution comprises a C-to-A substitution, in some embodiments, a nucleotide substitution comprises a C-to-T substitution. In some embodiments, a nucleotide substitution comprises a C-to-G substitution.
[0176] In some embodiments, a nucleotide insertion is at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 22 nucleotides, at least 24 nucleotides, at least 26 nucleotides, at least 28 nucleotides, at least 30 nucleotides, at least 32 nucleotides, at least 34 nucleotides, at least 36 nucleotides, at least 38 nucleotides, at least 40 nucleotides, at least 42 nucleotides, at least 44 nucleotides, at least 46 nucleotides, at least 48 nucleotides, or at least 50 nucleotides, in length. In some embodiments, a nucleotide insertion is from 1 to 2 nucleotides, from 1 to 3 nucleotides, from 1 to 4 nucleotides, from 1 to 5 nucleotides, form 2 to 5 nucleotides, from 3 to 5 nucleotides, from 3 to 6 nucleotides, from 3 to 8 nucleotides, from 4 to 9 nucleotides, from 5 to 10 nucleotides, from 6 to 11 nucleotides, from 7 to 12 nucleotides, from 8 to 13 nucleotides, from 9 to 14 nucleotides, from 10 to 15 nucleotides, from 11 to 16 nucleotides, from 12 to 17 nucleotides, from 13 to 18 nucleotides, from 14 to 19 nucleotides, from 15 to 20 nucleotides in length. In some embodiments, a nucleotide insertion is a single nucleotide insertion. In some embodiments, a nucleotide insertion comprises insertion of two nucleotides.
[0177] The editing template of a tagRNA may comprise one or more intended nucleotide edits, compared to the target gene to be edited. Position of the intended nucleotide edit(s) relevant to other components of the tagRNA, or to particular nucleotides (e.g., mutations) in the target gene may vary. In some embodiments, the nucleotide edit is in a region of the tagRNA corresponding to or homologous to the protospacer sequence. In some embodiments, thenucleotide edit is in a region of the tagRNA corresponding to a region of the target gene outside of the protospacer sequence.
[0178] In some embodiments, the position of a nucleotide edit incorporation in the target gene may be determined based on position of the protospacer adjacent motif (PAM). For instance, the intended nucleotide edit may be installed in a sequence corresponding to the protospacer adjacent motif (PAM) sequence. In some embodiments, a nucleotide edit in the editing template is at a position corresponding to the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit in the editing template is at a position corresponding to the 3’ most nucleotide of the PAM sequence, in some embodiments, position of an intended nucleotide edit in the editing template may be referred to by aligning the editing template with the partially complementary edit strand of the target gene, and referring to nucleotide positions on the editing strand where the intended nucleotide edit is incorporated, in some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream of the 5’ most nucleotide of the PAM sequence in the edit strand of the target gene. By 0 base pair upstream or downstream of a reference position, it is meant that the intended nucleotide is immediately upstream or downstream of the reference position. In some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to 16 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 3 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in is incorporated at a position corresponding to 4 nucleotides upstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit is incorporated at a position corresponding to 5 nucleotidesupstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, the nucleotide edit in the editing template is at a position corresponding to 6 nucleotides upstream of the 5’ most nucleotide of the PAM sequence.
[0179] In some embodiments, an intended nucleotide edit is incorporated at a position corresponding to about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides downstream of the 5’ most nucleotide of the PAM sequence in the edit strand of the target gene. In some embodiments, a nucleotide edit is incorporated at a position corresponding to about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to 16 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 3 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 4 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 5 nucleotides downstream of the 5’ most nucleotide of the PAM sequence. In some embodiments, a nucleotide edit is incorporated at a position corresponding to 6 nucleotides downstream of the 5’ most nucleotide of the PAM sequence.
[0180] In some embodiments, the position of a nucleotide edit incorporation in the target gene can be determined based on position of the nick site. In some embodiments, position ofan intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides apart from the nick site. In some embodiments, position of an intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39,40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides downstream of the nick site on the PAM strand (or the non-target strand, or the edit strand) of the double stranded target DNA. In some embodiments, position of the intended nucleotide edit in the editing template can be referred to by aligning the editing template with the partially complementary editing target sequence on the edit strand and referring to nucleotide positions on the editing strand where the intended nucleotide edit is incorporated. Accordingly, in some embodiments, a nucleotide edit in an editing template is at a position corresponding to a position about 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, or 150 nucleotides apart from the nick site. In some embodiments, a nucleotide edit in an editing template is at a position corresponding to a position about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, 40 to 50 nucleotides, 50 to 60 nucleotides, 60 to 70 nucleotides, 70 to 80 nucleotides, 80 to 90 nucleotides, 90 to 100 nucleotides, 100 to 110 nucleotides, 110 to 120 nucleotides, 120 to 130 nucleotides, 130 to 140 nucleotides, or 140 to 150 nucleotides apart from the nick site. In some embodiments, when referred to in the context of the PAM strand (or the non-target strand, or the edit strand), a nucleotide edit in an editing template is at a position corresponding to a position about 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 to 16 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides,12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, 20 to 30 nucleotides, 30 to 40 nucleotides, 40 to 50 nucleotides, 50 to 60 nucleotides, 60 to 70 nucleotides, 70 to 80 nucleotides, 80 to 90 nucleotides, 90 to 100 nucleotides, 100 to 110 nucleotides, 110 to 120 nucleotides, 120 to 130 nucleotides, 130 to 140 nucleotides, or 140 to 150 nucleotides downstream from the nick site. The relative positions of the intended nucleotide edit(s) and nick site may be referred to by numbers. For example, in some embodiments, the nucleotide immediately downstream of the nick site on a PAM strand (or the non-target strand, or the edit strand) may be referred to as at position 0. The nucleotide immediately upstream of the nick site on the PAM strand (or the non-target strand, or the edit strand) may be referred to as at position -1. The nucleotides downstream of position 0 on the PAM strand can be referred to as at positions +1, +2, +3, +4,... +n, and the nucleotides upstream of position -1 on the PAM strand may be referred to as at positions -2, -3, -4,.. -n. Accordingly, in some embodiments, the nucleotide in the editing template that corresponds to position 0 when the editing template is aligned with the partially complementary editing target sequence by complementarity can also be referred to as position 0 in the editing template, the nucleotides in the editing template corresponding to the nucleotides at positions +1, +2, +3, +4,..., +n on the PAM strand of the double stranded target DNA can also be referred to as at positions +1, +2, +3, +4,..., +n in the editing template, and the nucleotides in the editing template corresponding to the nucleotides at positions -1, -2, -3, -4,..., -n on the PAM strand on the double stranded target DNA may also be referred to as at positions -1, -2, -3, -4 -n on the editing template, even though when the tagRNA is viewed as a standalone nucleic acid, positions +1, +2, +3, +4,..., +n are 5' of position 0 and positions -1, -2, -3, -4,...-n are 3' of position 0 in the editing template. In some embodiments, an intended nucleotide edit is at position +n of the editing template relative to position 0. Accordingly, the intended nucleotide edit may be incorporated at position +n of the PAM strand of the double stranded target DNA (and subsequently, the target strand of the double stranded target DNA) by SyNTase editing. The number n may be referred to as the nick to edit distance.
[0181] When referred to within the tagRNA, positions of the one or more intended nucleotide edits may be referred to relevant to components of the tagRNA. For example, an intended nucleotide edit may be 5’ or 3’ to the FB sequence. In some embodiments, a tagRNA comprises the structure, from 5’ to 3’: a spacer, a gRNA core, an editing template, and a FB sequence. In some embodiments, the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11,12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides upstream to the 5’ most nucleotide of the FB sequence, in some embodiments, the intended nucleotide edit is 0 to 2 nucleotides, 0 to 4 nucleotides, 0 to 6 nucleotides, 0 to 8 nucleotides, 0 to 10 nucleotides, 2 to 4 nucleotides, 2 to 6 nucleotides, 2 to 8 nucleotides, 2 to 10 nucleotides, 2 to 12 nucleotides, 4 to 6 nucleotides, 4 to 8 nucleotides, 4 to 10 nucleotides, 4 to 12 nucleotides, 4 to 14 nucleotides, 6 to 8 nucleotides, 6 to 10 nucleotides, 6 to 12 nucleotides, 6 to 14 nucleotides, 6 tol6 nucleotides, 8 to 10 nucleotides, 8 to 12 nucleotides, 8 to 14 nucleotides, 8 to 16 nucleotides, 8 to 18 nucleotides, 10 to 12 nucleotides, 10 to 14 nucleotides, 10 to 16 nucleotides, 10 to 18 nucleotides, 10 to 20 nucleotides, 12 to 14 nucleotides, 12 to 16 nucleotides, 12 to 18 nucleotides, 12 to 20 nucleotides, 12 to 22 nucleotides, 14 to 16 nucleotides, 14 to 18 nucleotides, 14 to 20 nucleotides, 14 to 22 nucleotides, 14 to 24 nucleotides, 16 to 18 nucleotides, 16 to 20 nucleotides, 16 to 22 nucleotides, 16 to 24 nucleotides, 16 to 26 nucleotides, 18 to 20 nucleotides, 18 to 22 nucleotides, 18 to 24 nucleotides, 18 to 26 nucleotides, 18 to 28 nucleotides, 20 to 22 nucleotides, 20 to 24 nucleotides, 20 to 26 nucleotides, 20 to 28 nucleotides, or 20 to 30 nucleotides upstream to the 5’ most nucleotide of the FB sequence.
[0182] The corresponding positions of the intended nucleotide edit incorporated in the target gene may also be referred to based on the nicking position (i.e., the nick site) generated by a SyNTase editor based on sequence homology and complementarity. For example, in some embodiments, the distance between the intended nucleotide edit to be incorporated into the target gene and the nick site (also referred to as the “nick to edit distance”) may be determined by the position of the nick site and the position of the nucleotide(s) corresponding to the intended nucleotide edit(s), for example, by identifying sequence complementarity between the spacer and the search target sequence and sequence complementarity between the editing template and the editing target sequence. In some embodiments, the position of the nucleotide edit can be in any position downstream of the nick site on the edit strand (or the PAM strand) generated by the SyNTase editor, such that the distance between the nick site and the intended nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. In some embodiments, the position of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides upstream of the nick site on the edit strand. In some embodiments, the position of the nucleotide edit is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides downstream of the nick site on the edit strand. In some embodiments, the position of the nucleotide edit is 0 base pair from the nick site on the edit strand, that is, the editing position is at the same position as the nick site. As used herein, the distance between the nick site and the nucleotide edit, for example, where the nucleotide edit comprises aninsertion or deletion, refers to the 5’ most position of the nucleotide edit for a nick that creates a 3’ free end on the edit strand (i.e., the “near position” of the nucleotide edit to the nick site). Similarly, as used herein, the distance between the nick site and a PAM position edit, for example, where the nucleotide edit comprises an insertion, deletion, or substitution of two or more contiguous nucleotides, refers to the 5’ most position of the nucleotide edit and the 5’ most position of the PAM sequence.
[0183] In some embodiments, the editing template extends beyond a nucleotide edit to be incorporated into the target gene sequence. For example, in some embodiments, the editing template comprises at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, or 80 nucleotides 3’ to the nucleotide edit to be incorporated into the target gene sequence. In some embodiments, the editing template comprises 1 to 80 nucleotides 3’ to the nucleotide edit to be incorporated into the target gene sequence.
[0184] In some embodiments, the editing template can comprise a second editing sequence comprising a second mutation relative to a target sequence. The second mutation can be designed to mutate or otherwise silence a PAM sequence such that a corresponding nucleic acid guided nuclease or CRISPR nuclease is no longer able to cleave the target sequence. In some embodiments, this mutation or silencing of a PAM can serve as a method for selecting transformants in which the first editing sequence has been incorporated. In some embodiments, the mutation is in at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleic acids in a PAM motif.
[0185] The editing template of a tagRNA may encode a new single stranded DNA (e.g., by reverse transcription) to replace a editing target sequence in the target gene. In some embodiments, the editing target sequence in the edit strand of the target gene is replaced by the newly synthesized strand, and the nucleotide edit(s) are incorporated into the region of the target gene.
[0186] The tagRNAs may be modified in one or more ways to improve their overall stability and / or performance in editing methods of the disclosure (e.g., SyNTase editing). The tagRNA can comprise 3’ mN*mN*mN*N and 5’ mN*mN*mN* modifications, where m indicates that the nucleotide contains a 2’-0-Me modification and a * indicates the presence of a phosphorothioate bond. The tagRNA can comprise a structural motif at the 3’ terminus selected from the group consisting of: an inverted-dT, a prequeosinel-1 riboswitch aptamer (evopreQl) and variants thereof, a frameshifting pseudoknot from Moloney murine leukemia virus (MMLV) (mpknot), G-quadruplexes, hairpin structures, xrRNA, and a P4-P6 domain of the group I intron. In some embodiments, the structural motif is evopreQl or a variant thereof comprising anucleotide sequence selected from SEQ ID NOs: 24-30. In some embodiments, the structural motif is evopreQl or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 24-30 or a sequence having one, two or three mismatches relative a sequence selected from SEQ ID NOs: 24-30. In some embodiments, the structural motif is evopreQl or a variant thereof comprising a nucleotide sequence selected from SEQ ID NOs: 24-30 or a sequence at least 85% identical (85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) to a nucleotide sequence selected from SEQ ID NOs: 24-30. In some embodiments, the evopreQl is trimmed (T evopreQl).
[0187] In some embodiments, appending one or more RNA structural motifs to a tagRNA can protect against degradation of the tagRNA. Such RNA structural motifs can include, but are not limited to (i) a prequeosinel-1 riboswitch aptamer (evopreQl) and variants thereof, (ii) a frameshifting pseudoknot from Moloney murine leukemia virus (MMLV), hereafter referred to as “mpknot,” and variants thereof (iii) G-quadruplexes, (iv) hairpin structures (e.g., 15-bp hairpins), (v) xrRNA, and (vi) a P4-P6 domain of the group I intron.
[0188] The present disclosure provides modified tagRNAs with improved properties, including but not limited to, increased stability and cellular lifespan, and improved binding affinity for an editor (e.g., SyNTase editor, RT editor). These modified tagRNAs result in improved genome editing as demonstrated by increase editing efficiency at a wide variety of genomic sites. In some embodiments, by appending certain nucleic acid structural motifs to terminus of the extension arm of a tagRNA, including but limited to, a prequeosim-1 riboswitch aptamer (“evopreQi-1”) or variant thereof, a pseudoknot from the MMLV viral genome (“evopreQl- 1”) or variant thereof, a modified tRNA used by MMLV RT as a primer for reverse transcription or variant thereof, and a G quadruplex or variant thereof, a consistent increase in editing activity can be achieved. In some embodiments, the 3’ terminus of the tagRNA comprises a evopreQl aptamer or variant thereof, comprising or consisting of a sequence of any one of: TTGACGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 24), CGCGAGTCTAGGGGATAACGCGTTAAACTTCCTAGAAGGCGGTT (SEQ ID NO: 25), CGCGGATCTAGATTGTAACGCGTTAAACCATCTAGAAGGCGGTT (SEQ ID NO: 26), CGCGTCGCTACCGCCCGGCGCGTTAAACACACTAGAAGGCGGTT (SEQ ID NO: 27), CGCGGTTCTATCTAGTTACGCGTTAAACCAACTAGAA (SEQ ID NO: 28), TTGACGCGCTTCTATCTAGTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 29), and TTGACGCGGTTCTATCTACTTACGCGTTAAACCAACTAGAAA (SEQ ID NO: 30). For RNA sequences, T can be U.
[0189] In one embodiment, the modified tagRNAs include a nucleic acid moiety at the 3' end of the tagRNA. Optionally, the 3' end of the tagRNA is fused to the nucleic acid moietythrough a nucleotide linker. In various embodiments, it will be appreciated that a wide variety of nucleotide sequences will work reasonably well for each genomic target site. Linker length can also be variable. In some cases, linkers ranging in length from 3-18 nucleotides can be used. In other cases, the linker may be at least 3 nucleotides, at least 4 nucleotides, at least 5 nucleotides, at least 6 nucleotides, at least 7 nucleotides, at least 8 nucleotides, at least 9 nucleotides, at least 10 nucleotides, at least 11 nucleotides, at least 12 nucleotides, at least 13 nucleotides, at least 14 nucleotides, at least 15 nucleotides, at least 16 nucleotides, at least 17 nucleotides, at least 18 nucleotides, at least 19 nucleotides, at least 20 nucleotides, at least 21 nucleotides, at least 22 nucleotides, at least 23 nucleotides, at least 24 nucleotides, at least 25 nucleotides, at least 26 nucleotides, at least 27 nucleotides, at least 28 nucleotides, at least 29 nucleotides, or at least 30 nucleotides.
[0190] In general, the nucleic acid moieties that may be used to modify a tagRNA, for example, by attaching it to the 3' end of a tagRNA, may include any nucleic acid moiety, including, for instance, a nucleic acid molecule comprising or which forms a double-helix moiety, toeloop moiety, hairpin moiety, stem-loop moiety, pseudoknot moiety, aptamer moiety, G quadraplex moiety, tRNA moiety, or a ribozyme moiety. The nucleic acid moiety may be characterized as forming a secondary nucleic acid structure, a tertiary nucleic acid structure, or a quadruple nucleic acid structure. In other words, the nucleic acid moiety may form any two dimensional or three dimensional structure known to be formed by such structures. The nucleic acid moiety may be DNA or RNA.
[0191] In some embodiments, a tagRNA or an enhancer guide RNA can be chemically synthesized, or can be assembled or cloned and transcribed from a DNA sequence, e.g., a plasmid DNA sequence, or by any RNA oligonucleotide synthesis method known in the art. In some embodiments, DNA sequence that encodes a tagRNA (or enhancer guide RNA (egRNA)) can be designed to append one or more nucleotides at the 5' end or the 3' end of the tagRNA (or egRNA) encoding sequence to enhance tagRNA transcription. For example, in some embodiments, a DNA sequence that encodes a tagRNA (or an egRNA) can be designed to append a nucleotide G at the 5' end. Accordingly, in some embodiments, the tagRNA (or egRNA) can comprise an appended nucleotide G at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or egRNA) can be designed to append a sequence that enhances transcription, e.g., a Kozak sequence, at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or egRNA) can be designed to append the sequence CACC or CCACC at the 5' end. Accordingly, in some embodiments, the tagRNA (or egRNA) can comprise an appended sequence CACC or CCACC at the 5' end. In some embodiments, a DNA sequence that encodes a tagRNA (or egRNA) can be designed to append the sequence TTT, TTTT, TTTTT, TTTTTT, TTTTTTT at 3' end.Accordingly, in some embodiments, the tagRNA (or egRNA) can comprise an appended sequence UUU, UUUU, UUUUU, UUUUUU, or UUUUUUU at the 3' end.
[0192] Disclosed herein include, e.g., SyNTase editing systems. In some embodiments, the SyNTase editing system comprises: any of the tagRNAs disclosed herein, or a nucleic acid encoding the tagRNA; and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor. In some embodiments, the intended nucleotide edit incorporation rate of the SyNTase editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).
[0193] Disclosed herein include editing complexes (e.g., SyNTase editing complexes). In some embodiments, the SyNTase editing complex comprises: (i) any of the tagRNA disclosed herein and a SyNTase editor comprising a DNA endonuclease domain and a DNA polymerase domain; or (ii) any of the SyNTase editing systems of the disclosure. In some embodiments, the intended nucleotide edit incorporation rate of the SyNTase editing complex is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).Enhancer gRNA (egRNA)
[0194] Disclosed herein include editing systems (e.g., SyNTase editing, RT editing systems). In some embodiments, the SyNTase editing system comprises: any of the tagRNAs disclosed herein, or a nucleic acid encoding the tagRNA; and an SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor. The SyNTase editing system can comprise: an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises an egRNA spacer that is complementary to a second search target sequence in the target gene.
[0195] In some embodiments, an SyNTase editing system or composition further comprises an enhancer guide polynucleotide, such as an enhancer guide RNA (egRNA). Withoutwishing to be bound by any particular theory, the non-edit strand of a double stranded target DNA in the target gene may be nicked by a CRISPR-Cas nickase directed by an egRNA. In some embodiments, the nick on the non-edit strand directs endogenous DNA repair machinery to use the edit strand as a template for repair of the non-edit strand, which may increase efficiency of SyNTase editing. In some embodiments, the non-edit strand is nicked by an SyNTase editor localized to the non-edit strand by the egRNA. Accordingly, also provided herein are tagRNA systems comprising at least one tagRNA and at least one egRNA.
[0196] In some embodiments, the egRNA is a guide RNA which contains a variable spacer sequence and a guide RNA scaffold or core region that interacts with the DNA binding domain, a DNA endonuclease domain, e.g., Cas9 of the SyNTase editor, in some embodiments, the egRNA comprises a spacer sequence (referred to herein as an ng spacer, or a second spacer) that is substantially complementary to a second search target sequence (or ng search target sequence), which is located on the edit strand, or the non-target strand. Thus, in some embodiments, the egRNA search target sequence recognized by the egRNA spacer and the search target sequence recognized by the spacer sequence of the tagRNA are on opposite strands of the double stranded target DNA of target gene, e.g., the ATP7B gene, PAH gene, or SERPINA1 gene. In some embodiments, an egRNA spacer sequence is complementary to, and may hybridize with the second search target sequence only after an intended nucleotide edit has been incorporated on the edit strand, by the editing template of a tagRNA.
[0197] In some embodiments, the egRNA search target sequence is located on the nontarget strand, within 10 base pairs to 100 base pairs of an intended nucleotide edit incorporated by the tagRNA on the edit strand, in some embodiments, the egRNA target search target sequence is within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp of an intended nucleotide edit incorporated by the tagRNA on the edit strand. In some embodiments, the 5’ ends of the egRNA search target sequence and the tagRNA search target sequence are within 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp apart from each other. In some embodiments, the 5’ ends of the egRNA search target sequence and the tagRNA search target sequence are within 10 bp, 20 bp, 30 bp, 40 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 91 bp, 92 bp, 93 bp, 94 bp, 95 bp, 96 bp, 97 bp, 98 bp, 99 bp, or 100 bp apart from each other.
[0198] In some embodiments, an egRNA spacer sequence is complementary to, and may hybridize with the second search target sequence only after an intended nucleotide edit has been incorporated on the edit strand, by the editing template of a tagRNA. In some embodiments, the egRNA comprises a spacer sequence that matches only the edit strand after incorporation of the nucleotide edits, but not the endogenous target gene sequence on the edit strand. Accordingly,in some embodiments, an intended nucleotide edit is incorporated within the egRNA search target sequence. In some embodiments, the intended nucleotide edit is incorporated within about 1-10 nucleotides of the position corresponding to the PAM of the egRNA search target sequence.
[0199] The editing system (e.g., SyNTase editing, RT editing systems) can comprise: an egRNA, or a nucleic acid encoding the egRNA, wherein the egRNA comprises an egRNA spacer that is complementary to a second search target sequence in the target gene. The second search target sequence can be on the second strand of the target gene. The egRNA spacer can be from 16 to 22 nucleotides in length. In some embodiments, the length of the spacer varies from about 10 to about 100 nucleotides. In some embodiments, the spacer is 16 nucleotides, 17 nucleotides, 18 nucleotides, 19 nucleotides, 20 nucleotides, 21 nucleotides, 22 nucleotides, 23 nucleotides, 24 nucleotides, or 25 nucleotides in length. In some embodiments, the spacer is from 15 nucleotides to 30 nucleotides in length, 15 to 25 nucleotides in length, 18 to 22 nucleotides in length, 10 to 20 nucleotides in length, or 20 to 30 nucleotides in length. In some embodiments, the spacer is 16 to 22 nucleotides in length. In some embodiments, the spacer is 16 to 20 nucleotides in length. In some embodiments, the spacer is 17 to 18 nucleotides in length. In some embodiments, the egRNA spacer is 20 nucleotides in length. In some embodiments, the egRNA comprises a scaffold sequence.
[0200] In some embodiments, the intended nucleotide edit incorporation rate of the SyNTase editing system is greater than at least about 30%, about 40%, about 50%, about 60%, about 70%, or about 80% (e.g., about 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values).
[0201] The Tables below provide exemplary sequences related to the compositions and methods of the disclosure.Table 2: Exemplary Editor Linker SequencesSEQIDName Shorthand Sequence AA ResiduesNO GS3 (GGGGS)3 GGGGSGGGGSGGGGS 31 GS4 (GGGGS)4 GGGGSGGGGSGGGGSGGGGS 32 GS5 (GGGGS)5 GGGGSGGGGSGGGGSGGGGSGGGGS 33 GS6 (GGGGS)6 GGGGSGGGGSGGGGSGGGGSGGGGSGGGGS 34 GS7 (GGGGS)7 GGGGSGGGGSGGGGSGGGGSGGGGSGGGGSGGGGS 35 EAK3 (EAAAK)3 EAAAKEAAAKEAAAK 36 EAK4 (EAAAK)4 EAAAKEAAAKEAAAKEAAAK 37 EAK5 (EAAAK)5 EAAAKEAAAKEAAAKEAAAKEAAAK 38 EAK6 (EAAAK)6 EAAAKEAAAKEAAAKEAAAKEAAAKEAAAK 39EAK7 (EAAAK)7 EAAAKEAAAKEAAAKEAAAKEAAAKEAAAKEAAAK 40AEAK4- A(EAAAK)4ALEA(EA AEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAKEA 41 ALEA AAK)4A AAKEAAAKAGS- GSA(EAAAK)4ALEA( GSAEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAAAK 42 AEAK4- EAAAK)4ASG EAAAKEAAAKASGALEA GGS- GGGGS(EAAAK)4AL GGGGSEAAAKEAAAKEAAAKEAAAKALEAEAAAKEAA 43 AEAK4- EA AKEAAAKEAAAKSGGGGALEA (EAAAAK)4SGGGGAEAK AEAAAKEAAAKA AEAAAKEAAAKA 44 AEAK2 (AEAAAKEAAAKA)2 AEAAAKEAAAKA AEAAAKEAAAKA 45 AEAK3 (AEAAAKEAAAKA)3 AEAAAKEAAAKAAEAAAKEAAAKAAEAAAKEAAAKA 46 GGS- GGGGS(EAAAK)3GG GGGGSEAAAKEAAAKEAAAKGGGGS 47 EAK3- GGSSGG GGS- GGGGS(EAAAK)4GG GGGGSEAAAKEAAAKEAAAKEAAAKGGGGS 48 EAK4- GGSSGG GGS- GGGGS(EAAAK)5GG GGGGSEAAAKEAAAKEAAAKEAAAKEAAAKGGGGS 49 EAK5- GGSSGG GGS- GGGGS(EAAAK)3GG GGGGSEAAAKEAAAKEAAAKGGSGGSGEAAAKEAAAK 50 EAK3- SGGSG(EAAAK)3GG EAAAKGGGGSGGS GGS GSA-F SGAGSAAGSGEFGS GSAGSAAGSGEF 51 GSA-F-flip SGAGSAAGSGEFEG GSAGSAAGSGEFEGSGSAASGASG 52SGSAASGAGSK19D KESGSVSSEQLAQFR KESGSVSSEQLAQFRSLD 53SLDE14T EGKSSGSGSESKST EGKSSGSGSESKST 54 PAPAP PAPAP PAPAP 55 PAPAP-2 (PAPAP)2 PAPAPPAPAP 56 PAPAP-3 (PAPAP)3 PAPAPPAPAPPAPAP 57 PAPAP-4 (PAPAP)4 PAPAPPAPAPPAPAPPAPAP 58 PAPAP-5 (PAPAP)5 PAPAPPAPAPPAPAPPAPAPPAPAP 59 PAPAP-6 (PAPAP)6 PAPAPPAPAPPAPAPPAPAPPAPAPPAPAP 60 PAPAP-7 (PAPAP)7 PAPAPPAPAPPAPAPPAPAPPAPAPPAPAPPAPAP 61 PAPAP-8 (PAPAP)8 PAPAPPAPAPPAPAPPAPAPPAPAPPAPAPPAPAPPAPAP 62 GS- SGGGGSSGSETPGTSESATPESSGGGGS 63 XTEN16- GS SGGSx2- SGGSSGGSSGSETPGTSESATPESSGGSSGGS 18 XTEN16- SGGSx2GGS- SGGGGSAEAAAKEAAAKEAAAKEAAAKALEAEAAAKE 64 AEAK- AAAKEAAAKEAAAKASGGGGSALEA modl-gsGS-NLS- SGGSSGGSKRTADGSEFESPKKKRKVSGGSSGGS 65GS 4-gsTable 3A: Exemplary Cas9 SequencesDESCRIPTION SEQ ID NOSschlCas9 N583A 66SrolCas9 N584A 67Sha4Cas9 N583A 68SsuCas9 N622A 69evoCjCas9 H559A 70iSpyMac H840A 71Ssi5Cas9 N580A 72Ssi8Cas9 N580A 73Ssci4Cas9 N580A 74ShylCas9 N582A 75Sag3Cas9 N582A 76SlutrlCas9 N583A 77Ssch3Cas9 N583A 78SpRYCas9 H840A 79SpRYcCas9 H849A 80Sma2Cas9 N622A 81SsaCas9 N628A 82EvoSluCas9 N582A 83SluCas9 N582A 84EvosRGN3.1 N585A 85EvoNme2Cas9 H588A 86EvoAnaCas9 H582A 87SpCas9 R221K N394K H840A 133Table 3B: Exemplary Cas9 SequencesNAME SPECIES (native or Length SEQ derived) ID NO StlCas9-sim3_MetCas9_MBW4258562.1_Methanobacterium-sp- Methanobacterium sp 1096 88 YSL 01SsaCas9-siml ARC34060.1 Streptococcus-equinus Streptococcus equinus 1128 89 SsaCas9-siml_SsaCas9_AR00_CCB95040.1_Streptococcus- Streptococcus salivarius 1127 90 salivarius-JIM8777 01SsaCas9-siml_SsaCas9_AR01_ancestral-seq-recon_Streptococcus- Streptococcus salivarius 1127 91 salivarius 01SsaCas9-siml_SsaCas9_AR02_ancestral-seq-recon_Streptococcus- Streptococcus salivarius 1127 92 salivarius 01SsaCas9-siml WP 003097512.1 Streptococcus-vestibularis Streptococcus vestibularis 1128 93 SsaCas9-siml WP 014634195.1 Streptococcus-salivarius Streptococcus salivarius 1127 94 SsaCas9-siml WP 037611471.1 Streptococcus-sp-SR4 Streptococcus sp SR4 1128 95 SsaCas9-siml WP 048787825.1 Streptococcus Streptococcus 1122 96 SsaCas9-siml WP 048790568.1 Streptococcus-salivarius Streptococcus salivarius 1139 97 SsaCas9-siml WP 064519118.1 Streptococcus-vestibularis Streptococcus vestibularis 1128 98 SsaCas9-siml WP 145516278.1 Streptococcus-salivarius Streptococcus salivarius 1121 99 SsaCas9-siml WP 227278337.1 Streptococcus-vestibularis Streptococcus vestibularis 1122 100 SsaCas9-sim2 clustered ARC34060.1 Streptococcus-equinus Streptococcus equinus 1128 101 SsaCas9-sim2 clustered HEM3702072.1 Streptococcus-suis Streptococcus suis 1122 102 SsaCas9-sim2 clustered HEM4744366.1 Streptococcus-suis Streptococcus suis 1122 103 SsaCas9-sim2 clustered MDU6911040.1 Streptococcus-salivarius Streptococcus salivarius 1127 104 SsaCas9-sim2 clustered WP 002299171.1 Streptococcus-mutans Streptococcus mutans 1125 105 SsaCas9-sim2_clustered_WP_021002704.1 Streptococcus- Streptococcus intermedius 1125 106 intermediusSsaCas9-sim2 clustered WP 084910940.1 Streptococcus-oralis Streptococcus oralis 1129 107 SsaCas9-sim2_clustered_WP_115265868.1 Streptococcus- Streptococcus 1130 108 macedonicus macedonicusSsaCas9-sim2 clustered WP 126440388.1 Streptococcus-equinus Streptococcus equinus 1130 109 SsaCas9-sim2_clustered_WP_155974233.1 Streptococcus- Streptococcus 1125 110 ruminantium ruminantiumSsaCas9-sim2 clustered WP 198459777.1 Streptococcus-oralis Streptococcus oralis 1122 111 SsaCas9-sim2 clustered WP 219967735.1 Streptococcus-gordonii Streptococcus gordonii 1136 112 SsaCas9-sim2_clustered_WP_223322622.1_Streptococcus- Streptococcus sanguinis 1121 113 sanguinisSsaCas9-sim2 clustered WP 257737268.1 Streptococcus-uberis Streptococcus uberis 1122 114 SsaCas9-sim2 clustered WP 268734931.1 Streptococcus-cristatus Streptococcus cristatus 1127 115 SsaCas9-sim2_clustered_WP_270615522.1_Streptococcus- Streptococcus koreensis 1123 116koreensisSsaCas9-sim2_clustered_WP_281734300.1 Streptococcus- Streptococcus lutetiensis 1123 117 lutetiensisSsaCas9-sim2_clustered_WP_311154762.1 uncultured- Streptococcus sp 1123 118 Streptococcus-spSsaCas9-sim3_Structure-Ssa-Streptococcus- Streptococcus cristatus 1130 119 cristatus A0A3R9LSK4 STRCR Streptococcus-cristatus 01SsaCas9-sim3_structure-Ssa-Streptococcus- Streptococcus massiliensis 1126 120 massiliensis A0A380KY87 9STRE Streptococcus-massiliensis 01SsaCas9-sim3_structure-Ssa-Streptococcus- Streptococcus vestibularis 1128 121 vestibularis A0A3S4QG93 Streptococcus-vestibularis 01SsaCas9-sim3_structure-Stl_A0A7X2QGI6_Streptococcus- Streptococcus uberis 1122 122 uberis 01SsaCas9-sim3_structure-Stl-Enterococcus- Enterococcus asini 1123 123 asini A0A413AJC9 9ENTE Enterococcus-asini 01SsaCas9-sim3 stracture-Stl -Enterococcus-asini MCD5030292.1 Enterococcus asini 1123 124 SsaCas9-sim3 structure-Stl -Enterococcus-asini RGW10367.1 Enterococcus asini 1123 125 SsaCas9-sim3 structure-Stl-Enterococcus-asini WP 231453593.1 Enterococcus asini 1130 126 SsaCas9-sim3 structure-Stl-Enterococcus-asini WP 259305198.1 Enterococcus asini 1130 127 SsaCas9-sim3 structure-Stl-Enterococcus-asini WP 303220258.1 Enterococcus asini 1131 128 SsaCas9-sim3 structure-Stl-Enterococcus-asini WP 311832475.1 Enterococcus asini 1130 129 SsaCas9-sim3 structure-Stl-Enterococcus-asini WP 311835806.1 Enterococcus asini 1130 130 SsaCas9-sim3 structure-Stl-Enterococcus-asini WP 311877876.1 Enterococcus asini 1130 131SsaCas9-sim3 structure-Stl -Enterococcus-asini WP 311896923.1 Enterococcus asini 1130 132Table 4 A: Exemplary mRNA Sequences Related to Editors DESCRIPTION SEQ ID NO1SegPolyA 1 134SegPolyA 2 135SegPolyA 3 136SegPolyA 4 137120 PolyA 138130 PolyA 1393'UTR 140WPRE 141ek5 142STH 143STH-2 1445' UTR 145Exemplary editor mRNA sequence 1463'UTR-WPRE-SegPolyA 1 1473'UTR-WPRE-SegPolyA 2 1483'UTR-WPRE-SegPolyA 3 1493'UTR-WPRE-SegPolyA 4 1503'UTR-WPRE-120 PolyA 1513'UTR-WPRE-130 PolyA 1523'UTR-STH 1533'UTR-eK5-120 PolyA 1543'UTR-eK5-130 PolyA 1553'UTR-eK5-SegPolyA 1 1563'UTR-eK5-SegPolyA 2 1573'UTR-eK5-SegPolyA 3 1583'UTR-eK5-SegPolyA 4 159eK5-120 PolyA 160eK5-130 PolyA 161eK5-SegPolyA 1 162eK5-SegPolyA 2 163eK5-SegPolyA 3 164eK5-SegPolyA 4 1653'UTR-STH-120 PolyA 1663'UTR-STH-130 PolyA 1673'UTR-STH-SegPolyA 1 1683'UTR-STH-SegPolyA 2 1693'UTR-STH-SegPolyA 3 1703'UTR-STH-SegPolyA 4 171Comp 14 transcript 172Full-length segment from humanMALAT1 (HTH) 173'For RNA, T is U; U are T in the sequence listing.Table 4B: Exemplary mRNA ElementsName Element Sequence SEQ Position ID NO SAA1 5UTR AGGGACCCGCAGCTCAGCTACAGCACAGATCAGCACC 174 APOC3 5UTR CTGCTCAGTTCATCCCTAGAGGCAGCTGCTCCAGGAACAGAGGTGCC 175 ALB 5UTR CTAGCTTTTCTCTTCTGTCAACCCCACACGCCTTTGGCACA 176 HP 5UTR ACTGGAAAAGATAGTGACCTTACCAGGGCCAAAGTTTGTAGACACAGG 177AATTACGAAATGGAGAAGGGGGAGAAGTGAGCTAGTGGCAGCATAAA AAGACCAGCAGATGCCCCACAGCACTGCTCTTCCAGAGGCAAGACCAA CCAAG ORM1 5UTR AGCACTGCCTGGCTCCACGTGCCTCCTGGTCTCAGT 178 APOA2 5UTR AGGCACAGACACCAAGGACAGAGACGCTGGCTAGGCCGCCCTCCCCA 179CTGTTACCAAC SAA2 5UTR ACTATAAATAGCAGCCACCTCTCCCTGGCAGACAGGGACCCGCAGCTC 180AGCTACAGCACAGATCAGCACC SERPINA1 5UTR CTCCTCAGCTTCAGGCACCACCACTGACCTGGGACAGTGAATCGACA 181 APOCI 5UTR AGGCGGTCAGGGGAAGGCTCAGGAGGAGGGAGATCAACATCAACCTG 182CCCCGCCCCCTCCCCAGCCTGATAAAGGTCCTGCGGGCAGGACAGGAC CTCCCAACCAAGCCCTCCAGCAAGGATTCAGAGTGCCCCTCCGGCCTC GCCAptl7_ 5UTR ACTCACTATTTGTTTTCGCGCCCAGTTGCAAAAAGTGTCG 183 customcustom 5UTR AGGATAATATACTTACATACTTACTAATTAATACTAAACTCAACGCCA 184CCSyn_5'UTR 5UTR AGGAAATAAGAGAGAAAAGAAGAGTAAGAAGAAATATAAGAGCCAC 185C SAA1 3UTR GCTTCCTCTTCACTCTGCTCTCAGGAGATCTGGCTGTGAGGCCCTCAGG 186GCAGGGATACAAAGCGGGGAGAGGGTACACAATGGGTATCTAATAAA TACTTAAGAGGTGGAA APOC3 3UTR GACCTCAATACCCCAAGTCCACCTGCCTATCCATCCTGCGAGCTCCTTG 187GGTCCTGCAATCTCCAGGGCTGCCCCTGTAGGTTGCTTAAAAGGGACA GTATTCTCAGTGCTCTCCTACCCCACCTCATGCCTGGCCCCCCTCCAGG CATGCTGGCCTCCCAATAAAGCTGGACAAGAAGCTGCTATGA ALB 3UTR CATCACATTTAAAAGCATCTCAGCCTACCATGAGAATAAGAGAAAGAA 188AATGAAGATCAAAAGCTTATTCATCTGTTTTTCTTTTTCGTTGGTGTAA AGCCAACACCCTGTCTAAAAAACATAAATTTCTTTAATCATTTTGCCTC TTTTCTCTGTGCTTCAATTAATAAAAAATGGAAAGAATCTAATAGAGT GGTACAGCACTGTTATTTTTCAAAGATGTGTTGCTATCCTGAAAATTCT GTAGGTTCTGTGGAAGTTCCAGTGTTCTCTCTTATTCCACTTCGGTAGA GGATTTCTAGTTTCTTGTGGGCTAATTAAATAAATCATTAATACTCTTC TAAGTTATGGATTATAAACATTCAAAATAATATTTTGACATTATGATAA TTCTGAATAAAAGAACAAAAACCA HP 3UTR TGCAAGGCTGGCCGGAAGCCCTTGCCTGAAAGCAAGATTTCAGCCTGG 189AAGAGGGCAAAGTGGACGGGAGTGGACAGGAGTGGATGCGATAAGAT GTGGTTTGAAGCTGATGGGTGCCAGCCCTGCATTGCTGAGTCAATCAATAAAGAGCTTTCTTTTGACCCA0RM1 3UTR CAGGACACAGCCTTGGATCAGGACAGAGACTTGGGGGCCATCCTGCCC 190 CTCCAACCCGACATGTGTACCTCAGCTTTTTCCCTCACTTGCATCAATA AAGCTTCTGTGTTTGGAACAGCTAA APOA2 3UTR AGTGTCCAGACCATTGTCTTCCAACCCCAGCTGGCCTCTAGAACACCC 191ACTGGCCAGTCCTAGAGCTCCTGTCCCTACCCACTCTTTGCTACAATAA ATGCTGAATGAATCCA SAA2 3UTR GCTTCCTCTTCACTCTGCTCTCAGGAGACCTGGCTATGAGGCCCTCGGG 192GCAGGGATACAAAGTTAGTGAGGTCTATGTCCAGAGAAGCTGAGATAT GGCATATAATAGGCATCTAATAAATGCTTAAGAGGTGGAA SERPINA1 3UTR CTGCCTCTCGCTCCTCAACCCCTCCCCTCCATCCCTGGCCCCCTCCCTG 193GATGACATTAAAGAAGGGTTGAGCTGGTCCCTGCCTGCATGTGACTGT AAATCCCTCCCATGTTTTCTCTGAGTCTCCCTTTGCCTGCTGAGGCTGT ATGTGGGCTCCAGGTAACAGTGCTGTCTTCGGGCCCCCTGAACTGTGTT CATGGAGCATCTGGCTGGGTAGGCACATGCTGGGCTTGAATCCAGGGG GGACTGAATCCTCAGCTTACGGACCTGGGCCCATCTGTTTCTGGAGGG CTCCAGTCTTCCTTGTCCTGTCTTGGAGTCCCCAAGAAGGAATCACAGG GGAGGAACCAGATACCAGCCATGACCCCAGGCTCCACCAAGCATCTTC ATGTCCCCCTGCTCATCCCCCACTCCCCCCCACCCAGAGTTGCTCATCC TGCCAGGGCTGGCTGTGCCCACCCCAAGGCTGCCCTCCTGGGGGCCCC AGAACTGCCTGATCGTGCCGTGGCCCAGTTTTGTGGCATCTGCAGCAA CACAAGAGAGAGGACAATGTCCTCCTCTTGACCCGCTGTCACCTAACC AGACTCGGGCCCTGCACCTCTCAGGCACTTCTGGAAAATGACTGAGGC AGATTCTTCCTGAAGCCCATTCTCCATGGGGCAACAAGGACACCTATT CTGTCCTTGTCCTTCCATCGCTGCCCCAGAAAGCCTCACATATCTCCGT TTAGAATCAGGTCCCTTCTCCCCAGATGAAGAGGAGGGTCTCTGCTTTG TTTTCTCTATCTCCTCCTCAGACTTGACCAGGCCCAGCAGGCCCCAGAA GACCATTACCCTATATCCCTTCTCCTCCCTAGTCACATGGCCATAGGCC TGCTGATGGCTCAGGAAGGCCATTGCAAGGACTCCTCAGCTATGGGAG AGGAAGCACATCACCCATTGACCCCCGCAACCCCTCCCTTTCCTCCTCT GAGTCCCGACTGGGGCCACATGCAGCCTGACTTCTTTGTGCCTGTTGCT GTCCCTGCAGTCTTCAGAGGGCCACCGCAGCTCCAGTGCCACGGCAGG AGGCTGTTCCTGAATAGCCCCTGTGGTAAGGGCCAGGAGAGTCCTTCC ATCCTCCAAGGCCCTGCTAAAGGACACAGCAGCCAGGAAGTCCCCTGG GCCCCTAGCTGAAGGACAGCCTGCTCCCTCCGTCTCTACCAGGAATGG CCTTGTCCTATGGAAGGCACTGCCCCATCCCAAACTAATCTAGGAATC ACTGTCTAACCACTCACTGTCATGAATGTGTACTTAAAGGATGAGGTT GAGTCATACCAAATAGTGATTTCGATAGTTCAAAATGGTGAAATTAGC AATTCTACATGATTCAGTCTAATCAATGGATACCGACTGTTTCCCACAC AAGTCTCCTGTTCTCTTAAGCTTACTCACTGACAGCCTTTCACTCTCCA CAAATACATTAAAGATATGGCCATCACCAAGCCCCCTAGGATGACACC AGACCTGAGAGTCTGAAGACCTGGATCCAAGTTCTGACTTTTCCCCCTG ACAGCTGTGTGACCTTCGTGAAGTCGCCAAACCTCTCTGAGCCCCAGT CATTGCTAGTAAGACCTGCCTTTGAGTTGGTATGATGTTCAAGTTAGAT AACAAAATGTTTATACCCATTAGAACAGAGAATAAATAGAACTACATT TCTTGCA APOCI 3UTR GGACCTGAAGGGTGACATCCCAGGAGGGGCCTCTGAAATTTCCCACAC 194CCCAGCGCCTGTGCTGAGGACTCCCTCCATGTGGCCCCAGGTGCCACC AATAAAAATCCTACAGAAAA3WJ1_ 3UTR TTGCCATGTGTATGTGGGTTTTTTTTTTCCCACATACTCTGATGATCCTT 195 customTTTTTTTTGGATCATTCATGGCAA3WJ2_ 3UTR TTGCCATGTGTATGTGGGAAAAAAAAAACCCACATACTCTGATGATCC 196 custom AAAAAAAAAAGGATCATTCATGGCAAHBB 3UTR gtgctggtctgtgtgctggcccatcactttggcaaagaattcaccccaccagtgcaggctgcctatcagaaagtggtg 197 Extended gctggtgtggctaatgccctggcccacaagtatcactaagctcgctttcttgctgtccaatttctattaaaggttcctttgtt ccctaagtccaactactaaactgggggatattatgaagggccttgagcatctggattctgcctaataaaaaacatttatttt cattgcahHBA 3UTR GCTGGAGCCTCGGTGGCCATGCTTCTTGCCCCTTGGGCCTCCCCCCAGC 198CCCTCCTCCCCTTCCTGCACCCGTACCCCCGTGGTCTTTGAATAAAGTC TGAGTGGGCGGCmHBA 3UTR GCGGCCGCTTAATTAAGCTGCCTTCTGCGGGGCTTGCCTTCTGGCCATG 199CCCTTCTTCTCTCCCTTGCACCTGTACCTCTTGGTCTTTGAATAAAGCCTGAGTAGGAAGTCTAGAES- 3UTR CTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGGTA 200 mtRNRl CCCCGAGTCTCCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTCCACC TGCCCCACTCACCACCTCTGCTAGTTCCAGACACCTCCCAAGCACGCA GCAATGCAGCTCAAAACGCTTAGCCTAGCCACACCCCCACGGGAAACA GCAGTGATTAACCTTTAGCAATAAACGAAAGTTTAACTAAGCTATACT AACCCCAGGGTTGGTCAATTTCGTGCCAGCCACACC EK5 AE1cactcctccatggtgatataaagaccacccacttccttcgggtgagcccctaagccatggttgtactgcactatcatcct 201 aagacggtccttcttcggatcgcaatctcaccctggtgccgcgcttccttcgggaactgcacccgcggaccagggcc gtctttgaacttttctaactgttcttacSTH AE1GTCATGAAGGTTTTTCTTTTCCTGAGAAAACAAgaaaTTGTTTTCTCAGGT 202TTTGCTTTTTGAAAAAAAAGCAAAA1AE (additional elements) can be attached either downstream of the 3 ’-UTR or the poly A tail
[0202] The Table below additionally discloses 5’-UTR sequence variants wherein the difference in sequences are not within the UTR itself but in the second codon, as shown in the fourth column. For CESvarl, CESvar2, and CESvar3, ‘gcc’ was inserted in the second codon position. Column 4 shows the sequence of the first three codons.Table 4C: Exemplary mRNA ElementsElement First 3 SEQ Name Position Sequence of UTRcodons ID NO CES1 203 5UTR a ATGgccAAA varl gACAGAGACCTCGCAGGCCCCGAGAACTGTCGCCCTTCCACGCES1 204 5UTRvar2 agacagagacctcgcaggccccgagaactgtcgcccttccacc ATGgccAAA CES1 205 5UTR agacagagacct ATGgccAAA var3 cgcaggccccgagaactgtcgccctgccaccCES1 206 5UTR AGACAGAGACCTCGCAGGCCCCGAGAACTGTCGCCCTTCCACG ATG— AAAvar4Table 5: Exemplary Reverse Transcriptase VariantsNAME SPECIES (native or derived) Length SEQ ID NO RT RTBat-A ARS01 RTEni-V-derived Eptesicus nilssonii, bat 672 610 RT RTBat-A ARS01 With-RT5m Prot-neg2 Eptesicus nilssonii, bat 668 611 RT RTEni With-RT5m Eptesicus-nilssonii Eptesicus nilssonii, bat 679 607 RT_RTEni_With-RT5m_Eptesicus-nilssonii_within- Eptesicus nilssonii, bat 672 609 protRT_RTEni_With-RT5m_Eptesicus-nilssonii_Prot- Eptesicus nilssonii, bat 668 608 neg2RT RTMbr-A Myotis-brandtii Myotis brandtii, bat 673 599 RT_RTMbr-A_Myotis-brandtii_With-RT5m_Prot- Myotis brandtii, bat 669 600 neg2RT RTMcr-A Melozone crissalis, bird 741 585 RT RTMcr-A Prot-neg2 Melozone crissalis, bird 675 586 RT_RTMcr-A_Melozone-crissalis_Prot-neg2_With- Melozone crissalis, bird 675 587 RT5mRT RTMda-A Myotis daubertonii, bat 673 602 RT RTMda-A Myotis-daubentonii With- Myotis daubentonii, bat 669 603 RT 5 m_Prot-neg2RT RTMge With-RT5m Melospiza georgiana, bird 685 594 RT RTMge Melospiza-georgiana With- Melospiza georgiana, bird 675 593 RT 5 m_Prot-neg2RT RTMge -A Melospiza georgiana, bird 713 579RT RTMge -A Prot-neg2 Melospiza georgiana, bird 671 580RT RTMge-A Prot-neg2 With-RT5m Melospiza georgiana, bird 671 581 RT RTMge-B Melospiza georgiana, bird 741 588 RT RTMge-C var-of-B Prot-neg2 Melospiza georgiana, bird 675 589 RT RTMge-C var-of-B With-RT5m Prot-neg2 Melospiza georgiana, bird 675 592 RT RTMge-C var-of-B Within-prot-cleavage Melospiza georgiana, bird 679 590 RT RTPku-A Pipistrellus-kuhlii Pipistrellus kuhlii, bat 673 604 RT_RTPku-A_Pipistrellus-kuhlii_With-RT5m_Prot- Pipistrellus kuhlii, bat 669 605 neg2RT RTZal-A Zonotrichia albicollis, bird 741 582 RT RTZal-A Prot-neg2 Zonotrichia albicollis, bird 675 583 RT RTZal-A Zonotrichia-albicollis Prot- Zonotrichia albicollis, bird 675 584 neg2 With-RT5mKlebsiella pneumoniae mut1Klebsiella pneumonia (added 425 613mutations E48K, S339K, Y395E)Burkholderia cepacia mut1Burkholderia cepacia (added mutations 423 612T47K, S87E, S338K, Y394E)Caldibacillus thermoamylovorans Caldibacillus thermoamylovorans 442 618 Caldibacillus thermoamylovorans mut1Caldibacillus thermoamylovorans 442 619(added mutations T53K, S357K,Y408E)Lysinibacillus sphaericus Lysinibacillus sphaericus 459 614 Lysinibacillus sphaericus mut1Lysinibacillus sphaericus (added 459 615mutations E63K, S374K, Y425E)Bacillus sp. Bacillus sp. 438 616 Bacillus sp. mut1Bacillus sp. (added mutations S3391, 438 617S351K, Y402E)Truncated RT from Myotis brandtii Myotis brandtii 492 601 Truncated RT from Pipistrellus kuhlii Pipistrellus kuhlii 492 606 Truncated RT from MMLV MMLV 492 597 pPG273 MMLV protneg RT 667 596 PG270 Woolly monkey RT 672 598 PG273 Melospiza consensus protneg RT 675 591 PG273 Avian protneg RT 694 595^‘mut” indicates mutants that have been designed
[0203] Provided in Table 6 are the sequences of exemplary 5’UTRs.Table 6: Exemplary 5’ UTR VariantsNAME SEQUENCE1Length SEQ ID NO NCA7d AGCAAAAAUCAAAAUCAAUCAUCAUCACAACAUCAACAAU 70 207CAAUCAUCAACACAUCAUCAAGACACCACC NCA7d-50nt AGCAUCACAACAUCAACAAUCAAUCAUCAACACAUCAUCA 50 208AGACACCACC HBB AGAUUUGCUUCUGACACAACUGUGUUCACUAGCAACCUCA 50 209AACAGACACC RV-UML-m002 AGGUCCGUUAUAUUAUUUAUCUUGCAGAUCAAACUUCAGA 50 210GAGGAGGGCC RV-UML-m314 AGGGAACAGUGGAGCACAAUAGUUGUCACAAACGCGUGAC 50 211CUGUACCUUA RV-UML-mO17 AGAGAAAGAGUUUCGAGAAAGUUGCGAUACACACCGACCU 50 212AAUCCGUGUC RV-UML-m349 AGACUGAGUAGCGUUGUCGAUUGCUGUUUUCAUAACCGUA 50 213CCGAGCCUGC RV-UML-m286 AGUGGAGAGCGCAAUCCAGUUUCCUUAGCAAAGCAUAUCA 50 214CCCUUCGGCU RV-UML-m302 AGUUAUUUGCGAUAAACAGUAGGGAGACAAUUGACCUGCA 50 215CAGCGUGGCC RV-UML-m318 AGCCCUUACACGAAUUCUCGAAAAUCGUGCUUCAAGAAAC 50 216GUACAACGGCRV-UML-m348 AGCCUUCGCACAAUACUCGCGCGUCCUAAGCAAAAUCGGA 50 217 ACAAGUACUG RV-UML-mO25 AGGCGAAGUCUUCGACUACCACUACCCUCGCCGUCGCGGU 50 218CUUUGCUGUC RV-UML-m006 AGGCCGAGUCAUUAUCAUACUUGGUUCAAUCUAUUGAACG 50 219AAGUAUUGCU RV-UML-m004 AGGGACCAAGAGUUCGAUACCUCAUCGAACUGCGAGUCAU 50 220AAGCAGGGCC RV-UML-m007 AGGGAUUUACCGAGAAUCUAGUUACAGAACAACCUCUUCC 50 221UAAAAGGGCC RV-UML-m290 AGGUUGCACUUUACCCUCCCUUCCCGUGGGAACGCAUUGC 50 222GAAAAUCGCA RV-UML-m038 AGCCGGAACACGGUAUCUGUGCGUUCGUAAUCAGUACUAA 50 223CUCUAUCCAG RV-UML-m291 AGUGGAGAUUCGGGAUCUGCAAAUCACCGAUAAAACAUCC 50 224CUUCGGGAGG RV-UML-mO91 AGAACGCGAGUUCCUUCUUCACUGGUACCUCACACCCAAU 50 225AGCAAUUCCA RV-UML-m280 AGACGGUAGACAAGGAAUCGAAAGACUGCUCCUCACUGCU 50 226UACGUUUACC RV-UML-mO78 AGGUGGUCUUCUUGUUGCUACAGGAACCCACGGACUUAUU 50 227CGAUAGCGAC RV-UML-m354 AGUGAGUCCCCGUCAACCGUAUCGGGAGGGCGAUCGACAA 50 228UUCUCUACCA RV-UML-mO76 AGUGCCGCAUUGUCUUCUUGUGUUCGCUUGCCGCCGACGA 50 229ACAACUGGAA RV-UML-m358 AGAAGGUUCCGCGACGCUCCCGCCUUCUAUCCCCUUCUUU 50 230AAAGUCGCUGTor RNA, T is U; U are T in the sequence listing
[0204] In some embodiments, the gRNAs or tagRNAs described herein can be produced by in vitro transcription (IVT), synthetic and / or chemical synthesis methods, or a combination thereof. One or more of enzymatic IVT, solid-phase, liquid-phase, combined synthetic methods, small region synthesis, and ligation methods can be utilized. In some embodiments, the gRNAs or tagRNAs are made using IVT enzymatic synthesis methods. Methods of making polynucleotides by IVT are known in the art and are described in WO2013 / 151666. Polynucleotides constructs and vectors can be used to in vitro transcribe a gRNA or tagRNA described herein.
[0205] In some embodiments, a nucleic acid encoding an editor (e.g., SyNTase editor, RT editor) is administered to the subject. In some embodiments, the nucleic acid can be generated by an in vitro transcription reaction. In some embodiments, generating in vitro transcribed RNA comprises incubating a linear DNA template with an RNA polymerase and a nucleotide mixture under conditions to allow (run-off) RNA in vitro transcription. The nucleotide mixture can be part of an in vitro transcription mix (IVT-mix). In some embodiments, the RNA polymerase is a T7 RNA polymerase.
[0206] The nucleotide mixture used in RNA in vitro transcription can additionally contain modified nucleotides as defined below. In some embodiments, the nucleotide mixture(e.g., the fraction of each nucleotide in the mixture) used for RNA in vitro transcription reactions can be optimized for the given RNA sequence (optimized NTP mix). Such methods are described, for example in WO2015 / 188933. RNA obtained by a process using an optimized NTP mix is, in some embodiments, characterized by reduced immune stimulatory properties.
[0207] In various embodiments, the nucleotide mixture in an in vitro transcription reaction comprises a cap analog. Accordingly, in some embodiments the cap analog is a capO, capl, cap2, a modified capO or a modified capl analog, or a capl analog as described below.
[0208] The term “cap analog” or “5 ’-cap structure” as used herein can refer to the 5’ structure of the RNA, particularly a guanine nucleotide, positioned at the 5 ’-end of an RNA, e.g, an mRNA. In some embodiments, the 5’-cap structure is connected via a 5 ’-5 ’-triphosphate linkage to the RNA. In some embodiments, a “5’-cap structure” or a “cap analogue” is not considered to be a “modified nucleotide” or “chemically modified nucleotides”. 5 ’-cap structures which may be suitable include capO (methylation of the first nucleobase, e.g, m7GpppN), capl (additional methylation of the ribose of the adjacent nucleotide of m7GpppN), cap2 (additional methylation of the ribose of the 2nd nucleotide downstream of the m7GpppN), cap3 (additional methylation of the ribose of the 3rd nucleotide downstream of the m7GpppN), cap4 (additional methylation of the ribose of the 4th nucleotide downstream of the m7GpppN), ARCA (anti-reverse cap analogue), modARCA (e.g., phosphothioate modARCA), inosine, Nl-methyl-guanosine, 2’-fluoro-guanosine, 7-deaza-guanosine, 8-oxo-guanosine, 2-amino-guanosine, LNA-guanosine, and 2-azido-guanosine.
[0209] A 5’-cap (capO or capl) structure can be formed in chemical RNA synthesis, using capping enzymes, or in RNA in vitro transcription (co-transcriptional capping) using cap analogs. The term “cap analog” as used herein can refer to a non-polymerizable di-nucleotide or tri -nucleotide that has cap functionality in that it facilitates translation or localization, and / or prevents degradation of the RNA when incorporated at the 5 ’ -end of the RNA. Non-polymerizable means that the cap analogue will be incorporated only at the 5 ’-terminus because it does not have a 5’ triphosphate and therefore cannot be extended in the 3 ’-direction by a template-dependent polymerase, (e.g., a DNA-dependent RNA polymerase). Examples of cap analogues include m7GpppG, m7GpppA, m7GpppC; unmethylated cap analogues (e.g., GpppG); dimethylated cap analogue (e.g., m2,7GpppG), trimethylated cap analogue (e.g. m2,2,7GpppG), dimethylated symmetrical cap analogues (e.g. m7Gpppm7G), or anti reverse cap analogues (e.g., ARCA; m7,2’OmeGpppG, m7,2’dGpppG, m7,3’OmeGpppG, m7,3’dGpppG and their tetraphosphate derivatives). Further cap analogues have been described previously, e.g, W02008 / 016473, W02008 / 157688, WO2009 / 149253, WO2011 / 015347, and WO2013 / 059475. Further suitable cap analogues in that context are described in, e.g., WO2017 / 066793, WO2017 / 066781,WO2017 / 066791, WO2017 / 066789, WO2017 / 053297, WO2017 / 066782, WO2018 / 075827 and WO2017 / 066797 wherein the disclosures relating to cap analogues are incorporated herewith by reference.
[0210] In some embodiments, a capl structure is generated using tri-nucleotide cap analogue as disclosed in WO2017 / 053297, WO2017 / 066793, WO2017 / 066781, WO2017 / 066791, WO2017 / 066789, WO2017 / 066782, WO2018 / 075827 and WO2017 / 066797. For example, any cap analog derivable from the structure disclosed in claim 1-5 of WO2017 / 053297 may be suitably used to co-transcriptionally generate a capl structure. In some embodiments, any cap analog derivable from the structure described in WO2018 / 075827 can be suitably used to co-transcriptionally generate a capl structure. In some embodiments, the capl analog is a capl trinucleotide cap analog. In some embodiments, the capl structure of the in vitro transcribed RNA is formed using co-transcriptional capping using tri -nucleotide cap analog m7G(5’)ppp(5')(2’OMeA)pG or m7G(5’)ppp(5’)(2’OMeG)pG. In some embodiments, the capl analog is m7G(5’)ppp(5’)(2’OMeA)pG.
[0211] In some embodiments, the RNA (e.g., mRNA) comprises a 5 ’-cap structure, e.g., a capl structure. In some embodiments, the 5’ cap structure can improve stability and / or expression of the mRNA. A capl structure comprising mRNA (produced by, e.g., in vitro transcription) has several advantageous features including an increased translation efficiency and a reduced stimulation of the innate immune system. In some embodiments, the in vitro transcribed RNA comprises at least one coding sequence encoding at least one peptide or protein. In some embodiments, the protein is an RNA-guided endonuclease. In some embodiments, the RNA-guided endonuclease is Cas9 or a derivative thereof.
[0212] In some embodiments, the mRNA can comprise at least one chemically modified nucleoside and / or nucleotide. In some embodiments, the chemically modified nucleoside and / or nucleotide is selected from pseudouridine, N1 -methylpseudouridine, and 5-m ethoxyuridine. In some embodiments, the chemically modified nucleoside is Nl-methylpseudouridine (e.g., 1 -methylpseudouridine). In some embodiments, at least about 80% or more (e.g., about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) of uridines in the mRNA are modified or replaced with N1 -methylpseudouridine. In some embodiments, 100% of the uridines (e.g., uracils) in the mRNA are modified or replaced with N 1 -methylpseudouridine.
[0213] In some embodiments, the disclosure provides an mRNA comprising a nucleotide sequence that is 100% identical to the nucleotide sequence of SEQ ID NO: 620 or 621, wherein 100% of the uridines (e.g., uracils) of the mRNA are modified or replaced with Nl-methylpseudouridine. Optimized mRNAs encoding, for example Cas9, are also described in US20210355463A1, which is hereby incorporated by reference in its entirety.Compositions and Therapeutic Applications
[0214] Provided herein also includes compositions. A composition can include one or more gRNA(s) (e.g., tagRNA and egRNA), and a polymerase-based editor, e.g., and editor (e.g., SyNTase editor, RT editor) protein described herein. A composition can include one or more gRNA(s) (e.g., tagRNA and egRNA), and a nucleic acid encoding a polymerase-based editor (e.g., SyNTase editor, RT editor) protein described herein. The compositions can include one or more of the editing complexes of the disclosure. Also provided herein are ribonucleoprotein (RNP) complexes comprising any of the editing complexes disclosed herein, or a component thereof. In some embodiments, the RNP comprises, e.g., an editor (e.g., a polypeptide or a polynucleotide encoding an editor) and a tagRNA. In some embodiments, the RNP comprises, e.g., an editor (e.g., a polypeptide or a polynucleotide encoding an editor) and an egRNA. In some embodiments, the RNP comprises, e.g., an editor (e.g., a polypeptide or a polynucleotide encoding an editor) and both a tagRNA and an egRNA.
[0215] Disclosed herein include lipid nanoparticles (LNPs). In some embodiments, the LNP comprises any of the editing systems disclosed herein, or a component thereof. The LNP can comprise the tagRNA and the nucleic acid encoding the editor (e.g., SyNTase editor, RT editor). The nucleic acid encoding the editor can be mRNA. The LNP can comprise the egRNA. The lipid nanoparticle can comprise one or more neutral lipids, charged lipids, ionizable lipids, steroids, and polymers conjugated lipids. The lipid nanoparticle can comprise cholesterol, a polyethylene glycol (PEG) lipid, or both.
[0216] In some embodiments, each of the components of an editing system (e.g., a SyNTase editing system) can be separately formulated into lipid nanoparticles, or are all coformulated into one lipid nanoparticle.
[0217] In some embodiments, editor mRNA can be formulated in a lipid nanoparticle, while gRNAs (e.g., tagRNA and / or egRNA) can be delivered in an AAV vector.
[0218] Options are available to deliver the polymerase-based editor (e.g., SyNTase editor, RT editor) as a DNA plasmid, as mRNA or as a protein. A guide RNA (e.g., tagRNA and / or egRNA) can be expressed from the same DNA, or can also be delivered as an RNA. The RNA can be chemically modified to alter or improve its half-life, or decrease the likelihood or degree of immune response. The endonuclease protein can be complexed with the gRNA prior to delivery.
[0219] Disclosed herein include isolated cells. In some embodiments, the isolated cell comprises any of the tagRNAs, the SyNTase editing systems, the SyNTase editing complexes, the RNPs, the LNPs, or the vectors provided herein. The cell can be a mammalian cell. In some embodiments, the cell can be a human cell. The cell can be a primary cell. The cell can be ahepatocyte. In some embodiments, the cell has or is suspected of having a mutation that can be corrected via gene editing. In some embodiments, the gene editing system is a polymerase-based editing (e.g., SyNTase editing, RT editing) system. The cell can be from a subject having Wilson’s disease. The cell can be from a subject having A1ATD disease or disorder. The cell can be from a subject having phenylketonuria or hyperphenylalaninemia. The subject can be a human.
[0220] Disclosed herein include compositions. In some embodiments, the composition is a pharmaceutical composition. In some embodiments, the pharmaceutical composition comprises: any of the tagRNAs, the SyNTase editing systems, the SyNTase editing complexes, the RNPs, the LNPs, the polynucleotides, the vectors, or the cells disclosed herein; and (ii) a pharmaceutically acceptable carrier.
[0221] A composition described above can further have one or more additional reagents, where such additional reagents are selected from a buffer, a buffer for introducing a polypeptide or polynucleotide into a cell, a wash buffer, a control reagent, a control vector, a control RNA polynucleotide, a reagent for in vitro production of the polypeptide from DNA, adaptors for sequencing and the like. A buffer can be a stabilization buffer, a reconstituting buffer, a diluting buffer, or the like. In some embodiments, a composition can also include one or more components that can be used to facilitate or enhance the on-target binding or the cleavage of DNA by the endonuclease, or improve the specificity of targeting.
[0222] One or more components of a composition can be formulated with pharmaceutically acceptable excipients such as carriers, solvents, stabilizers, adjuvants, diluents, etc., depending upon the particular mode of administration and dosage form. In some embodiments, guide RNA compositions are generally formulated to achieve a physiologically compatible pH, and range from a pH of about 3 to a pH of about 11, about pH 3 to about pH 7, depending on the formulation and route of administration. In some embodiments, the pH is adjusted to a range from about pH 5 to about pH 8.
[0223] Suitable excipients can include, for example, carrier molecules that include large, slowly metabolized macromolecules such as proteins, polysaccharides, polylactic acids, polyglycolic acids, polymeric amino acids, amino acid copolymers, and inactive virus particles. Other exemplary excipients include antioxidants (for example and without limitation, ascorbic acid), chelating agents (for example and without limitation, EDTA), carbohydrates (for example and without limitation, dextrin, hydroxyalkylcellulose, and hydroxyalkylmethylcellulose), stearic acid, liquids (for example and without limitation, oils, water, saline, glycerol and ethanol), wetting or emulsifying agents, pH buffering substances, and the like.
[0224] Physiologically tolerable carriers are well known in the art. Exemplary liquid carriers are sterile aqueous solutions that contain no materials in addition to the active ingredientsand water, or contain a buffer such as sodium phosphate at physiological pH value, physiological saline or both, such as phosphate-buffered saline. Aqueous carriers can contain more than one buffer salt, as well as salts such as sodium and potassium chlorides, dextrose, polyethylene glycol and other solutes. Liquid compositions can also contain liquid phases in addition to and to the exclusion of water. Exemplary of such additional liquid phases are glycerin, vegetable oils such as cottonseed oil, and water-oil emulsions. The amount of an active compound used in the cell compositions that is effective in the treatment of a particular disorder or condition will depend on the nature of the disorder or condition, and can be determined by standard clinical techniques.
[0225] The terms “stable” or “stability” as used herein can refer to the ability of the compounds herein described (e.g., the SyNTase editor or a nucleic acid encoding the SyNTase editor and / or gRNA (e.g., tagRNA and / or egRNA)) to maintain therapeutic efficacy (e.g., all or the majority of its intended biological activity and / or physiochemical integrity) over extended periods of time. The stability of one or more of the compounds described herein (e.g., the SyNTase editor or a nucleic acid encoding the SyNTase editor and / or gRNA(e.g., tagRNA and / or egRNA)) can be 2 weeks, 15 days, 16 days, 17 days, 18 days, 19 days, 20 days, 3 weeks, 22 days, 23 days, 24 days, 25 days, 26 days, 27 days, 28 days, 29 days, 30 days, 31 days, 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 13 months, 14 months, 15 months, 16 months, 17 months, 18 months, 19 months, 20 months, 21 months, 22 months, 23 months, 2 years, 3 years, or more than 3 years. The temperature of storage can vary. For example, the storage temperature can be, can be about, can be at least, or can be at least about -80°C, -65°C, -20°C, 5°C, or a number or range between any two of these values. In some embodiments, the storage temperature is less than or equal to -65°C.
[0226] In some embodiments, the compounds herein described (e.g., the SyNTase editor or a nucleic acid encoding the SyNTase editor and / or gRNA(e.g., tagRNA and / or egRNA)) of a composition can be delivered via transfection such as calcium phosphate transfection, DEAE-dextran mediated transfection, cationic lipid-mediated transfection, electroporation, electrical nuclear transport, chemical transduction, electrotransduction, Lipofectamine-mediated transfection, Effectene-mediated transfection, lipid nanoparticle (LNP)-mediated transfection, or any combination thereof. In some embodiments, the composition is introduced to the cells via lipid-mediated transfection using a lipid nanoparticle.
[0227] Disclosed herein include methods for editing a target gene. In some embodiments, the method comprises contacting the target gene with (i) any of the tagRNAs of the disclosure and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain or (ii) any of the SyNTase editing systems disclosed herein. Insome embodiments, the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the target gene, thereby editing the target gene.
[0228] Disclosed herein include methods for editing a target gene. In some embodiments, the method comprises contacting the target gene with any of the SyNTase editing complexes disclosed herein, wherein the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the target gene, thereby editing the target gene. In some embodiments, the method is performed in vivo, in vitro or ex vivo.
[0229] In some embodiments, the polymerase-based editor (e.g., SyNTase editor, RT editor) synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the target gene. The target gene can be in a cell. The cell can be a mammalian cell. The cell can be a human cell. The cell can be a primary cell. The cell can be a hepatocyte. The cell can be in a subject. The subject can be a human. In some embodiments, the cell has or is suspected of having a mutation that can be corrected via gene editing. In some embodiments, the gene editing system is a polymerase-based editing (e.g., SyNTase editing, RT editing) system. The cell can be from a subject having a disease or disorder. The cell can be from a subject having Wilson’s disease. The cell can be from a subject having A1ATD disease or disorder. The cell can be from a subject having phenylketonuria or hyperphenylalaninemia. The method can comprise administering the cell to the subject after incorporation of the intended nucleotide edit. Also provided herein are cells generated by any of the methods disclosed herein. Disclosed herein include populations of cells generated by any of the methods of the disclosure.
[0230] Disclosed herein include methods for treating a disease or disorder in a subject in need thereof via gene editing. In some embodiments, the gene editing system is a polymerase-based editing (e.g., SyNTase editing, RT editing) system. Disclosed herein include methods for treating Wilson’s disease in a subject in need thereof. Disclosed herein include methods for treating Al ATD disease in a subject in need thereof. Disclosed herein include methods for treating phenylketonuria or hyperphenylalaninemia in a subject in need thereof. In some embodiments, the method comprises administering to the subject (i) any of the tagRNAs of the disclosure and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain or (ii) any of the SyNTase editing systems disclosed herein. In some embodiments, the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the mutated gene in the subject, thereby treating the disease or disorder in the subject. In some embodiments, the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the ATP7B gene in the subject, thereby treating Wilson’s disease in the subject. In someembodiments, the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the SERPINA1 gene in the subject, thereby treating Al ATD disease or disorder in the subject. In some embodiments, the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the PAH gene in the subject, thereby treating phenylketonuria or hyperphenylalaninemia in the subject.
[0231] Disclosed herein include methods for treating a disease in a subject in need thereof. In some embodiments, the method comprises administering to the subject any of the SyNTase editing complexes, any of the RNPs, any of the LNPs, or any of the pharmaceutical compositions disclosed herein. In some embodiments, the tagRNA directs the SyNTase editor to incorporate the intended nucleotide edit in the target gene in the subject, thereby treating the disease or disorder in the subject. The subject can be a human.
[0232] In some embodiments, editing efficiency of the SyNTase editing compositions and methods described herein can be measured by calculating the percentage of edited target genes in a population of cells introduced with the SyNTase editing composition. In some embodiments, the editing efficiency is determined after 1 hour, 2 hours, 6 hours, 12 hours, 24 hours, 36 hours, 48 hours, 3 days, 4 days, 5 days, 7 days, 10 days, or 14 days of exposing a target gene (e.g., ATP7B gene, PAH gene, or SERPINA1 gene within the genome of a cell) to a SyNTase editing composition (e.g., LNP). In some embodiments, the population of cells introduced with the SyNTase editing composition is ex vivo. In some embodiments, the population of cells introduced with the SyNTase editing composition is in vitro. In some embodiments, the population of cells introduced with the SyNTase editing composition is in vivo. In some embodiments, the SyNTase editing methods disclosed herein have an editing efficiency of at least about 1%, at least about 5%, at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, or at least about 99% relative to a suitable control.
[0233] In some embodiments, the SyNTase editing compositions provided herein are capable of incorporated one or more intended nucleotide edits without generating a significant proportion of indels. The term “indel(s)”, as used herein, refers to the insertion or deletion of a nucleotide base within a polynucleotide, for example, a target gene. Such insertions or deletions can lead to frame shift mutations within a coding region of a gene. Indel frequency of editing can be calculated by methods known in the art. In some embodiments, the methods disclosed herein can have an indel frequency of less than 20%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, or less than 1%. In some embodiments, any number of indels is determined after at least 1 hour, at least 2 hours, at least 6 hours, at least 12 hours, at least 24 hours, at least 36 hours, at least 48 hours, atleast 3 days, at least 4 days, at least 5 days, at least 7 days, at least 10 days, or at least 14 days of exposing a target gene (e.g., ATP7B gene, PAH gene, or SERPINA1 gene within the genome of a cell) to a SyNTase editing composition.
[0234] In some embodiments, the SyNTase editing methods disclosed hereto have an editing efficiency of at least about 95% and an indel frequency of less than 1% to a target cell, e.g., a human primary cell or hepatocyte, to some embodiments, the SyNTase editing methods disclosed hereto have an editing efficiency of at least about 95% and an indel frequency of less than 0.5% to a target cell, e.g., a human primary cell or hepatocyte, to some embodiments, the SyNTase editing methods disclosed hereto have an editing efficiency of at least about 95% and an indel frequency of less than 0.1% to a target cell, e.g., a human primary cell or hepatocyte, to some embodiments, any number of indels is determined after at least 1 hour, at least 2 hours, at least 6 hours, at least 12 hours, at least 24 hours, at least 36 hours, at least 48 hours, at least 3 days, at least 4 days, at least 5 days, at least 7 days, at least 10 days, or at least 14 days of exposing a target gene (e.g., ATP7B gene, PAH gene, or SERPINA1 gene within the genome of a cell) to a SyNTase editing composition, to some embodiments, the editing efficiency is determined after 1 hour, 2 hours, 6 hours, 12 hours, 24 hours, 36 hours, 48 hours, 3 days, 4 days, 5 days, 7 days, 10 days, or 14 days of exposing a target gene (e.g., a ATP7B gene, PAH gene, or SERPINA1 gene within the genome of a cell) to a SyNTase editing composition.
[0235] In some embodiments, the target tissue for the compositions and methods described herein is liver tissue. In some embodiments, the target cells for the compositions and methods described herein is hepatocyte.
[0236] In some embodiments, the pharmaceutical composition thereof can be administered by aerosol delivery, nasal delivery, vaginal delivery, rectal delivery, buccal delivery, ocular delivery, local delivery, topical delivery, intraci sternal delivery, intraperitoneal delivery, oral delivery, intramuscular injection, intravenous injection, subcutaneous injection, intranodal injection, intratumoral injection, intraperitoneal injection, and / or intradermal injection, or any combination thereof. The administration can be local or systemic. The systemic administration includes enteral and parenteral administration. In some embodiments, more than one administration can be employed to achieve the desired level of gene expression over a period of various intervals, e.g., daily, weekly, monthly, or yearly.
[0237] The pharmaceutical composition thereof can be administered to a subject in need thereof at a pharmaceutically effective amount. The term “pharmaceutically effective amount” as used herein means that the amount of the pharmaceutical composition that will elicit a desired therapeutic effect and / or biological or medical responses of a tissue, system, animal or human. The administration can result in a desired correction in target gene such restoration ofwild-type activity of the protein.
[0238] In certain embodiments, disease models for screening of the SyNTase editing system and guide RNAs are used to demonstrate efficacy and tolerability. In some embodiments, efficacy and tolerability are demonstrated in vitro. In some embodiments, efficacy and tolerability are demonstrated in vivo. In some embodiments, the in vitro system is a cell line. In certain embodiments, the cell line is a human hepatocyte cell line, for instance, Huh7 cell line wherein the specific disease mutation has been incorporated. In certain embodiments, the in vitro cell line are primary hepatocytes, for instance primary murine hepatocytes and primary human hepatocytes. In certain embodiments, the in vivo system are humanized mouse and rat models, wherein the specific disease mutation has been incorporated. Such disease models are well established in literature wherein the data may be extrapolated to efficacy and tolerability in humans.Compositions and methods related to Alpha- 1 Antitrypsin Deficiency (AATD)
[0239] In some embodiments, there are provided compositions and methods for, e.g., treating Alpha- 1 antitrypsin deficiency (AATD). AATD is a genetic disorder that may result in lung disease and / or liver disease. Onset of lung problems is typically between 20 and 50 years of age. This may result in shortness of breath, wheezing, or an increased risk of lung infections. Complications may include chronic obstructive pulmonary disease (COPD), cirrhosis, neonatal jaundice, or panniculitis.
[0240] AATD is due to a mutation in the SERPINA1 gene. Risk factors for lung disease include tobacco smoking and environmental dust. The underlying mechanism involves unblocked neutrophil elastase and buildup of abnormal A1AT in the liver. Provided herein are methods and compositions for editing of the SERPINA1 gene (to correct the disease mutation to wild-type) that encodes Al AT serine protease inhibitor for treating alphal antitrypsin deficiency (AATD) disease. Editing of the S allele mutation or the Z allele mutation in SERPINA1 gene are contemplated herein. The Z allele mutation on exon 5 of SERPINA1 is an E342K mutation. E342K mutation leads to misfolding of the AAT protein leading to polymers and liver damage. Contemplated are methods for a single nucleotide (T to C) correction leading to a K342E correction in the SERPINA1 gene. The SERPINA1 -editing SyNTase compositions and methods described herein can achieve protective AAT levels across a wide range of AATD patients, including those of varying genotypes. In some embodiments the SERPINA1 -editing SyNTase compositions and methods can yield 1.1-fold, 1.2-fold, 1.3-fold, 1.4-fold, 1.5-fold, 1.6-fold, 1.7-fold, 1.8-fold, 1.9-fold, 2.0-fold, 2.1-fold, 2.2-fold, 2.3-fold, 2.4-fold, 2.5-fold, 2.6-fold, 2.7-fold, 2.8-fold, 2.9-fold, 3.0-fold, 3.1-fold, 3.2-fold, 3.3-fold, 3.4-fold, 3.5-fold, 3.6-fold, 3.7-fold, 3.8-fold, 3.9-fold, 4.0-fold, 4.1-fold, 4.2-fold, 4.3-fold, 4.4-fold, 4.5-fold, 4.6-fold, 4.7-fold, 4.8-fold,4.9-fold, 5.0-fold, or a number or a range between any two of these values, serum AAT upregulation. SERPINA1 -editing SyNTase compositions and methods provided herein can achieve serum AAT levels of at least about 10.0 pM, 10.2 pM, 10.4 pM, 10.6 pM, 10.8 pM, 11.0 pM, 11.2 pM, 11.4 pM, 11.6 pM, 11.8 pM, 12.0 pM, 12.2 pM, 12.4 pM, 12.6 pM, 12.8 pM, 13.0 pM, 13.2 pM, 13.4 pM, 13.6 pM, 13.8 pM, 14.0 pM, 14.2 pM, 14.4 pM, 14.6 pM, 14.8 pM, 15.0 pM, 15.2 pM, 15.4 pM, 15.6 pM, 15.8 pM, 16.0 pM, 16.2 pM, 16.4 pM, 16.6 pM, 16.8 pM, 17.0 pM, 17.2 pM, 17.4 pM, 17.6 pM, 17.8 pM, 18.0 pM, 18.2 pM, 18.4 pM, 18.6 pM, 18.8 pM, 19.0 pM, 19.2 pM, 19.4 pM, 19.6 pM, 19.8 pM, 20.0 pM, 20.2 pM, 20.4 pM, 20.6 pM, 20.8 pM, 21.0 pM, 21.2 pM, 21.4 pM, 21.6 pM, 21.8 pM, 22.0 pM, or a number or a range between any two of these values. Accordingly, the SERPINA1 -editing SyNTase compositions and methods provided herein can achieve total serum AAT levels at or beyond the clinical protective thresholds. While a ll pM threshold has been historically used as a clinical threshold, given the variability of response, even if the mean level achieved for a population is, e.g., 12.4 pM, a significant proportion of the population can still fall below the 11 pM threshold. Additionally, the notion that there is a threshold level of protection is not grounded in any specific evidence. The 11 pM value is based on differences observed between SZ and ZZ genotype patients; however, patients may be below 11 pM and have no disease, and may be above 11 pM and still develop disease. The 11 pM threshold was originally used as entry criterion for early clinical trials and later consideration for replacement therapy. Genotype, particularly ZZ, can be a much more important factor in predicting risk of disease progression. In some embodiments, in the absence of clear evidence for a protective threshold in patients with COPD, achieving up to normal levels (>20 pM) would be in the “safest” zone since diagnosis of Al ATD can be made even with levels above 11 pM.
[0241] In some embodiments, the editing system (e.g., SyNTase editing system) comprises: a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises: a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule; an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; and a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor, wherein the tagRNA comprises a sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 712; and the SyNTase editor comprises an amino acid sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%,-n-95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 723.
[0242] Provided herein are methods for treating or preventing AATD in a subject in need thereof. In some embodiments, the method comprises administering to the subject an SyNTase editor or SyNTase editing system of the disclosure, an RNP of the disclosure, an LNP of the disclosure, a pharmaceutical composition of the disclosure, a cell of the disclosure, or a population of cells of the disclosure, thereby treating or preventing or reversing the AATD in the subject. In some embodiments, the tagRNA comprises a sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 712. In some embodiments, the SyNTase editor comprises an amino acid sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% (e.g., 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or a range between any two of these values) identical to the sequence of SEQ ID NO: 723. In some embodiments, the method is performed in vivo or in vitro or ex vivo.EXAMPLES
[0243] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure.Example 1Synthetic Nucleotide Template Polymerases
[0244] Described in this Example are Synthetic Nucleotide Template Polymerases (SyNTases, also referred to herein as a synthetic polymerase editor or SPE).
[0245] A natural reverse transcriptase (RT) is an RNA-dependent DNA polymerase, and a natural polymerase is a DNA-dependent DNA polymerase. This means that they either use RNA or DNA as templates for polymerization. Many gene editing technologies utilize Cas9-polymerase (e.g., Cas9-RT) fusion proteins.
[0246] Described herein is a new class of enzymes that use unnatural template nucleicacids (fully or partially chemically modified templates). Since these enzymes use a novel template, they are referred to herein as Synthetic Nucleotide Template Polymerases (SyNTases). The SyNTases’ use of fully modified chemical template significantly increases the stability of, e.g., tagRNAs in vivo, and can enhance potency and stability of the drug.
[0247] The MMLV RT editor fails to read through modified template RNA, even at high doses, with activity far lower than that observed using the standard unmodified template (FIG. 1). While the Evo Tfl editor demonstrates activity in read-through of fully chemically modified template RNA, its efficiency falls short compared to using unmodified template RNA (FIG. 2).
[0248] Twelve potential beneficial mutations for MMLV RT were identified by sequence alignment with EvoTfl RT (FIG. 3A). 10 out of 12 mutations identified through sequence alignment also align well in structural alignment (FIG. 3B).Table 7: Exemplary RT SequencesName mRNA SEQUENCE (SEQ ID NO) AA SEQUENCE (SEQ ID NO) EVO TF1 235 232TF1 RT 236 233MMLV RT 237 234
[0249] K103R and I125V are considered more conservative substitutions because both Tfl RT and MMLV RT naturally contain K and I at the aligned positions. In the evolution of Tfl into EvoTfl, K was replaced with R and I with V. Therefore, introducing K103R and I125V in MMLV RT may increase the likelihood of achieving a similar effect to that observed in EvoTfl RT. To further investigate, K103R and I125V were fixed while scanning the other 10 candidate beneficial mutations individually.Table 8: First Round SyNTase EngineeringName Description mRNA SEQUENCE AA SEQUENCE (SEQ ID NO) (SEQ ID NO) pNS39 MMLV RT editor 250 238pEY034 pNS39+K103R+I125V 251 239pEY037 pNS39+K103R+I125V+C157Q 252 240pEY038 PNS39+K103R+I125V+R278L 253 241pEY039 pNS39+K103R+I125V+P297Q 254 242pEY040 PNS39+K103R+I125V+K306Q 255 243pEY041 PNS39+K103R+I125V+E372V 256 244pEY042 pNS39+K103R+I125V+T420E 257 245pEY043 PNS39+K103R+I125V+S67T 258 246pEY044 PNS39+K103R+I125V+E69V 259 247pEY045 pNS39+K103R+I125V+Q84G 260 248pEY046 pNS39+K103R+I125V+G490N 261 249
[0250] The C157Q mutation in the MMLV RT editor pEY37 enables read-through of the fully modified template, but pEY37 still exhibits 2-fold lower efficiency compared to the control (FIG. 4A-FIG. 4B).
[0251] As shown in FIG. 5A-FIG. 5B, the C157Q mutation stabilizes the dNTP pocket by forming three hydrogen bonds with DI 53. This likely stabilizes the enzyme’s structure, allowing it to physically accommodate chemically modified RNA templates (e.g., 2'-O-Methyl, 2'-Fluoro) that are rigid or bulky. By reducing steric clashes, the mutation ensures the templateprimer duplex is properly positioned in the active site. Once positioned, the YVDD (SEQ ID NO: 265) motif catalyzes DNA strand extension using magnesium ions, retaining its essential role in forming phosphodiester bonds.
[0252] Collectively incorporating experimentally validated beneficial mutations in MMLV RT, along with structure-guided mutations predicted to stabilize interactions with the substrate DNA / RNA by forming new hydrogen bonds, SyNTases as shown in Table 9A-Table 9B were designed and tested.Table 9A: First Round SyNTase Engineering (Part 2)Name Description RationalepEY056 V101RD200C V101R improves electrostatic interactions, D200C improves thermal stabilitypEY047 T128N D200C V223Y Evolved mutations with ability to readthrough complex template structures pEY048 T128N D200C V223YV101R Combinations of pEY056 and pEY047 pEY049 K103R I125V T128N D200C V223Y Combinations of pEY048 and K103R and V101R I125VpEY051 T128N D200C V223YV101R Add K397R on pEY048 to enhance polar K397R interactions / form H-bondpEY052 T128N D200C V223YV101R Add K193R on pEY048 to enhance polar K193R interactions / form H-bondpEY053 T128N D200C V223YV101R Add D114N on pEY048 to enhance polar D114N interactions / form H-bondpEY055 T128N D200C V223YV101R Add E302T on pEY048 to enhance polarE302T interactions / form H-bondTable 9B: First Round SyNTase Engineering (Part 2)Name Description mRNA SEQUENCE AA SEQUENCE (SEQ ID NO) (SEQ ID NO) pEY056 V101RD200C 274 266pEY047 T128N D200C V223Y 275 267pEY048 T128N D200C V223YV101R 276 268pEY049 K103R I125V T128N D200C V223 Y V101R 277 269pEY051 T128N D200C V223Y V101R K397R 278 270pEY052 T128N D200C V223 Y V101R K193R 279 271pEY053 T128N D200C V223Y V101RD114N 280 272pEY055 T128N D200C V223Y V101R E302T 281 273
[0253] As shown in FIG. 6, structure-guided engineering shows E302T and K397R form H-bond with DNA / RNA substrate. K193, K397 and E302 are within 3A of the DNA / RNA substrate. Both K193R and K397R form two hydrogen bonds (H-bonds) to the RNA backbone, enhancing electrostatic interactions with the phosphate groups and reducing RNA flexibility for proper template alignment. E302T replaces a negatively charged glutamate with a neutral threonine, eliminating electrostatic repulsion with the DNA backbone while its hydroxyl group forms one H-bond, fine-tuning DNA positioning. Collectively, these mutations optimize charge balance and backbone rigidity, stabilizing the enzyme-substrate complex and boosting polymerase activity through enhanced processivity and catalytic alignment.
[0254] As shown in FIG. 7, pEY47, pEY051, and pEY055 show improved editing efficiency over pNS039 with a fully modified template, but significantly lower efficiency compared to pNS039 with an unmodified template.
[0255] Several SyNTases were tested using partially or fully modified templates. As shown in FIG. 8, pEY037, pEY047, pEY051, pEY052, pEY055 are top candidates that improve read through of modified templates.
[0256] The top MMLV RT editors selected from the first-round screening were pEY037, pEY047, pEY052, and pEY055. The L435K and V223M mutations are included to enhance activity at low dNTP concentrations in non-dividing cells. Additionally, single mutations of key amino acids, such as C157Q and N200C, were also designed and tested (Table 10 and FIG.9).Table 10: Second Round SyNTase EngineeringName Description mRNA AA SEQUENCE SEQUENCE(SEQ ID NO) (SEQ ID NO) pNS055 N200C 316 282 pNS056 V223M L435K 317 283 pNS057 N200C V223M L435K 318 284 pNS058 C157Q V223ML435K 319 285 pNS059 C157Q N200C V223M L435K 320 286 pNS060 C157Q 321 287 pNS061 C157Q N200C 322 288 pNS062 K103R I125V C157Q N200C 323 289 pNS063 K103RI125V C157Q K193R 324 290 pNS064 K103RI125V C157Q E302T 325 291 pNS065 K103R I125V C157Q N200C K193R 326 292 pNS066 K103R I125V C157Q N200C E302T 327 293 pNS067 K103R I125V C157Q K193R E302T 328 294 pNS068 K103R I125V C157Q T128N N200C V223Y 329 295 pNS069 K103R I125V C157Q T128N N200C V223M 330 296pNS070 K103R I125V C157Q N200C K193R E302T 331 297pNS071 K103R I125V C157Q T128N N200C V223M L435K 332 298 pNS072 T128N N200C V223M 333 299 pNS073 C157Q T128N N200C V223Y 334 300 pNS074 C157Q T128N N200C V223M 335 301 pNS075 C157Q T128N N200C V223Y L435K 336 302 pNS076 C157Q T128N N200C V223M L435K 337 303 pNS077 C157Q K193R 338 304 pNS078 C157Q E302T 339 305 pNS079 C157Q K193R E302T 340 306 pNS080 C157Q N200C K193R 341 307 pNS081 C157Q N200C E302T 342 308 pNS082 C157Q N200C K193R E302T 343 309 pNS083 C157Q K193R T128N N200C V223Y 344 310 pNS084 C157Q K193R T128N N200C V223M 345 311 pNS085 C157Q E302T T128N N200C V223Y 346 312 pNS086 C157Q E302T T128N N200C V223M 347 313 pNS087 C157Q K193R E302T T128N N200C V223Y 348 314pNS088 C157Q K193R E302T T128N N200C V223M 349 315
[0257] Top SyNTase variants, such as pNS073, pNS074, and pNS075, can read through a fully modified RNA template with efficiency comparable to that of the positive control, which uses the MMLV RT editor (pNS039) with an unmodified RNA template. In contrast, the MMLV RT editor itself fails to read through the fully modified RNA template, losing nearly all editing efficiency and performing at the level of the negative control. A ~2-fold increase in SyNTase efficiency is achieved compared to the previous best SyNTase, pEY037. A single C157Q mutation contributes 25-75% of the SyNTase efficiency observed in all C157Q-containing constructs for reading through fully modified RNA, however, other mutations such as N200C, E302T, K193R, I125V, K103R etc. are needed for full read-through. FIG. 10-FIG. 11 display results using unmodified or modified template with leading SyNTases (Table 11-Table 12). FIG.12 displays data related to relative efficiencies of each of the tested variants disclosed herein using either an unmodified or completely modified template. This screen identified additional mutations such as V223M / Y, L435K, T128N that help with read-through of chemically modified templates.Table 11: Top 10 SyNTase Variants (Unmodified Template)Name MutationspNS081 C157Q N200C E302TpNS079 C157Q K193R E302TpNS064 K103RI125V C157Q E302TpNS058 C157Q V223ML435KpNS066 K103R I125V C157Q N200C E302TpNS077 C157Q K193RpNS062 K103R I125V C157Q N200CpNS074 C157Q T128N N200C V223MpNS055 N200CpNS071 K103R I125V C157Q T128N N200CV223M L435KTable 12: Top 10 SyNTase Variants (Modified Template)Name MutationspNS073 C157Q T128N N200C V223YpNS075 C157Q T128N N200C V223Y L435KpNS085 C157Q E302T T128N N200C V223YpNS081 C157Q N200C E302TpNS074 C157Q T128N N200C V223MpNS076 C157Q T128N N200C V223M L435KpNS080 C157Q N200C K193RpNS061 C157Q N200CpNS084 C157Q K193R T128N N200C V223MpNS060 C157Q
[0258] In summary, pNS081 and pNS074 are among the top 10 ranked hits across three screens: in Huh7 with a fully modified template, in PHH with a fully modified template, and in PHH with an unmodified template. pNS061, pNS073, pNS074, pNS075, pNS076, pNS081, and pNS085 are among the top 10 ranked hits across two screens, using fully modified templates in both Huh7 and PHH.Table 13: Overall Top HitsName MutationspNS073 C157Q T128NN200C V223YpNS075 C157Q T128NN200C V223YL435KpNS085 C157Q E302T T128N N200C V223YpNS081 C157Q N200C E302TpNS074 C157Q T128NN200C V223MpNS076 C 157Q T 128N N200C V223M L435KpNS061 C157Q N200C
[0259] Further engineered Evo Tfl editors (Table 14) still exhibit low editing efficiency with fully chemically modified template RNA (FIG. 13).Table 14: Engineered Evo TF1 EditorsName Description mRNA AA SEQUENCE SEQUENCE(SEQ ID NO) (SEQ ID NO) pAM174 EvoTfl control 372 350 pAM184 evoTFl Ctruncl3 373 351 pAM185 evoTFl Ctrunc23 374 352 pAM216 evoTFl Ctruncl3 N18 375 353 pAM217 evoTFl Ctrunc23 N18 376 354 pAM218 pAM185+Q195C 377 355 pAM219 pAM185+L429K 378 356pAM220 pAM185+Q195C+M214Y 379 357pAM221 p AM 185 +P 130N+M214Y+Q 195 C 380 358 pAM222 pAM 185+P 130N+M214Y+Q 195C+L429K 381 359 pAM223 pAM185+K122R+K321R 382 360 pAM224 pAM185+Q195C+K122R+K321R 383 361 pAM225 pAM185+L429K+K122R+K321R 384 362 pAM226 pAM185+Q195C+M214Y+K122R+K321R 385 363 pAM227 p AM 185 +P 130N+M214Y+Q 195 C+K 122R+K321 R 386 364 pAM228 p AM 185 +P 130N+M214Y+Q 195 C+L429K+K 122R+K321 R 387 365 pAM229 p AM217+K 122R+K321 R 388 366 pAM230 pAM217+Q195C+K122R+K321R 389 367 pAM231 p AM217+L429K+K 122R+K321 R 390 368 pAM232 p AM217+Q 195 C+M214Y+K 122R+K321 R 391 369 pAM233 p AM217+P 130N+M214Y+Q 195 C+K 122R+K321 R 392 370pAM234 p AM217+P 130N+M214Y+Q 195 C+L429K+K 122R+K321 R 393 371
[0260] Screen of template partially or fully chemically modified tagRNA variants in Huh7 cells using EvoTfl RT editor and MMLV RT editor was performed (FIG. 14). Sequences of tagRNAs are shown in Table 15 below.Table 15: tagRNA modification screen with EvoTFlNAME SEQUENCE1-2SEQ ID NO unmodified mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 394 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGrU rCrGrArUrG*mG*mU*mCOme3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 395 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGrU rCrG / i2FA / rUmG*mG*mU*mCAltl-lF_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 396 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCrUmCrG mUrCmG / i2FA / mUrG*mG*mU*mCAlt2-lF_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 397 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCmUrCmGr UmCrG / i2FA / rUmG*mG*mU*mCfull-mod_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 398 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUmCmG / i2FA / mUmG*mG*mU*mCminl_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 399 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUmCrG / i2FA / mUmG*mG*mU*mCmin2_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 400 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUrCrG / i2FA / mUmG*mG*mU*mCmin3_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 401 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCmGrUrCrG / i2FA / mUmG*mG*mU*mCmin4_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 402 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCr GrUrCrG / i2FA / mUmG*mG*mU*mCmin5_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 403 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUrCrGr UrCrG / i2FA / mUmG*mG*mU*mCmin6_Ome3-lF-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 404 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCrUrCrGr UrCrG / i2FA / mUmG*mG*mU*mCFB Sonly_Ome3 - 1F-2 mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 405 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGrU rCrG / i2FA / mUmG*mG*mU*mCOme3_No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 406 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGrU rCrGrArUmG*mG*mU*mCAltl - IF No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 407 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCrUmCrG mUrCmGrAmUrG*mG*mU*mCA112- IF No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 408 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCmUrCmGr UmCrGmArUmG*mG*mU*mCfull-mod No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 409 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUmCmGmAmUmG*mG*mU*mCminl No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 410 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUmCrGmAmUmG*mG*mU*mCmin2_No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 411 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUrCrGmAmUmG*mG*mU*mCmin3_No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 412 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GrUrCrGmAmUmG*mG*mU*mCmin4_No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 413 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCr GrUrCrGmAmUmG*mG*mU*mCmin5_No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 414 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUrCrGr UrCrGmAmUmG*mG*mU*mCmin6_No-fluoro mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 415 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCrUrCrGr UrCrGmAmUmG*mG*mU*mCFBSonly Ome mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 416 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGrUrCrGmAmUmG*mG*mU*mCAltl - IF fullyMod mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 417 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmC / i2FU / m C / i2FG / mU / i2FC / mG / i2FA / mU / i2FG / *mG*mU*mCAlt2-lF_fullyMod mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 418 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FC / mU / i2 FC / mG / i2FU / mC / i2FG / mA / i2FU / mG*mG*mU*mCminl fullyMod mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 419 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmUmC / i2FG / / i2FA / mUmG*mG*mU*mCmin2_fullyMod mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 420 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm GmU / i2FC / / i2FG / / i2FA / mUmG*mG*mU*mCmin3_fullyMod mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 421 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmCm G / i2FU / / i2FC / / i2FG / / i2FA / mUmG*mG*mU*mCfull-mod_purines mC*mA*mU*rArArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 422 mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmCmUmC / i 2FG / mUmC / i2FG / / i2FA / mUmG*mG*mU*mC1For RNA, T is U; U are T in the sequence listing.2For the sequences: *, phosphorothioate modification; mN,2’0-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A, C, T / U, G.Example 2mRNAs Encoding Editors
[0261] Described in this Example are combinations of lead SyNTase, SsaCas9, Linkers and UTR for Lead Editor designs. Sequence information for these components can be found in Tables 1-12 above.Table 16: Exemplary mRNAs5’ and 3’ UTRs SsaCas9 Linkers RTsm007 and HBA Ssa H605A and Linker 15 pNS074 (SynRT74) N628A (EAAAK)x3 T128N, N200C,V223M, C157Q CESvar4 and AES- Ssa AR_0012 N628A Linker 16 pNS081 (SynRT81) mtRNRl (PAPAP)x3 N200C, C157Q,E302TCESvar4 and HBB- Ssa AR_0012 H605Aext and N628AALB and SAA2Table 17A: mRNA ConstructsName Description mRNA AA SEQUENCE SEQUENCE(SEQ ID NO) (SEQ ID NO) pEY167 Syn 5'UTR SSaAR012 Ng22A Linkl5 SynRT EY037 hHBA 472 423 pEY176 m007 SsaAR012 N622A Linkl5 SynRT074 HBA pA4 473 424 pEY177 m007 SsaAR012 N622A Linkl6 SynRT074 HBA pA4 474 425 pEY178 m007 SsaAR012 N622A Linkl5 SynRT081 HBA pA4 475 426pEY179 m007 SsaAR012 N622A Linkl6 SynRT081 HBA pA4 476 427pEY180 m007 SsaAR012 H599A N622A Linkl5 SynRT074 HBA pA4 477 428 pEY181 m007 SsaAR012 H599A N622A Linkl6 SynRT074 HBA pA4 478 429 pEY182 m007 SsaAR012 H599A N622A Linkl5 SynRT081 HBA pA4 479 430 pEY183 m007 SsaAR012 H599A N622A Linkl6 SynRT081 HBA pA4 480 431 pEY184 m007 nSsaCas9 H605A N628A Linkl5 SynRT074 HBA pA4 481 432 pEY185 m007 nSsaCas9 H605A N628A Linkl6 SynRT074 HBA pA4 482 433 pEY186 m007 nSsaCas9 H605A N628A Linkl5 SynRT081 HBA pA4 483 434 pEY187 m007 nSsaCas9 H605A N628A Linkl6 SynRT081 HBA pA4 484 435 pEY188 CESvar4 SsaAR012 N622A Linkl5 SynRT074 AES-mtRNRl pA4 485 436 pEY189 CESvar4 SsaAR012 N622A Linkl6 SynRT074 AES-mtRNRl pA4 486 437 pEY190 CESvar4 SsaAR012 N622A Linkl5 SynRT081 AES-mtRNRl pA4 487 438 pEY191 CESvar4 SsaAR012 N622A Linkl6 SynRT081 AES-mtRNRl pA4 488 439 pEY192 CESvar4_SsaAR012_H599A_N622A_Linkl5_SynRT074_AES- 489 440 mtRNRl pA4pEY193 CESvar4_SsaAR012_H599A_N622A_Linkl6_SynRT074_AES- 490 441 mtRNRl pA4pEY194 CESvar4_SsaAR012_H599A_N622A_Linkl5_SynRT081_AES- 491 442 mtRNRl pA4pEY195 CESvar4_SsaAR012_H599A_N622A_Linkl6_SynRT081_AES- 492 443 mtRNRl pA4pEY196 CESvar4_nSsaCas9_H605A_N628A_Linkl5_SynRT074_AES- 493 444 mtRNRl pA4pEY197 CESvar4_nSsaCas9_H605A_N628A_Linkl6_SynRT074_AES- 494 445 mtRNRl pA4pEY198 CESvar4_nSsaCas9_H605A_N628A_Linkl5_SynRT081_AES- 495 446 mtRNRl pA4pEY199 CESvar4_nSsaCas9_H605A_N628A_Linkl6_SynRT081_AES- 496 447 mtRNRl pA4pEY200 CESvar4 SsaAR012 N622A Linkl5 SynRT074 HBBext pA4 497 448 pEY201 CESvar4 SsaAR012 N622A Linkl6 SynRT074 HBBext pA4 498 449 pEY202 CESvar4 SsaAR012 N622A Linkl5 SynRT081 HBBext pA4 499 450 pEY203 CESvar4 SsaAR012 N622A Linkl6 SynRT081 HBBext pA4 500 451 pEY204 CESvar4 SsaAR012 H599A N622A Linkl5 SynRT074 HBBext pA4 501 452 pEY205 CESvar4 SsaAR012 H599A N622A Linkl6 SynRT074 HBBext pA4 502 453 pEY206 CESvar4 SsaAR012 H599A N622A Linkl5 SynRT081 HBBext pA4 503 454 pEY207 CESvar4 SsaAR012 H599A N622A Linkl6 SynRT081 HBBext pA4 504 455 pEY208 CESvar4 nSsaCas9 H605A N628A Linkl5 SynRT074 HBBext pA4 505 456 pEY209 CESvar4 nSsaCas9 H605A N628A Linkl6 SynRT074 HBBext pA4 506 457 pEY210 CESvar4 nSsaCas9 H605A N628A Linkl5 SynRT081 HBBext pA4 507 458 pEY211 CESvar4 nSsaCas9 H605A N628A Linkl6 SynRT081 HBBext pA4 508 459 pEY212 ALB SsaAR012 N622A Linkl5 SynRT074 SAA2 pA4 509 460 pEY213 ALB SsaAR012 N622A Linkl6 SynRT074 SAA2 pA4 510 461 pEY214 ALB SsaAR012 N622A Linkl5 SynRT081 SAA2 pA4 511 462 pEY215 ALB SsaAR012 N622A Linkl6 SynRT081 SAA2 pA4 512 463 pEY216 ALB SsaAR012 H599A N622A Linkl5 SynRT074 SAA2 pA4 513 464 pEY217 ALB SsaAR012 H599A N622A Linkl6 SynRT074 SAA2 pA4 514 465 pEY218 ALB SsaAR012 H599A N622A Linkl5 SynRT081 SAA2 pA4 515 466 pEY219 ALB SsaAR012 H599A N622A Linkl6 SynRT081 SAA2 pA4 516 467 pEY220 ALB nSsaCas9 H605A N628A Linkl5 SynRT074 SAA2 pA4 517 468 pEY221 ALB nSsaCas9 H605A N628A Linkl6 SynRT074 SAA2 pA4 518 469 pEY222 ALB nSsaCas9 H605A N628A Linkl5 SynRT081 SAA2 pA4 519 470pEY223 ALB nSsaCas9 H605A N628A Linkl6 SynRT081 SAA2 pA4 520 471Table 17B: tagRNA SequencesNAME SEQUENCE1-2SEQ ID NOSsa_s22_29dl_F6E7 HiA*mU*mA*rArGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmU 527 IF TtoC mUmGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmAr GrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGr UrCrG / i2FA / rUrG*mG*mU*mCSsa_s21_vl7dl_29_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmU 528 F6E7_TtoC mGmUmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGr ArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrCrUrCrGrUr CrGrArUrG*mG*mU*mC1For RNA, T is U; U are T in the sequence listing.2For the sequences: *, phosphorothioate modification; mN,2’0-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A, C, T / U, G.
[0262] Screening top AATD mRNA with fully modified tagRNA was performed in Huh7 (FIG. 25). Shown in FIG. 26 is data related to the best mRNA components from the screen in Huh7 with fully modified template. Shown in FIG. 27-FIG. 28 are results from screening top 48 AATD mRNA with standard tagRNA in Huh7 at different concentrations of editor (e.g., SyNTase editor). Screening top 48 AATD mRNA with standard tagRNA in PHH cells is shown in FIG. 29.In vivo studies with pEY167
[0263] pEY167 RT was selected from RT Engineering round 1 above, which showed improvements in RT read through (K103RI125V C157Q). Experiments were performed in female NSG-PiZ mice. Result is shown in FIG. 30. Results from a dose response curve performed in male NSG-PiZ mice is shown in FIG. 31. In these experiments, the tagRNA did not contain any modifications.Example 3Testing in the Context of Wilson’s Disease
[0264] In this Example, mutations to enable read-through of fully modified editing templates were tested in two reverse transcriptases: truncated MMLV RT and truncated Pipistrellus kuhlii (Pku) RT (Table 18A-Table 18B). Mutations were mapped from MMLVRT to Pku RT.Table 18A: Exemplary Syntase RT SequencesDescription mRNA SEQUENCE AA SEQUENCE(SEQ ID NO) (SEQ ID NO)MMLV-trunc 622 529Pku-trunc 623 530MMLV-trunc-C157Q 624 531MMLV-trunc-K103R I125V C157Q 625 532Pku-trunc-C154Q 626 533Pku-trunc-KlOOR I122V C154Q 627 534Table 18B: Exemplary Syntase SequencesDescription mRNA SEQUENCE AA SEQUENCE(SEQ ID NO) (SEQ ID NO)MMLV-trunc 634 628Pku-trunc 635 629MMLV-trunc-C157Q 636 630MMLV-trunc -K103R I125V C157Q 637 631Pku-trunc-C154Q 638 632Pku-trunc-KlOOR I122V C154Q 639 633
[0265] A range of editing template modification patterns were tested (2’OMe substitution of RNA): Fully modified, Every other nucleotide, Every two out of three nucleotides, No modification (all RNA), Every nucleotide in the Editing Template one-by-one (“walk”). See, e.g., Table 19A-Table 19B below. These constructs were tested in Huh7 H1069Q cells for correction of the Wilson’s disease H1069Q mutation.Table 19 A: tagRNA SequencesNAME Desc. SEQUENCE1-2SEQ ID NO VC_gPL2 All mod mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrU rUrUrUrArGrA 535 Al mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmU mGmCmGmU mGmAm AmCm AmCm CmCmCmUmUmG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Every other mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 536 _A2 walk 1 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmU mGmCmGrU mGr AmArCm ArCmCrC mCrUmUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Every other mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 537 _A3 walk 2 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmU mGmCrGmU rGmAr AmCr AmCrCmC rCmUrUmG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 2 / 3 walk 1 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrAr ArGrU rUrUrUrArGrA 538 _A4 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmU mGmCmGmU rGm AmArCmAmCrCm CmCrUmUmG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mU VC_gPL2 2 / 3 walk 2 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrAr ArGrU rUrUrUrArGrA 539 _A5 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmU mGmCmGrU mGm Ar AmCmArCmCm CrCmUmUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mU VC_gPL2 2 / 3 walk 3 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrAr ArGrU rUrUrUrArGrA 540 _A6 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmU mGmCrGmU mGr Am AmCr AmCmCr CmCmUrUmG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mU VC_gPL2 None mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 541 _A7 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCr AmCmCmGr ArGrU rCrGrGmUmGmCrGrUrGr Ar ArCrArCrCrCrCrUrUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walkl mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 542 _A8 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCmGrUrGrArArCrArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk2 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 543 _A9 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGmUrGrArArCrArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk3 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 544 _A10 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUmGrArArCrArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk4 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 545 All mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGmArArCrArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk5 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 546 _A12 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrAmArCrArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk6 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 547 Bl mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArAmCrArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk7 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 548 _B2 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCmArCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk8 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 549 _B3 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrAmCrCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walk9 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 550 _B4 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrArCmCrCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 WalklO mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 551 _B5 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrArCrCmCrCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walkl 1 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 552 _B6 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrArCrCrCmCrU rUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walkl2 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 553 _B7 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrArCrCrCrCmUrUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walkl 3 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 554 _B8 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrArCrCrCrCrU mUrG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mUVC_gPL2 Walkl4 mU*mU*mG*rGrUrGrArCrUrGrCrCrArCrGrCrCrCrArArGrUrUrUrUrArGrA 555 _B9 mGmCmUmAmGmAmAmAmUmAmGrCrArArGrUrUrArArArArUrArArGrG rCrUrArGrUrCrCrGrUrUrArUrCrArAmCmUmUmGmAmAmAmAmAmGrUr GmGrCrAmCmCmGrArGrUrCrGrGmUmGmCrGrUrGrArArCrArCrCrCrCrUr UmG / i2FG / rGrCmGmU / i2FG / mGmCmU*mU*mU*mU1For RNA, T is U; U are T in the sequence listing.2For the sequences: *, phosphorothioate modification; mN,2’0-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A, C, T / U, G.Table 19B: Editing Template SequencesNAME Description SEQUENCE1-2SEQ ID NO VC gPL2 Al All mod mGmU mGmAm AmCmAmCmCmCmCmU mU mG 556VC gPL2 A2 Every other walk 1 mGrU mGrAm ArCmArCmCrCmCrU mU rG 557VC gPL2 A3 Every other walk 2 rGmU rGm ArAmCr AmCrCmCrCmU rU mG 558VC gPL2 A4 2 / 3 walk 1 mGmU rGmAm ArCm AmCrCmCmCrU mU mG 559VC gPL2 A5 2 / 3 walk 2 mGrU mGmAr AmCmArCmCmCrCmU mU rG 560VC gPL2 A6 2 / 3 walk 3 rGmU mGr AmAmCr AmCmCrCmCmU rU mG 561VC gPL2 A7 None rGrUrGrArArCrArCrCrCrCrUrUrG 562VC gPL2 A8 Walkl mGrUrGrArArCrArCrCrCrCrUrUrG 563VC gPL2 A9 Walk2 rGmUrGrArArCrArCrCrCrCrUrUrG 564VC gPL2 A10 Walk3 rGrUmGrArArCrArCrCrCrCrUrUrG 565VC gPL2 All Walk4 rGrUrGmArArCrArCrCrCrCrUrUrG 566VC gPL2 A12 Walk5 rGrUrGrAmArCrArCrCrCrCrUrUrG 567VC gPL2 Bl Walk6 rGrUrGrArAmCrArCrCrCrCrUrUrG 568VC gPL2 B2 Walk7 rGrUrGrArArCmArCrCrCrCrUrUrG 569VC gPL2 B3 Walk8 rGrUrGrArArCrAmCrCrCrCrUrUrG 570VC gPL2 B4 Walk9 rGrUrGrArArCrArCmCrCrCrUrUrG 571VC gPL2 B5 WalklO rGrUrGrArArCrArCrCmCrCrUrUrG 572VC gPL2 B6 Walkl 1 rGrUrGrArArCrArCrCrCmCrUrUrG 573VC gPL2 B7 Walkl2 rGrUrGrArArCrArCrCrCrCmUrUrG 574VC gPL2 B8 Walkl3 rGrUrGrArArCrArCrCrCrCrUmUrG 575VC gPL2 B9 Walkl4 rGrUrGrArArCrArCrCrCrCrUrUmG 5761For RNA, T is U; U are T in the sequence listing.2For the sequences: *, phosphorothioate modification; mN,2’0-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A, C, T / U, G.
[0266] The C 157Q mutation in MMLV RT and the C 154Q mutation in Pku RT enable read-through of editing templates fully substituted with 2’0Me (FIG. 32-FIG. 33).Example 4Editing Systems for AATD
[0267] Described in this Example are designs and data of editing systems (e.g., SyNTase editing) for editing SERPINA1 gene (e.g., for treatment of AATD).
[0268] The following study (shown in FIG. 34A-FIG. 34C) used a tagRNA with the same nucleotide sequence in its spacer, scaffold, ET and FBS but different chemical modification patterns in its ET and FBS (FIG. 34A). In FIG. 34B-FIG. 34C, an editing system comprising SynRT81 (SynRT81 editor) was used. Experiments were performed in Huh7 cells. The results show that in both instances, the SynRT81 editor could read the chemically modified template andin some cases, such as, ET5F pattern, the efficacy is enhanced relative to control guide containing standard modifications.
[0269] Next, an in vivo study was performed with female NSG-PiZ mice with 0.1 mpk of pEY190 mRNA (comprising SynRT81 polymerase) and the tagRNA shown in FIG. 35. The study identified an ET5F chemical modification pattern that improved editing compared to the control (designated P6R10). The cm51 pattern, which is fully chemically modified, also showed some measurable activity in vivo suggesting RT81 can read through chemical modification in vivo as well.
[0270] Next, scaffold modification optimization studies were performed. FIG. 36A shows editing efficacy of tagRNAs containing scaffolds listed in Table 21. FIG. 36B-FIG. 36E shows the four designs that were considered for further evaluation as highlighted (the stars) in FIG. 36 A. The variants have extension of the repeat: anti -repeat (loop) region with strong base pairing, and show higher efficacy in vitro relative to v29 scaffold.
[0271] These four selected scaffold variants were then tested in vivo in female NSG-PiZ mice at 0.1 mpk. (FIG. 37) v29_10, and v29_26 scaffold variants improved editing efficacy in vivo by 2-3 fold.
[0272] FIG. 38 shows data from an in vivo screen for selection of the scaffold and chemical modifications. The data show that combining ET5F chemical modifications pattern (5 2’ fluoro modification) in the editing template of the tagRNAs with the two scaffold variants v29_10 and v29_26 enhanced editing in vivo even further. Comparison of RNA vs DNA editing (FIG. 39) is also shown, where editing is higher at the RNA level compared to DNA, however, the trends are the same. Overall, s21_v29_26_dl_ET5F tagRNA performed the best and was selected as lead.
[0273] Next, an in vitro mRNA screen for editing in Huh7 cells was performed. FIG.40A shows combinations of four 5’ and 3’ UTRs, three SsaCas9 variants, two linkers, and two Syntases (Syn74 and Synt81). These were tested with fully modified tagRNAs in Huh7 cells (FIG.40B) and standard tagRNA in PHH cells (FIG. 40C). Additional information regarding the mRNA sequences tested can be found, for example, in Table 17 A. The mRNAs that performed best with both fully modified tagRNAs and standard tagRNA in Huh7 and PHH cells were selected for in vivo evaluation.
[0274] These mRNAs were then tested in vivo (FIG. 41). It was observed that pEY190, pEY191 and pEY214 were the top performers. In general, SynRT81 performed better than SynRT74, and single nickase mutant (N622A) SsaCas9 performed better than the double nickase mutant (N622A, H599A). Overall pEY190 performed best in this in vivo screen and therefore was chosen for further optimization. Next, the impact of codon optimization was tested in the contextof mRNA pEY192. pEY192 was codon optimized by two different algorithms for high expression in liver. The codon optimized constructs (pAM246 and pAM250) showed higher editing in vitro in PHH cells (FIG. 42). These mRNAs were further tested in vivo in NSG-PiZ mice (FIG. 43), and pAM246 is shown to perform better than pEY192.
[0275] pEY 190 and pEY 194 were then codon optimized using the same algorithm that generated pAM246, and the resulting mRNAs are pAM320 and pAM321. These mRNAs were also tested in vivo for DNA and RNA editing. As expected, the codon optimizations improved editing efficiencies in vivo by up to 3 -fold (FIG. 44 and 45).
[0276] Two additional mouse studies were conducted with the SsaCas9 editor pAM320 and tagRNA.Study 1
[0277] mRNA encoding SsaCas9 RT editor (pAM320) and tagRNA was mixed at a 1:1 w / w ratio in an aqueous buffer. A total RNA dose of 0.5, 0.25, 0.1 and 0.05 mg / kg of the formulated LNP was delivered intravenously to NSG-PiZ mice carrying the E342K (PiZ) mutation. Seven days post-administration, the mice were sacrificed, and liver samples were collected for analysis. Genomic DNA and RNA were then extracted from the liver, and primers targeting the E342K locus were used to amplify this region. The resulting amplicons were analyzed via short-read sequencing using an Illumina MiSeq. Successful editing was indicated by the conversion of a T / A nucleotide to a C / G nucleotide.
[0278] Results (FIG. 53A) show up to 40% correction of the target E342K with the top dose of 0.5 mg / kg, and 20% correction at the lowest dose of 0.05 mg / kg. RNA editing levels show around 90% and 70% correction, respectively.Study 2: Longitudinal Study
[0279] A total RNA dose of 0.25 mg / kg of the formulated LNP was delivered intravenously to six NSG-PiZ mice carrying the E342K (PiZ) mutation, and six mice were injected with PBS. Serum collection started at seven days post-administration, followed by weekly or biweekly collection for 3 months. Serum was then used for a human Alpha- 1 -Antitrypsin ELISA assay to determine levels of total AAT in treated and untreated mice.
[0280] Results (FIG. 53B) show durable upregulation of total serum AAT levels after 9 weeks, with upto 11 -fold increase between treated and untreated mice and sustained levels of serum AAT over 10,000 ug / mL.
[0281] These data together demonstrate that SyNTase editors enable highly efficient correction of SERPINA1-E342K in the liver, resulting in potent and durable upregulation of serum AAT, approaching saturating levels of liver editing at very low doses as evidenced by >90% editing at the mRNA level..
[0282] Table 20 below shows components for combining best RT (SyNTase), Cas9 (Ssa), linkers and UTR from lead Editor.Table 20: mRNA Screen5’ and5’ 3’ UTRs SsaCas9 Linkers RTsm007 and HBA wtSsa H599A and Linker 15 pNS074N628A (EAAAK)x3 T128N, N200C,V223M, C157Q CESvar4 and AES- SsaAR_0012 Linker 16 pNS081 mtRNRl N622A (PAPAP)x3 N200C, C157Q, E302TCESvar4 and HBB- SsaAR_0012ext H599A and N622AALB and SAA2
[0283] Shown in Table 21 below are SEQ ID NOs of scaffolds tested.Table 21: Scaffolds TestedNAME SEQ ID NOv29 640v29- UUCG 6411 6422 6433 6444 6455 6466 6477 6488 6499 65010 65111 65212 65313 65414 65515 65616 65717 65818 65919 66020 66121 66222 66323 66424 66525 66626 66727 66828 66929 67030 67131 67232 67333 67434 67535 67636 67737 67838 67939 68040 68141 68242 68343 68444 685
[0284] Shown in Table 22 below are tagRNAs tested. For tagRNA sequences, the spacer sequences are underlined, and the extension arms are bold. The ET of the extension arm is bold underline. Table 23 A-Table 23B show sequences for editing of SERPINA1 gene.Table 22: tagRNAS TestedNAME SEQUENCE1-2SEQ ID NOEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 686 6R10_cm51 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2FU / mU / i2FC / / i2FU / m C / i2FG / / i2FU / mC / i2FG / / i2FA / mU / i2FG / *mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 687 6R10_cm53 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrC688mAmU / i2FU / mUmU / i2FC / mUmC Zi2FG / mUmC / i2FG / mAmU / i2FG / *mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 688 6R10_cm67 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUmUrCrUmCrGrUmCrG rAmUrG*mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 689 6R10_cm69 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUmUmUrCmUmCrGmUmC rGmAmUrG*mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 690 6R10_cm72 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCrUmCrGmUrC mGrAmUrG*mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 691 6R10_cm74 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCmUrCrGmUrC rGmArUrG*mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 692 6R10_cm75 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCrUmCrGrUmC rGrAmUrG*mG*mU*mCEvo_s21_29dl_P mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 693 6R10_cm77 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCmUmCrGmUmCrGmAmUrG*mG*mU*mCSsa_s21_29dl_P6 mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 694 R10_cm02 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUrCrUrCrGrUrC / i2FG ZZi2FAZmUZi2FGZ*mG*mU*mCSsa_s21_29dl_P6 mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 695 R10_cm03 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCrUrCrGrUrC / i 2FG / / i2FA / mU / i2FG / *mG*mU*mCSsa_s21_29dl_P6 mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 696 R10_cm83 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUmUrCrUmC / i2FG / rUm C / i2FG / / i2FA / mU / i2FG / *mG*mU*mCSsa_s21_29dl_P6 mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 697 R10_cm90 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCrUrC / i2FG / rUr C / i2FG / / i2FA / mUmG*mG*mU*mCSsa_s21_29dl_P6 mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 698 R10_cm92 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCmUmC / i2FG / r UrC / i2FG / / i2FA / mUmG*mG*mU*mCSsa_s21_v29_0_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 699 D1 P6R10 UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUrCrUrCrGrUrCrGrA rUrG*mG*mU*mCSsa_P6R10_s21_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 700 29dl_EtaltPlF UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / rUrU / i2FC / rUrC / i2FGZ rUrC / i2FG / / i2FA / rUrG*mG*mU*mCSsa_P6R10_s21_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 701 29dl_ET5F UmArCrUmCmCmGmAmAmAmGmGmAmArGrCrUrArCrArAmArGrArUrArArG rGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2FU / / i2FU / / i2FC / / i2FU ZrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 702 l_P6R10_cm00 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUr CrUrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 703 l_P6R10_cm02 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUr CrUrCrGrUrCZi2FGZZi2FAZmUZi2FGZ*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 704 l_P6R10_cm03 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUm UmCrUrCrGrUrCZi2FGZZi2FAZmUZi2FGZ*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 705 l_P6R10_cm83 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUmU rCrUmCZi2FGZrUmCZi2FGZZi2FAZmUZi2FGZ*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 706 l_P6R10_cm90 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUm UmCrUrCZi2FGZrUrCZi2FGZZi2FAZmUmG*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 707 l_P6R10_cm92 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUm UmCmUmCZi2FGZrUrCZi2FGZZi2FAZmUmG*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 708 1 P6R10 ET5F U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU r ArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2 FU / / i2FU / / i2FC / / i2FU / rCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_v29_09_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 709 D45 P6R10 U m ArCrU mCmCmGmCmGm AmAm AmGmCmGmGmAm ArGrCrU rArCr Ar Am Ar GrAmUmAmArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUrCr UrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_v29_10_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 710 D45 P6R10 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU rArCr A rAmArGrAmUmAmArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUr UrCrUrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_v29_26_ mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 711 D45 P6R10 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrAmUmAmArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU rUrUrUrCrUrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 712 1 P6R10 ET5F U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2 FU / / i2FU / / i2FU / / i2FC / / i2FU / rCrGrUrCrGrArUrG*mG*mU*mC Ssa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 713 l_P6R10_cm02_ U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ET5F ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2FU / / i2FU / / i2FC / / i2FU / rCrGrUrC / i2FG / / i2FA / mU / i2FG / *mG*mU*mC Ssa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 708 1 P6R10 ET5F U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU rArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2 FU / / i2FU / / i2FC / / i2FU / rCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 714 l_P6R10_cm02_ U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU rArCr A ET5F rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2FU / / i2FU / / i2FC / / i2FU / rCrGrUrC / i2FG / / i2FA / mU / i2FG / *mG*mU*mC Ssa_s21_29_10_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 702 l_P6R10_cm00 U m ArCrU mCmCmCmCmCmGmAmAm AmGmGmGmGmGmAm ArGrCrU rArCr A rAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrUrUrUr CrUrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 715 l_P6R10_cm00 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrU rUrUrCrUrCrGrUrCrGrArUrG*mG*mU*mCSsa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 716 l_P6R10_cm02 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrU rUrUrCrUrCrGrUrC / i2FG / / i2FA / mU / i2FG / *mG*mU*mCSsa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 717 l_P6R10_cm03 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUm UmUmUmCrUrCrGrUrC / i2FG / / i2FA / mU / i2FG / *mG*mU*mC Ssa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 718 l_P6R10_cm83 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUrU rUmUrCrUmC / i2FG / rUmC / i2FG / / i2FA / mU / i2FG / *mG*mU*mC Ssa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 719 l_P6R10_cm90 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUmUmUmUmCrUrC / i2FG / rUrC / i2FG / / i2FA / mUmG*mG*mU*mCSsa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 720 l_P6R10_cm92 U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmUm UmUmUmCmUmC / i2FG / rUrC / i2FG / / i2FA / mUmG*mG*mU*mC Ssa_s21_29_26_d mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUmUmGm 712 1 P6R10 ET5F U m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGmAmArGrCrU r ArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrArArAmUrCmAmU / i2FU / / i2FU / / i2FU / / i2FC / / i2FU / rCrGrUrCrGrArUrG*mG*mU*mC'For RNA, T is U; U are T in the sequence listing.2For the sequences: *, phosphorothioate modification; mN,2’0-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A, C, T / U, G.Table 23 A: Sequences for Editing SERPINA1 GeneNAME Description SEQUENCE1-2SEQID NO tagRNASsa_s21_v29 Full mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrArGrUrUrUmUm 712 _26_dl_P6R Sequence U mGmU m ArCrU mCmCmCmCmCmGmGmAm AmAmCmGmGmGmGmGm 10 ET5F, AmArGrCrUrArCrArAmArGrArUrArArGrGmCmUrUmCrArUrGrCrCrGrAr Benchling ArAmUrCmAmU / i2FU / / i2FU / / i2FU / / i2FC / / i2FU / rCrGrUrCrGrArUrG*mG* ID: mU*mCtagRNA00188, Othername (CMC):AATD-TAG- 00188s21 spacer spacer mU*mA*mA*rGrGrCrUrGrUrGrCrUrGrArCrCrArUrCrGrA 721 v29_26_dl Scaffold rGrU rU rU mU mU mGmU mArCrU mCmCmCmCmCmGmGmAmAmAmCmG 722 mGmGmGmGmAmArGrCrUrArCrArAmArGrArUrArArGrGmCmUrUmCrA rUrGrCrCrGrArArAmUrCmAmUET / i2FU / / i2FU / / i2FU / / i2FC / / i2FU / rCrGrUrCrG 748 FBS rArUrG*mG*mU*mCmRNApAM320 Full mRNA AGACAGAGACCTCGCAGGCCCCGAGAACTGTCGCCCTTCCACGATG 723 AAGCGGACCGCTGACGGCAGCGAGTTCGAGTCCCCCAAGAAGAAG CGGAAGGTGTCCGACCTGGTGCTGGGCCTGGATATTGGGATCGGCT CCGTGGGCGTGGGGATCCTGAACAAGGTCACCGGGGAGATCATCCA CAAGAACAGCAGGATCTTCCCCGCCGCCCAGGCCGAGAATAATGTG GAGCGCCGCATTAATCGGCAGGGGCGGCGGCTGACCCGCCGGAAG AAGCACCGGCGGGTCAGGCTGAATCACCTGTTCGAGGAGTCCGGCC TGATCACCGATTTCACCAACGTGTCCATCAATCTGAATCCCTACCAG CTGAGGGTGAAGGGGCTGACTGATGAGCTGTCCAACGAGGAGCTGT TCATCGCCCTGAAGAACATGGTGAAGCATCGGGGCATCTCCTACCT GGACGATGCCTCCGATGACGGGAACAGCTCCGTGGGGGACTACGCT CAGATCGTGAAGGAGAACAGCAAGCAGCTGGAAACAAAGACACCC GGCCAGATCCAGCTGGAGCGGTACCAGAAGTACGGGCAGCTGAGG GGCGATTTCACAGTGGAGGAGGATGGCAAGAAGCACAGGCTGATC AATGTGTTTCCTACTTCCGCTTACCGGGCCGAGGCTCTGCGGATCCT GCAGACCCAGCAGGAGTTCAACCCTCAGATCACCGATGAGTTCATC AACTCCTATCTGCAGATCCTGACCGGTAAGCGGAAGTACTACCACG GCCCAGGCAATGAGAAGAGTCGGACAGACTATGGGAGGTTTCGGA CAGACGGTACCACCCTCGATAACATCTTCGGGATCCTGATCGGCAA GTGCACCTTCTATCCCGAGGAGTACAGGGCAGCCAAGGCCAGCTAC ACTGCCCAGGAGTTCAACCTGCTGAATGATCTGAACAATCTGACCG TGCCCACCGAAACAAAGAAGCTGAGCGAGGAGCAGAAGAATCAGA TCATCACCAACGTGAAGAACGAGAAGGCCATGGGGCCTGCCAAGCT GTTTAAGTACATCGCCAAGCTGCTGAGCTGCGATGTGGCCGATATC AAGGGGTACAGGATCGACAAGTCCGATAAGGCCGAGATCCACACC TTCGAGGCCTATCGGAAGATGAAGACCCTGGCCACCATTGATATCG AGCAGATGGAGCGGGAGAACCTGGACAAGCTGGCCTATGTGCTGACCCTGAACACCGAGAAGGAGGGCATCCAGGAGGCCCTGGAGCACGAGTTCGCCGATGGCACCTTCAGCCAGGAGCAGATCGACGAGCTCGT GCAGTTCAGGAAGGCCAACAGCTCCATCTTCGGCAAGGGCTGGCAC AGCTTCAGCGTGAAGCTGATGATGGAGCTGATCCCCGAGCTGTACG CCACTTCCGAGGAGCAGATGACTATCCTGACCCGGCTGGGAAAGCA GAAGACTACCTCAAGCTCCAACAAGACCAAGTACATCGATGAGAA GCAGCTGACCGAGGAGATCTACAACCCTGTGGCCGCCAAGAGCGTG CGGCAGGCCATCAAGATCGTCAACGCTGCCATCAAGGAGTACGGGG ATTTCGACAACATCGTGATCGAGATGGCCAGGGAAACAAACGAGG ATGATGAGAAGAAGGCCATCCAGAAGATCCAGAAGGCCAACAAGG ATGAGAAGGACGCCGCCATGCTGAAGGCTGCCAACCAGTACAACG GCAAGGCCGAGCTGCCCCACAGCGTGTTCCATGGCCACAAGCAGCT GGCCACCAAGATCAGGCTGTGGCACCAGCAGGGCGAGCGCTGTCTG TACACCGGCAAGACCATCAGCATCCACGACCTGATCAACAACAGCA ACCAGTTCGAGATCGACcacATCCTGCCTCTGAGCATCACCTTCGATG ACTCCCTGGCCAACAAGGTGCTGGTGTATGCCACCGCAGCTCAGGA GAAGGGCCAGCGCACCCCCTACCAGGCCCTGGACTCCATGGACGAC GCCTGGAGCTTCAGGGAGCTGAAGGCCTTCGTGCGGGACAGCAAGG CTCTGTCTAACAAGAAGAAGGAGTACCTGCTGACCGAGGAGGACAT CAGCAAGTTCGATGTGCGGAAGAAGTTCATCGAGCGGAACCTGGTG GACACCCGGTATGCCAGCAGGGTGGTGCTGAACGCCCTGCAGGAGC ATTTCCGGGTGCACAAGACAGACACCAAGGTGTCTGTGGTGCGCGG CCAGTTCACCTCCCAGCTGCGGCGGCACTGGGGCATTGAGAAGACC AGGGACACCTACCATCATCACGCTGTGGACGCCCTGATCATCGCTG CTTCTTCCCAGCTGAACCTGTGGAAGAAGCAGAAGAACACCCTGGT GTCCTACAGCGAGGATCAGCTGCTGGACATCGAAACAGGCGAGCTG ATCTCTGATGATGAGTACAAGGAGTCCGTGTTCAAGGCCCCATACC AGCACTTTGTGGACACCCTGAAGTCCAAGGAGTTTGAGGATAGCAT TCTGTTTTCCTACCAGGTGGATAGCAAGTTCAATCGGAAGATCTCCG ACGCCACCATCTATGCCACCCGGCAGGCCAAGGTGGGGAAGGACA AGAAGGACGAAACATACGTGCTGGGCAAGATCAAGGACATCTATA GCCAGACCGGCTATGATGCCTTCATCAAGATCTACAAGAAGGACAA GTCCAAGTTCCTGATGTACCGCCATGATCCTCAGACCTTCGAGAAG GTGATTGAGCCCATCCTGGAGAACTACCCCAACAAGGAGCTGAACG AGAAGGGCAAGGAGGTGCCCTGCAATCCCTTCCTGAAGTACAAGGA GGATCATGGCTACATCAGGAAGTACAGCAAGAAGGGGAACGGCCC AGAGATCAAGTCCCTGAAGTACTACGACAGCAAGCTGGGGAATCAC ATCGACATCACCCCTAAGAACAGCAACAACAAGGTGGTGCTGCAGA GCGTGAGCCCCTGGCGGGCTGACGTGTACTTCAACAAGACCACCGG GAAGTACGAGATCCTGGGCCTGAAGTACGCCGACCTGAAGTTCGAG AAGGGAACCGGAACATACAAGATCTCCGAGGAGAAGTACAACGAC ATCAAGATCAAGGAGGGGGTGGACTCCGACTCCGAGTTCAAGTTCA CCCTGTACAAGAATGATCTGCTGCTGATCAAGGACACCGAAACAAA GGAGCAGCAGCTGTTCCGGTTCCTGTCTCGGACCATGCCCAACGTG AAGCACTACGTGGAGCTGAAGCCTTACGACAAGCAGAAGTTTGACG ACAATGAGGAGCTGATCAAGATCCTGGGGATCGTGGCCAAGGGCG GCCAGTGCAAGAAGGGCGTGTCCAAGCCCAACATCAGCATCTACAA GATCCGGACTGATGTGCTGGGCAACCAGCACATCATCAAGAATGAG GGCGACAAGCCCAAGCTGGACTTCGAGGCCGCCGCCAAGGAGGCC GCCGCCAAGGAGGCGGCCGCCAAGACCCTGAATATCGAGGATGAG TACCGTCTGCACGAAACATCCAAGGAGCCCGACGTGAGTCTGGGCT CCACATGGCTGTCCGACTTTCCTCAGGCCTGGGCCGAAACAGGGGG GATGGGCCTGGCCGTGCGCCAGGCCCCCCTGATCATCCCCCTGAAG GCCACCTCCACCCCTGTGAGCATCAAGCAGTACCCCATGTCCCAGG AGGCTCGGCTGGGCATCAAGCCCCACATCCAGCGGCTGCTGGATCA GGGGATCCTGGTGCCCTGCCAGAGCCCCTGGAACACCCCTCTGCTG CCCGTGAAGAAGCCTGGTACCAACGACTACAGGCCAGTGCAGGACC TGAGGGAGGTGAACAAGAGGGTGGAGGACATCCACCCTaccGTGCCT AACCCTTACAACCTGCTGTCTGGCCTGCCCCCCAGCCACCAGTGGTA CACCGTGCTGGATCTGAAGGATGCCTTTTTCCAGCTGCGGCTGCATC CCACCAGTCAGCCCCTGTTCGCCTTCGAGTGGAGGGATCCAGAGAT GGGCATCAGCGGGCAGCTGACCTGGACCCGGCTGCCCCAGGGCTTC AAGAACAGCCCCACCCTGTTCTGCGAGGCCCTGCATCGGGATCTGGCCGACTTCCGGATTCAGCACCCCGACCTGATCCTGCTGCAGTACgtgGATGATCTGCTGCTGGCCGCCACCAGCGAGCTGGACTGCCAGCAGGG CACCCGGGCCCTGCTGCAGACCCTGGGCAACCTGGGCTACCGGGCC TCCGCCAAGAAGGCCCAGATCTGCCAGAAGCAGGTGAAGTACCTGG GCTACCTGCTGAAGGAGGGGCAGCGGTGGCTGACCGAGGCCAGGA AGGAAACAGTGATGGGCCAGCCTACCCCAAAGACCCCTCGGCAGCT GAGGaccTTTCTGGGGAAGGCTGGCTTCTGCCGGCTGTTTATTCCTGG CTTCGCAGAGATGGCTGCCCCTCTGTACCCCCTGACCAAGCCTGGC ACCCTGTTCAACTGGGGCCCCGACCAGCAGAAGGCCTACCAGGAGA TCAAGCAGGCCCTGCTGACCGCCCCAGCCCTGGGCCTGCCTGATCT GACCAAGCCCTTCGAGCTGTTCGTGGACGAGAAGCAGGGCTATGCT AAGGGGGTGCTGACCCAGAAGCTGGGCCCTTGGCGGAGGCCCGTG GCCTACCTGTCCAAGAAGCTGGACCCCGTGGCAGCCGGCTGGCCTC CTTGCCTCAGGATGGTGGCCGCCATCGCCGTCCTGACCAAGGACGC CGGCAAGCTGACCATGGGCCAGCCCCTGGTGATCctgGCTCCCCACG CCGTGGAGGCCCTGGTGAAGCAGCCACCCGACCGGTGGCTGTCCAA CGCCAGGATGACCCACTACCAGGCCCTGCTGCTGGACACCGACCGC GTCCAGTTTGGCCCTGTGGTGGCCCTGAACCCCGCCACTCTGCTGCC ACTGCCCGAGGAGGGCAGTGGCGGCTCCAAGCGCACTGCCGACGG CAGTGAGTTTGAGAGCCCCAAGAAGAAGCGGAAGGTGGGGTCAGG GCCCGCCGCCAAGAGGGTGAAGCTGGACTAATAGTGACTGGTACTG CATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGGTACCCCGAG TCTCCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTCCACCTGCCC CACTCACCACCTCTGCTAGTTCCAGACACCTCCCAAGCACGCAGCA ATGCAGCTCAAAACGCTTAGCCTAGCCACACCCCCACGGGAAACAG CAGTGATTAACCTTTAGCAATAAACGAAAGTTTAACTAAGCTATAC TAACCCCAGGGTTGGTCAATTTCGTGCCAGCCACACCAAAAAAAAA AAAAAATAAAAAAAAAAAAAAAAAAAAAAAAAAATTAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAATTTTAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA CESvar4 5' UTR AGACAGAGACCTCGCAGGCCCCGAGAACTGTCGCCCTTCCACG 724 AES- 3' UTR CTGGTACTGCATGCACGCAATGCTAGCTGCCCCTTTCCCGTCCTGGG 200 mtRNRl TACCCCGAGTCTCCCCCGACCTCGGGTCCCAGGTATGCTCCCACCTC CACCTGCCCCACTCACCACCTCTGCTAGTTCCAGACACCTCCCAAGC ACGCAGCAATGCAGCTCAAAACGCTTAGCCTAGCCACACCCCCACG GGAAACAGCAGTGATTAACCTTTAGCAATAAACGAAAGTTTAACTA AGCTATACTAACCCCAGGGTTGGTCAATTTCGTGCCAGCCACACCpolyAv4 polyA AAAAAAAAAAAAAAATAAAAAAAAAAAAAAAAAAAAAAAAAAAT 137TAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAATTTTAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA CDS ATGAAGCGGACCGCTGACGGCAGCGAGTTCGAGTCCCCCAAGAAG 725AAGCGGAAGGTGTCCGACCTGGTGCTGGGCCTGGATATTGGGATCG GCTCCGTGGGCGTGGGGATCCTGAACAAGGTCACCGGGGAGATCAT CCACAAGAACAGCAGGATCTTCCCCGCCGCCCAGGCCGAGAATAAT GTGGAGCGCCGCATTAATCGGCAGGGGCGGCGGCTGACCCGCCGG AAGAAGCACCGGCGGGTCAGGCTGAATCACCTGTTCGAGGAGTCCG GCCTGATCACCGATTTCACCAACGTGTCCATCAATCTGAATCCCTAC CAGCTGAGGGTGAAGGGGCTGACTGATGAGCTGTCCAACGAGGAG CTGTTCATCGCCCTGAAGAACATGGTGAAGCATCGGGGCATCTCCT ACCTGGACGATGCCTCCGATGACGGGAACAGCTCCGTGGGGGACTA CGCTCAGATCGTGAAGGAGAACAGCAAGCAGCTGGAAACAAAGAC ACCCGGCCAGATCCAGCTGGAGCGGTACCAGAAGTACGGGCAGCT GAGGGGCGATTTCACAGTGGAGGAGGATGGCAAGAAGCACAGGCT GATCAATGTGTTTCCTACTTCCGCTTACCGGGCCGAGGCTCTGCGGA TCCTGCAGACCCAGCAGGAGTTCAACCCTCAGATCACCGATGAGTT CATCAACTCCTATCTGCAGATCCTGACCGGTAAGCGGAAGTACTAC CACGGCCCAGGCAATGAGAAGAGTCGGACAGACTATGGGAGGTTT CGGACAGACGGTACCACCCTCGATAACATCTTCGGGATCCTGATCG GCAAGTGCACCTTCTATCCCGAGGAGTACAGGGCAGCCAAGGCCAG CTACACTGCCCAGGAGTTCAACCTGCTGAATGATCTGAACAATCTG ACCGTGCCCACCGAAACAAAGAAGCTGAGCGAGGAGCAGAAGAAT CAGATCATCACCAACGTGAAGAACGAGAAGGCCATGGGGCCTGCCAAGCTGTTTAAGTACATCGCCAAGCTGCTGAGCTGCGATGTGGCCGATATCAAGGGGTACAGGATCGACAAGTCCGATAAGGCCGAGATCC ACACCTTCGAGGCCTATCGGAAGATGAAGACCCTGGCCACCATTGA TATCGAGCAGATGGAGCGGGAGAACCTGGACAAGCTGGCCTATGTG CTGACCCTGAACACCGAGAAGGAGGGCATCCAGGAGGCCCTGGAG CACGAGTTCGCCGATGGCACCTTCAGCCAGGAGCAGATCGACGAGC TCGTGCAGTTCAGGAAGGCCAACAGCTCCATCTTCGGCAAGGGCTG GCACAGCTTCAGCGTGAAGCTGATGATGGAGCTGATCCCCGAGCTG TACGCCACTTCCGAGGAGCAGATGACTATCCTGACCCGGCTGGGAA AGCAGAAGACTACCTCAAGCTCCAACAAGACCAAGTACATCGATGA GAAGCAGCTGACCGAGGAGATCTACAACCCTGTGGCCGCCAAGAG CGTGCGGCAGGCCATCAAGATCGTCAACGCTGCCATCAAGGAGTAC GGGGATTTCGACAACATCGTGATCGAGATGGCCAGGGAAACAAAC GAGGATGATGAGAAGAAGGCCATCCAGAAGATCCAGAAGGCCAAC AAGGATGAGAAGGACGCCGCCATGCTGAAGGCTGCCAACCAGTAC AACGGCAAGGCCGAGCTGCCCCACAGCGTGTTCCATGGCCACAAGC AGCTGGCCACCAAGATCAGGCTGTGGCACCAGCAGGGCGAGCGCT GTCTGTACACCGGCAAGACCATCAGCATCCACGACCTGATCAACAA CAGCAACCAGTTCGAGATCGACcacATCCTGCCTCTGAGCATCACCTT CGATGACTCCCTGGCCAACAAGGTGCTGGTGTATGCCACCGCAGCT CAGGAGAAGGGCCAGCGCACCCCCTACCAGGCCCTGGACTCCATGG ACGACGCCTGGAGCTTCAGGGAGCTGAAGGCCTTCGTGCGGGACAG CAAGGCTCTGTCTAACAAGAAGAAGGAGTACCTGCTGACCGAGGA GGACATCAGCAAGTTCGATGTGCGGAAGAAGTTCATCGAGCGGAAC CTGGTGGACACCCGGTATGCCAGCAGGGTGGTGCTGAACGCCCTGC AGGAGCATTTCCGGGTGCACAAGACAGACACCAAGGTGTCTGTGGT GCGCGGCCAGTTCACCTCCCAGCTGCGGCGGCACTGGGGCATTGAG AAGACCAGGGACACCTACCATCATCACGCTGTGGACGCCCTGATCA TCGCTGCTTCTTCCCAGCTGAACCTGTGGAAGAAGCAGAAGAACAC CCTGGTGTCCTACAGCGAGGATCAGCTGCTGGACATCGAAACAGGC GAGCTGATCTCTGATGATGAGTACAAGGAGTCCGTGTTCAAGGCCC CATACCAGCACTTTGTGGACACCCTGAAGTCCAAGGAGTTTGAGGA TAGCATTCTGTTTTCCTACCAGGTGGATAGCAAGTTCAATCGGAAG ATCTCCGACGCCACCATCTATGCCACCCGGCAGGCCAAGGTGGGGA AGGACAAGAAGGACGAAACATACGTGCTGGGCAAGATCAAGGACA TCTATAGCCAGACCGGCTATGATGCCTTCATCAAGATCTACAAGAA GGACAAGTCCAAGTTCCTGATGTACCGCCATGATCCTCAGACCTTC GAGAAGGTGATTGAGCCCATCCTGGAGAACTACCCCAACAAGGAG CTGAACGAGAAGGGCAAGGAGGTGCCCTGCAATCCCTTCCTGAAGT ACAAGGAGGATCATGGCTACATCAGGAAGTACAGCAAGAAGGGGA ACGGCCCAGAGATCAAGTCCCTGAAGTACTACGACAGCAAGCTGGG GAATCACATCGACATCACCCCTAAGAACAGCAACAACAAGGTGGTG CTGCAGAGCGTGAGCCCCTGGCGGGCTGACGTGTACTTCAACAAGA CCACCGGGAAGTACGAGATCCTGGGCCTGAAGTACGCCGACCTGAA GTTCGAGAAGGGAACCGGAACATACAAGATCTCCGAGGAGAAGTA CAACGACATCAAGATCAAGGAGGGGGTGGACTCCGACTCCGAGTTC AAGTTCACCCTGTACAAGAATGATCTGCTGCTGATCAAGGACACCG AAACAAAGGAGCAGCAGCTGTTCCGGTTCCTGTCTCGGACCATGCC CAACGTGAAGCACTACGTGGAGCTGAAGCCTTACGACAAGCAGAA GTTTGACGACAATGAGGAGCTGATCAAGATCCTGGGGATCGTGGCC AAGGGCGGCCAGTGCAAGAAGGGCGTGTCCAAGCCCAACATCAGC ATCTACAAGATCCGGACTGATGTGCTGGGCAACCAGCACATCATCA AGAATGAGGGCGACAAGCCCAAGCTGGACTTCGAGGCCGCCGCCA AGGAGGCCGCCGCCAAGGAGGCGGCCGCCAAGACCCTGAATATCG AGGATGAGTACCGTCTGCACGAAACATCCAAGGAGCCCGACGTGA GTCTGGGCTCCACATGGCTGTCCGACTTTCCTCAGGCCTGGGCCGAA ACAGGGGGGATGGGCCTGGCCGTGCGCCAGGCCCCCCTGATCATCC CCCTGAAGGCCACCTCCACCCCTGTGAGCATCAAGCAGTACCCCAT GTCCCAGGAGGCTCGGCTGGGCATCAAGCCCCACATCCAGCGGCTG CTGGATCAGGGGATCCTGGTGCCCTGCCAGAGCCCCTGGAACACCC CTCTGCTGCCCGTGAAGAAGCCTGGTACCAACGACTACAGGCCAGT GCAGGACCTGAGGGAGGTGAACAAGAGGGTGGAGGACATCCACCCTaccGTGCCTAACCCTTACAACCTGCTGTCTGGCCTGCCCCCCAGCCACCAGTGGTACACCGTGCTGGATCTGAAGGATGCCTTTTTCCAGCTGCGGCTGCATCCCACCAGTCAGCCCCTGTTCGCCTTCGAGTGGAGGGA TCCAGAGATGGGCATCAGCGGGCAGCTGACCTGGACCCGGCTGCCC CAGGGCTTCAAGAACAGCCCCACCCTGTTCTGCGAGGCCCTGCATC GGGATCTGGCCGACTTCCGGATTCAGCACCCCGACCTGATCCTGCT GCAGTACgtgGATGATCTGCTGCTGGCCGCCACCAGCGAGCTGGACT GCCAGCAGGGCACCCGGGCCCTGCTGCAGACCCTGGGCAACCTGGG CTACCGGGCCTCCGCCAAGAAGGCCCAGATCTGCCAGAAGCAGGTG AAGTACCTGGGCTACCTGCTGAAGGAGGGGCAGCGGTGGCTGACCG AGGCCAGGAAGGAAACAGTGATGGGCCAGCCTACCCCAAAGACCC CTCGGCAGCTGAGGaccTTTCTGGGGAAGGCTGGCTTCTGCCGGCTG TTTATTCCTGGCTTCGCAGAGATGGCTGCCCCTCTGTACCCCCTGAC CAAGCCTGGCACCCTGTTCAACTGGGGCCCCGACCAGCAGAAGGCC TACCAGGAGATCAAGCAGGCCCTGCTGACCGCCCCAGCCCTGGGCC TGCCTGATCTGACCAAGCCCTTCGAGCTGTTCGTGGACGAGAAGCA GGGCTATGCTAAGGGGGTGCTGACCCAGAAGCTGGGCCCTTGGCGG AGGCCCGTGGCCTACCTGTCCAAGAAGCTGGACCCCGTGGCAGCCG GCTGGCCTCCTTGCCTCAGGATGGTGGCCGCCATCGCCGTCCTGACC AAGGACGCCGGCAAGCTGACCATGGGCCAGCCCCTGGTGATCctgGC TCCCCACGCCGTGGAGGCCCTGGTGAAGCAGCCACCCGACCGGTGG CTGTCCAACGCCAGGATGACCCACTACCAGGCCCTGCTGCTGGACA CCGACCGCGTCCAGTTTGGCCCTGTGGTGGCCCTGAACCCCGCCACT CTGCTGCCACTGCCCGAGGAGGGCAGTGGCGGCTCCAAGCGCACTG CCGACGGCAGTGAGTTTGAGAGCCCCAAGAAGAAGCGGAAGGTGG GGTCAGGGCCCGCCGCCAAGAGGGTGAAGCTGGACProteinFull Protein MKRTADGSEFESPKKKRKVSDLVLGLDIGIGSVGVGILNKVTGEIIHKN 726 Sequence SRIFPAAQAENNVERRINRQGRRLTRRKKHRRVRLNHLFEESGLITDFT NVSINLNPYQLRVKGLTDELSNEELFIALKNMVKHRGISYLDDASDDG NSSVGDYAQIVKENSKQLETKTPGQIQLERYQKYGQLRGDFTVEEDGK KHRLINVFPTSAYRAEALRILQTQQEFNPQITDEFINSYLQILTGKRKYY HGPGNEKSRTDYGRFRTDGTTLDNIFGILIGKCTFYPEEYRAAKASYTA QEFNLLNDLNNLTVPTETI< I< LSEEQI< NQIITNVI< NEI< AMGPAI< LFI< YI AKLLSCDVADIKGYRIDKSDKAEIHTFEAYRKMKTLATIDIEQMERENL DKLAYVLTLNTEKEGIQEALEHEFADGTFSQEQIDELVQFRKANSSIFG KGWHSFSVKLMMELIPELYATSEEQMTILTRLGKQKTTSSSNKTKYIDE KQLTEEIYNPVAAKSVRQAIKIVNAAIKEYGDFDNIVIEMARETNEDDE KKAIQKIQKANKDEKDAAMLKAANQYNGKAELPHSVFHGHKQLATKI RLWHQQGERCLYTGKTISIHDLINNSNQFEIDHILPLSITFDDSLANKVL VYATAAQEKGQRTPYQALDSMDDAWSFRELKAFVRDSKALSNKKKE YLLTEEDISKFDVRKKFIERNLVDTRYASRWLNALQEHFRVHKTDTK VSVVRGQFTSQLRRHWGIEKTRDTYHHHAVDALIIAASSQLNLWKKQ KNTLVSYSEDQLLDIETGELISDDEYKESVFKAPYQHFVDTLKSKEFED SILFSYQVDSKFNRKISDATIYATRQAKVGKDKKDETYVLGKIKDIYSQ TGYDAFIKIYKKDKSKFLMYRHDPQTFEKVIEPILENYPNKELNEKGKE VPCNPFLKYKEDHGYIRKYSKKGNGPEIKSLKYYDSKLGNHIDITPKNS NNKVVLQSVSPWRADVYFNKTTGKYEILGLKYADLKFEKGTGTYKISE EKYNDIKIKEGVDSDSEFKFTLYKNDLLLIKDTETKEQQLFRFLSRTMP NVKHYVELKPYDKQKFDDNEELIKILGIVAKGGQCKKGVSKPNISIYKI RTDVLGNQHIIKNEGDKPKLDFEAAAKEAAAKEAAAKTLNIEDEYRLH ETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPLIIPLKATSTPVS IKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLLPVKKPGTNDY RPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYTVLDLKDAFFQ LRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNSPTLFCEALHRD LADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALLQTLGNLGYRA SAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMGQPTPKTPRQLR TFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPDQQKAYQEIKQ ALLTAP ALGLPDLTKPFELFVDEKQGYAKGVLTQKLGPWRRPVAYLS KKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVILAPHAVEAL VKQPPDRWLSNARMTHYQALLLDTDRVQFGPWALNPATLLPLPEEG SGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLDBPSV40 N-term NLS MKRTADGSEFESPKKKRKV 2nSsaCas9_A Cas9 SDLVLGLDIGIGSVGVGILNKVTGEIIHKNSRIFPAAQAENNVERRINRQ 727 R00 12 GRRLTRRKKHRRVRLNHLFEESGLITDFTNVSINLNPYQLRVKGLTDEL (N622A) SNEELFIALKNMVKHRGISYLDDASDDGNSSVGDYAQIVKENSKQLET KTPGQIQLERYQKYGQLRGDFTVEEDGKKHRLINVFPTSAYRAEALRIL QTQQEFNPQITDEFINSYLQILTGKRKYYHGPGNEKSRTDYGRFRTDGT TLDNIFGILIGKCTFYPEEYRAAKASYTAQEFNLLNDLNNLTVPTETKK LSEEQKNQIITNVKNEKAMGPAKLFKYIAKLLSCDVADIKGYRIDKSDK AEIHTFEAYRKMKTLATIDIEQMERENLDKLAYVLTLNTEKEGIQEALE HEFADGTFSQEQIDELVQFRKANSSIFGKGWHSFSVKLMMELIPELYAT SEEQMTILTRLGKQKTTSSSNKTKYIDEKQLTEEIYNPVAAKSVRQAIKI VNAAIKEYGDFDNIVIEMARETNEDDEKKAIQKIQKANKDEKDAAML KAANQYNGKAELPHSVFHGHKQLATKIRLWHQQGERCLYTGKTISIH DLINNSNQFEIDHILPLSITFDDSLANKVLVYATAAQEKGQRTPYQALD SMDDAWSFRELKAFVRDSKALSNKKKEYLLTEEDISKFDVRKKFIERN LVDTRYASRWLNALQEHFRVHKTDTKVSVVRGQFTSQLRRHWGIEK TRDTYHHHAVDALIIAASSQLNLWKKQKNTLVSYSEDQLLDIETGELIS DDEYKESVFKAPYQHFVDTLKSKEFEDSILFSYQVDSKFNRKISDATIY ATRQAKVGKDKKDETYVLGKIKDIYSQTGYDAFIKIYKKDKSKFLMYR HDPQTFEKVIEPILENYPNKELNEKGKEVPCNPFLKYKEDHGYIRKYSK KGNGPEIKSLKYYDSKLGNHIDITPKNSNNKVVLQSVSPWRADVYFNK TTGKYEILGLKYADLKFEKGTGTYKISEEKYNDIKIKEGVDSDSEFKFT LYKNDLLLIKDTETKEQQLFRFLSRTMPNVKHYVELKPYDKQKFDDNE ELIKILGIVAKGGQCKKGVSKPNISIYKIRTDVLGNQHIIKNEGDKPKLD FLinker 15 Linker EAAAKEAAAKEAAAK 36 SynRT81 RT TLNIEDEYRLHETSKEPDVSLGSTWLSDFPQAWAETGGMGLAVRQAPL 728 IIPLKATSTPVSIKQYPMSQEARLGIKPHIQRLLDQGILVPCQSPWNTPLL PVKKPGTNDYRPVQDLREVNKRVEDIHPTVPNPYNLLSGLPPSHQWYT VLDLKDAFFQLRLHPTSQPLFAFEWRDPEMGISGQLTWTRLPQGFKNS PTLFCEALHRDLADFRIQHPDLILLQYVDDLLLAATSELDCQQGTRALL QTLGNLGYRASAKKAQICQKQVKYLGYLLKEGQRWLTEARKETVMG QPTPKTPRQLRTFLGKAGFCRLFIPGFAEMAAPLYPLTKPGTLFNWGPD QQKAYQEIKQALLTAPALGLPDLTKPFELFVDEKQGYAKGVLTQKLGP WRRPVAYLSKKLDPVAAGWPPCLRMVAAIAVLTKDAGKLTMGQPLVI LAPHAVEALVKQPPDRWLSNARMTHYQALLLDTDRVQFGPWALNP ATLLPLPEEG BPSV40+C- C-term NLS SGGSKRTADGSEFESPKKKRKVGSGPAAKRVKLD 729 myc1For RNA, T is U; U are T in the sequence listing.2For the sequences: *, phosphorothioate modification; mN, 2’0-methyl modification; rN, unmodified ribonucleotide; i2FN, Int Fluoro modification. N=A, C, T / U, G.Table 23B: Sequences for Editing SERPINA1 Gene (tagRNA Sequences without Modifications)NAME SEQ ID NO or SequenceSsa s21 v29 26 dl P6R10 ET5F 745s21 spacer 746v29 26 dl 667ET 747FBS AUGGUCExample 5In vivo Editing of SERPINA1 in Rat
[0285] Described in this Example is proof of concept for editing of SERPINA1 in a rat disease model using the systems and methods disclosed herein. FIG. 46A depicts a non-limiting exemplary schematic related to the generation of a rat disease model bearing the PiZ mutation, showing the use of a targeting vector to replace a wild-type allele with the targeted allele (~4517bpof rat sequence was replaced with 660 Ibp human sequence). This design is based on human transcript-202 (NM_000295.5NP_000286.3) and rat transcript-202 (NM_022519.2 NP _071964.2).
[0286] Heterozygous or homozygous humanized Wistar rats carrying the E342K(PiZ) mutation (Biocytogen) were bled for serum collection 14 days prior to LNP injection. mRNAs encoding the SsaCas9 RT editor (pEY190 [SEQ ID NO: 487 and 438], pEY192 [SEQ ID NO: 489 and 440] or pEY200 [SEQ ID NO: 497 and 448]) and tagRNA (SEQ ID Nos: 528, 699, and 712) were mixed at a 1: 1 w / w ratio in an aqueous buffer. A total RNA dose of either 2 mg / kg (1 mg / kg each of editor and tagRNA) or 1 mg / kg (0.5 mg / kg each) was delivered intravenously via tail vein. Serum was collected 24 hours post-injection for liver function tests (LFTs). Seven days after dosing, rats were sacrificed, and liver and serum samples were collected. Genomic DNA and total RNA were extracted from the liver, and primers targeting the E342K locus were used for amplicon sequencing on an Illumina MiSeq. Editing efficiency was determined by the conversion of a T / A nucleotide to a C / G nucleotide. Total AAT levels in serum were quantified by ELISA at pre-dose and Day 7 post-dose.
[0287] FIGS. 46B-FIG. 52B depict data from dose response studies in this PiZ rat model showing the relationship between precise editing and increasing doses of the total RNA encapsulated in LNP (tagRNA and the mRNA encoding the editor). All animals tolerated the LNP doses with no mortality. LFTs were elevated in all treated rats compared to pre-dose levels, with higher elevations observed at 2 mg / kg. At 2 mg / kg, DNA editing was comparable between pEY192 and pEY200 (45.4% vs 45.0%), whereas at 1 mg / kg, pEY200 outperformed pEY192 (41.8% vs 15.1%), approaching its 2 mg / kg efficiency. RNA editing followed a similar trend. Predose AAT levels were comparable across groups (4.5-4.9 pmol / L). At 2 mg / kg, both constructs achieved similar AAT levels (16.6 vs 17.6 pmol / L). At 1 mg / kg, pEY200 achieved protective AAT levels (>11 pmol / L; 15.9 pmol / L), while pEY192 did not (8.4 pmol / L). AAT levels strongly correlated with DNA editing (R2= 0.97).AATD correction with lead mRNA and tagRNA in PiZ++rats
[0288] mRNAs encoding SsaCas9 RT editor (pAM320) and tagRNA (SEQ ID NO: 712) were mixed at a 1: 1 w / w ratio in an aqueous buffer and injected in the homozygous P. E366K Hu Ki Wistar custom rat model.
[0289] A total RNA dose of 0.5 mg / kg, 0.25 mg / kg, 0.1 mg / kg or 0.05 mg / kg for total RNA, formulated in the LNPs, were delivered intravenously via tail vein (n=3 animals per group). Seven days post-administration, the rats were sacrificed via CO2, liver and serum samples were collected for analysis. Genomic DNA was then extracted from the liver using quick-DNA / RNA MagBead kit following the manufacturer protocol (Zymo research cat# R2130). ), and primers(forward primer: GAGACTCAGAGAAAACATGGGA (SEQ ID NO: 749), reverse primer: CTAAAGGCGACCAATGAACAA (SEQ ID NO: 750)) targeting the E342K locus were used to amplify this region. The resulting amplicons were analyzed via short-read sequencing using an Illumina MiSeq. Successful editing was indicated by the conversion of a T / A nucleotide to a C / G nucleotide. ELISA for total AAT levels was quantified on pre-dose and seven days post injection serum collections using Human Alpha- 1 -Antitrypsin ELISA Kit, following the manufacturer protocol (Fortis life sciences cat# E88-122). LC-MS was performed by Inotiv, Inc on serum samples to quantify mutated Z-AAT (AVLTIDK, SEQ ID NO: 751), Wild Type M-AAT (AVLTIDEK, SEQ ID NO: 752) and total AAT (VVNPTQK, SEQ ID NO: 753) peptides. Briefly, serum was depleted of albumin and protein concentration was determined by BCA. 100 pg of protein was reduced, alkylated, and digested with trypsin. Isotopically labeled peptide standards, corresponding the sequences specific for the WT and Z alleles along with total SERPINA1 were added to all samples at 2 fmol / pg to allow quantification of the respective peptides. Peptides were fractionated by high pH reversed-phase fractionation and the fraction containing the peptides of interest were analyzed on a Thermo Exploris 480 mass spectrometer in PRM mode. An additional group (n=3) was dosed with 0.5 mg / kg mRNA (pAM320) and tagRNA in LNP, and a control group (n=2) received PBS. Both groups were bled weekly for serum collection to assess durability of M-AAT levels.
[0290] All animals tolerated LNP doses, and no deaths were observed. Dosedependent editing ranging from 66% to 18% with near saturating editing at 0.1 mg / kg (52.3%) was detected suggesting a direct correlation between dosage and editing efficiency. Very low indels were detected ranging from 0.8% at 0.5 mg / kg to 0.1% at 0.05 mg / kg. Pre-dose serum total AAT levels measured by ELISA were comparable across groups (11.9 to 16.0 pmol / L). AAT levels followed a similar trend, with dose-dependent decrease levels ranging from 68.9 pmol / L to 25.7 p...
Claims
WHAT IS CLAIMED IS:
1. An engineered polymerase comprising one or more amino acid substitutions as compared to a parent polymerase comprising a sequence selected from SEQ ID NOs: 232-234 and 579-619.
2. The engineered polymerase of claim 1, wherein the parent polymerase comprises the sequence of SEQ ID NO: 234 or SEQ ID NOs: 604-605.
3. The engineered polymerase of any one of claims 1-2, wherein each of the one or more amino acid substitutions is at an amino acid position functionally equivalent to S67, E69, Q84, V101, K103, D114, I125, T128, C157, K193, N200, V223, R278, P297, E302, K306, E372, K397, T420, L435, or G490 of SEQ ID NO: 234.
4. The engineered polymerase of any one of claims 1-3, wherein each of the one or more amino acid substitutions from the group comprising S67T, E69V, Q84G, V101R, K103R, D114N, I125V, T128N, C157Q, K193R, D / N200C, V223Y, V223M, R278L, P297Q, E302T, K306Q, E372V, K397R, T420E, L435K, or G490N of SEQ ID NO: 234.
5. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C157Q amino acid substitution relative to SEQ ID NO: 234.
6. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, and a V223Y amino acid substitution relative to SEQ ID NO: 234.
7. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, a V223Y amino acid substitution, and a L435K amino acid substitution relative to SEQ ID NO: 234.
8. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, a V223 Y amino acid substitution, and an E302T amino acid substitution relative to SEQ ID NO: 234.
9. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C157Q amino acid substitution, an N200C amino acid substitution, and an E302T amino acid substitution relative to SEQ ID NO: 234.
10. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C157Q amino acid substitution, a T128N amino acid substitution, an N200C amino acid substitution, and a V223M amino acid substitution relative to SEQ ID NO: 234.
11. The engineered polymerase of claim 4, wherein the engineered polymerase comprises a C154Q amino acid substitution, an I222V amino acid substitution, a K100R substitution, or any combination thereof, relative to any one of SEQ ID NOs: 604-605.
12. A polynucleotide encoding the engineered polymerase of any one of claims 1-11.
13. A Synthetic Nucleotide Template Polymerase (SyNTase) editor comprising: a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain,wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein,wherein the DNA polymerase domain comprises the engineered polymerase of any one of claims 1-11.
14. The SyNTase editor of claim 13, wherein the SyNTase editor is a reverse transcriptase (RT) editor.
15. A Synthetic Nucleotide Template Polymerase (SyNTase) editing system comprising:a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises:a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule;an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; anda Synthetic Nucleotide Template Polymerase (SyNTase) editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor,wherein the DNA polymerase domain comprises the engineered polymerase of any one of claims 1-11.
16. The SyNTase editing system of claim 15, wherein the SyNTase editing system is reverse transcriptase (RT) editing system.
17. The SyNTase editing system of any one of claims 15-16, wherein the tagRNA comprises a flap binding sequence at least partially complementary to the spacer.
18. The SyNTase editing system of any one of claims 15-17, wherein the scaffold sequence is between the spacer and the editing template.
19. The SyNTase editing system of any one of claims 15-18, wherein the tagRNA comprises from 5’ to 3’: the spacer, the scaffold sequence, the editing template, and the flap binding sequence.
20. The SyNTase editing system of any one of claims 15-19, wherein the spacer, the scaffold sequence, the editing template, and the flap binding sequence form a contiguous sequence in a single molecule.
21. The SyNTase editing system of any one of claims 15-20, wherein the editing template comprises an intended nucleotide edit compared to the double stranded target DNA.
22. The SyNTase editing system of any one of claims 15-21, wherein the tagRNA guides the SyNTase editor to incorporate the intended nucleotide edit into the double stranded target DNA when the tagRNA is contacted with the double stranded target DNA.
23. The SyNTase editing system of any one of claims 15-22, wherein the SyNTase editor synthesizes a single stranded DNA encoded by the editing template, wherein the single stranded DNA replaces the editing target sequence and results in incorporation of the intended nucleotide edit into a region corresponding to the editing target in the double stranded target DNA.
24. The SyNTase editing system of any one of claims 15-23, wherein the search target sequence is complementary to a protospacer sequence in the double stranded target DNA, and wherein the protospacer sequence is adjacent to a protospacer adjacent motif (PAM) in the double stranded target DNA.
25. The SyNTase editing system of any one of claims 15-24, wherein the tagRNA results in incorporation of a nucleotide edit in the PAM when contacted with the double stranded target DNA.
26. The SyNTase editing system of any one of claims 15-25, wherein:the spacer of the tagRNA is from 16 to 25 nucleotides in length, optionally 20 nucleotides in length or 21-23 nucleotides in length;the flap binding sequence is about 2 to 20 nucleotides in length, optionally about 8 to 16 nucleotides in length or 6 nucleotides in length; and / orthe editing template is about 4 to 30 nucleotides in length, optionally about 10 to 30 nucleotides in length, further optionally 6 to 9 nucleotides in length.
27. The SyNTase editing system of any one of claims 21-26, wherein the tagRNA results in incorporation of the intended nucleotide edit about 0 to 30 base pairs downstream of the nickase cleavage site.
28. The SyNTase editing system of any one of claims 21-27, wherein the intended nucleotide edit:-no-comprises a single nucleotide substitution compared to the region corresponding to the editing target in the double stranded target DNA;comprises an insertion compared to the region corresponding to the editing target in the double stranded target DNA, optionally an insertion of a nucleotide sequence at least 50, at least 45, at least 40, at least 35, at least 30, at least 25, at least 20, at least 15, at least 10, or at least 5, or at least 1 nucleotides in length; and / orcomprises a deletion compared to the region corresponding to the editing target in the double stranded target DNA.
29. The SyNTase editing system of any one of claims 15-28, wherein the editing template:comprises one or more silent nucleotide edits compared to the region corresponding to the editing target in the double stranded target DNA, optionally said silent nucleotide edits do not alter the amino acid sequence of the protein encoded by the double stranded target DNA, further optionally said one or more silent nucleotide edits comprise a substitution of 2 to 5 contiguous nucleotides; and / orcomprises a wild type double stranded target DNA sequence.
30. The SyNTase editing system of any one of claims 15-29, wherein the tagRNA results in correction of a mutation when contacted with the double stranded target DNA.
31. The SyNTase editing system of any one of claims 15-30, further comprising: an enhancer guide RNA (egRNA), or a nucleic acid encoding the egRNA, wherein the egRNA comprises a egRNA spacer that is complementary to a second search target sequence in the double stranded target DNA, optionally the egRNA comprises a scaffold sequence.
32. The SyNTase editor or SyNTase editing system of any one of claims 13-30, wherein the DNA endonuclease domain is a CRISPR associated (Cas) protein domain, optionally the Cas protein domain is Cas9.
33. The SyNTase editor or SyNTase editing system of claim 32, wherein the Cas protein domain has nickase activity.
34. The SyNTase editor or SyNTase editing system of any one of claims 32-33, wherein the Cas9 comprises a mutation in an HNH domain, optionally the Cas9 comprises an H840A mutation in the HNH domain.
35. The SyNTase editor or SyNTase editing system of any one of claims 32-33, wherein the Cas protein domain is a Casl2a, Casl2b, Casl2c, Casl2d, Casl2e, Casl4a, Casl4b, Cas 14c, Casl4d, Casl4e, Casl4f, Cas 14g, Casl4h, Casl4u, Cas, SschlCas9, SrolCas9, Sha4Cas9, SsuCas9, iSpyMacCas9, Ssi5Cas9, Ssi8Cas9, Ssci4Cas9, ShylCas9, Sag3Cas9,SlutrlCas9, Ssch3Cas9, SpRYCas9, SpRYcCas9, Sma2Cas9, SsaCas9, EvoCjCas9, iSpyMac, SsaCas9_AR12, or SveCas9, optionally the Cas protein domain is a Casl2b.
36. The SyNTase editor or SyNTase editing system of any one of claims 13-35, wherein the DNA polymerase domain, the DNA binding domain, and the DNA endonuclease domain are fused or linked to form a fusion protein.
37. The SyNTase editing system of any one of claims 15-36, wherein the scaffold sequence of the tagRNA comprises the sequence of any one of SEQ ID NOs: 640-685, optionally the scaffold sequence of the tagRNA comprises the sequence of any one of SEQ ID NOs: 640, 651, and 667.
38. The SyNTase editing system of any one of claims 15-37, wherein the editing template comprises: (i) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 14 2’ O-methyl RNA base(s); and / or (ii) at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, or 142’ Fluoro RNA base(s).
39. The SyNTase editing system of any one of claims 15-38, wherein all of the nucleotides of the editing template are chemically modified or wherein none of the nucleotides of the editing template are chemically modified, optionally the chemical modification is a 2’ O-methyl RNA or a 2’ Fluoro RNA base(s), optionally about 25%, about 50%, about 75%, or about 100% of the nucleotides of the editing template are chemically modified.
40. The SyNTase editing system of any one of claims 17-39, wherein all of the nucleotides of the flap binding sequence are chemically modified or wherein none of the nucleotides of the flap binding sequence are chemically modified, optionally the chemical modification is a 2’ O-methyl RNA or a 2’ Fluoro RNA base(s), optionally the flap binding sequences comprising one or more phosphorothioate linkages, further optionally about 25%, about 50%, about 75%, or about 100% of the nucleotides of the flap binding sequence are chemically modified.
41. The SyNTase editing system of any one of claims 17-40, wherein the first five nucleotides of the editing template each comprise a 2’ Fluoro RNA base, and the last 3 nucleotides of the flap binding sequence each comprise a 2’ O-methyl RNA and a phosphorothioate linkage, from 5’ to 3’.
42. The SyNTase editing system of any one of claims 15-41, wherein the tagRNA comprises the sequence of any one of SEQ ID NOs: 686-720, optionally the tagRNA comprises the sequence of SEQ ID NO: 712.
43. The SyNTase editor or SyNTase editing system of any one of claims 13-42, wherein the SyNTase editor comprises an amino acid sequence that is at least 80% identical to any one of SEQ ID NOs: 238-249, 266-273, 282-315, 350-371, 423-471, and 628-633.
44. The SyNTase editing system of any one of claims 15-43, wherein:the tagRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 712; andthe SyNTase editor comprises an amino acid sequence at least 80% identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723.
45. A SyNTase editing system for editing SERPINA1 gene comprising:a template armed guide RNA (tagRNA), or a nucleic acid encoding the tagRNA, wherein the tagRNA comprises:a spacer that is complementary to a search target sequence on a first strand of a nucleic acid molecule;an editing template that comprises a region of complementarity to an editing target sequence on a second strand of the nucleic acid molecule; and a scaffold sequence that associates with a SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain; anda SyNTase editor comprising a DNA binding domain, a DNA endonuclease domain and a DNA polymerase domain, or a nucleic acid encoding the SyNTase editor, wherein the tagRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 712, andwherein the SyNTase editor comprises an amino acid sequence at least 80% identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723.
46. A ribonucleoprotein (RNP) complex comprising the SyNTase editing system of any one of claims 15-45 or a component thereof.
47. A lipid nanoparticle (LNP) comprising the SyNTase editing system of any one of claims 15-45, or a component thereof.
48. The LNP of claim 47, comprising the tagRNA and the nucleic acid encoding the SyNTase editor, optionally the nucleic acid encoding the SyNTase editor is mRNA.
49. The LNP of any one of claims 47-48, further comprising the egRNA.
50. A polynucleotide encoding the SyNTase editor or SyNTase editing system of any one of claims 13-45.
51. The polynucleotide of claim 50, wherein the polynucleotide is an mRNA.
52. The polynucleotide of claim 51, wherein the mRNA comprises one or more of a 5'-cap structure, a 5’-UTR, a 3’-UTR, and a nuclear localization sequence (NLS).
53. The polynucleotide of any one of claims 51-52, wherein the mRNA comprises (i) a 5 '-cap, (ii) a 5 ’-untranslated region (UTR); (iii) an open reading frame (ORF) comprising a nucleotide sequence that encodes the SyNTase editor; and (iv) a 3' untranslated region (UTR).
54. The polynucleotide of any one of claims 51-53, wherein the mRNA has a structure comprising or consisting of 5’ - [5’UTR]-[NLS]-[nCas9]-[linker]-[RT]-[NLS]-[3’UTR and / or viral element]-[polyA sequence] - 3’.
55. The polynucleotide of any one of claims 51-54, wherein one or more nucleotides of the mRNA sequence are chemically modified.
56. The polynucleotide of any one of claims 51-55, wherein the mRNA comprises one or more additional segmented polyA sequences configured to increase mRNA stability and / or half-life.
57. The polynucleotide of any one of claims 51-56, wherein the mRNA comprises one or more viral element sequences, optionally 3’ of the 3 ’-UTR, optionally wherein:the viral element comprises a woodchuck hepatitis virus post-transcriptional regulatory element (WPRE) sequence, optionally comprising the sequence of SEQ ID NO: 141, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 141; and / orthe viral element comprises a eK5 sequence, optionally located 3’ of the 3 ’-UTR, optionally comprising the sequence of SEQ ID NO: 142, or a sequence that exhibits at least about 85% identity to SEQ ID NO: 142.
58. The polynucleotide of any one of claims 51-57, wherein the mRNA comprises one or more secondary structure motifs.
59. The polynucleotide of claim 58, wherein a secondary structure motif comprises a triple helix sequence, optionally a synthetic triple helix (STH) sequence, optionally the STH sequence comprises the sequence derived from a sequence element of a long non-coding RNA, optionally MALATE60. The polynucleotide of claim 59, wherein the STH sequence comprises any one of the sequences of SEQ ID NOs: 143-144 and 172-173, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 143-144 and 172-173.
61. The polynucleotide of any one of claims 51-60, wherein the mRNA comprises:a polyA sequence comprising any one of the sequences of SEQ ID NOs: 134- 139, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 134-139;one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 140-142, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 140-142;one or more 3’ UTR structural motifs selected from the group comprising any one of the sequences of SEQ ID NOs: 143-144, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 143-144; and / ora 5’UTR sequence comprising the sequence of any one of SEQ ID NOs: 145 and 207-230, or a sequence that exhibits at least about 85% identity to any one of SEQ ID NOs: 145 and 207-230.
62. The polynucleotide of any one of claims 51-61, wherein the mRNA comprises a 3’ UTR comprising any one of the sequences of SEQ ID NOs: 147-171, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 147-171.
63. The polynucleotide of any one of claims 51-62, wherein the mRNA comprises:one or more 5’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 174-185, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 174-185;one or more 3’ UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 186-200, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 186-200; and / orone or more additional elements selected from the group comprising any one of the sequences of SEQ ID NOs: 201-202, or a sequence that exhibits at least about 85% identity to any one of the sequences of SEQ ID NOs: 201-202, optionally the one or more additional elements are situated downstream of the 3 ’UTR and / or the polyA tail.
64. The polynucleotide of any one of claims 51-63, wherein the mRNA comprises one or more 5’-UTR element sequences selected from the group comprising any one of the sequences of SEQ ID NOs: 203-206, and wherein the second codon of the mRNA is a “gcc”.
65. The polynucleotide of any one of claims 51-64, wherein the mRNA is codon optimized for expression in human cells.
66. The polynucleotide of any one of claims 51-65, wherein the mRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723.
67. The polynucleotide of any one of claims 50-66, wherein the polynucleotide is operably linked to a regulatory element, optionally the regulatory element is an inducible regulatory element.
68. A vector comprising the polynucleotide of any one of claims 50-67.
69. The vector of claim 68, wherein the vector is an AAV vector.
70. An isolated cell comprising the SyNTase editor or SyNTase editing system of any one of claims 13-45, the RNP of claim 46, the LNP of any one of claims 47-49, the polynucleotide of any one of claims 50-67, or the vector of any one of claims 68-69.
71. The cell of claim 70, wherein the cell is a mammalian cell, optionally a human cell.
72. The cell of any one of claims 70-71, wherein the cell is a primary cell.
73. The cell of any one of claims 70-72, wherein the cell is a hepatocyte.
74. The cell of any one of claims 70-73, wherein the cell is from a subject having a disease or disorder, optionally P-thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, Wilson’s disease, phenylketonuria, or hyperphenylalaninemia, further optionally the subject is a human.
75. A pharmaceutical composition comprising:(i) the SyNTase editor or SyNTase editing system of any one of claims 13-45, the RNP of claim 46, the LNP of any one of claims 47-49, the polynucleotide of any one of claims 50-67, the vector of any one of claims 68-69, or the cell of any one of claims 70- 74; and(ii) a pharmaceutically acceptable carrier.
76. A method for editing a double stranded target DNA, the method comprising contacting the double stranded target DNA with the SyNTase editor or SyNTase editing system of any one of claims 13-45, thereby editing the double stranded target DNA.
77. The method of claim 76, wherein the double stranded target DNA is in a cell.
78. The method of claim 77, wherein the cell is a mammalian cell, optionally a human cell.
79. The method of any one of claims 77-78, wherein the cell is a primary cell.
80. The method of any one of claims 77-79, wherein the cell is a hepatocyte.
81. The method of any one of claims 77-80, wherein the cell is a stem cell; optionally, an embryonic stem cell, an induced pluripotent stem cell, or an adult stem cell.
82. The method of any one of claims 77-81, wherein the cell is in a subject, optionally the subject is a human.
83. The method of any one of claims 77-82, wherein the cell is from a subject having a disease or disorder, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof, optionally P-thalassemia, sickle cell disease, alphal antitrypsin deficiency disease, Wilson’s disease, phenylketonuria, or hyperphenylalaninemia.
84. The method of claim 83, further comprising administering the cell to the subject after incorporation of the intended nucleotide edit.
85. A cell generated by the method of any one of claims 77-84.
86. A population of cells generated by the method of any one of claims 77-84.
87. A method for treating or preventing a disease or disorder in a subject in need thereof, the method comprising administering to the subject the SyNTase editor or SyNTase editing system of any one of claims 13-45, the RNP of claim 46, the LNP of any one of claims 47-49, the pharmaceutical composition of claim 75, the cell of claim 85, or the population of cells of claim 86, thereby treating or preventing the disease or disorder in the subject, optionally the disease or disorder is selected from the group comprising an autoimmune disease, a neurological disease or disorder, a cancer, an inflammatory disease, a cardiovascular disease, an infectious disease, a genetic disease, a trinucleotide repeat expansion disorder, a metabolic disease, or any combination thereof.
88. The method of claim 87, wherein the disease or disorder comprises alpha-1 antitrypsin deficiency (AATD).
89. A method for treating or preventing or reversing AATD in a subject in need thereof, the method comprising administering to the subject the SyNTase editor or SyNTase editing system of any one of claims 13-45, the RNP of claim 46, the LNP of any one of claims 47-49, the pharmaceutical composition of claim 75, the cell of claim 85, or the population of cells of claim 86, thereby treating or preventing or reversing the AATD in the subject.
90. The method of claim 89, wherein:the tagRNA comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 712; andwherein the SyNTase editor comprises an amino acid sequence at least 80% identical to the sequence of SEQ ID NO: 726 and / or the nucleic acid encoding the SyNTase editor comprises a sequence at least 80% identical to the sequence of SEQ ID NO: 723.
91. The method of any one of claims 76-84 and 87-90, wherein the method is performed in vivo, in vitro or ex vivo.
92. The method of any one of claims 87-91, comprising administering to the subject a single dose of about 0.05 to 0.5 mg / kg of total nucleic acids comprising (a) the tagRNA and (b) the nucleic acid encoding the SynTase editor.
93. The method of any one of claims 88-92, wherein the method provides a SERPINA1 gene editing efficiency of at least 20% or the method provides a SERPINA1 gene editing efficiency of at least 80%.
94. The method of any one of claims 88-93, wherein the method provides a SERPINA1 mRNA editing efficiency of at least 80%, at least 85%, at least 90%, at least 95%, or at least 99%.
95. The method of any one of claims 88-94, wherein a clinically relevant increase in M-AAT protein levels is detected in the serum of the subject following the administration.
96. The method of any one of claims 88-95, wherein:an at least five-fold increase in total serum AAT protein levels is detected in the subject following the administration; and / orthe ratio of M-AAT: Z-AAT in the serum of the subject is at least 99% following the administration.
97. The method of any one of claims 88-96, wherein an increase in total serum AAT protein levels is detected in the subject at least seven days after the administration, optionally at least 7 weeks, further optionally at least 9 weeks.
98. The method of any one of claims 95-97, wherein a linear correlation is observed between SERPINA1 gene and / or mRNA editing efficiency and the increase in AAT protein levels in the subject.
99. The method of any one of claims 95-98, wherein the M-AAT is capable of functional rescue in a human neutrophil elastase inhibition assay.