RBM20 gene editing systems and methods of use
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENAYA THERAPEUTICS INC
- Filing Date
- 2026-01-30
- Publication Date
- 2026-08-06
Smart Images

Figure US2026013369_06082026_PF_FP_ABST
Abstract
Description
Attorney Docket No.:TYA-072WORBM20 GENE EDITING SYSTEMS AND METHODS OF USECROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 752,386 filed January 31, 2025, U.S. Provisional Application No. 63 / 804,445 filed May 12, 2025, and U.S. Provisional Application No. 63 / 947,933 filed December 23, 2025, the contents of each of which are incorporated herein by reference in their entireties.STATEMENT REGARDING SEQUENCE LISTING
[0002] This application contains a Sequence Listing which has been submitted electronically in XML format. The Sequence Listing XML is incorporated herein by reference. Said XML file, created on January 27, 2026, is named TYA-072WO_SL.xml, is 366,552 bytes in size, and is being submitted electronically via USPTO Patent Center.TECHNICAL FIELD
[0003] In some aspects, the present disclosure relates to compositions and methods for editing the RNA binding motif protein 20 (RBM20) gene in humans. In some embodiments, the present disclosure relates to genetically-modified, non-human animals comprising a humanized RBM20 gene, and methods of use thereof.BACKGROUND
[0004] RBM20 is a regulator of mRNA splicing, and certain variants in RBM20 are associated with dilated cardiomyopathy (DCM). RBM20 mutations that cluster within an arginine- serine-rich (RS) domain of RBM20 (also called a “mutation hotspot”) have been linked to aggressive forms of DCM with poor clinical outcomes. Therapeutic strategies for treating DCM are needed, and gene editing is one potential strategy.
[0005] There remains a need in the art for gene editing therapies targeting the RBM20 gene. However, the DNA sequence of murine Rbm20 differs from that of human RBM20. Since gene editing systems are DNA sequence-dependent, this difference in sequence hinders testing of human RBM20-targeting gene editing systems in murine models. There is a need for murine models in which the mouse Rbm20 gene is replaced with a human DNA sequence (i.e., humanized). Murine models inAttorney Docket No.:TYA-072WO which the Rbm20 mutation hotspot is humanized would be of particular value for testing of human RBM20-targeting gene editing systems.SUMMARY
[0006] In some aspects, the present disclosure provides a genetically-modified, non-human animal comprising a humanized RNA binding motif protein 20 (RBM20) gene, wherein nucleic acids 250-264 of SEQ ID NO: 1 within the Rbm20 gene are replaced with a human RBM20 nucleic acid sequence. In some embodiments, the non-human animal is a mouse. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) at amino acid position corresponding to position 634 of human RBM20 protein. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q).
[0007] In some aspects, the present disclosure provides a genetically-modified, non-human animal comprising a humanized RNA binding motif protein 20 (RBM20) gene, wherein at least 15 consecutive nucleic acids between 230-284 of SEQ ID NO: 1 are replaced with a human RBM20 nucleic acid sequence. In some embodiments, the non-human animal is a mouse. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 6. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 7. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 8. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 9. In some embodiments, the human RBM20 nucleic acid sequence replaces at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85. at least 90, at least 95, at least 100, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, or at least 220 nucleotides upstream and / or downstream of nucleic acids 250-264 of SEQ ID NO: 1 within the murine Rbm20 gene.
[0008] In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4 and: (a) a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 4; and / or (b) a sequence comprising any one of “GGTGA” or SEQAttorney Docket No.:TYA-072WO ID NOs: 50-88 downstream of SEQ ID NO: 4. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5 and: (a) a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 5: and / or (b) a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 5. In some embodiments, the human RBM20 nucleic acid sequence replaces nucleic acids 31-294 of SEQ ID NO: 1 within the murine Rbm20 gene. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 89. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 90.
[0009] In some embodiments, the non-human animal comprises a first allele comprising a first humanized RBM20 gene and a second allele comprises a murine Rbm20 gene. In some embodiments, the non-human animal comprises a first allele comprising a first humanized RBM20 gene and a second allele comprises a second humanized RBM20 gene. In some embodiments, the first humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) at an amino acid position corresponding to position 634 of human RBM20 protein. In some embodiments, the non-human animal comprises a first allele comprising a first humanized RBM20 gene that encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q).
[0010] In some embodiments, a human RBM20-targeting guide RNA specifically binds to the humanized RBM20 gene. In some embodiments, humanized RBM20 gene allows targeting of a nucleic acid-guided nuclease to the RBM20 nucleotide sequence using a human RBM20-targeting guide RNA. In some embodiments, endogenous and human RBM20 sequences encode a humanized RBM20 protein with substantially similar activity to a wild-type RBM20 protein. In some embodiments, the animal has substantially similar ejection fraction (EF). left ventricle internal diameter at systolic stage (LVID;s). and / or left ventricle internal diameters at diastolic stage (LVID;d) to a non-human animal comprising two wild-type RBM20 alleles. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q) (an “R634Q mutation”), and wherein the humanized RBM20 protein comprising an R634Q mutation has altered activity compared to a wild-type RBM20 protein. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, and wherein the animal hasAttorney Docket No.:TYA-072WO reduced ejection fraction (EF), increased left ventricle internal diameters at systolic stage (LVID;s), and / or increased left ventricle internal diameter at diastolic stage (LVID;d) compared to a non-human animal comprising two wild-type RBM20 alleles. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the animal has increased levels of Nppa transcripts, Nppb transcripts, natriuretic peptide A, and / or natriuretic peptide B.
[0011] In some aspects, the present disclosure provides a method of determining the efficacy of a gene editing system targeting the human RBM20 gene, comprising: (a) contacting a cell of the non-human animal of any one of claims 1-22 expressing a humanized RNA binding motif protein 20 (RBM20) gene with the gene editing system, wherein the gene editing system comprises a guide RNA comprising a spacer sequence that binds to a target sequence of the humanized RBM20 gene and an RNA-guided nuclease. In some embodiments, the method further comprises: (b) analyzing the humanized RBM20 gene by NGS to determine the presence or absence of editing events in the humanized RBM20 gene, wherein the gene editing system is determined to be efficacious by the presence of editing events; and / or (c) monitoring the heart functions of the animal, wherein the gene editing system is determined to be efficacious by the improvement of heart functions. In some embodiments, cell is in vivo or ex vivo. In some embodiments, guide RNA comprises any of SEQ ID NOs: 95, 96, 98-103, 195, 199, and 200. In some embodiments, the guide RNA comprises the sequence of SEQ ID NO: 96 and / or SEQ ID NO: 103. In some embodiments, the guide RNA comprises the sequence of SEQ ID NO: 96 and / or SEQ ID NO: 199. In some embodiments, the guide RNA comprises the sequence of SEQ ID NO: 109. In some embodiments, the guide RNA comprises the sequence of SEQ ID NO: 187.
[0012] In some aspects, the present disclosure provides a population of cardiomyocytes isolated from a non-human animal described herein.
[0013] In some aspects, the present disclosure provides a guide RNA polynucleotide comprising (a) a spacer sequence; (b) a scaffold sequence; and (c) a template + primer binding site sequence, wherein the spacer sequence binds a target sequence within a human RBM20 gene or a humanized RBM20 gene. In some embodiments, the target sequence is within a mutant human RBM20 gene or a mutant humanized RBM20 gene. In some embodiments, the human RBM20 gene or the humanized RBM20 gene comprises the sequence of SEQ ID NO: 5. In some embodiments, the human RBM20 gene or the humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R)Attorney Docket No.:TYA-072WO to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q) (an “R634Q mutation”). In some embodiments, the human RBM20 gene or the humanized RBM20 gene comprises the sequence of SEQ ID NO: 7 or SEQ ID NO: 9. In some embodiments, the spacer sequence comprises the nucleotide sequence of any one of SEQ ID NOs: 95. 96, and 195. In some embodiments, the spacer sequence comprises the sequence of SEQ ID NO: 96. In some embodiments, the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 97. In some embodiments, the template + primer binding site sequence comprises the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200. In some embodiments, the template + primer binding site sequence comprises the nucleotide sequence of SEQ ID NO: 103. In some embodiments, the template + primer binding site sequence comprises the nucleotide sequence of SEQ ID NO: 199. In some embodiments, the guide RNA comprises: (a) a spacer comprising the nucleotide sequence of SEQ ID NO: 95 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 98; (b) a spacer comprising the nucleotide sequence of SEQ ID NO: 95 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 99; (c) a spacer comprising the nucleotide sequence of SEQ ID NO: 95 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 100; (d) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 101; (e) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 102; (f) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 103; or (g) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 199. In some embodiments, the guide RNA comprises a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 103. In some embodiments, the guide RNA comprises a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 199. In some embodiments, the guide RNA further comprises a linker and a 3’ motif, wherein the linker and 3’ motif comprise the nucleotide sequence of any one of SEQ ID NOs: 126 and 202-204. In some embodiments, the guide RNA comprises the nucleotide sequence of any one of SEQ ID NOs: 104-109, 187, and 188. In some embodiments, the guide RNA comprises the nucleotide sequence of SEQ ID NO: 109. In some embodiments, the guide RNA comprises the nucleotide sequence of SEQ ID NO: 187.Attorney Docket No.:TYA-072WO
[0014] In some aspects, the present disclosure provides a gene editing system comprising a guide RNA described herein and an effector protein. In some embodiments, the effector protein is a nucleic acid-guided nuclease. In some embodiments, the nucleic acid-guided nuclease is a Cas nuclease. In some embodiments, the nucleic acid-guided nuclease is an RNA-guided nickase. In some embodiments, the effector protein comprises an RNA-guided nickase and a reverse transcriptase. In some embodiments, the effector protein is a fusion protein comprising an RNA-guided nickase and a reverse transcriptase. In some embodiments, the fusion protein comprises: (a) an N-terminal fragment comprising an N-terminal fragment of an RNA-guided nickase and a split intein domain; and (b) a C-terminal fragment comprising a C-terminal fragment of the RNA-guided nickase, a reverse transcriptase, and a split intein domain. In some embodiments, the system further comprises a nicking guide RNA (ngRNA). In some embodiments, the ngRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 110-117 and 194. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 113. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 194. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 118-125 and 193. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 121. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 193.
[0015] In some aspects, the present disclosure provides a vector comprising a nucleotide sequence encoding (a) a gRNA described herein or (b) a gene editing system described herein or a portion thereof. In some embodiments, the vector is an AAV viral vector.
[0016] In some aspects, the present disclosure provides a method for expressing a polynucleotide in a target cell, comprising contacting the target cell with a gene editing system described herein or transducing the target cell with one or more vectors described herein. In some embodiments, the target cell is a cardiac cell, a muscle cell, an induced pluripotent stem cell-derived cardiomyocyte (iPSC-CM), a cardiomyocyte, or an iPSC. In some embodiments, the target cell comprises at least one mutant human RBM20 allele or at least one mutant humanized RBM20 allele. In some embodiments, the method results in editing of the mutant human RBM20 allele or mutant humanized RBM20 allele. In some embodiments, the mutant human RBM20 allele or mutant humanized RBM20 allele encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q)Attorney Docket No.:TYA-072WO (an “R634Q mutation”). In some embodiments, the method does not, or substantially does not, edit the wild-type human RBM20 allele or wild-type humanized RBM20 allele.
[0017] In some aspects, the present disclosure provides a method for preventing and / or treating a disease or condition in a subject in need thereof, comprising administering to the subject a gene editing system described herein or one or more vectors described herein, wherein the disease or condition is a cardiac pathology. In some embodiments, the subject comprises at least one mutant human RBM20 allele or at least one mutant humanized RBM20 allele. In some embodiments, the method results in editing of the mutant human RBM20 allele or mutant humanized RBM20 allele. In some embodiments, the mutant human RBM20 allele or mutant humanized RBM20 allele encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q) (an “R634Q mutation”). In some embodiments, the method does not, or substantially does not, edit the wild-type human RBM20 allele or wild-type humanized RBM20 allele.
[0018] In some aspects, the present disclosure provides a method for identifying guide RNAs suitable for gene editing systems, wherein the guide RNAs allow efficient editing of a human RBM20 gene or a humanized RBM20 gene In some embodiments, the guide RNAs provide editing with an efficiency of at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%. or at least 70%. In some embodiments, the guide RNAs are pegRNAs and / or ngRNAs.
[0019] In some aspects, the present disclosure provides a nicking guide RNA (ngRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 110-117 and 194. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 113. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 194. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 118-125 and 193. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 121. In some embodiments, the ngRNA comprises the nucleotide sequence of SEQ ID NO: 193.
[0020] In some aspects, the present disclosure provides a gene editing system comprising: (a) a first expression cassette comprising: (i) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein; and (ii) a second polynucleotide encoding a nicking guide RNA (ngRNA)Attorney Docket No.:TYA-072WO comprising the nucleotide sequence of any one of SEQ ID NOs: 110-117 and 194; and b) a second expression cassette comprising: (iii) a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein; and (iv) a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200.
[0021] In some aspects, the present disclosure provides a gene editing system comprising: (a) a first expression cassette comprising: (i) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein; and (ii) a second polynucleotide encoding a nicking guide RNA (ngRNA) comprising the nucleotide sequence of SEQ ID NO: 113; and b) a second expression cassette comprising: (iii) a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein; and (iv) a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of SEQ ID NO: 103.
[0022] In some embodiments, the ngRNA comprises the sequence of any one of SEQ ID NOs: 118-125 and 193, and the pegRNA comprises the sequence of any one of SEQ ID NOs: 104-109, 187, and 188. In some embodiments, the ngRNA comprises the sequence of SEQ ID NO: 121, and the pegRNA comprises the sequence of SEQ ID NO: 109. In some embodiments, the N-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 129 and the C-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 132.
[0023] In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 177. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 178. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 180. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 181.
[0024] In some embodiments, the first expression cassette comprises the sequence of SEQ IDAttorney Docket No.:TYA-072WO NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 189. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 190. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 191. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 192.
[0025] In some aspects, the present disclosure provides a gene editing system comprising a first expression cassette comprising SEQ ID NO: 176 and a second expression cassette comprising SEQ ID NO: 177 or SEQ ID NO: 178.
[0026] In some aspects, the present disclosure provides a gene editing system comprising a first expression cassette comprising SEQ ID NO: 179 and a second expression cassette comprising SEQ ID NO: 180 or SEQ ID NO: 181.
[0027] In some aspects, the present disclosure provides a gene editing system comprising a first expression cassette comprising SEQ ID NO: 176 and a second expression cassette comprising SEQ ID NO: 189 or SEQ ID NO: 190.
[0028] In some aspects, the present disclosure provides a gene editing system comprising a first expression cassette comprising SEQ ID NO: 179 and a second expression cassette comprising SEQ ID NO: 191 or SEQ ID NO: 192.
[0029] In some aspects, the present disclosure provides a system comprising: (a) a first vector comprising a first expression cassette comprising the sequence of SEQ ID NO: 176; and (b) a second vector comprising a second expression cassette comprising the sequence of any one of SEQ ID NOs: 177, 178, 189, and 190. In some aspects, the present disclosure provides a system comprising: (a) a first vector comprising a first expression cassette comprising the sequence of SEQ ID NO: 179; and (b) a second vector comprising a second expression cassette comprising the sequence of any one of SEQ ID NOs: 180, 181. 191, and 192. In some embodiments, the first vector and the second vector are viral vectors. In some embodiments, the viral vectors are AAV vectors.
[0030] In some aspects, the present disclosure provides a cell comprising a guide RNA described herein, a gene editing system described herein or a portion thereof, a vector describedAttorney Docket No.:TYA-072WO herein, or a system described herein or a portion thereof. In some embodiments, the cell is a cardiac cell, a muscle cell, an iPSC-CM, a cardiomyocyte, or an iPSC.
[0031] In some aspects, the present disclosure provides a cell produced by a method for expressing a polynucleotide in a cell described herein. In some embodiments, the cell is a cardiac cell, a muscle cell, an iPSC-CM, a cardiomyocyte, or an iPSC.
[0032] In some aspects, provided herein is a gene editing system capable of editing a target genomic locus and achieving self-inactivation. In some embodiments, the system comprises one or more polynucleotides encoding a prime editor protein or fragments thereof, wherein at least one polynucleotide comprises a self-inactivation site positioned upstream of or within a coding region. In some embodiments, the system further comprises a polynucleotide encoding a pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site. In some embodiments, editing of the self-inactivation site by the pegRNA introduces one or more protein coding changes that reduce or eliminate expression or function of the prime editor protein.
[0033] In some embodiments, the one or more polynucleotides encoding the prime editor protein or fragments thereof is located in a single vector. In some embodiments, the one or more polynucleotides encoding the prime editor protein or fragments thereof and the polynucleotide encoding the pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site are located in the same vector. In some embodiments, the one or more polynucleotides encoding the prime editor protein or fragments thereof and the polynucleotide encoding the pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site are located in at least two separate vectors.
[0034] In some embodiments, the self-inactivation site encodes a peptide sequence of 1, 2, 3, 4. 5, 6, 7. 8, 9, 10, 11. 12. 13. 14. 15. 16. 17. 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29. 30. 31. 32.33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70. 71. 72. 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, or 100 amino acids that does not substantially impair expression or function of the prime editor protein prior to editing.
[0035] In some embodiments, the one or more protein coding changes comprise a frameshift mutation, a premature stop codon, an amino acid change, or a combination thereof. In someAttorney Docket No.:TYA-072WO embodiments, editing of the self-inactivation site reduces or eliminates expression or function of the prime editor protein by at least 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%. 68%, 69%, 70%, 71%. 72%, 73%, 74%, 75%, 76%, 77%, 78%.79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.
[0036] In some embodiments, the self-inactivation site is positioned within the coding region within 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, or 500 nucleotides downstream of a start codon.
[0037] In some embodiments, the one or more polynucleotides encoding the prime editor protein or fragments thereof comprise a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein, and a second polynucleotide encoding a C-terminal fragment of the fusion protein comprising a C-terminal fragment of the RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein.
[0038] In some embodiments, the first polynucleotide comprises a first self-inactivation site. In some embodiments, the second polynucleotide comprises a second self-inactivation site. In some embodiments, the first polynucleotide and / or the second polynucleotide comprises a self-inactivation site. In some embodiments, both the first polynucleotide and the second polynucleotide comprise selfinactivation sites.
[0039] In some embodiments, the first polynucleotide and the polynucleotide encoding the pegRNA are packaged in a first AAV vector, and the second polynucleotide is packaged in a second AAV vector. In some embodiments, the first AAV vector and the second AAV vector are delivered to cardiac tissue. In some embodiments, the target genomic locus is the RBM20 locus.
[0040] In some embodiments, a self-inactivating gene editing system described herein comprises a first expression cassette comprising a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein, and a second polynucleotide encoding a nicking guide RNA (ngRNA), and a second expression cassette comprising a third polynucleotide encoding a C-terminalAttorney Docket No.:TYA-072WO fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein, and a fourth polynucleotide encoding a prime editing guide RNA (pegRNA), wherein at least one of the first polynucleotide or the third polynucleotide comprises a self-inactivation site that is recognized and edited by the pegRNA, and wherein editing of the self-inactivation site introduces one or more protein coding changes that reduce or eliminate expression or function of the gene editing system.
[0041] In some embodiments, the self-inactivating gene editing system comprises a first expression cassette comprising a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein, and a second polynucleotide encoding a nicking guide RNA (ngRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 110, 111, 112, 113, 114, 115, 116, 117, and 194, and a second expression cassette comprising a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein, and a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 98, 99, 100, 101.102, 103, 199, and 200. In some embodiments, at least one of the first polynucleotide or the third polynucleotide comprises a self-inactivation site that is recognized and edited by the pegRNA. In some embodiments, editing of the self-inactivation site introduces one or more protein coding changes that reduce or eliminate expression or function of the gene editing system. In some embodiments, the ngRNA comprises the sequence of any one of SEQ ID NOs: 118, 119, 120, 121, 122, 123, 124, 125, and 193. In some embodiments, the pegRNA comprises the sequence of any one of SEQ ID NOs: 104, 105, 106, 107, 108, 109, 187, and 188. In some embodiments, the N-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 129. In some embodiments, the C-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 132. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 177. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 178. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 180. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 181. In someAttorney Docket No.:TYA-072WO embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 189. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 190. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 191. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 192.
[0042] In some embodiments, provided herein is a system comprising a first vector comprising a first expression cassette and a second vector comprising a second expression cassette, wherein at least one of the first expression cassette or the second expression cassette comprises a self-inactivation site that is recognized and edited by the pegRNA, wherein editing of the self-inactivation site introduces one or more protein coding changes that reduce or eliminate expression or function of the gene editing system.
[0043] In some embodiments, provided is a system comprising a first vector comprising a first expression cassette and a second vector comprising a second expression cassette. In some embodiments, the first expression cassette comprises a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein, and a second polynucleotide encoding a nicking guide RNA (ngRNA). In some embodiments, the second expression cassette comprises a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein, and a fourth polynucleotide encoding a prime editing guide RNA (pegRNA). In some embodiments, at least one of the first expression cassette or the second expression cassette comprises a self-inactivation site that is recognized and edited by the pegRNA. In some embodiments, editing of the self-inactivation site introduces one or more protein coding changes that reduce or eliminate expression or function of the gene editing system.
[0044] In some embodiments of the system of the disclosure, the first vector is a first adeno-associated virus (AAV) vector and the second vector is a second AAV vector. In some embodiments, the first AAV vector and the second AAV vector are delivered to cardiac tissue. In someAttorney Docket No.:TYA-072WO embodiments, the cardiac tissue is heart tissue. In some embodiments, the system is for use in editing a target genomic locus in the cardiac tissue. In some embodiments, the target genomic locus is the RBM20 locus. In some embodiments, the ngRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 110, 111, 112, 113, 114, 115, 116, 117, and 194. In some embodiments, the pegRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 98, 99, 100, 101, 102, 103, 199, and 200.
[0045] In some embodiments of the system of the disclosure, the ngRNA comprises the sequence of any one of SEQ ID NOs: 118, 119, 120, 121, 122, 123, 124, 125, and 193. In some embodiments, the pegRNA comprises the sequence of any one of SEQ ID NOs: 104, 105, 106, 107, 108, 109, 187, and 188. In some embodiments, the N-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 129. In some embodiments, the C-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 132.
[0046] In some embodiments of the system of the disclosure, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 177. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 178. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 180. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 181. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 189. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 190. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 191. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 192.
[0047] In some embodiments, the gene editing system or system described herein is for use in treating a genetic disorder in a subject.Attorney Docket No.:TYA-072WO BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIGs. 1A-1D provide a schematic (FIG. 1A) and DNA and protein sequences of mouse Rbm20, human RBM20, and human RBM20R634Q(FIG. 1B) at the mutation hotspot region, an alignment of the mouse wildtype, humanized wildtype knock-in, and humanized R634Q knock-in alleles of RBM20 (FIG. 1C), and survival data for mice either heterozygous or homozygous for humanized RBM20 wild-type or R634Q mutant alleles (FIG. ID).
[0049] FIGs.2A-2G provide ejection fractions (EF), left ventricle internal diameters at systolic stage (LVID;s), and left ventricle internal diameters at diastolic stage (LVID;d) at various time points for wild-type mice and mice either heterozygous or homozygous for mouse or humanized wild-type or mutant RBM20. FIGs. 2A-2D provide data at various time points (FIGs. 2A and 2C and at 26 weeks of age (FIGs. 2B and 2D) for wildtype mice and mouse Rbm20R636Q / +mice (FIGs. 2A-2B) and for wildtype mRbm20+ / +mice, heterozygous humanized wildtype hRBM20+ / mRbm20+mice, and heterozygous humanized mutant hRBM20R634Q / mRbm20+mice (FIGs.2C-2D). FIGs.2E-2G provide data at various time points (FIG. 2E) and at 8 weeks of age (FIG. 2F) for heterozygous-humanized wildtype hRBM20+ / mRbm20+mice, homozygous-humanized hRBM20+ / hRBM20+mice, humanized heterozygous-mutant hRBM20R634Q / hRBM20+mice, and humanized homozygous-mutant hRBM20R634Q / hRBM20R634Qmice. FIG. 2G provides expression levels of heart failure marker genes Nppa and Nppb.
[0050] FIGs.3A-3F provide relative expression data indicating splicing of Ttn (FIGs.3A, 3C, and 3E) and Camk2d (FIG.3B, 3D, and 3F) for wildtype mice and mouse Rbm20R636Q / +mice (FIGs.3A-3B), for wildtype mRbm20+ / +mice and heterozygous humanized mutant hRBM20R634Q / mRbm20+mice (FIGs. 3C-3D), and for hRBM20+ / mRbm20+mice, hRBM20+ / hRBM20+mice. hRBM20R634Q / hRBM20+heterozygous mutant mice, and hRBM20R634Q / hRBM20R634Qhomozygous mutant mice (FIGs.3E-3F).
[0051] FIGs.4A-4H provide comparative data for editing efficiency of a human RBM20 locus using prime editing systems comprising combinations of pegRNAs and ngRNAs in HEK293T cells.
[0052] FIGs. 5A-5E provide comparative data for editing efficiency of a human RBM20 locus using prime editing systems comprising combinations of pegRNAs and ngRNAs in HEK293T cells.
[0053] FIG. 6 provides screening results for exemplary combinations of exemplary pegRNAsAttorney Docket No.:TYA-072WO and exemplary ngRNAs in mouse embryonic fibroblasts (MEFs) isolated from humanized RBM20 mice carrying the humanized wildtype hRBM20 and mutant hRBM20R634Qalleles.
[0054] FIGs. 7A-7B provide data showing in vivo gene editing efficiencies of two prime editing drugs targeting the human RBM20 sequence in humanized RBM20 mice (FIG.7A) and dosedependent efficacy for one of the prime editing drugs (FIG. 7B).
[0055] FIGs. 8A-8G show dual-AAV based prime editing treatment improving cardiac function in a humanized RBM20 cardiomyopathy mouse model. FIG. 8A provides a schematic overview of the experimental design. FIG. 8B shows ejection fractions (EF), left ventricle internal diameters at diastolic stage (LVID;d), and left ventricle internal diameters at systolic stage (LVID;s) measured at pre-treatment baseline and post-injection time points. FIG. 8C shows ejection fractions (EF), left ventricle internal diameters at diastolic stage (LVID;d), and left ventricle internal diameters at systolic stage (LVID;s) measured before dosing and at 16-week post-injection time point. FIG.8D and FIG. 8E plot gene editing results at the humanized RBM20 locus from heart RNA samples.FIG. 8F and FIG.8G show expression levels of mRNA isoforms of Ttn gene and Camk2d gene.
[0056] FIG. 9 depicts comparative editing efficiency data for prime editing of a human RBM20 locus using various pegRNA and ngRNA combinations screened in HEK293T cells, ng-26 and ng-36 were identified as high-performing nicking guide RNAs.
[0057] FIGs. 10A-10D depict an exemplary self-inactivating prime editing system design.FIG. 10A provides a generic schematic of a self-inactivating prime editing system comprising a selfinactivation site (SI site) in the PE-N cassette. FIG. 10B provides a specific example of a selfinactivating prime editing system in which the self-inactivation site is inserted after the start codon of the PE-N cassette. The self-inactivation site is recognized and edited by the pegRNA targeting the mouse Rbm20 locus, where editing events at the self-inactivation site can introduce amino acid changes, frameshift mutations, and / or a premature stop codon, thereby abolishing protein coding capability and inactivating expression of the prime editor protein. FIGs. 10C and 10D provide in vivo study results demonstrating that the self-inactivating prime editing system achieves editing at both the target genomic locus and the self-inactivation site.DETAILED DESCRIPTION
[0058] The present technology relates to genetically-modified, non-human animals comprisingAttorney Docket No.:TYA-072WO a humanized RNA binding motif protein 20 (RBM20) gene, guide RNAs (gRNAs) for targeting pathogenic variants of the human RBM20 gene, gene editing systems comprising such gRNAs, and methods of use thereof.
[0059] In some embodiments, the gRNA is a gRNA for targeting the human RBM20 gene, a humanized RBM20 gene, or a human RBM20 nucleic acid sequence. In some embodiments, the gRNA comprises a spacer sequence that binds to a target sequence, wherein the target sequence is a human RBM20 nucleic acid sequence. In some embodiments, the human RBM20 nucleic acid sequence is within a human RBM20 gene or humanized RBM20 gene. In some embodiments, the human RBM20 gene or humanized RBM20 gene is a mutant allele.
[0060] In some embodiments, the human RBM20 gene or humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of the human RBM20 protein (R634Q). In some embodiments, the human RBM20 gene or humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to the human RBM20 protein. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of any one of SEQ ID NOs: 5, 7, and 9.
[0061] In some aspects, the present disclosure provides a gene editing system comprising any guide RNA described herein and an RNA-guided nuclease. In some embodiments, the RNA-guided nuclease is a Cas nuclease. In some embodiments, the gene editing system comprises one or more vectors comprising one or more expression cassettes. In some embodiments, the expression cassettes comprise one or more promoters. Non-limiting examples of promoters include a modified TNNT2 promoter and a Pol III promoter. In some embodiments, the one or more vectors are adeno-associated virus (AAV) vectors.
[0062] In some aspects, the present disclosure provides genetically-modified, non-human animals comprising a humanized RBM20 gene, wherein nucleic acids 250-264 of SEQ ID NO: 1 within the murine Rbm20 gene are replaced with a human RBM20 nucleic acid sequence. In some embodiments, the human RBM20 nucleic acid sequence encodes an arginine (R) at an amino acid position corresponding to position 634 of the human RBM20 protein. In some embodiments, the human RBM20 protein comprises or consists of the amino acid sequence of SEQ ID NO: 93. In some embodiments, the human RBM20 nucleic acid sequence encodes an R634Q mutation, wherein theAttorney Docket No.:TYA-072WO numbering is according to the human RBM20 protein having the amino acid sequence of SEQ ID NO: 93. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R636Q mutation, wherein the numbering is according to the murine RBM20 protein having the amino acid sequence of SEQ ID NO: 91. In some embodiments, the human RBM20 nucleic acid sequence comprises at least 15 consecutive nucleic acids between 230-284 of SEQ ID NO: 1 that are replaced with a human RBM20 nucleic acid sequence. In some embodiments, the human RBM20 nucleic acid sequence comprises a sequence disclosed herein. In some embodiments, the human RBM20 nucleic acid sequence replaces at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, or at least 220 nucleotides upstream and / or downstream of nucleic acids 250-264 of SEQ ID NO: 1. In some embodiments, a human RBM20-targeting guide RNA specifically binds to the humanized RBM20 gene. In some embodiments, the humanized RBM20 gene allows targeting of a nucleic acid-guided nuclease to the RBM20 nucleotide sequence using a human RBM20-targeting guide RNA. In some embodiments, the present disclosure provides methods of use thereof. In some aspects, the present disclosure provides methods of use thereof. In some aspects, the present disclosure provides methods of determining the efficacy of gene editing systems targeting the human RBM20 gene. In some aspects, the present disclosure provides a population of cardiomyocytes isolated thereof.Definitions
[0063] Unless the context indicates otherwise, the features of the invention can be used in any combination. Any feature or combination of features set forth can be excluded or omitted. Certain features of the invention, which are described in separate embodiments may also be provided in combination in a single embodiment. Features of the invention, which are described in a single embodiment may also be provided separately or in any suitable sub-combination. All combinations of the embodiments are disclosed herein as if each and every combination were individually disclosed. All sub-combinations of the embodiments and elements are disclosed herein as if every such subcombination were individually disclosed.
[0064] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.Attorney Docket No.:TYA-072WO The detailed description is divided into sections only for the reader’s convenience and disclosure found in any section may be combined with that in another section. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the exemplary methods and materials are now described. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. Reference to a publication is not an admission that the publication is prior art.
[0065] The singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. For example, reference to “a recombinant AAV virion” includes a plurality of such virions and reference to “the cardiac cell” includes one or more cardiac cells.
[0066] The conjunction “and / or” means both “and” and “or,” and lists joined by “and / or” encompasses all possible combinations of one or more of the listed items.
[0067] The use of numerical values in the various quantitative values specified in this application, unless expressly indicated otherwise, are stated as approximations as though the minimum and maximum values within the stated ranges were both preceded by the word "about." It is to be understood, although not always explicitly stated, that all numerical designations are preceded by the term “about.” It is to be understood that such range format is used for convenience and brevity and should be understood flexibly to include numerical values explicitly specified as limits of a range, but also to include all individual numerical values or sub-ranges encompassed within that range as if each numerical value and sub-range is explicitly specified. For example, a ratio in the range of about 1 to about 200 should be understood to include the explicitly recited limits of about 1 and about 200, but also to include individual ratios, such as about 2, about 3, and about 4, and sub-ranges, such as about 10 to about 50, about 20 to about 100, and so forth. It also is to be understood, although not always explicitly stated, that the reagents described herein are merely exemplary and that equivalents of such are known in the art.
[0068] The term, “protospacer adjacent motif (PAM),” as used herein, refers to a nucleotide sequence found in a target nucleic acid that directs an effector protein to modify the target nucleic acid at a specific location. In some instances, a PAM is required for a complex of an effector protein and a gRNA to hybridize to and modify the target nucleic acid. In some instances, the complex does not require a PAM to modify the target nucleic acid. One example of a PAM sequence is NTTN,Attorney Docket No.:TYA-072WO where N can be any nucleic acid.
[0069] The term, “template sequence,” as used herein, refers to a portion of a pegRNA that contains a desired nucleotide modification relative to a target sequence or portion thereof. By way of non-limiting example, the desired edit may comprise one or more nucleotide insertions, deletions or substitutions relative to a target sequence or portion thereof. In some embodiments, it is identical to, complementary to, or reverse complementary to a target sequence or portion thereof. In some embodiments, the template sequence is identical to or reverse complementary to a sequence of the target nucleic acid that is adjacent to a nick site of a target site to be edited, with the exception that it includes one or more desired edits. The template sequence can be identical to or reverse complementary to at least a portion of the target sequence with the exception of at least one nucleotide.
[0070] The terms, “primer binding site (PBS),” as used herein, refer to a portion of a pegRNA and serves to bind to a primer sequence of the target nucleic acid. In some embodiments, the PBS binds to a primer sequence in the target nucleic acid that is formed after the target nucleic acid is cleaved by an effector protein. In some embodiments, the PBS is linked to the 3’ end of a pegRNA. In some embodiments, the PBS is located at the 5’ end of a pegRNA.
[0071] ‘Primer sequence” as used herein refers to a portion of the target nucleic acid that is capable of hybridizing with the primer binding site (PBS) portion of a pegRNA that is generated after cleavage of the target nucleic acid by an effector protein described herein.
[0072] The term “trans-activating RNA (tracrRNA),” as used herein, refers to a nucleic acid that comprises a first sequence that is capable of being non-covalently bound by an effector protein, and a second sequence that hybridizes to a repeat portion of a crRNA, which may be referred to as a repeat hybridization sequence.
[0073] The terms, “CRISPR RNA” or “crRNA,” as used herein, refer to a type of gRNA, wherein the nucleic acid is RNA comprising a spacer sequence, that hybridizes to a target sequence of a target nucleic acid, and a repeat sequence that interacts with an effector protein. In some instances, the repeat sequence is bound by the effector protein. In some instances, the repeat sequence hybridizes to a portion of a tracrRNA, wherein the tracrRNA forms a complex with the effector protein.
[0074] The term, “nickase” as used herein refers to an enzyme that possesses catalytic activityAttorney Docket No.:TYA-072WO for single stranded nucleic acid cleavage of a double stranded nucleic acid. A nickase cleaves a phosphodiester bond between two nucleotides of only one strand of dsDNA.
[0075] The term, “nuclease activity,” is used to refer to catalytic activity that results in nucleic acid cleavage (e.g., ribonuclease activity (ribonucleic acid cleavage), or deoxyribonuclease activity (deoxyribonucleic acid cleavage), etc.).
[0076] The term "humanized' refers to a nucleic acid or protein whose sequence (i.e., nucleotide or amino acid sequence) includes one or more sequences that correspond substantially or identically with the sequences of a particular gene or protein found in nature in a non-human animal (i.e.., one or more endogenous sequences), and also include one or more sequences that differ from that found in the particular gene or protein (i.e., one or more heterologous sequences) and instead correspond more closely with comparable sequences found in a corresponding human gene or protein (i.e., one or more human sequences). In some embodiments, a humanized gene comprises at least a portion of a DNA sequence of a human gene. In some embodiment, a humanized gene comprises an entire DNA sequence of a human gene. In some embodiments, a humanized protein comprises a sequence having a portion that appears in a human protein. In some embodiments, a humanized protein comprises an entire sequence of a human protein and is expressed from an endogenous locus of a non-human animal that corresponds to the homolog or ortholog of the human gene.
[0077] The term “variant” refers to a protein or nucleic acid having one or more genetic changes (e.g., insertions, deletions, substitutions, or the like) that returns all or substantially all of the functions of the reference protein or nucleic acid. For example, a variant of a therapeutic protein retains the same or substantially the same activity and / or provides the same or substantially the same therapeutic benefit to a subject in need thereof. A variant of a promoter sequence retains the ability to initiate transcription at the same, substantially the same, or an increased level as the reference promoter, and retains the same or substantially the same cell type specificity. In particular embodiments, polynucleotides variants have at least or about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity to a reference sequence. In particular embodiments, protein variants have at least or about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%. 91%, 92%, 93%, 94%, 95%. 96%, 97%, 98%, 99% or 100% sequence identity to aAttorney Docket No.:TYA-072WO reference sequence.
[0078] The term “vector” refers to a macromolecule or complex of molecules comprising a polynucleotide or protein to be delivered to a cell. In some embodiments, a “vector” refers to a DNA construct containing a nucleic acid molecule that is operably linked to a suitable control sequence capable of effecting the expression of the nucleic acid molecule in a suitable host. Such control sequences may include a promoter to effect transcription, an optional operator sequence to control such transcription, a sequence encoding suitable mRNA ribosome binding sites, and sequences which control termination of transcription and translation. The vector may be a plasmid, a phage particle, a virus, or simply a potential genomic insert. Once transformed into a suitable host, the vector may replicate and function independently of the host genome, or may, in some instances, integrate into the genome itself.
[0079] The term “wild-type” or “WT” refers to the naturally-occurring polynucleotide sequence encoding a protein, or a portion thereof, or protein sequence, or portion thereof, respectively, as it normally exists in vivo in a normal or healthy subject.
[0080] “AAV” is an abbreviation for adeno-associated virus. The term covers all subtypes of AAV. except where a subtype is indicated, and to both naturally occurring and recombinant forms. The abbreviation “rAAV” refers to recombinant adeno-associated virus. “AAV” includes AAV or any subtype. “AAV5” refers to AAV subtype 5. “AAV9” refers to AAV subtype 9. The genomic sequences of various serotypes of AAV, as well as the sequences of the native inverted terminal repeats (ITRs), Rep proteins, and capsid subunits may be found in the literature or in public databases such as GenBank. See, e.g., GenBank Accession Numbers NC_002077 (AAV1), AF063497 (AAV1), NC_001401 (AAV2). AF043303 (AAV2). NC_001729 (AAV3). NC_001829 (AAV4). U89790 (AAV4), NC_006152 (AAV5), AF513851 (AAV7), AF513852 (AAV8), NC_006261 (AAV8), and AY530579 (AAV9). Publications describing AAV include Srivistava et al. (1983) J. Virol. 45:555; Chiorini et al. (1998) J. Virol. 71:6823; Chiorini et al. (1999) J. Virol. 73:1309; Bantel-Schaal et al. (1999) J. Virol. 73:939; Xiao etal. (1999) J. Virol. 73:3994; Muramatsu et al. (1996) Virol. 221:208; Shade et al. (1986) J. Virol. 58:921; Gao et al. (2002) Proc. Nat. Acad. Sci. USA 99: 11854; Moris et al. (2004) Virology 33:375-383; Int’l Pat. Publ Nos. WO2018 / 222503A1, WO2012 / 145601A2, W02000 / 028061A2, WO1999 / 61601A2, and WO1998 / 11244A2; U. S. Pat. Appl. Nos. 15 / 782,980 and 15 / 433,322; and U. S. Pat. Nos. 10,036.016, 9,790,472. 9.737,618, 9,434.928, 9.233,131.Attorney Docket No.:TYA-072WO 8,906,675, 7,790,449, 7,906,111, 7,718,424, 7,259,151, 7,198,951, 7,105,345, 6,962,815, 6,984,517, and 6,156,303.
[0081] An “AAV vector” or “rAAV vector” as used in the art to refer either to the DNA packaged into in the rAAV virion or to the rAAV virion itself, depending on context. As used herein, unless otherwise apparent from context, rAAV vector refers to a nucleic acid (typically a plasmid) comprising a polynucleotide sequence capable of being packaged into an rAAV virion, but with the capsid or other proteins of the rAAV virion. Generally an rAAV vector comprises a heterologous polynucleotide sequence (i.e., a polynucleotide not of AAV origin) and one or two AAV inverted terminal repeat sequences (ITRs) flanking the heterologous polynucleotide sequence. Only one of the two ITRs may be packaged into the rAAV and yet infectivity of the resulting rAAV virion may be maintained. See Wu et al. (2010) Mol Ther. 18:80. An rAAV vector may be designed to generate either single-stranded (ssAAV) or self-complementary (scAAV). See McCarty D. (2008) Mo. Ther.16:1648-1656; W02001 / 11034; W02001 / 92551; WO2010 / 129021.
[0082] An “rAAV virion” refers to an extracellular viral particle including at least one viral capsid protein (e.g. VP1) and an encapsulated rAAV vector (or fragment thereof), including the capsid proteins.
[0083] The term “inverted terminal repeats” or “ITRs” as used herein refers to AAV viral ciselements named so because of their symmetry. These elements are essential for efficient multiplication of an AAV genome.
[0084] The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of tissue culture, immunology, molecular biology, cell biology and recombinant DNA, which are within the skill of the art. See, e.g., Sambrook and Russell eds. (2001) Molecular Cloning: A Laboratory Manual, 3rdedition; Ausubel et al. eds. (2007) Current Protocols in Molecular Biology; Methods in Enzymology (Academic Press, Inc., N. Y.); MacPherson et al. (1991) PCR 1: A Practical Approach (IRL Press at Oxford University Press); MacPherson et al. (1995) PCR 2: A Practical Approach; Harlow and Lane eds. (1999) Antibodies, A Laboratory Manual; Freshney (2005) Culture of Animal Cells: A Manual of Basic Technique, 5thedition; Gait ed. (1984) Oligonucleotide Synthesis; U. S. Pat. No. 4,683.195; Hames and Higgins eds. (1984) Nucleic Acid Hybridization; Anderson (1999) Nucleic Acid Hybridization; Hames and Higgins eds. (1984) Transcription and Translation; IRL Press (1986) Immobilized Cells and Enzymes; Perbal (1984) AAttorney Docket No.:TYA-072WO Practical Guide to Molecular Cloning; Miller and Calos eds. (1987) Gene Transfer Vectors for Mammalian Cells (Cold Spring Harbor Laboratory); Makrides ed. (2003) Gene Transfer and Expression in Mammalian Cells; Mayer and Walker eds. (1987) Immunochemical Methods in Cell and Molecular Biology (Academic Press, London); Herzenberg et al. eds (1996) Weir’s Handbook of Experimental Immunology; Manipulating the Mouse Embryo: A Laboratory Manual, 3rdedition (2002) Cold Spring Harbor Laboratory Press; Sohail (2004) Gene Silencing by RNA Interference: Technology and Application (CRC Press); and Sell (2013) Stem Cells Handbook.
[0085] The term “isolated” means separated from constituents, cellular and otherwise, in which the virion, cell, tissue, polynucleotide, peptide, polypeptide, or protein is normally associated in nature. For example, an isolated cell is a cell that is separated from tissue or cells of dissimilar phenotype or genotype.
[0086] As used herein, “sequence identity” or “identity” refers to the percentage of number of amino acids that are identical between a sequence of interest and a reference sequence. Generally identity is determined by aligning the sequence of interest to the reference sequence, determining the number of amino acids that are identical between the aligned sequences, dividing that number by the total number of amino acids in the reference sequence, and multiplying the result by 100 to yield a percentage. Sequences can be aligned using various computer programs, such BLAST, available at ncbi.nlm.nih.gov. Other techniques for alignment are described in Methods in Enzymology, vol. 266: Computer Methods for Macromolecular Sequence Analysis ( 1996); and Meth. Mol. Biol. 70: 173- 187 (1997); J. Mol. Biol. 48: 44. Skill artisans are capable of choosing an appropriate alignment method depending on various factors including sequence length, divergence, and the presence of absence of insertions or deletions with respect to the reference sequence.
[0087] A “gene” refers to a polynucleotide containing at least one open reading frame that is capable of encoding a particular protein after being transcribed and translated. For avoidance of doubt, the term “gene” as used herein encompasses the coding and non-coding strand of DNA at a particular genetic locus. A “gene product” is a molecule resulting from expression of a particular gene. Gene products may include, without limitation, a polypeptide, a protein, an aptamer, an interfering RNA, or an mRNA. Gene-editing systems (e.g. a prime editing system) may be described as one gene product or as the several gene products required to make the system (e.g. a Cas protein, a reverse transcriptase, and a guide RNA).Attorney Docket No.:TYA-072WO
[0088] A “control element” or “control sequence” is a nucleotide sequence involved in an interaction of molecules that contributes to the functional regulation of a polynucleotide, including replication, duplication, transcription, splicing, translation, or degradation of the polynucleotide. The regulation may affect the frequency, speed, or specificity of the process, and may be enhancing or inhibitory in nature. Control elements include transcriptional regulatory sequences such as promoters and / or enhancers.
[0089] A “promoter” is a DNA sequence capable under certain conditions of binding RNA polymerase and initiating transcription of a coding region usually located downstream (in the 3’ direction) from the promoter.
[0090] “Operatively linked” or “operably linked” refers to a juxtaposition of genetic elements, wherein the elements are in a relationship permitting them to operate in the expected manner. For instance, a promoter is operatively linked to a coding region if the promoter helps initiate transcription of the coding sequence. There may be intervening residues between the promoter and coding region so long as this functional relationship is maintained.
[0091] The term “expression cassette” refers to a polynucleotide cassette comprising a coding sequence which encodes a gene product of interest used to effect the expression of the gene product in target cells. Unless otherwise specified, the expression cassette of an AAV vector includes only the polynucleotides between (and not including) the ITRs.
[0092] The terms “upstream” and “upstream end” refer to a portion of a polynucleotide that is, with reference to a specified site or sequence, 5' to the specified site or sequence on the sense strand (or coding strand) of the polynucleotide; and 3' to the specified site or sequence on the antisense strand of the polynucleotide. The terms “downstream” and “downstream end” refer to a portion of a polynucleotide that is, with reference to a specified site or sequence, 3' to the specified site or sequence on the sense strand (or coding strand) of the polynucleotide; and 5' to the specified site or sequence on the antisense strand of the polynucleotide.
[0093] “Heterologous” means derived from a genotypically distinct entity from that of the rest of the entity to which it is being compared. For example, a polynucleotide introduced by genetic engineering techniques into a plasmid or vector derived from a different species is a heterologous polynucleotide. A promoter removed from its native coding sequence and operatively linked to aAttorney Docket No.:TYA-072WO coding sequence with which it is not naturally found linked is a heterologous promoter. Thus, for example, an rAAV that includes a heterologous nucleic acid is an rAAV that includes a nucleic acid not normally included in a naturally- occurring AAV.
[0094] The terms “genetic alteration” and “genetic modification” (and grammatical variants thereof), are used interchangeably herein to refer to a process wherein a genetic element (e.g., a polynucleotide) is introduced into a cell other than by mitosis or meiosis. The element may be heterologous to the cell, or it may be an additional copy or improved version of an element already present in the cell. Genetic alteration may be effected, for example, by transfecting a cell with a polynucleotide through any process known in the art, such as electroporation, calcium phosphate precipitation, or contacting with a polynucleotide-liposome complex. Genetic alteration may also be effected, for example, by transduction or infection with a vector.
[0095] A cell is said to be “stably” altered, transduced, genetically modified, or transformed with a polynucleotide sequence if the sequence is available to perform its function during extended culture of the cell in vitro. Generally, such a cell is “heritably” altered (genetically modified) in that a genetic alteration is introduced which is also inheritable by progeny of the altered cell.
[0096] The term “transfection” is as used herein refers to the uptake of an exogenous nucleic acid molecule by a cell. A cell has been “transfected” when exogenous nucleic acid has been introduced inside the cell membrane. A number of transfection techniques are generally known in the art. See, e.g., Graham et al. (1973) Virology, 52:456, Sambrook et al. (1989) Molecular Cloning, a laboratory manual, Cold Spring Harbor Laboratories, New York, Davis et al. (1986) Basic Methods in Molecular Biology, Elsevier, and Chu et al. (1981) Gene 13:197. Such techniques can be used to introduce one or more exogenous nucleic acid molecules into suitable host cells.
[0097] The term “transduction” is as used herein refers to the transfer of an exogenous nucleic acid into a cell by a recombinant virion, in contrast to “infection” by a wild-type virion. When infection is used with respect to a recombinant virion, the terms “transduction” and “infectious” are synonymous, and therefore “infectivity” and “transduction efficiency” are equivalent and can be determined using similar methods.
[0098] Unless otherwise specified, all medical terminology is given the ordinary meaning of the term used by medical professionals as, for example, in Harrison ’s Principles of Internal Medicine.Attorney Docket No.:TYA-072WO 15ed., which is incorporated by reference in its entirety for all purposes, in particular the chapters on cardiac or cardiovascular diseases, disorders, conditions, and dysfunctions.
[0099] “Administration,” “administering” and the like, when used in connection with a composition of the invention refer both to direct administration (administration to a subject by a medical professional or by self-administration by the subject) and / or to indirect administration (prescribing a composition to a patient). Typically, an effective amount is administered, which amount can be determined by one of skill in the art. Any method of administration may be used. Administration to a subject can be achieved by, for example, intravenous, intra-arterial, intramuscular, intravascular, or intramyocardial delivery.
[0100] As used herein the term “effective amount” and the like in reference to an amount of a composition refers to an amount that is sufficient to induce a desired physiologic outcome (e.g., reprogramming of a cell or treatment of a disease). An effective amount can be administered in one or more administrations, applications or dosages. Such delivery is dependent on a number of variables including the time period which the individual dosage unit is to be used, the bioavailability of the composition, the route of administration, etc. It is understood, however, that specific amounts of the compositions for any particular subject depends upon a variety of factors including the activity of the specific agent employed, the age, body weight, general health, sex, and diet of the subject, the time of administration, the rate of excretion, the composition combination, severity of the particular disease being treated and form of administration.
[0101] The terms “individual,” “subject,” and “patient” are used interchangeably herein, and refer to a mammal, including, but not limited to, human and non-human primates (e.g., simians); mammalian sport animals (e.g., horses); mammalian farm animals (e.g., sheep, goats, etc.); mammalian pets (e.g., dogs, cats, etc.); and rodents (e.g., mice, rats, etc.). In some embodiments, the subject is a genetically-modified, non-human animal described herein.
[0102] As used herein the term “cardiac cell” refers to any cell present in the heart that provides a cardiac function, such as heart contraction or blood supply, or otherwise serves to maintain the structure of the heart. Cardiac cells as used herein encompass cells that exist in the epicardium, myocardium or endocardium of the heart. Cardiac cells also include, for example, cardiac muscle cells or cardiomyocytes, and cells of the cardiac vasculatures, such as cells of a coronary artery or vein. Other non-limiting examples of cardiac cells include epithelial cells, endothelial cells,Attorney Docket No.:TYA-072WO fibroblasts, cardiac stem or progenitor cells, cardiac conducting cells and cardiac pacemaking cells that constitute the cardiac muscle, blood vessels and cardiac cell supporting structure. Cardiac cells may be derived from stem cells, including, for example, embryonic stem cells or induced pluripotent stem cells.
[0103] The term “cardiomyocyte” or “cardiomyocytes” as used herein refers to sarcomerecontaining striated muscle cells, naturally found in the mammalian heart, as opposed to skeletal muscle cells. Cardiomyocytes are characterized by the expression of specialized molecules, e.g., proteins like myosin heavy chain, myosin light chain, cardiac a-actinin. The term “cardiomyocyte” as used herein is an umbrella term comprising any cardiomyocyte subpopulation or cardiomyocyte subtype, e.g., atrial, ventricular and pacemaker cardiomyocytes.
[0104] The term “cardiomyocyte-like cells” is intended to mean cells sharing features with cardiomyocytes, but which may not share all features. For example, a cardiomyocyte-like cell may differ from a cardiomyocyte in expression of certain cardiac genes.
[0105] Unless stated otherwise, the abbreviations used throughout the specification have the following meanings: AAV, adeno-associated virus, rAAV, recombinant adeno-associated virus; kg, kilogram; μg, microgram; μl, microliter; mg, milligram; ml, milliliter; min, minute; PBS, phosphate buffered saline; sec, second.RNA Binding Motif Protein 20 (RBM20) Gene Editing Systems
[0106] In some embodiments, the present disclosure provides gene editing systems that are capable of modifying pathogenic variants of the human RBM20 gene when introduced into a cell. RNA-binding motif protein 20 (RBM20) is an RNA binding protein that regulates splicing (UniProt #Q5T481). In some embodiments, the gene editing system is a base editing system. In some embodiments, the gene editing system is a prime editing system.gRNAs
[0107] In some embodiments, the present disclosure provides guide nucleic acids for targeting pathogenic variants of the human RBM20 gene and gene editing systems comprising the same. In some embodiments, the guide nucleic acid is a guide RNA (gRNA). In general, a gRNA comprises a protein binding segment that is capable of being non-covalently bound by an effector protein and a spacer sequence that hybridizes to a target nucleic acid sequence in the gene of interest (e.g., a targetAttorney Docket No.:TYA-072WO sequence in the human RBM20 gene). It is understood that guide nucleic acids may comprise DNA, RNA, or a combination thereof (e.g., RNA with a thymine base). Guide nucleic acids may include a chemically modified nucleobase or phosphate backbone. In some embodiments, the gRNA specifically targets a human RBM20 nucleic acid sequence.Spacer sequences
[0108] The spacer sequence comprises a nucleotide sequence that binds to a target nucleic acid sequence. In some embodiments, the spacer sequence binds to a target sequence on the coding strand of a dsDNA molecule. In some embodiments, the spacer sequence binds to a target sequence on the template strand of a dsDNA molecule. The spacer sequence can function to direct the guide nucleic acid to the target nucleic acid for detection and / or modification of the gene of interest. The spacer sequence may bind to a target sequence that is adjacent to (e.g., is 5’ or 3’ to) a PAM that is recognizable by an effector protein of interest.
[0109] As such, the spacer sequence of a gRNA interacts with a target nucleic acid in a sequence- specific manner via hybridization (i.e., base pairing) and determines the location within the target nucleic acid that the gRNA will bind. The spacer sequence of a gRNA can be modified e.g., by genetic engineering) to hybridize to a desired sequence within a target nucleic acid sequence. In some embodiments, the spacer sequence is between about 15 and about 25 nucleotides in length. In some embodiments, the spacer sequence is about 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 nucleotides in length. In some embodiments, the spacer sequence is about 19 nucleotides in length. In some embodiments, the spacer sequence is about 20 nucleotides in length. In some embodiments, the spacer sequence is about 21 nucleotides in length.
[0110] In some embodiments, the spacer sequence binds to a target sequence, wherein the target sequence is a human RBM20 nucleic acid sequence. In some embodiments, the target sequence is located on the coding strand of a dsDNA molecule comprising a human RBM20 gene or a humanized RBM20 gene. In some embodiments, the target sequence is located on the template strand of a dsDNA molecule comprising a human RBM20 gene or a humanized RBM20 gene. In some embodiments, the human RBM20 gene or humanized RBM20 gene is a wild-type human RBM20 allele. In some embodiments, the spacer sequence binds to a target sequence of a wild-type human RBM20 allele, a wild-type humanized RBM20 allele, or a wild-type human RBM20 nucleic acid sequence (hereafter called a “wild-type human RBM20 target sequence”). In some embodiments, the human RBM20 geneAttorney Docket No.:TYA-072WO or humanized RBM20 gene is a mutant human RBM20 allele. In some embodiments, the spacer sequence binds to a target sequence of a mutant human RBM20 allele (e.g., a pathogenic mutant human RBM20), a mutant humanized RBM20 allele, or a mutant human RBM20 nucleic acid sequence (hereafter called a “mutant human RBM20 target sequence”). In some embodiments, the human RBM20 gene or humanized RBM20 gene encodes an RBM20 protein comprising arginine (R) at an amino acid position corresponding to position 634 of the human RBM20 protein. In some embodiments, the human RBM20 protein comprises or consists of the amino acid sequence of SEQ ID NO: 93. In some embodiments, the human RBM20 nucleic acid sequence comprises an R634Q mutation, wherein the numbering is according to the human RBM20 protein of SEQ ID NO: 93. Illustrative examples of spacer sequences that bind to a target sequence of the human RBM20 gene are show in Table 1 below.Table 1: Exemplary spacer sequences targeting the RBM20 geneReference RNA Sequence SEQ ID NO: peg-1 ctcacagatatggcccagaa 95 peg-28 ctcacagatatggcccagaa 95 peg-32 ctcacagatatggcccagaa 95 peg- 186 gggagagtgaccggctcac 96 peg- 189 gggagagtgaccggctcac 96 peg-224 gggagagtgaccggctcac 96 peg-262 gggagagtgaccggctcac 96 peg-263 ggagagtgaccggctcac 195
[0111] In some embodiments, the spacer sequence comprises or consists a nucleotide sequence having at least 80% (e.g., at least 80%. at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity to the nucleotide sequence of any one of SEQ ID NOs: 95, 96, and 195. In some embodiments, the spacer sequence comprises or consists the nucleotide sequence of any one of SEQ ID NOs: 95, 96, and 195, with up to 1, up to 2, up to 3, up to 4, or more nucleotide mismatches or substitutions. In some embodiments, the spacer sequence comprises or consists the nucleotide sequence of any one of SEQ ID NOs: 95, 96. and 195. In some embodiments, the spacer sequence comprises or consists of the sequence of SEQ ID NO: 95. In some embodiments, the spacer sequence comprises or consists of the sequence of SEQ ID NO: 96. In some embodiments,Attorney Docket No.:TYA-072WO the spacer sequence comprises or consists of the sequence of SEQ ID NO: 195.Scaffold sequences
[0112] In some embodiments, the gRNA comprises a scaffold sequence that is capable of being non-covalently bound by an effector protein. In some embodiments, the scaffold sequence comprises one or more hairpin or stem-loop structures that are recognized by an effector protein. The scaffold sequence can comprise a repeat sequence and / or a tracrRNA. In some embodiments, the scaffold sequence comprises a tracrRNA and a repeat sequence. An exemplary scaffold sequences described herein is 5’ - gttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgc - 3’ (SEQ ID NO: 97), comprising a repeat sequence and a tracr sequence.tracrRNA sequences
[0113] Guide nucleic acids (e.g., gRNAs) described herein may comprise a tracr sequence. In general, the tracrRNA sequence non-covalently binds to an effector protein. In some embodiments, the tracrRNA forms a secondary structure, for example in a cell, and an effector protein binds the secondary structure.
[0114] A tracrRNA sequence may also comprise or form a secondary structure (e.g., one or more hairpin loops) that facilitates the binding of an effector protein to a guide nucleic acid and / or modification activity of an effector protein on a target nucleic acid (e.g., a hairpin region). In some embodiments, a tracrRNA sequence comprises a stem-loop structure comprising a stem region and a loop region. In some embodiments, the stem region is 4 to 8 linked nucleotides in length. In some embodiments, the stem region is 5 to 6 linked nucleotides in length. In some embodiments, the stem region is 4 to 5 linked nucleotides in length. An effector protein may interact with a tracrRNA sequence comprising a single stem region or multiple stem regions. In some embodiments, a tracrRNA sequence comprises 1, 2, 3, 4, 5 or more stem regions.Nucleic acid linkers
[0115] In some embodiments, a guide nucleic acid for use with compositions, systems, and methods described herein comprises one or more linkers, or a nucleic acid encoding one or more linkers. In some embodiments, the guide nucleic acid comprises at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten linkers. In some embodiments, the guide nucleic acid comprises one, two, three, four, five, six, seven,Attorney Docket No.:TYA-072WO eight, nine, or ten linkers. In some embodiments, the guide nucleic acid comprises two or more linkers. In some embodiments, at least two or more linkers are the same. In some embodiments, at least two or more linkers are not the same.
[0116] In some embodiments, a linker comprises one to ten, one to seven, one to five, one to three, two to ten, two to eight, two to six, two to four, three to ten, three to seven, three to five, four to ten, four to eight, four to six, five to ten, five to seven, six to ten, six to eight, seven to ten, or eight to ten linked nucleotides. In some embodiments, the linker comprises one, two, three, four, five, six, seven, eight, nine, or ten linked nucleotides.
[0117] In some embodiments, a guide nucleic acid comprises one or more linkers connecting one or more repeat sequences. In some embodiments, the guide nucleic acid comprises one or more linkers connecting one or more repeat sequences and one or more spacer sequences. In some embodiments, the guide nucleic acid comprises at least two repeat sequences connected by a linker. Repeat sequences
[0118] Guide nucleic acids (e.g., gRNAs) described herein may comprise one or more repeat sequences. In some embodiments, a repeat sequence comprises a nucleotide sequence that is not complementary to a target sequence of a target nucleic acid. In some embodiments, a repeat sequence comprises a nucleotide sequence that may interact with an effector protein. In some embodiments, a repeat sequence includes a nucleotide sequence that is capable of forming a guide nucleic acideffector protein complex (e.g., a RNP complex).
[0119] In some embodiments, a repeat sequence is adjacent to a spacer sequence. In some embodiments, a repeat sequence is followed by a spacer sequence in the 5’ to 3’ direction. In some embodiments, a repeat sequence is adjacent to a tracrRNA sequence. In some embodiments, a repeat sequence is 3’ to a tracrRNA sequence. In some embodiments, a tracrRNA sequence is followed by a repeat sequence, which is followed by a spacer sequence in the 5’ to 3’ direction. In some embodiments, a repeat sequence is linked to a spacer sequence and / or a tracrRNA sequence. In some embodiments, a guide nucleic acid comprises a repeat sequence linked to a spacer sequence, which may be a direct link or by any suitable linker, examples of which are described herein.
[0120] In some embodiments, the spacer sequence(s) and the repeat sequence(s) of the guide nucleic acid are present within the same polynucleotide molecule. In some embodiments, a spacerAttorney Docket No.:TYA-072WO sequence is adjacent to a repeat sequence. In some embodiments, a spacer sequence follows a repeat sequence in a 5’ to 3’ direction. In some embodiments, a spacer sequence precedes a repeat sequence in a 5’ to 3’ direction. In some embodiments, the spacer(s) and repeat sequence(s) are linked directly to one another. In some embodiments, a linker is present between the spacer(s) and repeat sequence(s). Linkers may be any suitable linker. In some embodiments, the spacer sequence(s) and the repeat sequence(s) of the guide nucleic acid are present in separate polynucleotide molecules, which are joined to one another by base pairing interactions.
[0121] In some embodiments, the repeat sequence comprises two sequences that are reverse complementary to each other and hybridize to form a double stranded RNA duplex (dsRNA duplex). In some embodiments, the two sequences are not directly linked and hybridize to form a stem loop structure. In some embodiments, the dsRNA duplex comprises 5, 10, 15, 20 or 25 base pairs (bp). In some embodiments, not all nucleotides of the dsRNA duplex are paired, and therefore the duplex forming sequence may include a bulge. In some embodiments, the repeat sequence comprises a hairpin or stem-loop structure, optionally at the 5’ portion of the repeat sequence. In some embodiments, a strand of the stem portion comprises a sequence and the other strand of the stem portion comprises a sequence that is, at least partially, reverse complementary. In some embodiments, such sequences may have 65% to 100% complementarity (e.g.. 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% complementarity). In some embodiments, a guide nucleic acid comprises nucleotide sequence that when involved in hybridization events may hybridize over one or more segments such that intervening or adjacent segments are not involved in the hybridization event (e.g., a bulge, a loop structure or hairpin structure, etc.).Single Nucleic Acid Systems
[0122] In some embodiments, compositions, systems and methods described herein comprise a single nucleic acid system comprising a guide nucleic acid or a nucleotide sequence encoding the guide nucleic acid, and one or more effector proteins or a nucleotide sequence encoding the one or more effector proteins. An exemplary guide nucleic acid for a single nucleic acid system is a crRNA or a sgRNA.crRNA
[0123] In some embodiments, guide nucleic acid comprises a crRNA comprising a spacerAttorney Docket No.:TYA-072WO sequence(s) and a repeat sequence(s) present within the same polynucleotide molecule. Tn some embodiments, the spacer sequence is adjacent to the repeat sequence. In some embodiments, the spacer sequence follows the repeat sequence in a 5’ to 3’ direction. In some embodiments, the spacer sequence precedes the repeat sequence in a 5’ to 3’ direction. In some embodiments, the spacer(s) and repeat sequence(s) are linked directly to one another. In some embodiments, a linker is present between the spacer(s) and repeat sequence(s). Linkers may be any suitable linker.
[0124] In some embodiments, a crRNA is useful as a single nucleic acid system for compositions, methods, and systems described herein or as part of a single nucleic acid system for compositions, methods, and systems described herein. In some embodiments, a crRNA is useful as part of a single nucleic acid system for compositions, methods, and systems described herein. In such embodiments, a single nucleic acid system comprises a guide nucleic acid comprising a crRNA wherein, a repeat sequence of a crRNA is capable of connecting a crRNA to an effector protein. In some embodiments, a single nucleic acid system comprises a guide nucleic acid comprising a crRNA linked to another nucleotide sequence that is capable of being non-covalently bound by an effector protein.
[0125] In some embodiments, a crRNA is sufficient to form a complex with an effector protein (e.g., to form an RNP) through the repeat sequence and direct the effector protein to a target nucleic acid sequence through the spacer sequence. In some embodiments, the repeat sequence in the crRNA polynucleotide hybridizes with a tracr sequence present in a separate polynucleotide. In some embodiments, the hybridization with the tracr sequences permits formation of an RNP complex with an effector protein.
[0126] A crRNA may include deoxyribonucleosides, ribonucleosides, chemically modified nucleosides, or any combination thereof. In some embodiments, a crRNA comprises about: 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, or 60 linked nucleotides. In some embodiments, a crRNA comprises at least: 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60 linked nucleotides. In some embodiments, the length of the crRNA is about 20 to about 120 linked nucleotides. In some embodiments, the length of a crRNA is about 20 to about 100, about 30 to about 100, about 40 to about 100, about 40 to about 90, about 40 to about 80, about 40 to about 70, about 40 to about 60, about 40 to about 50, about 50 to about 90, about 50 to about 80, about 50 to aboutAttorney Docket No.:TYA-072WO 70, or about 50 to about 60 linked nucleotides. In some embodiments, the length of a crRNA is about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70 or about 75 linked nucleotides.sgRNA
[0127] In some embodiments, a guide nucleic acid comprises a single guide RNA (sgRNA). In some embodiments, the guide nucleic acid is a sgRNA. The combination of a spacer sequence (e.g., a nucleotide sequence that hybridizes to a target sequence in a target nucleic acid) with a scaffold sequence may be referred to herein as a single guide RNA (sgRNA), wherein the spacer sequence and the scaffold sequence are covalently linked. In some embodiments, the spacer sequence and scaffold sequence are linked by a phosphodiester bond. In some embodiments, the spacer sequence and scaffold sequence are linked by one or more linked nucleotides. In some embodiments, a guide nucleic acid may comprise a spacer sequence, a repeat sequence, a scaffold sequence, or a combination thereof. In some embodiments, the scaffold sequence may comprise a portion of, or all of, a repeat sequence. In general, a sgRNA comprises a first region and a second region, wherein the first region comprises a scaffold sequence and the second region comprises a spacer sequence.
[0128] In some embodiments, a sgRNA comprises a scaffold sequence and a spacer sequence. In some embodiments, a scaffold sequence is 5’ to a spacer sequence in an sgRNA. In some embodiments, the scaffold sequence is 3’ to a spacer sequence in an sgRNA. In some embodiments, a sgRNA comprises a linked scaffold sequence and spacer sequence. In some embodiments, a scaffold sequence and a spacer sequence are linked in an sgRNA directly (e.g., covalently linked, such as through a phosphodiester bond) In some embodiments, a scaffold sequence and a spacer sequence are linked in an sgRNA by any suitable linker, examples of which are provided herein.Dual Nucleic Acid Systems
[0129] In some embodiments, compositions, systems and methods described herein comprise a dual nucleic acid system comprising a crRNA or a nucleotide sequence encoding the crRNA, a tracrRNA or a nucleotide sequence encoding the tracrRNA, and one or more effector proteins or a nucleotide sequence encoding the one or more effector proteins, wherein the crRNA and the tracrRNA are separate, unlinked molecules, wherein a repeat hybridization region of the tracrRNA is capable of hybridizing with an equal length portion of the crRNA to form a tracrRNA-crRNA duplex, whereinAttorney Docket No.:TYA-072WO the equal length portion of the crRNA does not include a spacer sequence of the crRNA, and wherein the spacer sequence is capable of hybridizing to a target sequence of the target nucleic acid. In the dual nucleic acid system having a complex of the guide nucleic acid, tracrRNA, and the effector protein, the effector protein is transactivated by the tracrRNA. In other words, in a dual nucleic acid system, activity of the effector protein requires binding to a tracrRNA molecule.
[0130] In some embodiments, a repeat hybridization sequence is at the 3’ end of a tracrRNA. In some embodiments, a repeat hybridization sequence may have a length of about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 12, about 14, about 16, about 18, or about 20 linked nucleotides. In some embodiments, the length of the repeat hybridization sequence is 1 to 20 linked nucleotides. In some embodiments, systems, compositions, and methods comprise a crRNA or a use thereof. In general, a crRNA comprises a first region and a second region, wherein the first region of the crRNA comprises a repeat sequence, and the second region of the crRNA comprises a spacer sequence. In some embodiments, the repeat sequence and the spacer sequences are directly connected to each other (e.g., covalent bond (phosphodiester bond)). In some embodiments, the repeat sequence and the spacer sequence are connected by a linker.
[0131] In some embodiments, systems, compositions, and methods comprise a tracrRNA or a use thereof. In some embodiments, systems, compositions, and methods do not comprise a tracrRNA or a use thereof. A tracrRNA and / or tracrRNA-crRNA duplex may form a secondary structure that facilitates the binding of an effector protein to a tracrRNA or a tracrRNA-crRNA. In some embodiments, the secondary structure modifies activity of the effector protein on a target nucleic acid. In some embodiments, the secondary structure comprises a stem-loop structure comprising a stem region and a loop region. In some embodiments, the stem region is 4 to 8 linked nucleotides in length. In some embodiments, the stem region is 5 to 6 linked nucleotides in length. In some embodiments, the stem region is 4 to 5 linked nucleotides in length. In some embodiments, the secondary structure comprises a pseudoknot (e.g., a secondary structure comprising a stem at least partially hybridized to a second stem or half-stem secondary structure). An effector protein may recognize a secondary structure comprising multiple stem regions. In some embodiments, nucleotide sequences of the multiple stem regions are identical to one another. In some embodiments, the nucleotide sequences of at least one of the multiple stem regions is not identical to those of the others. In some embodiments, the secondary structure comprises at least two, at least three, at least four, or at leastAttorney Docket No.:TYA-072WO five stem regions. In some embodiments, the secondary structure comprises one or more loops. In some embodiments, the secondary structure comprises at least one, at least two, at least three, at least four, or at least five loops.Prime Editing gRNAs
[0132] Prime editing enables the replacement of genomic sequences without the need for a double-stranded DNA break and a donor template carrying the desired sequence changes. In brief, a reverse transcriptase is fused to a nucleic acid-guided nuclease and targets the genomic region guided by a prime editing guide RNA (pegRNA). The pegRNA directs the fusion protein to the target region; the edit encoded in the 3’ extension of the pegRNA is reverse transcribed to the 3’ end of the nicked genomic DNA strand. This leads to the generation of a 3’ flap containing the edited sequence and a 5’ flap of wild-type (WT) sequence surrounding the nicked site. These flaps are resolved via a flap excision and heteroduplex repair system. Excision of the 5’ flap and incorporation of the 3’ flap lead to an editing event (Anzalone et al., 2019).
[0133] In some embodiments, the gRNA is a prime editing guide RNA (pegRNA). In the context of prime editing, the strand to which the spacer sequence of the pegRNA binds is referred to as the “target strand”, and the strand to which the templated edit is incorporated (and not bound by the spacer sequence of the pegRNA) is referred to as the “non-target strand”.
[0134] PegRNAs comprise a spacer sequence that binds to a target nucleic acid on the target strand of a dsDNA, a scaffold sequence capable of binding the effector protein, a template sequence, and a primer binding site (PBS). The primer binding sequence then hybridizes with a sequence on the non-target strand of the dsDNA molecule (in this embodiment, the coding strand) and the template sequence is incorporated into the non-target strand through the actions of the reverse transcriptase fused to a Cas protein. In some embodiments, the spacer sequence may be any spacer sequence disclosed herein. In some embodiments, the scaffold sequence may be any scaffold sequence disclosed herein or known in the art.
[0135] The template sequence may comprise one or more nucleotides having a different nucleobase than that of a nucleotide at the corresponding position in the target nucleic acid when a spacer sequence of the guide RNA and the target sequence are aligned for maximum identity. As such, the template sequence can comprise a desired edit to be introduced at the target nucleic acidAttorney Docket No.:TYA-072WO sequence.
[0136] The template sequence may comprise one or more nucleotides having a different nucleobase than that of a nucleotide at the corresponding position in the target nucleic acid when a spacer sequence of the guide RNA and the target sequence are aligned for maximum identity. The one or more nucleotides may be contiguous. The one or more nucleotides may not be contiguous. The one or more nucleotides may each independently be selected from guanine, adenine, cytosine and thymine.
[0137] In some embodiments, the template sequence comprises or consists a nucleotide sequence having at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%. or 100%) identity to the nucleotide sequence of aggccgcggtctcgtagtccggtg (SEQ ID NO: 197), ggccgcggtctcgtagtccggtg (SEQ ID NO: 198), or aggccgcggtctcgtagtccggtc (SEQ ID NO: 201). In some embodiments, the template sequence comprises or consists the nucleotide sequence of any one of SEQ ID NOs: 197. 198. and 201. with up to 1, up to 2, up to 3, up to 4, up to 5, or more nucleotide mismatches or substitutions. In some embodiments, the template sequence comprises or consists the nucleotide sequence of any one of SEQ ID NOs: 197, 198, and 201.
[0138] In some instances, the primer binding site (PBS) hybridizes to a primer sequence on the non-target strand of the dsDNA molecule. In some embodiments, the spacer sequence of the pegRNA binds a target sequence on the target strand of the dsDNA molecule, and the PBS binds a primer sequence on the non-target strand of the target dsDNA molecule.
[0139] In some embodiments, the PBS is 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides long. In some embodiments, the PBS is at least 7, 8, 9. 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 3435, 36, 37, 38, 39, or 40 nucleotides long. In some embodiments, at least a portion of the PBS binds at least a portion of the target nucleic acid sequence. In some embodiments, the PBS comprises or consists a nucleotide sequence having at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity to the nucleotide sequence of agccggtcac (SEQ ID NO: 196). In some embodiments, the PBS comprises or consists the nucleotide sequence of SEQ ID NO: 196, with up to 1, up to 2, or more nucleotide mismatches or substitutions. In some embodiments, the PBS comprises or consists the nucleotide sequence of SEQ ID NO: 196.Attorney Docket No.:TYA-072WO
[0140] In some embodiments, the template sequence and PBS are described herein as a combined “template + primer binding site sequence” or “template + PBS sequence”. In some embodiments, the template + PBS sequence comprises or consists a nucleotide sequence having at least 80% (e.g., at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity to the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199. and 200. In some embodiments, the template + PBS sequence comprises or consists the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200, with up to 1, up to 2, up to 3, up to 4, up to 5, or more nucleotide mismatches or substitutions. In some embodiments, the template + PBS sequence comprises or consists the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200. In some embodiments, the template + PBS sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 103. In some embodiments, the template + PBS sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 199. Illustrative examples of template + PBS sequences are show in Table 2 below.
[0141] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of any one of SEQ ID NOs: 95. 96, and 195; a scaffold sequence; and a template + PBS sequence comprising or consisting of the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200. In some embodiments, the pegRNA comprises a combination of a specific spacer sequence and a specific template + PBS sequence. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 98. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 99. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 100. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 101. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence, and a template + PBS sequence comprising or consisting of theAttorney Docket No.:TYA-072WO nucleotide sequence of SEQ ID NO: 102. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 103. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 195, a scaffold sequence, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 200. Illustrative examples of combinations of pegRNA spacer sequences and template + PBS sequences are show in Table 2 below.Table 2: Exemplary combinations of pegRNA spacer sequences and template + PBS sequences Reference Spacer sequence SEQ ID NO: Template + PBS sequence SEQ ID NO: peg-1 ctcacagatatggcccagaa 95 tcaccggactacgagaccgcggtctttctgg 98gccatatcpeg-28 ctcacagatatggcccagaa 95 acgagaccgcggtctttctgggccatatctg 99 peg-32 ctcacagatatggcccagaa 95 agaccgcggtctttctgggccatatc 100 peg- 186 gggagagtgaccggctcac 96 aggccgcggtctcgtagtccggtcagccggt 101cacpeg- 189 gggagagtgaccggctcac 96 aggccgaggtctcgtagtccggtcagccggt 102cactcpeg-224 gggagagtgaccggctcac 96 aggccgcggtctcgaagtccggtcagccgg 103tcacpeg-262 gggagagtgaccggctcac 96 aggccgcggtctcgtagtccggtgagccggt 199cacpeg-263 ggagagtgaccggctcac 195 ggccgcggtctcgtagtccggtgagccggtc 200ac
[0142] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NOs: 95 or SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of any one of SEQ ID NOs: 98-103. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 98. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequenceAttorney Docket No.:TYA-072WO comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 99. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 100. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 101. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 102. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 103. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 195, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, and a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 200.
[0143] In some embodiments, the pegRNA comprises, from 5’ to 3’, a spacer sequence, a scaffold sequence, and a template + primer binding site. In some embodiments, the spacer sequence, the scaffold sequence, and the template sequence + primer binding site may each be any sequence described herein.
[0144] In some embodiments, the pegRNA further comprises a 3’ motif. Without wishing to be bound by theory, the 3’ motif may enhance stability of the 3' end of the pegRNA, avoiding degradation and leading to increased editing efficiency. In some embodiments, the 3’ motif is anyAttorney Docket No.:TYA-072WO suitable 3’ motif known in the art. Tn some embodiments, the 3’ motif is operably linked to the primer binding site via a linker. The linker may be any suitable linker known in the art. In some embodiments, the 3’ motif comprises or consists of the nucleotide sequence of SEQ ID NO: 175. In some embodiments, the linker comprises or consists of the nucleotide sequence of any one of: “actattca”, “ataataca”, “ataaacac”, “aaagtaag”, “aatattac”, and “aagaaatt”. In some embodiments, the linker and 3’ motif comprise or consists of the nucleotide sequence of SEQ ID NO: 126. In some embodiments, the linker and 3’ motif comprise or consists of the nucleotide sequence actattcacgcggttctatctagttacgcgttaaaccaactagaa (SEQ ID NO: 126). In some embodiments, the linker and 3’ motif comprise or consist of the nucleotide sequence aaagtaagcgcggttctatctagttacgcgttaaaccaactagaa (SEQ ID NO: 202). In some embodiments, the linker and 3’ motif comprise or consist of the nucleotide sequence aagtaattcgcggttctatctagttacgcgttaaaccaactagaa (SEQ ID NO: 203). In some embodiments, the linker and 3’ motif comprise or consist of the nucleotide sequence aggtaaagcgcggttctatctagttacgcgttaaaccaactagaa (SEQ ID NO: 204). In some embodiments, the pegRNA comprises, from 5’ to 3’, a spacer sequence, a scaffold sequence, a template sequence, a primer binding site, and a linker + 3’ motif.
[0145] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of any one of SEQ ID NOs: 95, 96, and 195; a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97; a template + PBS sequence comprising or consisting of the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200; and a linker + 3’ motif comprising or consisting of the nucleotide sequence of any one of SEQ ID NOs: 126 and 202-204. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 98, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 126. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 95, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 99, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 126. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of theAttorney Docket No.:TYA-072WO nucleotide sequence of SEQ ID NO: 95, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 100, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 125. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 101, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 126. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 102, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 126. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 103, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 126. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising SEQ ID NO: 202. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 195. a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 200, and a linker + 3’ motif comprising SEQ ID NO: 202. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising SEQ ID NO: 203. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 195. a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequenceAttorney Docket No.:TYA-072WO comprising or consisting of the nucleotide sequence of SEQ ID NO: 200, and a linker + 3’ motif comprising SEQ ID NO: 203. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising SEQ ID NO: 204. In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 195, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 200, and a linker + 3’ motif comprising SEQ ID NO: 204.
[0146] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96. a scaffold sequence, a template + PBS, and a linker + 3’ motif.
[0147] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96. a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS, and a linker + 3’ motif.
[0148] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif.
[0149] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96. a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif.
[0150] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96. a scaffold sequence, a template + PBS, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0151] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising orAttorney Docket No.:TYA-072WO consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0152] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0153] In some embodiments, the pegRNA comprises: a spacer sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 96, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203. In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS. and a linker + 3’ motif.
[0154] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif.
[0155] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0156] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0157] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif.
[0158] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif.Attorney Docket No.:TYA-072WO
[0159] Tn some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0160] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0161] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence, a template + PBS, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0162] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0163] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0164] In some embodiments, the pegRNA comprises: a spacer sequence, a scaffold sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 97, a template + PBS sequence comprising or consisting of the nucleotide sequence of SEQ ID NO: 199, and a linker + 3’ motif comprising or consisting of the nucleotide sequence of SEQ ID NO: 203.
[0165] In some embodiments, the pegRNA comprises, from 5’ to 3’, a spacer sequence, a scaffold sequence, and a template + primer binding site. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of any one of SEQ ID NOs: 104-109. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 104. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 105. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 106. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 107. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 108. In some embodiments, the pegRNA comprises or consists of theAttorney Docket No.:TYA-072WO nucleotide sequence of SEQ ID NO: 109. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 187. In some embodiments, the pegRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 188. Exemplary pegRNA sequences are provided in Table 3.Table 3: Exemplary pegRNA sequencesReference RNA Sequence SEQ ID NO: peg-1 ctcacagatatggcccagaagttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 104aaaagtggcaccgagtcggtgctcaccggactacgagaccgcggtctttctgggccatatcactattcacgc ggttctatctagttacgcgttaaaccaactagaapeg-28 ctcacagatatggcccagaagttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 105aaaagtggcaccgagtcggtgcacgagaccgcggtctttctgggccatatctgataatacacgcggttctatc tagttacgcgttaaaccaactagaapeg-32 ctcacagatatggcccagaagttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 106aaaagtggcaccgagtcggtgcagaccgcggtctttctgggccatatcataaacaccgcggttctatctagtta cgcgttaaaccaactagaapeg- 186 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaa 107aaagtggcaccgagtcggtgcaggccgcggtctcgtagtccggtcagccggtcacaaagtaagcgcggttc tatctagttacgcgttaaaccaactagaapeg- 189 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaa 108 aaagtggcaccgagtcggtgcaggccgaggtctcgtagtccggtcagccggtcactcaatattaccgcggtt ctatctagttacgcgttaaaccaactagaapeg-224 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaa 109aaagtggcaccgagtcggtgcaggccgcggtctcgaagtccggtcagccggtcacaagaaattcgcggttc tatctagttacgcgttaaaccaactagaapeg-262 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaa 187aaagtggcaccgagtcggtgcaggccgcggtctcgtagtccggtgagccggtcacaagtaattcgcggttct atctagttacgcgttaaaccaactagaapeg-263 ggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaa 188 aagtggcaccgagtcggtgcggccgcggtctcgtagtccggtgagccggtcacaggtaaagcgcggttctatctagttacgcgttaaaccaactagaa
[0166] In some embodiments, a second guide RNA (nicking guide RNA; ngRNA) is used to nick the non-target strand either upstream or downstream of the original target to render the excision repair in favor of 3’ flap incorporation (Anzalone et al., 2019; Yang et al., 2019, each incorporated by reference herein). Without wishing to be bound by theory, the ngRNA-mediated nicking of the non-target strand is thought to bias DNA repair mechanisms towards replacing the non-target strand with the edit encoded by the template sequence.
[0167] In some embodiments, the ngRNA comprises a ngRNA spacer sequence. The ngRNA spacer sequence comprises a nucleic acid sequence that binds a ngRNA target sequence. In some embodiments, the ngRNA spacer sequence binds the ngRNA target sequence on the non-target strandAttorney Docket No.:TYA-072WO of a dsDNA molecule. The ngRNA spacer sequence may bind a ngRNA target sequence that is adjacent to a PAM that is recognizable by an effector protein of interest.
[0168] In some embodiments, the ngRNA spacer sequence comprises or consists of the nucleotide sequence of any one of SEQ ID NOs: 110-117. In some embodiments, the ngRNA spacer sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 113. In some embodiments, the ngRNA spacer sequence comprises or consists of the nucleotide sequence of SEQ ID NO: 194. Illustrative examples of ngRNA spacer sequences that bind a non-target sequence of the human RBM20 gene are shown in Table 4 below.Table 4. Exemplary ngRNA spacer sequencesReference ngRNA spacer sequence SEQ ID NO:ng-2 ctacgagaccgcggtctttc 110ng-20 tgggacctcggggagagtgac 111ng-6 gggacctcggggagagtgac 112ng-26 agattctaaatcctgctcct 113ng-28 ggtctcgtagtccggtcagc 114ng-31 gattctaaatcctgctcct 115ng-35 ctgtgtgtctgtgtgtggg 116ng-30 tcagccggtcactctccccg 117ng-36 attctaaatcctgctcct 194
[0169] In some embodiments, the ngRNA further comprises a scaffold sequence. In some embodiments, the ngRNA comprises or consists of the nucleotide sequence of any one of SEQ ID NOs: 118-125 and 193. In some embodiments, the ngRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 121. In some embodiments, the ngRNA comprises or consists of the nucleotide sequence of SEQ ID NO: 193. Illustrative examples of ngRNA sequences that bind a nontarget sequence of the human RBM20 gene are show in Table 5 below.Table 5. Exemplary ngRNA sequencesReference RNA Sequence SEQ ID NO: ng-2 ctacgagaccgcggtctttcgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttg 118aaaaagtggcaccgagtcggtgcng-20 tgggacctcggggagagtgacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaact 119tgaaaaagtggcaccgagtcggtgcAttorney Docket No.:TYA-072WO Reference RNA Sequence SEQ ID NO: ng-6 gggacctcggggagagtgacgttttagagctagaaatagcaagttaaaataaggctagtccgtatcaactt 120gaaaaagtggcaccgagtcggtgcng-26 agattctaaatcctgctcctgtttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 121aaaagtggcaccgagtcggtgcng-28 ggtctcgtagtccggtcagcgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttg 122aaaaagtggcaccgagtcggtgcng-31 gatctaaatcctgctcctgtttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaa 123aaagtggcaccgagtcggtgcng-35 ctgtgtgtctgtgtgtggggttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaa 124aaagtggcaccgagtcggtgcng-30 tcagccggtcactctccccggttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttg 125aaaaagtggcaccgagtcggtgcng-36 attctaaatcctgctcctgttttagagctagaaatagcaagttaaaataaggctagtccgtatcaactgaaa 193aagtggcaccgagtcggtgcEffector Proteins
[0170] In some embodiments, the gene editing systems described herein comprise an effector protein. In some embodiments, the effector protein is a nucleic acid-guided nuclease. In some embodiments, the nucleic acid-guided nuclease is an RNA-guided nuclease.
[0171] In some embodiments, the nucleic acid-guided nuclease is a Cas nuclease. In some embodiments, the Cas nuclease is any Cas nuclease (e.g., any Cas9 nuclease) which is encoded by a gene equal to or less than 3.3 kb in size, equal to or less than 3.2 kb in size, equal to or less than 3.1 kb in size, equal to or less than 3 kb in size, equal to or less than 2.9 kb in size, or equal to or less than 2.8 kb in size. In some embodiments, the Cas nuclease is any Cas nuclease (e.g., any Cas9 nuclease) which has the protein size of equal to or less than 1,100 amino acids, equal to or less than 1,075 amino acids, equal to or less than 1,060 amino acids, equal to or less than 1,050 amino acids, equal to or less than 1,000 amino acids, equal to or less than 950 amino acids, or equal to or less than 900 amino acids.
[0172] In some embodiments, the Cas nuclease is a Cas9 nuclease. In some embodiments, the Cas9 is derived from. S', aureus, S. pyogenes, F. novicida, N. meningitidis, S. thermophilus, Acidaminococcus sp., G. stearothermophilus, N. mucosa, S. canis, Lachnospiraceae bacterium., S. sanguinis, N. subflava, S. epidermidis, S. agalactiae, L. monocytogenes, Actinomyces sp., S. anginosus, S. dysgalactiae, C. jejuni, B. thuringiensis, C. difficile, S. cristatus, S. mutans, or S. pneumonia. In some embodiments, the Cas9 endonuclease is a S. aureus Cas9 (SaCas9) or a variant thereof. In some embodiments, the Cas9 endonuclease is a S. pyogenes Cas9 (SpCas9) or a variantAttorney Docket No.:TYA-072WO thereof. In some embodiments, the Cas9 endonuclease is nSPCas9max.
[0173] In some embodiments, the Cas nuclease is a Casl2a nuclease. In some embodiments, the Casl2a is derived from Acidaminococcus sp., M. bovoculi, A. cellulolyticus, Prevotella sp., S. mutans, Acidovorax sp., A. acidocaldarius, P. aeruginosa, L. crispatus, M. osloensis, or K. oxytoca.
[0174] In some embodiments, the Cas nuclease is a Cas 12b nuclease. In some embodiments, the Cas 12b is derived from Lachnospiraceae bacterium, Ruminococcus sp., P. gingivalis, P. intermedia, Enterobacter sp., S. pneumoniae, Acidobacterium sp., B. fragilis, S. marcescens, B. uniformis, or B. thuringiensis.
[0175] In some embodiments, the Cas nuclease is a CasX nuclease. In some embodiments, the CasX is derived from Ruminococcus sp., Pseudomonas sp., or S. epidermidis.
[0176] In some embodiments, the Cas nuclease may comprise one or more mutations to alter their activity, specificity, recognition, and / or other characteristics. For example, the Cas nuclease may have one or more mutations that alter its fidelity to mitigate off-target effects (e.g., eSpCas9, SpCas9-HFl, HypaSpCas9, HeFSpCas9, and evoSpCas9 high-fidelity variants of SpCas9). For another example, the Cas nuclease may have one or more mutations that alter its PAM specificity.
[0177] In some embodiments, a coding sequence encoding a Cas nuclease can be codon optimized for expression in particular cells, such as eukaryotic cells. The eukaryotic cells may be those of or derived from a particular organism, such as a mammal, including but not limited to human, mouse, rat, rabbit, dog, or non-human primate. In general, codon optimization refers to a process of modifying a nucleic acid sequence for enhanced expression in the host cells of interest by replacing at least one codon of the native sequence with codons that are more frequently or most frequently used in the genes of that host cell while maintaining the native amino acid sequence. Various species exhibit particular bias for certain codons of a particular amino acid. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which is in turn believed to be dependent on, among other things, the properties of the codons being translated and the availability of particular transfer RNA (tRNA) molecules. The predominance of selected tRNAs in a cell is generally a reflection of the codons used most frequently in peptide synthesis. Accordingly, genes can be tailored for optimal gene expression in a given organism based on codon optimization.Attorney Docket No.:TYA-072WO
[0178] In some embodiments, the polynucleotide encoding a Cas9 nuclease comprises or consists of a sequence that shares at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%. at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 127. In some embodiments, the Cas9 nuclease used as described herein or encoded by the polynucleotides described herein comprises or consists of an amino acid sequence that shares at least 70%. at least 75%, at least 80%. at least 85%, at least 90%. at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 128. Illustrative examples of Cas nuclease sequences are shown in Table 6 below.Table 6. Example Cas nuclease sequencesName Sequence SEQ ID NO:SaCas9 atggccccaaagaagaagcggaaggtcggtatccacggagtcccagcagccaagcggaactacat 127 nucleotide cctgggcctggacatcggcatcaccagcgtgggctacggcatcatcgactacgagacacgggacgt sequence gatcgatgccggcgtgcggctgttcaaagaggccaacgtggaaaacaacgagggcaggcggagca agagaggcgccagaaggctgaagcggcggaggcggcatagaatccagagagtgaagaagctgct gttcgactacaacctgctgaccgaccacagcgagctgagcggcatcaacccctacgaggccagagt gaagggcctgagccagaagctgagcgaggaagagttctctgccgccctgctgcacctggccaagag aagaggcgtgcacaacgtgaacgaggtggaagaggacaccggcaacgagctgtccaccaaagag cagatcagccggaacagcaaggccctggaagagaaatacgtggccgaactgcagctggaacggct gaagaaagacggcgaagtgcggggcagcatcaacagattcaagaccagcgactacgtgaaagaag ccaaacagctgctgaaggtgcagaaggcctaccaccagctggaccagagcttcatcgacacctacat cgacctgctggaaacccggcggacctactatgagggacctggcgagggcagccccttcggctggaa ggacatcaaagaatggtacgagatgctgatgggccactgcacctacttccccgaggaactgcggagc gtgaagtacgcctacaacgccgacctgtacaacgccctgaacgacctgaacaatctcgtgatcacca gggacgagaacgagaagctggaatattacgagaagttccagatcatcgagaacgtgttcaagcagaa gaagaagcccaccctgaagcagatcgccaaagaaatcctcgtgaacgaagaggatattaagggcta cagagtgaccagcaccggcaagcccgagttcaccaacctgaaggtgtaccacgacatcaaggacat taccgcccggaaagagattattgagaacgccgagctgctggatcagattgccaagatcctgaccatct accagagcagcgaggacatccaggaagaactgaccaatctgaactccgagctgacccaggaagag atcgagcagatctctaatctgaagggctataccggcacccacaacctgagcctgaaggccatcaacct gatcctggacgagctgtggcacaccaacgacaaccagatcgctatcttcaaccggctgaagctggtg cccaagaaggtggacctgtcccagcagaaagagatccccaccaccctggtggacgacttcatcctga gccccgtcgtgaagagaagcttcatccagagcatcaaagtgatcaacgccatcatcaagaagtacgg cctgcccaacgacatcattatcgagctggcccgcgagaagaactccaaggacgcccagaaaatgatc aacgagatgcagaagcggaaccggcagaccaacgagcggatcgaggaaatcatccggaccaccg gcaaagagaacgccaagtacctgatcgagaagatcaagctgcacgacatgcaggaaggcaagtgc ctgtacagcctggaagccatccctctggaagatctgctgaacaaccccttcaactatgaggtggacca catcatccccagaagcgtgtccttcgacaacagcttcaacaacaaggtgctcgtgaagcaggaagaa aacagcaagaagggcaaccggaccccattccagtacctgagcagcagcgacagcaagatcagctacgaaaccttcaagaagcacatcctgaatctggccaagggcaagggcagaatcagcaagaccaagaaagagtatctgctggaagaacgggacatcaacaggttctccgtgcagaaagacttcatcaaccggaacAttorney Docket No.:TYA-072WO Name Sequence SEQ ID NO: ctggtggataccagatacgccaccagaggcctgatgaacctgctgcggagctacttcagagtgaaca acctggacgtgaaagtgaagtccatcaatggcggcttcaccagctttctgcggcggaagtggaagttt aagaaagagcggaacaaggggtacaagcaccacgccgaggacgccctgatcattgccaacgccga tttcatcttcaaagagtggaagaaactggacaaggccaaaaaagtgatggaaaaccagatgttcgagg aaaagcaggccgagagcatgcccgagatcgaaaccgagcaggagtacaaagagatcttcatcaccc cccaccagatcaagcacattaaggacttcaaggactacaagtacagccaccgggtggacaagaagc ctaatagagagctgattaacgacaccctgtactccacccggaaggacgacaagggcaacaccctgat cgtgaacaatctgaacggcctgtacgacaaggacaatgacaagctgaaaaagctgatcaacaagag ccccgaaaagctgctgatgtaccaccacgacccccagacctaccagaaactgaagctgattatggaa cagtacggcgacgagaagaatcccctgtacaagtactacgaggaaaccgggaactacctgaccaag tactccaaaaaggacaacggccccgtgatcaagaagattaagtattacggcaacaaactgaacgccc atctggacatcaccgacgactaccccaacagcagaaacaaggtcgtgaagctgtccctgaagcccta cagattcgacgtgtacctggacaatggcgtgtacaagttcgtgaccgtgaagaatctggatgtgatcaa aaaagaaaactactacgaagtgaatagcaagtgctatgaggaagctaagaagctgaagaagatcagc aaccaggccgagtttatcgcctccttctacaacaacgatctgatcaagatcaacggcgagctgtataga gtgatcggcgtgaacaacgacctgctgaaccggatcgaagtgaacatgatcgacatcacctaccgcg agtacctggaaaacatgaacgacaagaggccccccaggatcattaagacaatcgcctccaagaccc agagcattaagaagtacagcacagacattctgggcaacctgtatgaagtgaaatctaagaagcaccct cagatcatcaaaaagggcaaaaggccggcggccacgaaaaaggccggccaggcaaaaaagaaaa agtaaSaCas9 MAPKKKRKVGIHGVPAAKRNYILGLDIGITSVGYGIIDYETRDVI 128 amino acid DAGVRLFKEANVENNEGRRSKRGARRLKRRRRHRIQRVKKLLF sequence DYNLLTDHSELSGINPYEARVKGLSQKLSEEEFSAALLHLAKRR GVHNVNEVEEDTGNELSTKEQISRNSKALEEKYVAELQLERLKK DGEVRGSINRFKTSDYVKEAKQLLKVQKAYHQLDQSFIDTYIDL LETRRTYYEGPGEGSPFGWKDIKEWYEMLMGHCTYFPEELRSV KYAYNADLYNALNDLNNLVITRDENEKLEYYEKFQIIENVFKQK KKPTLKQIAKEILVNEEDIKGYRVTSTGKPEFTNLKVYHDIKDIT ARKEIIENAELLDQIAKILTIYQSSEDIQEELTNLNSELTQEEIEQIS NLKGYTGTHNLSLKAINLILDELWHTNDNQIAIFNRLKLVPKKV DLSQQKEIPTTLVDDFILSPVVKRSFIQSIKVINAIIKKYGLPNDIIIE LAREKNSKDAQKMINEMQKRNRQTNERIEEIIRTTGKENAKYLI EKIKLHDMQEGKCLYSLEAIPLEDLLNNPFNYEVDHIIPRSVSFD NS FNNKVL VKQEENS KKGNRTPFQ YLS S S D S KIS YETFKKHILNL AKGKGRISKTKKEYLLEERDINRFSVQKDFINRNLVDTRYATRG LMNLLRSYFRVNNLDVKVKSINGGFTSFLRRKWKFKKERNKGY KHHAEDALIIANADFIFKEWKKLDKAKKVMENQMFEEKQAESM PEIETEQEYKEIFITPHQIKHIKDFKDYKYSHRVDKKPNRELINDT LYSTRKDDKGNTLIVNNLNGLYDKDNDKLKKLINKSPEKLLMY HHDPQT YQKLKLIMEQ YGDEKNPLY KY YEETGN YLTK Y S KKD NGPVIKKIKYYGNKLNAHLDITDDYPNSRNKVVKLSLKPYRFDV YLDNGVYKFVTVKNLDVIKKENYYEVNSKCYEEAKKLKKISNQ AEFIASFYNNDLIKINGELYRVIGVNNDLLNRIEVNMIDITYREYLENMNDKRPPRIIKTIASKTQSIKKYSTDILGNLYEVKSKKHPQIIKAttorney Docket No.:TYA-072WO Name Sequence SEQ ID NO:KGKRPAATKKAGQAKKKK
[0179] Non-limiting examples of nucleic acid-guided nucleases include Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csnl and Csxl2), CaslO, CasX, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, homologs thereof, or modified versions thereof. These enzymes are known, for example, the amino acid sequence of S. pyogenes Cas9 protein may be found in the SwissProt database under accession number Q99ZW2. In some embodiments, the nucleic acid-guided nuclease is a S. aureus Cas9 (SaCas9) or a variant thereof. In some embodiments, the nucleic acid-guided nuclease is a 5. pyogenes Cas9 (SpCas9) or a variant thereof.
[0180] For the Cas nuclease to function, there must be a PAM immediately downstream of (3’ to) or upstream of (5’ to) the target sequence in the genomic DNA. Recognition of the PAM by the Cas nuclease is thought to destabilize the adjacent genomic sequence, allowing interrogation of the sequence by the gRNA and resulting in gRNA-DNA pairing when a matching sequence is present. The specific sequence of PAM varies depending on the species of the Cas gene. For example, the SpCas recognizes a PAM sequence of 5’-NGG-3’ or, at less efficient rates, 5’-NAG-3’, where N can be any nucleotide. For another example, the SaCas9 PAM sequence is NNGRR or for optimal on-target cutting is NNGRRT, wherein N can be any nucleotide, and R can be guanine or adenine. Other Cas nuclease variants with alternative PAMs have also been characterized and successfully used for genome editing. Therefore, the PAM sequence adjacent to the target sequence is an essential targeting component for the design of CRISPR / Cas9-mediated gene editing, and when designing gRNAs targeting a specific genomic locus, one skilled in the art needs to consider the availability and / or location of PAM sequences optimal for the Cas nuclease of choice at the target locus and, if necessary, select a different Cas nuclease / PAM sequence pair based on the DNA sequence at the target locus. PAMs for use with different Cas endonucleases are known in the art. Illustrative examples of Cas enzymes that can be used as described herein and PAMs for use with their respective Cas endonucleases are shown in Table 7 below (where M is adenine or cytosine; N is any nucleotide; R is guanine or adenine; V is guanine, cytosine, or adenine; W is adenine or thymine; and Y is cytosineAttorney Docket No.:TYA-072WO or thymine).Table 7. Exemplary Cas proteins and their corresponding PAM sequencesName Species PAM Sequence Size (amino acids) SpCas9 Streptococcus pyogenes NGG 1,368SaCas9 Staphylococcus aureus NNGRRT 1,053FnCas9 Francisella novicida NGG 1,244NmCas9 Neisseria meningitidis NNNNGATT 1,055StCas9 Streptococcus, thermophilus NGG 1,364AaCas9 Acidaminococcus sp NNNNRYAC 984GstCas9 Geobacillus NGG 1,245stearothermophilusNmuCas9 Neisseria mucosa NNAGAAW 1,054ScCas9 Streptococcus canis NGG 1,342LbCas9 Lachnospiraceae bacterium TTTA 1,305SsCas9 Streptococcus sanguinis NNGRRT 1,365NsuCas9 Neisseria subflava NNNNGMTT 1,056SeCas9 Staphylococcus epidermidis NNGS 1,054SagalCas9 Streptococcus agalactiae NNGRRT 1,363LmCas9 Listeria monocytogenes NGG 1,359AspCas9 Actinomyces sp. NNAGAAW 1,045SangCas9 Streptococcus anginosus NNNNGATT 1,056SdCas9 Streptococcus dysgalactiae NNGRRT 1,341CjCas9 Campylobacter jejuni NNNNRYAC 1,235BtCas9 Bacillus thuringiensis NGG 1,243CdCas9 Clostridium difficile NGG 1,406ScristCas9 Streptococcus cristatus NNNNGATT 1,056SmCas9 Streptococcus mutans NNGG 1,333AsCasl2a Acidaminococcus sp. TTTV 1,278LbCasl2b Lachnospiraceae bacterium TYCV 1,187MboCasl2a Moraxella bovoculi TTTV 1,286RspCasl2b Ruminococcus sp TYCV 1,185AccCasl2a Acidothermus cellulolyticus TTTV 1,285PspCasl2a Prevotella sp. TTTV 1,264Attorney Docket No.:TYA-072WO Name Species PAM Sequence Size (amino acids) PgiCasl2b Porphyromonas gingivalis TYCV 1,167SmaCasl2a Streptococcus inutans TTTV 1,259PinCasl2b Prevotella intermedia TYCV 1,188AspCasl2a Acidovorax sp. TTTV 1,284EspCasl2b Enterobacter sp. TYCV 1,296SpyCasl2b Streptococcus pneumoniae TYCV 1,157AacCasl2a Alicyclobacillus TTTV 1,285acidocaldariusAcbCasl2b Acidobacterium sp. TYCV 1,174PaeCasl2a Pseudomonas aeruginosa TTTV 1,284BfrCasl2b Bacteroides fragilis TYCV 1,181LcrCasl2a Lactobacillus crispatus TTTV 1,259SmaCasl2b Serratia marcescens TYCV 1,188MolCasl2a Moraxella osloensis TTTV 1,284BunCasl2b Bacteroides uniformis TYCV 1,197KoxCasl2a Klebsiella oxytoca TTTV 1,292BthCasl2b Bacillus thuringiensis TYCV 1,182RspCasX Ruminococcus sp. TTN 926PspCasX Pseudomonas sp. TTN 918SepCasX Staphylococcus epidermidis TTN 927
[0181] In some embodiments, the nucleic acid-guided nuclease is an RNA-guided nickase. Any suitable RNA-guided nickase may be used in the systems described herein. In various embodiments, the RNA-guided nickase may be any Class 2 CRISPR-Cas system having a nickase activity (e.g., only cleaves one strand of the target DNA sequence), including any type II, type V, or type VI CRISPR-Cas enzyme. In some embodiments, the RNA-guided nickase is an SaCas9 nickase. In some embodiments, the RNA-guided nickase is an SpCas9 nickase.
[0182] The RNA-guided nickase may also comprise nickase variants of Cas9 equivalents, including Casl2a (Cpfl), Casl2e (CasX), Casl2bl (C2cl), Casl2b2, Casl2c (C2c3), C2c4, C2c8, C2c5. C2cl0, C2c9 Casl3a (C2c2), Casl3d, Casl3c (C2c7), Casl3b (C2c6). and Casl3b. Further Cas-equivalents are described in Makarova et al., Science 2016; 353(6299) and Makarova et al., The CRISPR Journal, Vol.l. No.5, 2018, the contents of which are incorporated herein by reference.Attorney Docket No.:TYA-072WO For example, an aspartate-to-alanine substitution (D 10 A) in the RuvC I catalytic domain of Cas9 from S. pyogenes converts Cas9 from a nuclease that cleaves both strands to a nickase (cleaves a single strand). Other examples of mutations that render Cas9 a nickase include, without limitation, H840A, N854A, and N863A in reference to the canonical SpCas9 sequence, or to equivalent amino acid positions in other Cas9 variants or Cas9 equivalents.
[0183] In some embodiments, the system comprises a fusion protein comprising an RNA-guided nickase and a reverse transcriptase.
[0184] Reverse transcriptases are multi-functional enzymes typically with three enzymatic activities including RNA- and DNA-dependent DNA polymerization activity, and an RNaseH activity that catalyzes the cleavage of RNA in RNA-DNA hybrids. Some mutants of reverse transcriptases have disabled the RNaseH moiety. These enzymes synthesize complementary DNA (cDNA) using RNA as a template. More recently, mutants and fusion proteins have been created in the quest for improved properties such as thermostability, fidelity and activity. Any of the wild type, variant, and / or mutant forms of reverse transcriptase which are known in the art or which can be made using methods known in the art are contemplated herein.
[0185] Non-limiting examples of reverse transcriptases include Moloney Murine Leukemia Virus (M-MLV); Human Immunodeficiency Virus (HIV) reverse transcriptase and avian Sarcoma-Leukosis Virus (ASLV) reverse transcriptase, which includes but is not limited to Rous Sarcoma Virus (RSV) reverse transcriptase, Avian Myeloblastosis Virus (AMV) reverse transcriptase. Avian Erythroblastosis Virus (AEV) Helper Virus MCAV reverse transcriptase, Avian Myelocytomatosis Virus MC29 Helper Virus MCAV reverse transcriptase, Avian Reticuloendotheliosis Virus (REV-T) Helper Virus REV-A reverse transcriptase, Avian Sarcoma Virus UR2 Helper Virus UR2AV reverse transcriptase, Avian Sarcoma Virus Y73 Helper Virus YAV reverse transcriptase, Rous Associated Virus (RAV) reverse transcriptase, and Myeloblastosis Associated Virus (MAV) reverse transcriptase.
[0186] In some embodiments, the reverse transcriptase comprises one or more mutations or deletions in the RNase H domain. As mentioned above, one of the intrinsic properties of reverse transcriptases is the RNase H activity, which cleaves the RNA template of the RNA:cDNA hybrid concurrently with polymerization. The RNase H activity can be unnecessary for certain applications. The RNase H activity may also lower reverse transcription efficiency, presumably due to itsAttorney Docket No.:TYA-072WO competition with the polymerase activity of the enzyme. Thus, the present disclosure contemplates any reverse transcriptase variants that comprise a modified RNaseH activity or lack the RNase H domain.
[0187] In some embodiments, the reverse transcriptase is an M-MLV reverse transcriptase lacking the RNase H domain. In some embodiments the reverse transcriptase is an evolved Tfl reverse transcriptase. In some embodiments the reverse transcriptase is an evolved and engineered Tfl reverse transcriptase. In some embodiments the reverse transcriptase is an evolved and engineered M-MLV reverse transcriptase lacking the RNase H domain.
[0188] In some embodiments, the present disclosure provides a fusion protein comprising an RNA-guided nickase and a reverse transcriptase described herein. In some embodiments, the fusion protein comprises, from the N-terminus to the C-terminus, an RNA-guided nickase and a reverse transcriptase. In some embodiments, the fusion protein further comprises a linker.
[0189] In some embodiments, the fusion protein described herein may be divided into two or more fragments which become assembled inside the cell into the mature fusion protein. In some embodiments, the fusion protein is divided into an N-terminal fragment and a C-terminal fragment. In some embodiments, the fusion protein N-terminal fragment comprises an N-terminal fragment of the RNA-guided nickase disclosed herein and a split intein domain. In some embodiments, the fusion protein C-terminal fragment comprises a C-terminal fragment of the RNA-guided nickase, a reverse transcriptase, and a split intein domain.
[0190] In some embodiments, the N-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 129, and the C-terminal fragment of the fusion protein comprises an amino acid sequence selected from the group consisting of the sequences of SEQ ID NOs: 130-138, 182, and 183. In some embodiments, the N-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 129, and the C-terminal fragment of the fusion protein comprises an amino acid sequence selected from the group consisting of sequences of SEQ ID NOs: 130-132, 134, 136, and 138.
[0191] Once delivered or expressed within a cell, the split intein domains of the different fragments associate and bind to one another, and then undergo trans- splicing, which results in the excision of the split-intein domains from each of the fragments, and a concomitant formation of aAttorney Docket No.:TYA-072WO peptide bond between the fragments, thereby resulting in the mature version of the fusion protein comprising the RNA-guided nickase and the reverse transcriptase (the mature fusion protein is also referred to herein as a “prime editor”). Non-limiting examples of split inteins are described in Stevens etal., PNAS, 2017, Vol.114: 8538-8543; Iwai et al.,, FEBS Lett, 580: 1853-1858, each of which are incorporated herein by reference. Additional split intein sequences can be found, for example, in WO 2013 / 045632, WO 2014 / 055782, WO 2016 / 069774, and EP2877490. the contents each of which are incorporated herein by reference.Expression Cassettes and Vectors
[0192] In some embodiments, the present disclosure provides one or more expression cassettes encoding a gene editing system described herein. In some embodiments, the present disclosure provides a first expression cassette and a second expression cassette. In some embodiments, the first expression cassette comprises a first polynucleotide encoding an N-terminal fragment of an RNA-guided nickase described herein and an N-terminal fragment of a split-intein. In some embodiments, the first expression cassette further comprises a second polynucleotide encoding a first guide RNA. In some embodiments, the second expression cassette comprises a third polynucleotide encoding a C-terminal fragment of an RNA-guided nickase described herein and a C-terminal fragment of a split-intein. In some embodiments, the second expression cassette further comprises a fourth polynucleotide encoding a second guide RNA. In some embodiments, the first guide RNA is a ngRNA described herein and the second guide RNA is a pegRNA described herein. In some embodiments, the first guide RNA is a pegRNA described herein and the second guide RNA is a ngRNA described herein.
[0193] In some embodiments, the present disclosure provides a first expression cassette comprising a first promoter described herein operatively linked to a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase described herein and an N-terminal fragment of a split-intein. In some embodiments, the first expression cassette further comprises a second promoter operatively linked to a second polynucleotide encoding a nicking guide RNA (ngRNA) described herein. In some embodiments, the first expression cassette comprises (a) a first promoter described herein operatively linked to a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase described herein and an N-terminal fragment of a split-intein;Attorney Docket No.:TYA-072WO and (b) a second promoter operatively linked to a second polynucleotide encoding a ngRNA described herein. In some embodiments, the first expression cassette comprises a first truncated TNNT2 promoter operatively linked to a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein.
[0194] In some embodiments, the present disclosure further provides a second expression cassette comprising a third promoter described herein operatively linked to a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase described herein, a polymerase described herein, and a C-terminal fragment of a split-intein. In some embodiments, the second expression cassette further comprises a fourth promoter described herein operatively linked to a fourth polynucleotide encoding a prime editing guide RNA (pegRNA). In some embodiments, the second expression cassette comprises (a) a second truncated TNNT2 promoter operatively linked to a third polynucleotide encoding a C-terminal fragment of the fusion protein comprising a C-terminal fragment of the RNA-guided nickase, a reverse transcriptase, and a C-terminal fragment of the split-intein; and (b) a fourth promoter operatively linked to a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) described herein.
[0195] In some embodiments, the C-terminal fragment of the fusion protein further comprises an RNA binding domain of small RNA binding exonuclease protection factor La (e.g., as described in Yan et al. Nature. 2024; 628(8008):639-647.) In some embodiments, the small RNA binding exonuclease protection factor La is according to Gene ID: 6741 and / or UniProt ID P05455. In some embodiments, the effector proteins are arranged as Cas9 — La — reverse transcriptase.
[0196] In some embodiments, the first expression cassette comprises a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity to the sequence of any one of SEQ ID NOs: 143, 145, 146, 149, and 151. In some embodiments, the first expression cassette comprises the sequence of any one of SEQ ID NOs: 143, 145, 146, 149, and 151. In some embodiments, the first expression cassette further comprises a promoter described herein controlling expression of a ngRNA. In some embodiments, the first expression cassette is flanked by a left ITR on the 5’ end of the expression cassette and a right ITR on the 3’ end of the expression cassette. In some embodiments, the left ITR comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 141, or aAttorney Docket No.:TYA-072WO nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 141. In some embodiments, the right ITR comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 142, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 142.
[0197] In some embodiments, the second expression cassette comprises a sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 92%, at least 94%, at least 95%, at least 96%, at least 97%. at least 98%, or at least 99% identity to the sequence of any one of SEQ ID NOs: 144, 146, 148, 150, 152-160, 184, and 185. In some embodiments, the second expression cassette comprises the sequence of any one of SEQ ID NOs: 144, 146, 148, 150, 152-160, 184, and 185. In some embodiments, the second expression cassette further comprises a promoter described herein controlling expression of a pegRNA. In some embodiments, the second expression cassette is flanked by a left ITR on the 5’ end of the expression cassette and a right ITR on the 3’ end of the expression cassette. In some embodiments, the left ITR comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 141, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 141. In some embodiments, the right ITR comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 142, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 142.
[0198] In some embodiments, the first expression cassette comprises a nucleotide sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 147, and the second expression cassette comprises a nucleotide sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 148. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 147, and the second expression cassette comprises the sequence of SEQ ID NO: 148.
[0199] In some embodiments, the first expression cassette comprises a nucleotide sequence that is at least 70%. at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identicalAttorney Docket No.:TYA-072WO to the sequence of SEQ ID NO: 145, and the second expression cassette comprises a nucleotide sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence of any one of SEQ ID NOs: 146, 153, 154, 156, 158, and 160. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 147, and the second expression cassette comprises the sequence of any one of SEQ ID NOs: 146, 153, 154. 156, 158, and 160.
[0200] In some embodiments, the first expression cassette comprises the sequence “GGCCACC” between the promoter and the start codon. In some embodiments, the second expression cassette comprises the sequence “GGCCACC” between the promoter and the start codon.
[0201] In some embodiments, a junction sequence exists between two transcript units of an expression cassette. In some embodiments, the junction sequence comprises the sequence “CTCGAG”.
[0202] In some embodiments, the first and second expression cassettes are selected from those shown in Table 8. In the table, “Orientation” refers whether the polynucleotide encoding the N-terminal or C-terminal fragment of the fusion protein face the same direction (tail-to-head) or opposite direction (tail-to-tail) as the polynucleotide encoding the gRNA.
[0203] In some embodiments, the N-terminal PE protein is encoded by the nucleotide sequence of SEQ ID NO: 205. In some embodiments, the N-terminal PE protein is encoded by the nucleotide sequence of SEQ ID NO: 206.Table 8. Exemplary expression cassettesCassette Promoter sequence Split-PE protein 3'UTR+polyA Orientation Guide Full sequence sequence signal sequence RNA excluding ITRs (SEQ ID NO:) V4-PE-N pTNNT2 (3O4bp) N-terminal PE SEQ pZC659 SEQ ID Tail-to-tail ngRNA143 SEQ ID NO: 172 ID NO: 129 NO: 161V4-PE-C pTNNT2 (304bp) C-terminal PE SEQ pZC659 SEQ ID Tail-to-tail pegRNA144 SEQ ID NO: 172 ID NO: 130 NO: 161V13-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNASEQ ID NO: 170 ID NO: 129 NO: 165 145 V13-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA 146SEQ ID NO: 170 ID NO: 130 NO: 165V15-PE-N pTNNT2 (304bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA 147SEQ ID NO: 172 ID NO: 129 NO: 165V15-PE-C pTNNT2 (304bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNASEQ ID NO: 172 ID NO: 130 NO: 165 148 V16-PE-N pTNNT2 (304bp) N-terminal PE SEQ pZC692 SEQ ID Tail-to-tail ngRNA149SEQ ID NO: 172 ID NO: 129 NO: 166Attorney Docket No.:TYA-072WO Cassette Promoter sequence Split-PE protein 3'UTR+polyA Orientation Guide Full sequence sequence signal sequence RNA excluding ITRs (SEQ ID NO:) V16-PE-C pTNNT2 (304bp) C-terminal PE SEQ pZC692 SEQ ID Tail-to-tail pegRNA150 SEQ ID NO: 172 ID NO: 130 NO: 166V18-PE-N pTNNT2 (304bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-head ngRNA151 SEQ ID NO: 172 ID NO: 129 NO: 165V18-PE-C pTNNT2 (304bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-head pegRNA152 SEQ ID NO: 172 ID NO: 130 NO: 165V19-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V19-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA153 SEQ ID NO: 170 ID NO: 131 NO: 165V20-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V20-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA154 SEQ ID NO: 170 ID NO: 132 NO: 165V21-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V21-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA155 SEQ ID NO: 170 ID NO: 133 NO: 165V22-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V22-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA156 SEQ ID NO: 170 ID NO: 134 NO: 165V23-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V23-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA157 SEQ ID NO: 170 ID NO: 135 NO: 165V24-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V24-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA158 SEQ ID NO: 170 ID NO: 136 NO: 165V25-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V25-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNASEQ ID NO: 170 ID NO: 137 NO: 165 159 V26-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V26-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA160 SEQ ID NO: 170 ID NO: 138 NO: 165V27-PE-N pTNNT2 (400bp) N-terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V27-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA184 SEQ ID NO: 170 ID NO: 182 NO: 165V28-PE-N pTNNT2 (400bp) N- terminal PE SEQ pZC691 SEQ ID Tail-to-tail ngRNA145 SEQ ID NO: 170 ID NO: 129 NO: 165V28-PE-C pTNNT2 (400bp) C-terminal PE SEQ pZC691 SEQ ID Tail-to-tail pegRNA185SEQ ID NO: 170 ID NO: 183 NO: 165
[0204] In some embodiments, the present disclosure provides a gene editing system comprising: a) a first expression cassette comprising: (i) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein; and (ii) a second polynucleotide encoding a nicking guide RNA (ngRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 110-117; andAttorney Docket No.:TYA-072WO b) a second expression cassette comprising: (iii) a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein; and (iv) a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 98-103.
[0205] In some embodiments, the present disclosure provides a gene editing system comprising: a) a first expression cassette comprising: (i) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein; and (ii) a second polynucleotide encoding a nicking guide RNA (ngRNA) comprising the nucleotide sequence of SEQ ID NO: 113; and b) a second expression cassette comprising: (iii) a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein; and (iv) a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of SEQ ID NO: 103.
[0206] In some embodiments, the present disclosure provides a gene editing system comprising: a) a first expression cassette comprising: (i) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein; and (ii) a second polynucleotide encoding a nicking guide RNA (ngRNA) comprising the nucleotide sequence of SEQ ID NO: 193; and b) a second expression cassette comprising: (iii) a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein; and (iv) a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of SEQ ID NO: 187.
[0207] In some embodiments, the ngRNA comprises the sequence of any one of SEQ ID NOs: 118-125. In some embodiments, the ngRNA comprises the sequence of 121. In some embodiments, the pegRNA comprises the sequence of any one of SEQ ID NOs: 104-109, 187, and 188. In some embodiments, the pegRNA comprises the sequence of SEQ ID NO: 109. In some embodiments, the pegRNA comprises the nucleotide sequence of SEQ ID NO: 187.
[0208] In some embodiments, the first polynucleotide encodes a polypeptide comprising a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, atAttorney Docket No.:TYA-072WO least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 129. In some embodiments, the third polynucleotide encodes a polypeptide comprising a sequence having at least at least 70%, at least 75%. at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 132.
[0209] In some embodiments, the first expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 176, and the second expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 177 or SEQ ID NO: 178. In some embodiments, the first expression cassette comprises SEQ ID NO: 176. and the second expression cassette comprises the sequence of SEQ ID NO: 177. In some embodiments, the first expression cassette comprises SEQ ID NO: 176, and the second expression cassette comprises the sequence of SEQ ID NO: 178.
[0210] In some embodiments, the first expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 179, and the second expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 180 or SEQ ID NO: 181. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179, and the second expression cassette comprises the sequence of SEQ ID NO: 180. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179, and the second expression cassette comprises the sequence of SEQ ID NO: 181.
[0211] In some embodiments, the first expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%. at least 99%, or 100% identity to the sequence of SEQ ID NO: 176. and the second expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 189 or SEQ ID NO: 190. In some embodiments, the firstAttorney Docket No.:TYA-072WO expression cassette comprises the sequence of SEQ ID NO: 176, and the second expression cassette comprises the sequence of SEQ ID NO: 189. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 176. and the second expression cassette comprises the sequence of SEQ ID NO: 190.
[0212] In some embodiments, the first expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 179, and the second expression cassette comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 191 or SEQ ID NO: 192. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179, and the second expression cassette comprises the sequence of SEQ ID NO: 191. In some embodiments, the first expression cassette comprises the sequence of SEQ ID NO: 179, and the second expression cassette comprises the sequence of SEQ ID NO: 192.
[0213] The polynucleotides, expression cassettes, and / or vectors contemplated herein may be combined with other sequences, such as promoters, enhancers, untranslated regions (UTRs), introns, signal sequences, Kozak sequences, polyadenylation (poly(A)) signals, post-transcriptional regulatory elements, additional restriction enzyme sites, multiple cloning sites, internal ribosomal entry sites (IRES), recombinase recognition sites (e.g., LoxP, FRT, and Att sites), termination codons, transcriptional termination signals, polynucleotides encoding self-cleaving polypeptides, epitope tags, and / or any other regulatory elements as disclosed elsewhere herein or as known in the art. In some embodiments, the polynucleotides, expression cassettes, and / or vectors described herein may also contain a ribosome binding site for translation initiation, a transcription terminator, and / or polynucleotide sequences for amplifying expression. The expression cassette may be flanked by one or more inverted terminal repeats (ITRs). The ITRs in an expression cassette serve as markers used for viral packaging of the expression cassette. The expression cassette can be integrated into the host cell genome, thereby expressing the transgene within a host cell.
[0214] As used herein, the term “regulatory element” refers to those non-translated regions of the vector (e.g., origin of replication, selection cassettes, promoters, enhancers, translation initiation signals (Kozak sequence), introns, poly(A) sequences, 5' and 3' untranslated regions) which interactAttorney Docket No.:TYA-072WO with host cellular proteins to carry out transcription and translation. Such elements may vary in their strength and specificity. The transcriptional regulatory element may be functional in either a eukaryotic cell (e.g., a mammalian cell) or a prokaryotic cell (e.g., bacterial or archaeal cell). In some embodiments, a polynucleotide sequence described herein is operably linked to multiple control elements that allow expression of the polynucleotide in both prokaryotic and eukaryotic cells.Poly (A) sequences
[0215] In some embodiments, the vector and / or expression cassettes described herein further comprises one or more poly(A) sequences. The term “poly(A) sequence” as used herein denotes a DNA sequence which directs both the termination and polyadenylation of the nascent RNA transcript by RNA polymerase II. Polyadenylation sequences can promote mRNA stability by addition of a poly (A) tail to the 3’ end of the coding sequence and thus, contribute to increased translational efficiency. Cleavage and polyadenylation are directed by a poly(A) sequence in the RNA. The core poly(A) sequence for mammalian pre-mRNAs has two recognition elements flanking a cleavage-polyadenylation site. Typically, an almost invariant AAUAAA hexamer lies 20-50 nucleotides upstream of a more variable element rich in U or GU residues. Cleavage of the nascent transcript occurs between these two elements and is coupled to the addition of up to 250 adenosines to the 5’ cleavage product. In some embodiments, the core poly(A) sequence is an ideal poly(A) sequence (e.g., AATAAA, ATTAAA, AGTAAA). Non-limiting examples of poly(A) sequences include SV40 poly(A) sequence, bovine growth hormone (BGH) poly(A) sequence, rabbit P-globin poly(A) sequence (rPgpA), variants thereof, and other suitable heterologous or endogenous poly(A) sequences known in the art. Exemplary poly(A) sequences are provided in Table 9 below.Table 9. Exemplary poly(A) sequencesName Sequence SEQ ID NO: pZC659 taagcttggatccaagcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccg 161tgccttccttgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcat cgcattgtctgagtaggtgtcattctattctggggggtggggtggggcaggacagcaaggggg aggattgggaagacaatagcaggcatgctggggapZC688 taagcttgcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttcct 162tgaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtc tgagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgg gaagacaatagcaggcatgctggggapZC689 agcttgcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttg 163accctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgAttorney Docket No.:TYA-072WO Name Sequence SEQ ID NO:agtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattggga agacaatagcaggcatgctggggapZC690 ctgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaagg 164tgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcat tctattctggggggtggggtggggcaggacagcaagggggaggattgggaagacaatagca ggcatgctggggapZC691 ttgccagccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccac 165tgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggggg gtggggtggggcaggacagcaagggggaggattgggaagacaatagcaggcatgctgggg apZC692 gatgccttctctctccatctaccttccagtcaggatgacggtattatgcttcttggagtcttccaaac 166caccttccctcatctttcatcaatcattgtacagtttgtttacacacgtgcaatttgtttgtgcttctaat atttattgctttatactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttcctt gaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtct gagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattggg aagacaatagcaggcatgctggggapZC693 gatgccttctctctccatctaccttccagtcaggatgacggtattatgcttcttggagtcttccaaac 167caccttccctcatctttcatcaatcattgtacagtttgtttacacacgtgcaatttgtttgtgcttctaat atttattgctttataaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctgg ggggtggggtggggcaggacagcaagggggaggattgggaagacaatagcaggcatgctg gggapZC695 agatctgcctcgactgtgccttctagttgccagccatctgttgtttgcccctcccccgtgccttcctt 168gaccctggaaggtgccactcccactgtcctttcctaataaaatgaggaaattgcatcgcattgtct gagtaggtgtcattctattctggggggtggggtggggcaggacagcaagggggaggattgggaagacaatagcaggcatgctgggga
[0216] In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 161, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 161.
[0217] In some embodiments, the poly(A) sequence comprises a SV40 poly(A) sequence. In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 162, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 162.
[0218] In some embodiments, the poly(A) sequence comprises a SV40 poly(A) sequence. In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 163, or a nucleotide sequence that shares at least 80%, at leastAttorney Docket No.:TYA-072WO 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 163.
[0219] In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 164, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 164.
[0220] In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 165, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 165
[0221] In some embodiments, the poly(A) sequence comprises a SV40 poly(A) sequence. In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 166, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%. at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 166.
[0222] In some embodiments, the poly(A) sequence comprises a SV40 poly(A) sequence. In some embodiments, the poly(A) sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 167, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 167.
[0223] In some embodiments, the poly(A) sequence comprises, consists of. or consists essentially of the nucleotide sequence of SEQ ID NO: 168, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 168.ITRs
[0224] In some embodiments, the expression cassette is flanked by AAV inverted terminal repeats (ITRs) at the 5’ and 3’ ends. ITRs function as recognition sites for replication and markers used for viral packaging of the expression cassette. ITRs form T-shaped secondary structures by two adjacent inverted repeats separated by an unpaired nucleotide. ITRs are required for packaging theAttorney Docket No.:TYA-072WO expression cassette into an rAAV virion, which provide the function of expressing the transgene after a host cell is targeted by the rAAV virion. The ITRs contain tetranucleotide repeat motifs called Repbinding elements (RBE) that act as contact points for the Rep68 / 78 proteins encoded by the rep gene. The ITRs also contain a packaging signal for genome encapsidation, which directs 3’ genomic transport into preassembled capsids by Rep proteins. Any naturally occurring or synthetically derived ITRs described herein or known in the art can be used.
[0225] In some embodiments, the ITRs flanking the expression cassette are ITRs of the same AAV serotype as the Rep protein used in making the virions described herein. For example, where a Rep protein from AAV9 is used, the transgene expression cassette used in the expression system comprises ITRs from AAV9 as well. In another example, where a Rep protein from AAV2 is used, the transgene expression cassette used in the expression system comprises ITRs from AAV2 as well. In another example, where a Rep protein from AAV5 is used, the transgene expression cassette used in the expression system comprises ITRs from AAV5 as well. The ITRs may be of the same or different serotype as the capsid protein used in packaging the virion described herein.
[0226] In some embodiments, the first and / or second expression cassettes may each be flanked by one or more inverted terminal repeats (ITRs). In some embodiments, the first and second expression cassettes are each flanked by a left ITR on the 5’ end of the expression cassette and a right ITR on the 3’ end of the expression cassette. In some embodiments, the left ITR comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 141, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%. or 100% identity to the sequence of SEQ ID NO: 141. In some embodiments, the right ITR comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 142, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%. at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 142. In some embodiments, the first and second expression cassette are each flanked on the 5’ end of the expression cassette by a nucleotide sequence comprising the sequence of SEQ ID NO: 141 and flanked on the 3’ end of the expression cassette by a nucleotide sequence comprising the sequence of SEQ ID NO: 142. In some embodiments, the ITR sequences comprise, consist of, or consist essentially of any nucleotide sequence provided in Table 10 below.
[0227] In some embodiments, the expression cassettes further comprise a junction sequence. InAttorney Docket No.:TYA-072WO some embodiments, the expression cassette is flanked by one or both of a left TTR and junction sequence on the 5’ end of the expression cassette and a right ITR and junction sequence on the 3’ end of the expression cassette. In some embodiments, the left ITR and junction sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 139, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 139. In some embodiments, the right ITR and junction sequence comprises, consists of, or consists essentially of the nucleotide sequence of SEQ ID NO: 140, or a nucleotide sequence that shares at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 140. In some embodiments, the first and second expression cassette are each flanked on the 5’ end of the expression cassette by a nucleotide sequence comprising the sequence of SEQ ID NO: 139 and flanked on the 3’ end of the expression cassette by a nucleotide sequence comprising the sequence of SEQ ID NO: 140. In some embodiments, the ITR plus junction sequences comprise, consist of. or consist essentially of any nucleotide sequence provided in Table 10 below.Table 10. Exemplary ITR sequencesName Sequence SEQ ID NO:Left ITR ctgcgcgctcgctcgctcactgaggccgcccgggcaaagcccgggcgtcgggcgacctttggtcgccc 139 and ggcctcagtgagcgagcgagcgcgcagagagggagtggccaactccatcactaggggttcctggtacc functionRight gctagcaggaacccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgg 140 ITR and gcgaccaaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcag junctionLeft ITR ctgcgcgctcgctcgctcactgaggccgcccgggcaaagcccgggcgtcgggcgacctttggtcgccc 141 ggcctcagtgagcgagcgagcgcgcagagagggagtggccaactccatcactaggggttcctRight aggaacccctagtgatggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgacc 142ITR aaaggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcagSplit Prime Editing Systems
[0228] In some embodiments, the present disclosure provides a system comprising a first vector comprising a first expression cassette described herein and a second vector comprising a second expression cassette described herein.
[0229] The vector can be any viral vector or non-viral vector known in the art or describedAttorney Docket No.:TYA-072WO herein. In some embodiments, the vector is a viral vector. In some embodiments the viral vector is an adeno-associated virus vector (AAV), an adenoviral vector, a lentiviral vector, a retroviral vector, a herpes simplex virus vector (HSV), or a poxvirus vector.
[0230] As used herein, the term “retrovirus” or “retroviral” refers to an RNA virus that reverse transcribes its genomic RNA into a linear double-stranded DNA copy and subsequently covalently integrates its genomic DNA into a host genome. Retrovirus vectors are a common tool for gene delivery. Once the vims is integrated into the host genome, it is referred to as a “provirus.” The provirus serves as a template for RNA polymerase II and directs the expression of RNA molecules encoded by the vims. In some embodiments, a retroviral vector is altered so that it does not integrate into the host cell genome. Illustrative retro viruses include, but are not limited to, (1) genus gammaretrovirus, such as, Moloney murine leukemia virus (M-MuLV or M-MLV), Moloney murine sarcoma virus (MoMSV), murine mammary tumor virus (MuMTV). gibbon ape leukemia virus (GaLV), and feline leukemia virus (FLV); (2) genus spumavirus, such as, simian foamy virus; and (3) genus lentivims, such as, human immunodeficiency virus- 1 and simian immunodeficiency virus.
[0231] As used herein, the term “lentiviral” or “lentivirus” refers to a group (or genus) of complex retroviruses. Illustrative lentiviruses include but are not limited to, human immunodeficiency virus (HIV), including HIV type 1, and HIV type 2; visna-maedi virus (VMV) virus; caprine arthritis-encephalitis virus (CAEV); equine infectious anemia virus (EIAV); feline immunodeficiency virus (FIV); bovine immune deficiency virus (BTV); and simian immunodeficiency virus (SIV).
[0232] In some embodiments, the viral vector is an adenoviral vector. The genetic organization of adenovirus includes an approximate 36 kb, linear, double-stranded DNA virus, which allows substitution of large pieces of adenoviral DNA with foreign sequences up to 7 kb.
[0233] In some embodiments, the viral vector is an AAV vector, such as an AAV vector selected from the group consisting of serotype 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, rh.10, rh.20, rh.74, and a variant or chimeric AAV derived thereof. In some embodiments, the AAV expression vector is pseudotyped to enhance targeting. A pseudotyping strategy can promote gene transfer and sustain expression in a target cell type. For example, the AAV2 genome can be packaged into the capsid of another AAV serotype such as AAV5, AAV7, or AAV8, producing pseudotyped vectors such as AAV2 / 5, AAV2 / 7, and AAV2 / 8 respectively, as described in Balaji et al., J. Surg. Res. ep. (2013)Attorney Docket No.:TYA-072WO 184(1):691-698. Tn some embodiments, an AAV9 may be used to target expression in myofibroblastlike lineages, as described in Piras et al., Gene Therapy (2016) 23:469-478. In some embodiments, AAV1, AAV6. or AAV9 is used, and in some embodiments, the AAV is engineered, as described in Asokari et al., Hum. Gene Ther. Nov. (2013) 24(11):906-913; Pozsgai et al., Mol. Ther. (2017) 25(4): 855-869; Kotterman, M. A. and D. V. Schaffer, Nature Reviews Genetics (2014) 15:445-451; and US20160340393A1 to Schaffer et al. In some embodiments, the viral vector is AAV engineered to increase target cell infectivity as described in US20180066285A1. In some embodiments, the vector is an AAV9 vector.
[0234] In some embodiments, the vector is a non-viral vector. In some embodiments, the non-viral vector is a naked DNA (e.g., a DNA plasmid). In some embodiments, the non-viral vector is a plasmid. In some embodiments, the non-viral vector is a liposome or lipid vector comprising plasmid DNA and a lipid solution.
[0235] In some embodiments, the vector is a recombinant vector. In some embodiments, the viral vectors described herein are replication incompetent, in that it cannot independently further replicate and package its genome. For example, when a cardiac cell is targeted with a virion, the transgene is expressed in the targeted cardiac cell, however, since the targeted cardiac cell lacks packaging and accessory function genes, the virion is not able to replicate. In some embodiments, the viral vectors described herein are replication competent.
[0236] In some embodiments, the vectors described herein are capable of being delivered to both dividing and non-dividing cells. In some embodiments, the vectors described herein are capable of being delivered to non-dividing cells. In some embodiments, the vectors described herein are capable of being delivered to dividing cells.
[0237] In some of these embodiments, the vector is an AAV vector or a variant thereof. In some of these embodiments, the vector is an AAV9 vector or a variant thereof. In some of these embodiments, the vector is an AAV5 vector or a variant thereof. In some of these embodiments, the vector is an AAV2 vector or a variant thereof.
[0238] The capsid proteins of AAV largely determine the immunogenicity and tropism of AAV vectors. In some embodiments, the AAV is an AAV subtype 9 (AAV9). In some embodiments, AAV9 is a preferred AAV vector due to its ability to transduce the heart following systemic delivery. WhileAttorney Docket No.:TYA-072WO AAV9 can achieve moderate transduction of the heart, the majority of vector traffics to the liver. Moreover, in order to achieve therapeutic levels of transduction in the heart, relatively high systemic doses are required, potentially leading to systemic inflammation and in turn, toxicity.
[0239] Methods of introducing polynucleotides into a host cell are known in the art, and any known method can be used to introduce the polynucleotides described herein into a cell. Suitable methods include e.g., viral or bacteriophage infection, transfection, conjugation, protoplast fusion, lipofection, electroporation, calcium phosphate precipitation, polyethyleneimine (PEI)-mediated transfection, DEAE-dextran mediated transfection, liposome-mediated transfection, particle gun technology, calcium phosphate precipitation, direct micro injection, nanoparticle-mediated nucleic acid delivery, microfluidics delivery methods, and the like.
[0240] In some embodiments, the first vector comprising a first expression cassette comprising a first promoter described herein operatively linked to a polynucleotide encoding an N-terminal fragment of a fusion protein described herein. In some embodiments, the N-terminal fragment of a fusion protein comprises an N-terminal fragment of an RNA-guided nickase. In some embodiments, the N-terminal fragment of a fusion protein additionally comprises an N-terminal fragment of a split-intein. In some embodiments, the first expression cassette further comprises a second promoter operatively linked to a second polynucleotide encoding a nicking guide RNA (ngRNA).
[0241] In some embodiments, the second vector comprises a second expression cassette comprising a third promoter described herein operatively linked to a third polynucleotide encoding a C-terminal fragment of a fusion protein described herein. In some embodiments, the C-terminal fragment of a fusion protein comprises a C-terminal fragment of an RNA-guided nickase. In some embodiments, the C-terminal fragment of a fusion protein additionally comprises a reverse transcriptase. In some embodiments, the C-terminal fragment of a fusion protein additionally comprises a C-terminal fragment of a split-intein. In some embodiments, the second expression cassette further comprises a fourth promoter described herein operatively linked to a fourth polynucleotide encoding a prime editing guide RNA (gRNA). In some embodiments, the C-terminal fragment of the fusion protein further comprises an RNA binding domain of small RNA binding exonuclease protection factor La.
[0242] In some embodiments, the present disclosure provides a system comprising:Attorney Docket No.:TYA-072WO (a) a first vector comprising a first expression cassette comprising a first polynucleotide encoding an N-terminal fragment of an RNA-guided nickase described herein and an N-terminal fragment of a split-intein, and a second polynucleotide encoding a first guide RNA; and(b) a second vector comprising a third polynucleotide encoding a C-terminal fragment of an RNA-guided nickase described herein and a C-terminal fragment of a split-intein, and a fourth polynucleotide encoding a second guide RNA.
[0243] In some embodiments, the first guide RNA is a ngRNA described herein and the second guide RNA is a pegRNA described herein. In some embodiments, the first guide RNA is a pegRNA described herein and the second guide RNA is a ngRNA described herein.
[0244] In some embodiments, the present disclosure provides a system comprising(i) a first vector comprising a first expression cassette comprising:(a) a first modified TNNT2 promoter operatively linked to a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein;(ii) a second vector comprising a second expression cassette comprising:(a) a second modified TNNT2 promoter operatively linked to a second polynucleotide encoding a C-terminal fragment of the fusion protein comprising a C-terminal fragment of the RNA-guided nickase, a reverse transcriptase, and a C-terminal fragment of the split-intein; and(b) a third promoter operatively linked to a third polynucleotide encoding a prime editing guide RNA (pegRNA).
[0245] In some embodiments, the first and second modified TNNT2 promoters are independently selected from (a) all TNNT2 promoters described herein, including TNNT2 promoters consisting of the sequence of any one of SEQ ID NOs: 169-174, and (b) a TNNT2 promoter consisting of the sequence of SEQ ID NO: 170.
[0246] In some embodiments, the first expression cassette further comprises a fourth promoter operatively linked to a fourth polynucleotide encoding a nicking guide RNA (ngRNA). In some embodiments, the ngRNA causes the nickase to nick the non-edited strand and increase the efficiencyAttorney Docket No.:TYA-072WO of prime editing. Without wishing to be bound by theory, the ngRNA-mediated nicking of the nonedited strand is thought to bias DNA repair mechanisms towards replacing the non-edited strand.
[0247] In some embodiments, the third promoter and / or fourth promoter are Pol III promoters. Without wishing to be bound by theory, the vector embodiments described herein enable expressing the pegRNA from one vector and the ngRNA from a second vector. This enables the use of the same promoter to each gRNA. Expressing the two gRNAs from the same vector requires using different promoters for these two transcription units (e.g., a human U6 promoter and a murine U6 promoter), as use of two identical promoters in one cassette may cause cassette instability as a result of the homology regions between the two identical promoters. Splitting the gRNAs into two different vectors allows the use of the same human promoter to express both.
[0248] In some embodiments, the first and second vectors are adeno-associated virus (AAV) vectors.
[0249] In some embodiments, the first vector comprises a nucleotide sequence that is at least 70%, at least 80%, at least 85%. at least 90%, at least 95%, at least 99%. or 100% identical to the sequence of SEQ ID NO: 147, a promoter described herein operatively linked to an ngRNA, a left ITR disclosed herein, and a right ITR disclosed herein. In some embodiments, the second vector comprises a nucleotide sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 148, a promoter described herein operatively linked to an pegRNA, a left ITR disclosed herein, and a right ITR disclosed herein.
[0250] In some embodiments, the first vector comprises a nucleotide sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to the sequence of SEQ ID NO: 145, a promoter described herein operatively linked to an ngRNA, a left ITR disclosed herein, and a right ITR disclosed herein. In some embodiments, the second vector comprises a DNA sequence that is at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identical to a sequence selected from SEQ ID NOs: 146, 153, 154, 156, 158, and 160, a promoter described herein operatively linked to a pegRNA, a left ITR disclosed herein, and a right ITR disclosed herein.
[0251] In some embodiments, the system expresses the fusion protein comprising the RNA-guided nickase and the reverse transcriptase. In some embodiments, the system expresses the fusionAttorney Docket No.:TYA-072WO protein comprising the RNA-guided nickase, the reverse transcriptase, and the RNA binding domain of small RNA binding exonuclease protection factor La.
[0252] In some embodiments, the system comprises a first vector comprising a first expression cassette sequence having at least 70%. at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 179. In some embodiments, the system comprises a second vector comprising a second expression cassette sequence having at least 70%, at least 80%, at least 85%, at least 90%, at least 95%, at least 99%, or 100% identity to the sequence of SEQ ID NO: 180 or SEQ ID NO: 181. In some embodiments, the system comprises: (a) a first vector comprising a first expression cassette comprising the sequence of SEQ ID NO: 179; and (b) a second vector comprising a second expression cassette comprising the sequence of SEQ ID NO: 180 or SEQ ID NO: 181. In some embodiments, the first vector and the second vector are viral vectors. In some embodiments, the viral vectors are AAV vectors.Self- Inactivating Prime Editing Systems
[0253] In some embodiments, the present disclosure provides self-inactivating expression cassettes of a prime editing system. When a prime editing system is encoded by expression cassette(s) and delivered by adeno-associated virus (AAV) to organs such as the heart, its RNA and protein expression may persist indefinitely. Even after the desired editing is achieved at the target genomic locus, the expression cassette(s) may continue to produce guide RNAs and prime editor protein in the cell. This persistent expression may create safety and regulatory concerns, including the potential for continued off-target editing activity, immune responses to the foreign protein, and cellular stress from ongoing expression of the editing machinery. One approach to address this concern is to inactivate the expression of the prime editing system after the desired editing has been achieved. However, conventional methods for controlling transgene expression are not straightforward in the context of AAV-delivered systems, as AAV genomes typically persist as episomal elements in non-dividing cells. Methods to terminate AAV-mediated expression post-delivery remain limited.
[0254] Disclosed herein is a self-inactivating prime editing system design that allows the prime editor system to inactivate its own protein expression. The general design incorporates a selfinactivation site sequence, which encodes a short peptide, into one or more expression cassettes of the prime editing system. The self-inactivation site is designed to cause minimal harm to theAttorney Docket No.:TYA-072WO expression and function of the prime editing system in its unedited state. Importantly, the selfinactivation site may comprise a sequence that can be recognized and edited by the pegRNA component of the prime editing system, thereby transforming the self-inactivation site from a benign form (encoding a short peptide) into a form that disrupts protein coding (for example, by introducing a premature stop codon, frameshift mutation, and / or deleterious amino acid changes).
[0255] In certain embodiments, the prime editing system is designed to introduce one or more edits at a target genomic locus while simultaneously editing its own expression cassette(s). This selfinactivating capability is achieved by incorporating a self-inactivation site into one or more of the expression cassettes in the prime editing system.
[0256] In some embodiments, the presence of a self-inactivation site allows the expression and function of the prime editor protein prior to editing of the self-inactivation site. The self-inactivation site comprises a nucleic acid sequence that can be recognized and edited by the prime editing system itself. In some embodiments, the self-inactivation site is recognized by the pegRNA component of the prime editing system. In some embodiments, the pegRNA that recognizes and edits the selfinactivation site is the same pegRNA that targets the genomic locus of interest. In some embodiments, the pegRNA that recognizes and edits the self-inactivation site is a different pegRNA from the pegRNA that targets the genomic locus of interest.
[0257] In some aspects, provided herein is a gene editing system capable of editing a target genomic locus and achieving self-inactivation. In some embodiments, the system comprises one or more polynucleotides encoding a prime editor protein or fragments thereof, wherein at least one polynucleotide comprises a self-inactivation site positioned upstream of or within a coding region. In some embodiments, the system further comprises a polynucleotide encoding a pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site. In some embodiments, editing of the self-inactivation site by the pegRNA introduces one or more protein coding changes that reduce or eliminate expression or function of the prime editor protein.
[0258] In some embodiments, the one or more polynucleotides encoding the prime editor protein or fragments thereof is located in a single vector. In some embodiments, the one or more polynucleotides encoding the prime editor protein or fragments thereof and the polynucleotide encoding the pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site are located in the same vector. In some embodiments, the one or more polynucleotides encodingAttorney Docket No.:TYA-072WO the prime editor protein or fragments thereof and the polynucleotide encoding the pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site are located in at least two separate vectors.
[0259] In some embodiments, the editing event at the self-inactivation site leads to one or more protein coding changes. Such protein coding changes include, but are not limited to, amino acid changes, frameshift mutations, and premature stop codons. In some embodiments, the editing event at the self-inactivation site introduces one or more amino acid changes to the prime editor protein. In other embodiments, the editing event at the self-inactivation site introduces a frameshift mutation. In still other embodiments, the editing event at the self-inactivation site introduces a premature stop codon. In certain embodiments, the editing event at the self-inactivation site introduces a combination of amino acid changes, frameshift mutations, and premature stop codons.
[0260] In some embodiments, the protein coding changes resulting from editing of the selfinactivation site reduce the expression or function of the prime editor protein. In some embodiments, the protein coding changes reduce the expression of the prime editor protein by at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%. In some embodiments, the protein coding changes abolish the expression of the prime editor protein. In some embodiments, the protein coding changes reduce the function of the prime editor protein by at least 50%, at least 60%, at least 70%. at least 80%, at least 90%, at least 95%, or at least 99%. In certain embodiments, the protein coding changes abolish the function of the prime editor protein.
[0261] In some embodiments, prior to editing, the self-inactivation site encodes a peptide sequence that is incorporated into the prime editor protein without substantially impairing protein function. In some embodiments, the peptide encoded by the unedited self-inactivation site comprises 1 to 100 amino acids, e.g., 5 to 100 amino acids, 5 to 75 amino acids, 5 to 50 amino acids, 5 to 25 amino acids, 10 to 50 amino acids, 10 to 40 amino acids, 10 to 25 amino acids, or 15 to 30 amino acids.
[0262] In some embodiments, the prime editing system comprises a split prime editor system, wherein the self-inactivation site can be added to the PE-N expression cassette, the PE-C expression cassette, or both expression cassettes. In some embodiments, once the prime editing system comprises self-inactivation site(s) inserted into its PE-N cassette, PE-C cassette, or both cassettes, the prime editing system can reduce or abolish its own expression and function over time. In someAttorney Docket No.:TYA-072WO embodiments, this self-inactivation mechanism provides temporal control over prime editor activity and enhances the safety profile of the system for therapeutic applications.
[0263] In some embodiments, the prime editing system comprises a split prime editor system, wherein the self-inactivation site can be added to the PE-N expression cassette, the PE-C expression cassette, or both expression cassettes. In certain embodiments, the self-inactivation site is incorporated into the PE-N expression cassette. In other embodiments, the self-inactivation site is incorporated into the PE-C expression cassette. In still other embodiments, a self-inactivation site is incorporated into both the PE-N expression cassette and the PE-C expression cassette.
[0264] In embodiments where self-inactivation sites are incorporated into both the PE-N and PE-C expression cassettes, the self-inactivation site in PE-N and the self-inactivation site in PE-C can be the same or different. In some embodiments, the self-inactivation sites in PE-N and PE-C comprise the same nucleotide sequence. In other embodiments, the self-inactivation sites in PE-N and PE-C comprise different nucleotide sequences.
[0265] In some embodiments, the self-inactivation site is positioned within a coding sequence of the PE-N or PE-C expression cassette. In certain embodiments, the self-inactivation site is inserted after the start codon of the expression cassette. In other embodiments, the self-inactivation site is positioned within the first 100 nucleotides, first 200 nucleotides, first 300 nucleotides, first 400 nucleotides, or first 500 nucleotides downstream of the start codon.
[0266] In some embodiments, after being delivered to cells, the self-inactivating prime editing system achieves both targeted genome editing at the genomic locus of interest and self-inactivation through editing of its own expression cassette(s). In some embodiments, the self-inactivation capability provides temporal control over prime editor activity. In some embodiments, the selfinactivation capability reduces potential off-target editing events.Recombinant AAV Virions
[0267] In some embodiments, the gene editing system comprises at least one vector comprising an expression cassette encoding a nucleic acid-guided nuclease. In some embodiments, the gene editing system comprises at least one vector comprising an expression cassette encoding a guide RNA (e.g., a pegRNA). In some embodiments, the at least one vector is an AAV vector.
[0268] In some embodiments, the AAV is any AAV known in the art or described herein. InAttorney Docket No.:TYA-072WO some embodiments, the AAV is an AAV selected from the group consisting of serotype 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, rh.10, rh.20, rh.74, or a chimeric or variant AAV derived therefrom. In some embodiments, the AAV is AAV1. AAV2, AAV3. AAV4, AAV5. AAV6, AAV7. AAV8, AAV9. AAV10, AAV11, AAV12, AAVrh.10, AAVrh.20, AAVrh.74, or a variant thereof.
[0269] In some embodiments, the rAAV virus or virion comprises an AAV capsid protein and an expression cassette as described herein. Capsid proteins are structural proteins that make up the assembled icosahedral packaging of the rAAV virion that contains the expression cassette. Capsid proteins are classified by the serotype. Wild-type capsid serotypes in rAAV virions can be, for example, AAV1, AAV2, AAV3. AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV 12, AAVrh.10, AAVrh.20, or AAVrh.74. Engineered capsid types include chimeric capsids and mosaic capsids. Capsids are selected for rAAV virions based on their ability to transduce specific tissue or cell types.
[0270] Any capsid protein that can facilitate rAAV virion transduction into cardiac cells for delivery of a transgene, as described herein, can be used. Capsid proteins used in rAAV virions for transgene delivery to cardiac cells that result in high expression can be, for example. AAV4, AAV6, AAV7, AAV8, and AAV9. In some embodiments, the AAV capsid protein described herein is a wild-type AAV capsid protein from AAV serotype 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, rh.10, rh.20, rh.74, or a variant thereof. In some embodiments, the AAV is AAV9 or a variant thereof. In some embodiments, the AAV is AAV5 or a variant thereof. In some embodiments, the AAV is AAV2 or a variant thereof.
[0271] Artificial capsids, such as chimeric capsids generated through combinatorial libraries, can also be used for transgene delivery to cardiac cells that results in high expression. Other capsid proteins with various features can also be used in the rAAV virions of the disclosure. AAV vectors and capsids are provided in U. S. Pat. Pub. Nos. US10011640B2; US7892809B2, US8632764B2, US8889641B2, US9475845B2, US10889833B2, US 10480011B2, and US10894949B2, the entire contents of each of which are incorporated by reference herein; and Int’l Pat. Pub. Nos. WO2020198737A1, W02019028306A2, WO2016054554A1, WO2018152333A1, WO2017106236A1, WO2008124724A1, W02017212019A1, WO2020117898A1, WO2017192750A1, W02020191300A1, and W02017100671A1, the entire contents of each of which are incorporated by reference herein.Attorney Docket No.:TYA-072WO
[0272] In some embodiments, the AAV capsid protein is an engineered capsid protein, comprising two or more non-naturally occurring amino acid motifs (also called “variant amino acid sequences”) relative to a corresponding wild-type or parental capsid protein. The wild-type or parental capsid protein can be any wild-type, chimeric, or mosaic capsid protein as described herein or as known in the art, or a variant thereof. In some embodiments, the wild-type or parental capsid protein is an AAV9, AAV5, AAVrh.10, or AAVrh.74 capsid. Non-limiting examples of engineered capsid proteins are provided in Int’l Pat. Pub. Nos. WO2021216456A2 and W02023201207A1, the entire contents of each of which are incorporated by reference herein.
[0273] In some embodiments, the AAV capsid protein is an engineered capsid protein comprising the amino acid sequence set forth in SEQ ID NO. 186 (ZC755).
[0274] As described herein, a human RBM20-encoding gene may be any RBM20 gene sequence that occurs in a human. In some embodiments, the human RBM20-encoding gene is a “wild-type” allele. In some embodiments, the human RBM20-encoding gene is a mutant allele. In some embodiments, the mutant allele comprises a single nucleotide substitution, a deletion, an insertion, an inversion, or a complex rearrangement. In some embodiments, the mutant allele is pathogenic mutant.
[0275] As described herein, a pathogenic mutant is a mutant associated with altered activity of the encoded RBM20 protein. In some embodiments, altered activity of the encoded RBM20 protein comprises reduced activity. In some embodiments, altered activity of the encoded RBM20 protein comprises gain of one or more deleterious activities. Non-limiting examples of activities of the encoded RBM20 protein include binding of RNA and regulation of splicing. In some embodiments, a pathogenic mutant is associated with altered splicing of one or more genes (e.g., titin (TIN) and / or calcium / calmodulin-dependent protein kinase 2 (CAMK2). Gene splicing function can be measured by any method known in the art. In some embodiments, gene splicing is measured by determining the quantity (expression level) of at least one splice variant of a gene. In some embodiments, gene splicing is measured by determining the relative quantity (relative expression level) of at least one splice compared to a control gene. In some embodiments, altered gene splicing of an RBM20-encoding gene that is a mutant allele (e.g., a pathogenic mutant) results in increased expression of Ttn N2BA and decreased expression of Ttn N2B. In some embodiments, altered gene splicing of an RBM20-encoding gene that is a mutant allele (e.g., a pathogenic mutant) results in increased expression ofAttorney Docket No.:TYA-072WO Camk2d-A and decreased expression of Camk2d-B and / or Camk2d-C.
[0276] In some embodiments, the human RBM20 gene or the mouse Rbm20 gene comprises a genetic variant. Several genetic variants have been associated with cardiomyopathy. Non-limiting examples of cardiomyopathy include dilated cardiomyopathy, hypertrophic cardiomyopathy, restrictive cardiomyopathy, and arrhythmogenic cardiomyopathy. In some embodiments, the genetic variant is associated with dilated cardiomyopathy (DCM). In some embodiments, cardiomyopathy is associated with one or more altered heart functions. Cardiomyopathy, and / or one or more heart functions, may be measured by any method known in the art. Non-limiting examples of heart functions include ejection fraction (EF), left ventricle internal diameter at systolic stage (LVID;s), and / or left ventricle internal diameters at diastolic stage (LVID;d).
[0277] In some embodiments, the genetic variant comprises a mutation. In some embodiments, the mutation is located in the arginine-serine-rich (RS) domain of RBM20. In some embodiments, the RS domain (also called a “mutation hotspot”) comprises amino acid positions 636 to 640 of the mouse RBM20 protein. In some embodiments, the RS domain comprises amino acid positions 634 to 638 of the human RBM20 protein. In some embodiments, the genetic variant comprises an R636Q mutation in the mouse RBM20 or an R634Q mutation in the human RBM20.Non-Human Animals Comprising a Humanized RBM20 Gene
[0278] In some aspects, the present disclosure provides genetically-modified, non-human animals comprising a humanized RNA binding motif protein 20 (RBM20) gene. In some embodiments, the humanized RBM20 gene comprises nucleic acids of the murine Rbm20 gene (e.g., within SEQ ID NO: 1) that are replaced with a human RBM20 sequence.
[0279] As described herein, a human RBM20 sequence may be any RBM20 sequence that occurs in a human. In some embodiments, the human RBM20 sequence comprises a sequence that occurs in a “wild-type” allele. In some embodiments, the human RBM20 sequence comprises a sequence that occurs in a mutant allele. In some embodiments, the mutant allele comprises a single nucleotide substitution, a deletion, an insertion, an inversion, or a complex rearrangement. In some embodiments, the mutant allele is pathogenic mutant.
[0280] In some embodiments, the nucleic acids replaced within the murine Rbm20 gene comprise a nucleic acid sequence encoding the RS domain of the mouse RBM20 protein. In someAttorney Docket No.:TYA-072WO embodiments, the non-human animal comprises a humanized RBM20 gene wherein nucleic acids 250-264 of SEQ ID NO: 1 within the murine Rbm20 gene are replaced with a human RBM20 nucleic acid sequence.
[0281] In some embodiments, the human RBM20 nucleic acid sequence comprises a nucleic acid sequence encoding the RS domain of the human RBM20. In some embodiments, the human RBM20 sequence comprises the sequence of SEQ ID NO: 4.
[0282] In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising a mutation. In some embodiments, the mutation is located in the RS domain of the human RBM20 protein. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising arginine (R) at an amino acid position corresponding to position 634 of human RBM20 protein. In some embodiments, the human RBM20 nucleic acid sequence comprises an R634Q mutation, wherein the numbering is according to the human RBM20 protein of SEQ ID NO: 93. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R636Q mutation, wherein the numbering is according to the murine RBM20 protein of SEQ ID NO: 91. In some embodiments, the human RBM20 sequence comprises the sequence of SEQ ID NO: 5. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to the human RBM20 protein of SEQ ID NO: 93.
[0283] In some aspects, the present disclosure provides genetically-modified, non-human animals comprising a humanized RNA binding motif protein 20 (RBM20) gene, wherein at least 15 consecutive nucleic acids between 230-284 of SEQ ID NO: 1 are replaced with a human RBM20 nucleic acid sequence. In some embodiments, the at least 15 consecutive nucleic acids include nucleic acids 250-264. In some embodiments, the at least 15 consecutive nucleic acids include nucleic acids upstream of nucleic acids 250-264 and / or nucleic acids downstream of nucleic acids 250-264. In some embodiments, the at least 15 consecutive nucleic acids encode the RS domain of the human RBM20. Non-limiting examples of the at least 15 consecutive nucleic acids include sequences comprising nucleic acids 37-294, nucleic acids 37-264, or nucleic acids 250-294. Further non-limiting examples of the at least 15 consecutive nucleic acids include sequences comprising nucleic acids nucleic acids 50-264, nucleic acids 100-264, nucleic acids 150-264, nucleic acids 200-264, nucleic acids 210-264, nucleic acids 220-264, nucleic acids 230-264, nucleic acids 235-264, nucleic acids 240-264, nucleic acids 245-264, nucleic acids 246-264, nucleic acids 247-264, nucleic acids 248-264, nucleic acidsAttorney Docket No.:TYA-072WO 249-264, nucleic acids nucleic acids 250-264, nucleic acids 250-265, nucleic acids 250-266, nucleic acids 250-267, nucleic acids 250-268, nucleic acids 250-269, nucleic acids 250-274, nucleic acids 250-279, nucleic acids 250-284, or nucleic acids 250-294.
[0284] In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 6. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 7. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 8. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 9.
[0285] In some embodiments, the human RBM20 nucleic acid sequence replaces at least 25, at least 30, at least 35, at least 40. at least 45, at least 50, at least 60. at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, or at least 220 nucleotides upstream and / or downstream of nucleic acids 250-264 of SEQ ID NO: 1 within the murine Rbm20 gene.
[0286] In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4 and further comprises a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 4. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4 and further comprises a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 4. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4 and: (a) a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 4; and / or (b) a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 4.
[0287] In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5 and further comprises a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 5. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5 and further comprises a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 5. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5 and: (a) a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 5;Attorney Docket No.:TYA-072WO and / or (b) a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 5.
[0288] In some embodiments, the human RBM20 nucleic acid sequence replaces nucleic acids 31-294 of SEQ ID NO: 1 within the murine Rbm20 gene. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 89. In some embodiments, the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 90.
[0289] In some embodiments, the non-human animal comprises a first allele and a second allele, wherein at least the first allele comprises a humanized RBM20 gene. In some embodiments, the second allele comprises a humanized RBM20 gene. In some embodiments, the non-human animal comprises a first allele comprising a first humanized RBM20 gene and a second allele comprising a murine Rbm20 gene. In some embodiments, the non-human animal comprises a first allele comprising a first humanized RBM20 gene and a second allele comprising a second humanized RBM20 gene. In some embodiments, the first humanized RBM20 gene encodes an RBM20 protein comprising arginine (R) at an amino acid position corresponding to position 634 of human RBM20 protein. In some embodiments, the human RBM20 protein comprises or consists of the sequence of SEQ ID NO: 93. In some embodiments, the RBM20 protein comprising arginine (R) at an amino acid position corresponding to position 634 of human RBM20 protein comprises or consists of the sequence of SEQ ID NO: 94.
[0290] In some embodiments, the first humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q), wherein the numbering is according to the human RBM20 protein of SEQ ID NO: 93. In some embodiments, the first humanized RBM20 gene encodes an RBM20 protein comprising an R636Q mutation, wherein the numbering is according to the murine RBM20 protein of SEQ ID NO: 91. In some embodiments, the second humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q), wherein the numbering is according to the human RBM20 protein of SEQ ID NO: 93. In some embodiments, the second humanized RBM20 gene encodes an RBM20 protein comprising an R636Q mutation, wherein the numbering is according to the murine RBM20 protein of SEQ ID NO: 91.
[0291] In some embodiments, a human RBM20-targeting guide RNA specifically binds to theAttorney Docket No.:TYA-072WO humanized RBM20 gene. Non-limiting embodiments of / ? / < W20- targeting guide RNAs include any disclosed herein and any known in the art.
[0292] In some embodiments, the humanized RBM20 gene allows targeting of a nucleic acid-guided nuclease to the RBM20 nucleotide sequence using a human RBM20-targeting guide RNA (gRNA). The nucleic acid-guided nuclease may be any described herein or known in the art.
[0293] In some embodiments, the endogenous and human RBM20 sequences encode a humanized RBM20 protein with substantially similar activity to a wild-type RBM20 protein. As described herein, “substantially similar activity” refers to the ability to perform one or more functions of a wild-type RBM20 protein at a similar level and / or rate. Non-limiting examples of activities include binding of RNA and regulation of splicing.
[0294] In some embodiments, the non-human animal has substantially similar ejection fraction (EF), left ventricle internal diameter at systolic stage (LVID;s), and / or left ventricle internal diameters at diastolic stage (LVID;d) to a non-human animal comprising two wild-type RBM20 alleles. In some embodiments, the non-human animal has substantially similar ejection fraction (EF), left ventricle internal diameter at systolic stage (LVID;s), and / or left ventricle internal diameters at diastolic stage (LVID;d) to a non-human animal that is homozygous for the murine wild-type Rbm20 gene.
[0295] In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to human RBM20 protein, and the humanized RBM20 protein comprising an R634Q mutation has altered activity compared to a wild-type RBM20 protein. In some embodiments, the human RBM20 protein comprises or consists of the amino acid sequence of SEQ ID NO: 93. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R636Q mutation, wherein the numbering is according to murine RBM20 protein. In some embodiments, the murine RBM20 protein comprises or consists of the amino acid sequence of SEQ ID NO: 91. In some embodiments, altered activity of the encoded RBM20 protein comprises reduced activity. In some embodiments, altered activity of the encoded RBM20 protein comprises gain of one or more deleterious activities.
[0296] In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to human RBM20 protein, and the non-human animal has reduced ejection fraction (EF), increased left ventricle internal diametersAttorney Docket No.:TYA-072WO at systolic stage (LVID;s), and / or increased left ventricle internal diameter at diastolic stage (LVID;d) compared to a non-human animal comprising two wild-type RBM20 alleles. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to human RBM20 protein, and the non-human animal has reduced ejection fraction (EF), increased left ventricle internal diameters at systolic stage (LVID;s), and / or increased left ventricle internal diameter at diastolic stage (LVID;d) compared to a non-human animal that is homozygous for the murine wild-type Rbm20 gene. In some embodiments, the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to human RBM20 protein, and wherein the animal has increased levels of Nppa transcripts, Nppb transcripts, natriuretic peptide A, and / or natriuretic peptide B.
[0297] In some embodiments, the genetically-modified, non-human animal is a mammal. In some embodiments, the genetically-modified, non-human animal is a mouse, a rat, a sheep, a pig, a dog, a cow, a rabbit, a dog, or a non-human primate. In some embodiments, the genetically-modified, non-human animal is a mouse.
[0298] In some embodiments, the genetically-modified, non-human animal is a mouse of any genetic background or strain known in the art. In some embodiments, the genetically-modified, non-human animal is a mouse of a C57BL strain, a BALB strain, a 129 strain, an A / I strain, a CBA strain, or a DBA strain. Non-limiting examples of C57BL strains are C57BL / A, C57BL / An, C57BL / GrFa, C57BL / KaLwN, C57BL / 6, C57BL / 6I, C57BL / 6N, C57BL / 6ByJ, C57BL / 6NJ, C57BL / 10, C57BL / 10ScSn, C57BL / 10Cr, and C57BL / 01a. Non-limiting examples of BALB strains areBALB / c and BALB / cI. Non-limiting examples of 129 strains are 129P1. 129P2, 129P3, 129X1, 129S1 (e.g., 129S1 / SV, 129Sl / SvIm), 129S2, 129S4, 129S5, 129S9 / SvEvH, 129 / SvIae, 129S6 (129 / SvEvTac), 129S7, 129S8, 129T1, and 129T2. In some embodiments, a mouse of the present invention is a mix of two aforementioned strains.
[0299] In some embodiments, the genetically modified mouse is used to screen gRNAs for efficacy in editing the mutant human RBM20 gene. In some embodiments, the present disclosure provides methods for identifying guide RNAs suitable for gene editing systems described herein, wherein the method comprises administering the guide RNA and / or gene editing system to a genetically-modified, non-human animal described herein. In some embodiments, the methods identify guide RNAs that allow efficient editing of a human RBM20 gene or humanized RBM20 geneAttorney Docket No.:TYA-072WO when the guide RNA is used as part of a gene editing system described herein. In some embodiments, the guide RNAs provide editing with an efficiency of at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60% or 70%.
[0300] In some embodiments, the methods identify guide RNAs that allow efficient editing of a mutant human RBM20 allele or humanized RBM20 allele but do not allow efficient editing of a wild-type human RBM20 allele or humanized RBM20 allele. In some embodiments, the guide RNAs provide editing of a pathogenic mutant human RBM20 allele or humanized RBM20 allele with an efficiency of at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60% or 70%, but does not, or substantially does not, edit the wild-type human RBM20 allele or humanized RBM20 allele.
[0301] In some embodiments, the guide RNAs identified by the method are pegRNAs and / or ngRNAs. Any suitable method known in the art may be used to identify such guide RNAs. In some embodiments, the method comprises contacting a cell with the guide RNA and an effector protein described herein, a gene editing system described herein, or a vector described herein, amplifying the target region DNA, and sequencing by NGS to quantify editing.Methods of Determining the Efficacy of a Gene Editing System
[0302] In some aspects, the present disclosure provides a method of determining the efficacy of a gene editing system targeting the human RBM20 gene. In some embodiments, the method of determining the efficacy of a gene editing system targeting the human RBM20 gene comprises contacting a cell of a non-human animal comprising a humanized RNA binding motif protein 20 (RBM20) gene, as described herein, with the gene editing system, wherein the gene editing system comprises a guide RNA comprising a spacer sequence that binds to a target sequence of the humanized RBM20 gene and an RNA-guided nuclease. In some embodiments, the method further comprises analyzing the humanized RBM20 gene to determine the presence or absence of editing events in the humanized RBM20 gene, wherein the gene editing system is determined to be efficacious by the presence of editing events. The presence or absence of editing events can be determined by any method known in the art. In some embodiments, the presence or absence of editing events is determined by next- generation sequencing (NGS). In some embodiments, the cell is in vivo. In some embodiments the cell is ex vivo.. In some embodiments, the method further comprises measuring in the non-human animal one or more heart functions associated with cardiomyopathy. Heart functions associated with cardiomyopathy may be measured by any method known in the art. Non-limitingAttorney Docket No.:TYA-072WO examples of heart functions include ejection fraction (EF), left ventricle internal diameter at systolic stage (LVID;s), and / or left ventricle internal diameters at diastolic stage (LVID;d).
[0303] In some embodiments, the method of determining the efficacy of a gene editing system targeting the human RBM20 gene comprises administering a guide RNA and an RNA-guided nuclease to a genetically-modified, non-human animal described herein. In some embodiments, the administering comprises administering an effective amount of the guide RNA and RNA-guided nuclease. In some embodiments, the guide RNA and RNA-guided nuclease are administered in one or more vectors. In some embodiments, the one or more vectors are AAV vectors. In some embodiments, the method the method further comprises analyzing the humanized RBM20 gene to determine the presence or absence of editing events in the humanized RBM20 gene, wherein the gene editing system is determined to be efficacious by the presence of editing events. The presence or absence of editing events can be determined by any method known in the art. In some embodiments, the presence or absence of editing events is determined by next-generation sequencing (NGS). In some embodiments, the method further comprises measuring in the non-human animal one or more heart functions associated with cardiomyopathy. Heart functions associated with cardiomyopathy may be measured by any method known in the art. Non-limiting examples of heart functions include ejection fraction (EF), left ventricle internal diameter at systolic stage (LVID;s), and / or left ventricle internal diameters at diastolic stage (LVID;d).
[0304] In some embodiments, the guide RNA specifically binds the humanized RBM20 gene.
[0305] In some aspects, the present disclosure provides a population of cells isolated from any genetically-modified, non-human animal described herein. In some embodiments, the population of cells is a population of cardiomyocytes. In some aspects, the present disclosure provides a cell isolated from any genetically-modified, non-human animal described herein. In some embodiments, the cell is an embryonic stem cell.Methods of Use
[0306] In some embodiments, the present disclosure provides methods for expressing a polynucleotide in a target cell. The method may comprise, for example, transducing a target cell with one or more vectors described herein. A target cell can be, for example and without limitation, a cardiac cell, a muscle cell, an induced pluripotent stem cell-derived cardiomyocyte (iPSC-CM), aAttorney Docket No.:TYA-072WO cardiomyocyte.
[0307] In some embodiments, the one or more vectors encode an effector protein and a gRNA specific for a human RBM20 gene or a humanized RBM20 gene. In some embodiments, the target cell comprises at least one human RBM20 gene or at least one humanized RBM20 gene. In some embodiments, the target cell comprises one human RBM20 gene or one humanized RBM20 gene. In some embodiments, the target cell comprises two human RBM20 genes or two humanized RBM20 genes. In some embodiments, a method of expressing an effector protein and / or a gRNA specific for a human RBM20 gene or a humanized RBM20 gene in a target cell comprises transducing a target cell or population of target cells with one or more AAV vectors or AAV virions described herein.
[0308] In some embodiments, the target cell comprises at least one target sequence. In some embodiments, the target sequence is within a human RBM20 gene or a humanized RBM20 gene described herein. In some embodiments, the human RBM20 gene or the humanized RBM20 gene is a mutant human RBM20 allele or a mutant humanized RBM20 allele. In some embodiments, the human RBM20 gene or the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, wherein the numbering is according to human RBM20 protein.
[0309] In some embodiments, the method results in editing in the mutant human RBM20 allele or mutant humanized RBM20 allele. In some embodiments, the editing has an efficiency of at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60% or 70%. In some embodiments, the editing has an efficiency of at least 25%. In some embodiments, the editing has an efficiency of at least 30%. In some embodiments, the editing has an efficiency of at least 35%. In some embodiments, the editing has an efficiency of at least 40%. In some embodiments, the editing has an efficiency of at least 45%. In some embodiments, the editing has an efficiency of at least 50%. In some embodiments, the editing has an efficiency of at least 55%. In some embodiments, the editing has an efficiency of at least 60%. In some embodiments, the editing has an efficiency of at least 65%. In some embodiments, the editing has an efficiency of at least 70%. In some embodiments, the editing has an efficiency of at least 75%. In some embodiments, the editing has an efficiency of at least 80%. In some embodiments, the editing has an efficiency of at least 85%. In some embodiments, the editing has an efficiency of at least 90%. In some embodiments, the editing has an efficiency of at least 95%.
[0310] In some embodiments, the method does not, or substantially does not, edit or form editing events in a wild-type human RBM20 allele or wild-type humanized RBM20 allele.Attorney Docket No.:TYA-072WO Methods of Treatment
[0311] In some embodiments, the present disclosure provides methods for preventing and / or treating a disease or condition in a subject in need thereof, comprising administering to the subject a gene editing system described herein or a vector encoding the same, wherein the disease or condition is a cardiac pathology caused by or associated with a mutation in the RBM20 gene.
[0312] Subjects who are suitable for the compositions and / or methods of the present technology include individuals (e.g., mammalian subjects, such as humans, non-human primates, domestic mammals, experimental non-human mammalian subjects such as mice, rats, etc.) having a heterozygous mutation (e.g., a dominant negative mutation) in the RBM20 gene. In some embodiments, the subject is a human.
[0313] In some embodiments, the method for preventing and / or treating a disease or condition caused by or associated with a mutant human RBM20 allele in a subject in need thereof comprises editing of the RBM20 gene.
[0314] In some embodiments, the disease or condition caused by or associated with a mutation (e.g., a dominant negative mutation) in the RBM20 gene is a heart disease. In some embodiments, the heart disease is cardiomyopathy. In some embodiments, the heart disease is dilated cardiomyopathy (DCM). In some embodiments, the heart disease is arrhythmia, including, for example, ventricular arrhythmia and malignant ventricular arrhythmia. In some embodiments, the subject has or is at risk of having a heart disease (e.g., cardiomyopathy) caused by or associated with a mutant human RBM20 allele. In some embodiments, the mutant allele is a pathogenic mutant. In some embodiments, the mutant RBM20 allele encodes an RBM20 protein comprising an R634Q mutation.
[0315] In some embodiments, the pharmaceutical composition for use in the method for preventing and / or treating a disease or condition caused by or associated with a mutant RBM20 allele (e.g., rAAV vectors, viruses, or virions) can be administered to the subject by systemic application (such as parenteral application), for example, by intravenous (e.g., by IV infusion), intra-arterial, or intraperitoneal delivery. In some embodiments, pharmaceutical compositions of the gene editing systems described herein or vectors encoding the same (e.g., rAAV vectors, viruses, or virions) can be delivered by direct administration to the heart tissue.Attorney Docket No.:TYA-072WO
[0316] In some embodiments, the pharmaceutical compositions of the gene editing systems described herein or vectors encoding the same (e.g., rAAV vectors, viruses, or virions) can be delivered by intracoronary administration. In some embodiments, the administration is by antegrade epicardial coronary artery infusion, e.g., a single infusion over a 10-minute period in a cardiac catheterization laboratory after angiography (percutaneous intracoronary delivery without vessel balloon occlusion) with the use of standard 5F or 6F guide or diagnostic catheters.
[0317] In some embodiments, the pharmaceutical compositions of the gene editing systems described herein or vectors encoding the same (e.g., rAAV vectors, viruses, or virions) can be delivered by direct injection into the heart or cardiac catheterization, or by intracardiac catheter delivery via retrograde coronary sinus infusion (RCSI).
[0318] When direct injection is used, it may be performed either by open-heart surgery or by minimally invasive surgery. In some cases, the pharmaceutical compositions of the gene editing systems described herein or vectors encoding the same (e.g., rAAV vectors, viruses, or virions) can be delivered to the pericardial space by injection or infusion.
[0319] In some embodiments, the amount, concentration, and volume of the pharmaceutical composition administered to a subject can be controlled and / or optimized to substantially improve the functional parameters of the heart while mitigating adverse side effects.
[0320] The amount of the composition administered to myocardial tissue can also be an amount required to result in the detectable expression of a gene editing system in the heart; preserve and / or improve contractile function; reduce or eliminate ventricular arrhythmia; and / or delay the emergence of cardiomyopathy or reverse the pathological course of the disease.
[0321] In some embodiments, the compositions and / or methods disclosed herein result in reduced expression of the mutant RBM20 gene and / or restoration of function of the RBM20 gene and gene product in a cardiac cell or tissue (such as heart) of the subject being treated.
[0322] In some embodiments, the method comprises administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of about IxlO8genome copies per milliliter (GC / mL), about 5xl08GC / mL, about IxlO9GC / mL, about 5xl09GC / mL, about IxlO10GC / mL, about 5xlO10GC / mL, about IxlO11GC / mL, about 5xlOnGC / mL, about IxlO12GC / mL, about 5xl012GC / mL, about 5xl013GC / mL, about IxlO14GC / mL, or about 5xl014GC / mL of theAttorney Docket No.:TYA-072WO rAAV vector, virus, or virion.
[0323] In some embodiments, the method comprises intravenously administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of about 3xl012GC / mL, about 3xl013GC / mL, about IxlO14GC / mL, or about 3xl014GC / mL of the rAAV vector, virus, or virion.
[0324] In some embodiments, the method comprises administering, by localized delivery to the heart, an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of about 3×1011GC / mL, about 3×1012GC / mL, about 1×1013GC / mL, or about 3×1013GC / mL of the rAAV vector, virus, or virion.
[0325] In some embodiments, the method comprises administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of about IxlO8viral genomes per milliliter (vg / mL), about 5xl08vg / mL, about IxlO9vg / mL, about 5xl09vg / mL, about IxlO10vg / mL, about 5xlO10vg / mL, about IxlO11vg / mL, about 5xl0nvg / mL, about IxlO12vg / mL, about 5xl012vg / mL, about 5xl013vg / mL, about IxlO14vg / mL, or about 5xl014vg / mL of the rAAV vector, virus, or virion.
[0326] In some embodiments, the method comprises administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of less than about IxlO15viral genomes per milliliter (vg / mL), less than about 5x1014vg / mL, less than about IxlO14vg / mL, less than about 5xl013vg / mL, less than about IxlO13vg / mL, less than about 5xl012vg / mL, less than about IxlO12vg / mL, less than about 5×1011vg / mL, or less than about IxlO11vg / mL of the rAAV vector, virus, or virion.
[0327] In some embodiments, the method comprises administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of less than about IxlO14viral genomes per milliliter (vg / mL) or less than about IxlO13vg / mL of the rAAV vector, virus, or virion.
[0328] In some embodiments, the method comprises administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of from about IxlO11viral genomes per milliliter (vg / mL) to about IxlO15vg / mL, from about IxlO11vg / mL to about IxlO14vg / mL, from about IxlO12vg / mL to about IxlO14vg / mL, from about IxlO12vg / mL to about IxlO13vg / mL, or from about IxlO12vg / mL to about 6xl013vg / mL of the rAAV vector, virus, or virion.Attorney Docket No.:TYA-072WO
[0329] In some embodiments, the method comprises administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at any dose or dose range of the disclosure between the values referenced herein.
[0330] In some embodiments, the method comprises intravenously administering an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of about lxlO12vg / mL, about 3xlO12vg / mL, about 6xl012vg / mL, or about 9xl012vg / mL of the rAAV vector, virus, or virion.
[0331] In some embodiments, the method comprises administering, by localized delivery to the heart, an rAAV vector, virus, or virion encoding the gene editing systems described herein at a dose of about lxlO12vg / mL, about 3xlO12vg / mL, about 6xl012vg / mL, or about 9xl012vg / mL of the rAAV vector, virus, or virion.
[0332] Genome copies per milliliter can be determined by quantitative polymerase change reaction (qPCR) using a standard curve generated with a reference sample having a known concentration of the polynucleotide genome of the virus. For AAV, the reference sample used is often the transfer plasmid used in generation of the rAAV virion, but other reference samples may be used.
[0333] Alternatively, the concentration of a viral vector can be determined by measuring the titer of the vector on a cell line. Viral titer is typically expressed as viral particles (vp) per unit volume (e.g., vp / mL). In various embodiments, the pharmaceutical composition comprises about IxlO8viral particles per milliliter (vp / mL), about 5xl08vp / mL, about IxlO9vp / mL, about 5xl09vp / mL, about IxlO10vp / mL, about 5xlO10vp / mL, about IxlO11vp / mL, about 5×1011vp / mL, about IxlO12vp / mL, about 5xl012vp / mL, about 5xl013vp / mL, about IxlO14vp / mL, or about 5xl014of the rAAV vector, virus, or virion encoding the gene editing systems described herein.
[0334] The vector, virus, or virion encoding the gene editing systems described herein administered to the subject can be traced by a variety of methods. For example, recombinant viruses labeled with or expressing a marker (such as green fluorescent protein, or beta-galactosidase) can readily be detected. The recombinant viruses may be engineered to cause the target cell to express a marker protein, such as a surface-expressed protein or a fluorescent protein. Alternatively, the infection of target cells with recombinant viruses can be detected by their expression of a cell markerAttorney Docket No.:TYA-072WO that is not expressed by the animal employed for testing (for example, a human-specific antigen when injecting cells into an experimental animal). The presence and phenotype of the target cells can be assessed by fluorescence microscopy (e.g., for green fluorescent protein, or beta-galactosidase), by immunohistochemistry (e.g., using an antibody against a human antigen), by ELISA (using an antibody against a human antigen), or by RT-PCR analysis using primers and hybridization conditions that cause amplification to be specific for RNA indicative of a cardiac phenotype.
[0335] In some embodiments, the vector, virus, or virion encoding the gene editing systems described herein, or a pharmaceutical composition containing the same, is administered to the subject once a day, twice a day, three times a day, or four times a day for a period of about 1 day, about 2 days, about 3 days, about 5 days, about 7 days, about 10 days, about 2 weeks, about 3 weeks, about 4 weeks, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, about 7 months, about 8 months, about 9 months, about 10 months, about 11 months, about 1 year, about 2 years, about 3 years, about 4 years, about 5 years, or more than about 5 years. In some embodiments, the vector, virus, or virion encoding the gene editing systems described herein, or a pharmaceutical composition containing the same, is administered every day, every other day, 3 times a week, every third day, weekly, biweekly (i.e., every other week), every third week, monthly, every other month, every third month, every fourth month, every fifth month, every sixth month, every ninth month, every year, every 18 months, every 2 years, every 5 years, every 10 years, or every 20 years. In some embodiments, the dose regimens listed above could be repeated after a period of about 1 week, about 1 month, about 2 months, about 3 months, about 4 months, about 5 months, about 6 months, about 7 months, about 8 months, about 9 months, about 10 months, about 11 months, about 1 year, about 2 years, about 3 years, about 4 years, about 5 years, or more than about 5 years. In some embodiments, the schedule of administration is a hybrid of these periods, for example, the vector, virus, or virion encoding the gene editing systems described herein, or a pharmaceutical composition containing the same, is administered a number of times a week and that pattern is repeated a number of times a month every month or every second, third, fourth, fifth, or sixth month, and treatment according to that pattern is continued for part of a year to several years, as set out above. In some embodiments, the vector, virus, or virion encoding the gene editing systems described herein, or a pharmaceutical composition containing the same, is administered in a cycle of a number or administrations over a week or two weeks, and the cycle is repeated at spaced intervals over a number of months or years, as set out above. In some embodiments treatment is continued until disease isAttorney Docket No.:TYA-072WO eliminated, until no further improvement is achieved, or as long as the disease does not progress. In some embodiments, a disease or condition within a subject to be treated can be monitored to evaluate the effectiveness of the treatment using any appropriate method known to a skilled artisan.
[0336] In some embodiments, the vector, virus, or virion encoding the gene editing systems described herein, or a pharmaceutical composition containing the same, is administered over a predetermined time period. Alternatively, the vector, virus, or virion encoding the gene editing systems described herein, or a pharmaceutical composition containing the same, is administered until a particular therapeutic benchmark is reached. In some embodiments, the methods provided herein further include a step of evaluating one or more therapeutic benchmarks in the subject to determine whether to continue administration of the treatment.
[0337] In some embodiments, the method results in an editing efficiency (e.g., knocking down, knocking out, or otherwise altering the expression of the mutant allele) of, or an efficiency of indel formation (e.g., forming a frameshift mutation) in the mutant human RBM20 allele of at least 20%. at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100%. In some embodiments, the method does not, or substantially does not, edit, or form editing events in, the wild-type RBM20 allele.
[0338] In some embodiments, the method results in a change in splicing of one or more genes (e.g., titin (TTN) and / or calcium / calmodulin-dependent protein kinase 2 (CAMK2). Gene splicing can be measured by any method known in the art. In some embodiments, gene splicing is measured by determining the quantity (expression level) of at least one splice variant of a gene. In some embodiments, gene splicing is measured by determining the relative quantity (relative expression level) of at least one splice compared to a control gene.
[0339] In some embodiments, the method restores or improves cardiac function in the subject, and / or restores or improves contractile function of the heart in the subject.
[0340] In some embodiments, the method improves ejection fraction in the subject, for example, increases the ejection fraction by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%. at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%.Attorney Docket No.:TYA-072WO at least 150%, or at least 200% in the subject with a mutant human RBM20 allele.
[0341] In some embodiments, the method reduces left ventricular hypertrophy, left ventricular mass, and / or left ventricular wall thickness in the subject. In some embodiments, the method improves left ventricular relaxation and / or left ventricular filling pressure in the subject.
[0342] In some embodiments, the method decreases or prevents an increase in left ventricle internal dimension (LVID) in the subject, for example, decreases or prevents an increase in LVID (mm) by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, or at least 200% in the subject with a mutant human RBM20 allele.
[0343] In some embodiments, the method decreases or prevents an increase in left ventricle (LV) mass in the subject, for example, decreases or prevents an increase in LV mass (as measured in mg / g of body weight) by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%. at least 45%, at least 50%. at least 55%, at least 60%. at least 65%, at least 70%. at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, or at least 200% in the subject with a mutant human RBM20 allele.
[0344] In some embodiments, the method increases or prevents a decrease in stroke volume in the subject, for example, increases or prevents a decrease in stroke volume by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, or at least 200% in the subject with a mutant human RBM20 allele.
[0345] In some embodiments, the method increases or prevents a decrease in R amplitude in the subject, for example, increases or prevents a decrease in R amplitude (mV) by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%. at least 60%, at least 65%. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, or at least 200% in the subject with a mutant human RBM20 allele.
[0346] In some embodiments, the method decreases or prevents an increase in Q wave, R wave,Attorney Docket No.:TYA-072WO and S wave (QRS) interval in the subject, for example, decreases or prevents an increase in QRS interval by at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%. at least 50%, at least 55%. at least 60%, at least 65%. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 100%, at least 150%, or at least 200% in the subject with a mutant human RBM20 allele.
[0347] In some embodiments, the method increases life span or prevents an increase in mortality over time in the subject with a mutant human RBM20 allele.EXAMPLESExample 1. Development of humanized RBM20 mouse models
[0348] The mouse Rbm20R636Qmodel exhibits dilated cardiomyopathy and allows testing the efficacy of gene editing drugs targeting mouse Rbm20. However, mouse Rbm20 and human RBM20 have different DNA sequences, which hinders testing of human RBM20- targeting gene editing systems in the mouse model because gene editing is DNA sequence-dependent. To enable in vivo efficacy and editing test of human RBM20 gene editing systems, humanized RBM20 mouse models were designed in which the mouse Rbm20 DNA sequence around the mutation hotspot is replaced by the corresponding human RBM20 DNA sequence (FIG. 1A). An example of gene editing systems is prime editing (PE).
[0349] The mutation hotspot in the RBM20 gene is located in exon 9 at about nucleic acid positions 20 to 34 (both human RBM20 and murine Rbm20\ The DNA sequences, protein sequences, and amino acid numbering of mouse Rbm20, human RBM20, and human RBM20R634Qat the mutation hotspot region are shown in FIG. 1B. Human RBM20 R634 corresponds to mouse Rbm20 R636, and the arginine (R) residue mutated to glutamine (Q) is located at position 636 of the mouse genome in the humanized RBM20 mouse models. However, to better distinguish endogenous mouse Rbm20 and humanized RBM20, this residue is referred to as R634 in the models described herein (e.g., human nomenclature and number). Two humanized RBM20 knock-in alleles were constructed: a wildtype knock-in (hRBM20) and a R634Q mutant knock-in hRBM20R634Q). The DNA sequence alignment of mouse Rbm20 wildtype allele, humanized wildtype knock-in allele, and humanized R634Q knock-in allele is shown in FIG.1C, with the humanized region highlighted and the mutation hotspot region outlined in red. The G> A point mutation of the R634Q allele is marked by an arrowhead.Attorney Docket No.:TYA-072WO
[0350] Mice with various genotypes carrying hRBM20+and / or hRBM20R634Qwere generated: one carrying one wild-type humanized RBM20 allele and one wild-type mouse Rbm20 allele (hRBM20+ / mRbm20+), a second carrying two wild-type humanized RBM20 alleles (hRBM20+ / hRBM20+). a third carrying one R634Q-mutant humanized RBM20 allele and one wildtype mouse Rbm20 allele (hRBM20R634Q / mRbm20+a fourth carrying one wild-type humanized RBM20 allele and one humanized R634Q-mutant RBM20 allele hRBM20R634Q / hRBM20+) and a fifth carrying two humanized R634Q-mutant RBM20 alleles (hRBM20R634Q / hRBM20R634Q)). Exemplary genotypes of humanized RBM20 mice (including mutants and controls) and wild-type naive mice are provided in Table 11.Table 11. Exemplary mouse genotypesFirst Allele Second Allele DescriptionhRBM20+mRBM20+One wild-type humanized RBM20 allele;one wild-type mouse Rbm20 allele hRBM20+hRBM20+Two wild-type humanized RBM20 alleles hRBM20R634QmRBM20+One R634Q-mutant humanized RBM20 allele;one wild-type mouse Rbm20 allele hRBM20R634QhRBM20+One wild-type humanized RBM20 allele;one R634Q-mutant humanized RBM20 allele hRBM20R634QhRBM20R634QTwo R634Q-mutant humanized RBM20 alleles mRBM20+mRBM20+Two wild-type mouse Rbm20 alleles
[0351] Survival rates of hRBM20+ / mRbm20+and hRBM20+ / hRBM20+control mice, as well as hRBM20R634Q / hRBM20+heterozygous mutant mice and hRBM20R634Q / hRBM2(634Qhomozygous mutant mice, were monitored (FIG. ID). Mortality observed in the hRBM20+ / hRBM20+control group was due to severe fighting wounds and was not cardiac -related. No mortality was observed in the heterozygous hRBM20R634Q / hRBM20+mutant group at 22 weeks of age. Homozygous hRBM20R634Q / hRBM20R634QRBM20 mutant mice showed dramatically reduced lifespan.Example 2. Humanized hRBM20 wildtype and hRBM20R634Qmimic heart function phenotypes of mouse mRbm20 wildtype and mRbm20R636Q
[0352] The ejection fractions (EF), left ventricle internal diameters at systolic stage (LVID;s),Attorney Docket No.:TYA-072WO and left ventricle internal diameters at diastolic stage (LVID;d) of wildtype mice and mouse Rbm20R636Q / +mice at various time points are shown in FIG. 2A. Ejection fractions, left ventricle internal diameters at systolic stage (LVID;s), and left ventricle internal diameters at diastolic stage (LVID;d) of wildtype mice and mouse Rbm20R636Q / +mice at 26 weeks of age are shown in FIG. 2B.
[0353] Corresponding EF, LVID;s and LVID;d values for wildtype mRbm20+ / +mice, heterozygous humanized wildtype hRBM20+ / mRbm20+mice, and heterozygous humanized mutant hRBM20R634Q / mRbm20+mice are shown in FIG.2C and FIG.2D. Humanized WT and mutant alleles were shown to mimic the phenotype of their corresponding mouse orthologous alleles.
[0354] The ejection fractions (EF), left ventricle internal diameters at systolic stage (LVID;s), and left ventricle internal diameters at diastolic stage (LVID;d) of hRBM20+ / mRbm20+control mice, hRBM20+ / hRBM20+control mice, hRBM20R634Q / hRBM20+heterozygous mutant mice, and hRBM20R634Q / hRBM20R634Qhomozygous mutant mice at various time points are shown in FIG. 2E.
[0355] The ejection fractions (EF), left ventricle internal diameters at systolic stage (LVID;s), and left ventricle internal diameters at diastolic stage (LVID;d) of hRBM20+ / mRbm20+control mice. hRBM20+ / hRBM20+control mice, hRBM20R634Q / hRBM20+heterozygous mutant mice and hRBM20R634Q / hRBM20R634Qhomozygous mutant mice at 8 weeks of age are shown in FIG.2F.
[0356] Natriuretic peptide A (ANP, also called atrial natriuretic peptide), and natriuretic peptide B (BNP, also called brain natriuretic peptide), encoded by Nppa and Nppb, respectively, are important biomarkers in clinical cardiology. Nppa and Nppb are induced by cardiac stress and the levels of atrial natriuretic factor and brain natriuretic peptide are found to be increased in heart failure (HF) proportionally to HF severity. Relative expression levels of heart failure marker genes Nppa and Nppb of hRBM20+ / mRbm20+control mice, hRBM20+ / hRBM20+control mice, hRBM20R634Q / hRBM20+heterozygous mutant mice, and hRBM20R634Q / hRBM20R634Qhomozygous mutant mice at 8 weeks of age are shown in FIG.2G.
[0357] As shown in FIG.2E and FIG.2F, hRBM20R634Q / hRBM20+heterozygous mutant mice developed a moderate DCM phenotype, characterized by a moderate decline in ejection fraction (EF) and increased left ventricle internal diameters at systolic stage (LVID;s), and left ventricle internal diameters at diastolic stage (LVID;d). In comparison, hRBM20R634Q / hRBM20R634Qhomozygous mutant mice exhibited a much more severe phenotype, featuring a dramatic EF drop and significantAttorney Docket No.:TYA-072WO left ventricular enlargement (FIG.2E and FIG.2F), as well as dramatically elevated Nppa and Nppb expression FIG.2G.Example 3. Humanized hRBM20R634Qmimic gene splicing phenotypes of mouse mRbm20R636Q
[0358] RBM20 protein is a splicing factor regulating the splicing of multiple genes in cardiomyocytes. FIGs. 3A and 3B show that compared to wildtype mice, adult Rbm20R636Q / +mice exhibit abnormal splicing of Ttn (FIG. 3A) and Camk2d (FIG. 3B) genes. Similarly, humanized mutant hRBM20R634Q / mRbm20+mice show abnormal splicing of Ttn (FIG. 3C) and Camk2d (FIG.3D) genes. FIG. 3E and FIG. 3F shows the splicing patterns of Ttn and Camk2d in 8-week-old hRBM2(T / mRbm20+control mice, hRBM2(r / hRBM2(T control mice, hRBM20R634Q / hRBM20+heterozygous mutant mice, and hRBM20R634Q / hRBM20R634Qhomozygous mutant mice. Both hRBM2()K634< J / hRBM2(T heterozygous mutant mice and hRBM2()K634< J / hR M2()Kf,34< Jhomozygous mutant mice show abnormal splicing of Ttn and Camk2d. This observation suggests that the humanized mouse model described herein recapitulates gene splicing phenotypes of the endogenous mouse model.Example 4. Assessment of ngRN As and pegRNAs for prime editing at human RBM20 locus
[0359] A prime editing system requires a prime editing guide RNA (pegRNA) to provide a spacer sequence (which directs the editor to the target locus), a reverse transcription primer binding site (PBS), and a reverse transcription template sequence (RTT). Optionally, a second guide RNA, nicking guide RNA (ngRNA) is included in the prime editing system to direct (through the spacer sequence of ngRNA) the editor to generate a second nick site near the editing locus. Without wishing to be bound by theory, the second nicking events may boost prime editing efficiency.
[0360] To identify pegRNA and ngRNA sequences that facilitate editing within a human RBM20 locus, an in vitro study was performed in HEK293T cells using a split-prime editing (split PE) gene editing system. There are many variables in pegRNA and ngRNA design, including but not limited to the sequence and length of pegRNA spacer, the sequence and length of PBS, the sequence and length of RTT, the sequence and length of ngRNA spacer, etc. Plasmids encoding prime editing machinery and DNA fragments expressing pegRNA and ngRNA were delivered to HEK293T cells via chemical transfection. The pegRNA + ngRNA designs screened in this system install one or more silent mutations (mutations to a gene's DNA sequence that have no effect on the amino acid sequenceAttorney Docket No.:TYA-072WO coded for by that gene) on RBM20 relative to the wildtype. This design allows the efficiencies of guide RNA pairs to be determined by quantifying the frequencies of editing events in HEK293T cells that carry wildtype RBM20 sequence. Frequencies of editing events were measured by next generation sequencing (NGS) of DNA extracted from cells.
[0361] FIGs. 4A-4H and FIGs. 5A-5E show quantification of editing event frequencies in HEK293T cells transfected with different combinations of pegRNAs and ngRNAs. In FIGs.4A-4H and FIGs. 5A-5E, each panel represents one set of results comparing prime editing efficiencies of multiple pairs of pegRNA + ngRNA. Light color on the heatmap represents higher efficiency, dark stands for lower efficiency. Based on these the identification of pegRNA + ngRNA candidates were identified.
[0362] In a similar screening approach, additional pegRNA and ngRNA combinations were evaluated for their efficiency in editing the human RBM20 locus in HEK293T cells (FIG.9). Multiple pegRNA and ngRNA pairs were systematically screened to identify optimal combinations for RBM20 editing. The screening identified ng-26 (ngRNA comprising the spacer sequence of SEQ ID NO: 113) and ng-36 (ngRNA comprising the spacer sequence of SEQ ID NO: 194) as high-performing ngRNAs, demonstrating superior editing efficiency compared to other ngRNA candidates tested in this study.
[0363] Exemplary pegRNA and ngRNA candidates are listed in Table 12 and Table 13 below.Table 12. Exemplary pegRNA candidatesName Sequence SEQ ID NO: peg-1 ctcacagatatggcccagaagttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 104 aaaagtggcaccgagtcggtgctcaccggactacgagaccgcggtctttctgggccatatcactattcacgc ggttctatctagttacgcgttaaaccaactagaapeg-28 ctcacagatatggcccagaagttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 105 aaaagtggcaccgagtcggtgcacgagaccgcggtctttctgggccatatctgataatacacgcggttctat ctagttacgcgttaaaccaactagaapeg-32 ctcacagatatggcccagaagttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 106 aaaagtggcaccgagtcggtgcagaccgcggtctttctgggccatatcataaacaccgcggttctatctagtt acgcgttaaaccaactagaapeg- 186 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 107 aaaagtggcaccgagtcggtgcaggccgcggtctcgtagtccggtcagccggtcacaaagtaagcgcgg ttctatctagttacgcgttaaaccaactagaapeg- 189 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 108Attorney Docket No.:TYA-072WO Name Sequence SEQ ID NO: aaaagtggcaccgagtcggtgcaggccgaggtctcgtagtccggtcagccggtcactcaatattaccgcg gttctatctagttacgcgttaaaccaactagaapeg-224 gggagagtgaccggctcacgttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttga 109 aaaagtggcaccgagtcggtgcaggccgcggtctcgaagtccggtcagccggtcacaagaaattcgcggttctatctagttacgcgttaaaccaactagaaTable 13. Exemplary ngRNA candidatesName spacer sequence full ngRNA sequence SEQ ID NO: ng-2 ctacgagaccgcggtct ctacgagaccgcggtctttcgttttagagctagaaatagcaagttaaaataaggct 118 ttc agtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-20 tgggacctcggggaga tgggacctcggggagagtgacgttttagagctagaaatagcaagttaaaataag 119 gtgac gctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-6 gggacctcggggaga gggacctcggggagagtgacgttttagagctagaaatagcaagttaaaataag 120 gtgac gctagtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-26 agattctaaatcctgctc agattctaaatcctgctcctgttttagagctagaaatagcaagttaaaataaggcta 121 Ct gtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-28 ggtctcgtagtccggtc ggtctcgtagtccggtcagcgttttagagctagaaatagcaagttaaaataaggc 122 age tagtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-31 gattctaaatcctgctcc gattctaaatcctgctcctgttttagagctagaaatagcaagttaaaataaggctag 123 t tccgttatcaacttgaaaaagtggcaccgagtcggtgcng-35 ctgtgtgtctgtgtgtgg ctgtgtgtctgtgtgtggggttttagagctagaaatagcaagttaaaataaggcta 124 g gtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-30 tcagccggtcactctcc tcagccggtcactctccccggttttagagctagaaatagcaagttaaaataaggc 125 ccg tagtccgttatcaacttgaaaaagtggcaccgagtcggtgcng-36 attctaaatcctgctcct attctaaatcctgctcctgttttagagctagaaatagcaagttaaaataaggctagt 193ccgttatcaacttgaaaaagtggcaccgagtcggtgcExample 5. Assessment of ngRN As and pegRNAs for prime editing of humanized RBM20 gene in mouse embryonic fibroblasts (MEFs)
[0364] To identify pegRNA and ngRNA sequences, and combinations thereof, that facilitate editing of the humanized mutant hRBM20R634Qallele sequence, an in vitro study was performed in mouse embryonic fibroblasts (MEFs) isolated from humanized RBM20 mice carrying the humanized wildtype hRBM20 and mutant hRBM20R634Qalleles. Plasmids encoding prime editing machinery and pegRNA + ngRNA candidates were delivered to MEFs via chemical transfection. Frequencies of editing events were measured by next generation sequencing (NGS) of DNA extracted from cells.FIG. 6 shows the relative editing efficiency comparison of nine combinations of a pegRNA and a ngRNA. Light color on the heatmap represents higher efficiency, dark stands for lower efficiency.Attorney Docket No.:TYA-072WO Among these nine pairs, the combination of peg- 186 and ng-26 showed the highest efficiency. Exemplary combinations of pegRNA and ngRNA are shown in Table 14 below.Table 14. Exemplary combinations of a pegRNA and an ngRNApegRNA name pegRNA SEQ ID NO: ngRNA name ngRNA SEQ ID NO: peg-1 104 ng-2 118peg-28 105 ng-2 118peg-32 106 ng-2 118peg-1 104 ng-20 119peg- 186 107 ng-26 121peg- 186 107 ng-28 122peg- 186 107 ng-31 123peg- 186 107 ng-35 124peg- 189 108 ng-26 121peg- 189 108 ng-31 123peg- 189 108 ng-35 124peg-224 109 ng-26 121peg-224 109 ng-30 125peg-224 109 ng-31 123Example 6. Assessment of in vivo editing efficiencies of RBM20-targeting primer editors in humanized RBM20 mice
[0365] RBM20- targeting prime editing drugs were administered at various dosages to hRBM20R634Q / mRbm20+mice via retro-orbital injection. At 3 weeks post-injection, RNA was extracted from the whole heart and the percentage of humanized RBM20 transcript that carry the desired edit was measured by NGS analysis.
[0366] FIG.7A shows in vivo testing results of two RBM20- targeting prime editing drugs, PE-hRBM20-4.9 and PE-hRBM20-4.10. PE-hRBM20-4.9 and PE-hRBM20-4.10 both utilize an AAV9 capsid variant comprising expression cassettes according to the V20 expression cassette design, in which the first expression cassette is V20-PE-N (SEQ ID NO: 145) and the second expression cassette is V20-PE-C (SEQ ID NO: 154). In both PE-hRBM20-4.9 and PE-hRBM20-4.10. the ngRNA is ng-26 (SEQ ID NO: 121). In PE-hRBM20-4.9, the pegRNA is peg-186 (SEQ ID NO: 107), and PE-hRBM20-4.10 the pegRNA is peg-224 (SEQ ID NO: 109). The AAV capsid protein ZC755 (SEQ ID NO. 186) was used in this AAV9 capsid variant. The features and sequences of p PE-hRBM20-4.9 and PE-hRBM20-4.10 are shown below in Table 15.Attorney Docket No.:TYA-072WO Table 15. Exemplary human or humanized RBM20 gene editing systemsName First cassette Second cassetteN-terminal Full C-terminal Full ngRNA pegRNAfragment sequence fragment sequence PE- V20 PE-N V20 PE-C peg- 186SEQ ID NO: SEQ ID NO: 0 (SEQ ID ng-26 (SEQhRBM2 - (SEQ ID (SEQ IDID NO: 121) 176 177 4.9 NO: 145) NO: 154) NO: 107)PE- peg-224 V20 PE-N SEQ ID NO: V20 PE-Cng-26 (SEQ SEQ ID NO: hRBM20- (SEQ ID (SEQ ID (SEQ IDID NO: 121) 176 178 4.10 NO: 145) NO: 154) NO: 109)
[0367] FIG. 7B shows in vivo testing results of PE-hRBM20-4.9 at three dosages. The editing efficiency increases substantially from 1.5E13 vg / kg per vector dosage to 3E13 vg / kg per vector, but only slightly from 3E13 vg / kg per vector to 5E13 vg / kg per vector.
[0368] These data demonstrate that the humanized RBM20 mouse models described herein are effective for testing of prime editing drugs in vivo. These data also demonstrate that the PE-hRBM20-4.9 and PE-hRBM20-4.10 prime editing drugs effectively mediate in vivo editing of the human RBM20 sequence.Example 7. Dual- AAV based prime editing treatment improves cardiac function in humanized RBM20 cardiomyopathy mouse model
[0369] Dual-AAV based prime editing drugs PE-hRBM20-4.9, PE- hRB M20-4.ll, and PE-hRBM20-4.12 were tested in this study. PE-hRBM20-4.9, PE-hRBM20-4.11, and PE-hRBM20-4.12 all utilize an AAV9 capsid variant comprising expression cassettes according to the V20 expression cassette design, in which the first expression cassette is V20-PE-N (SEQ ID NO: 145) and the second expression cassette is V20-PE-C (SEQ ID NO: 154). Upon co-delivery to cells, the PE-N and PE-C components reconstitute to form a functional prime editor. The AAV capsid protein ZC755 (SEQ ID NO. 186) was used in the V20 expression cassette design. The features and sequences of PE-hRBM20-4.9, PE-hRBM20-4.11, and PE-hRBM20-4.12 are shown below in Table 16. The combinations of pegRNA and ngRNA used in this study are: peg-186 (SEQ ID NO: 107) and ng-26 (SEQ ID NO: 121); peg-262 (SEQ ID NO: 187) and ng-26 (SEQ ID NO: 121); peg-263 (SEQ ID NO: 188) and ng-26 (SEQ ID NO: 121). Additional exemplary combinations of a pegRNA and an ngRNA are shown below in Table 17. For example, an exemplary combination of ng-36 and peg-262 comprises a first cassette comprising the sequence set forth in SEQ ID NO: 207 and a second cassette comprising the sequence set forth in SEQ ID NO: 189.Attorney Docket No.:TYA-072WO Table 16. Exemplary human or humanized RBM20 gene editing systemsName First cassette Second cassetteN-terminal Full C-terminal Full ngRNA pegRNAfragment sequence fragment sequence PE- V20 PE-N V20 PE-C peg- 186SEQ ID NO: SEQ ID NO: h (SEQ ID ng-26 (SEQRBM20- (SEQ ID (SEQ IDID NO: 121) 176 177 4.9 NO: 145) NO: 154) NO: 107)PE- V20 PE-N peg-26Q SEQ ID V20 PE-C 2ng-26 (SE NO: SEQ ID NO: hRBM20- (SEQ ID (SEQ ID (SEQ ID4.11 ID NO: 121) 176 189 NO: 145) NO: 154) NO: 187PE- V20 PE-N V20 PE-Cng-26 (SEQ SEQ ID NO: peg- SEQ ID NO: hRBM20- (SEQ ID (SEQ ID 263(SEQ IDID NO: 121) 176 1904.12 NO: 145) NO: 154) NO: 188Table 17. Exemplary combinations of a pegRNA and an ngRNApegRNA name pegRNA SEQ ngRNA name ngRNA SEQ ngRNA Spacer ID NO: ID NO: SEQ ID NO: peg- 186 107 ng-2 118 110 peg- 186 107 ng-20 119 111 peg- 186 107 ng-6 120 112 peg- 186 107 ng-30 125 117 peg- 186 107 ng-36 193 194 peg-262 187 ng-2 118 110 peg-262 187 ng-20 119 111 peg-262 187 ng-6 120 112 peg-262 187 ng-26 121 113 peg-262 187 ng-28 122 114 peg-262 187 ng-31 123 115 peg-262 187 ng-35 124 116 peg-262 187 ng-30 125 117 peg-262 187 ng-36 193 194 peg-263 188 ng-2 118 110 peg-263 188 ng-20 119 111 peg-263 188 ng-6 120 112 peg-263 188 ng-26 121 113 peg-263 188 ng-28 122 114 peg-263 188 ng-31 123 115 peg-263 188 ng-35 124 116 peg-263 188 ng-30 125 117peg-263 188 ng-36 193 194
[0370] As shown in FIG. 8A, PE-N and PE-C vectors were co-injected to 3.5-week-oldAttorney Docket No.:TYA-072WO humanized hRBM20R634Q / hRBM20+mice through RO injection at 3E13 vg / kg per vector (6E13 vg / kg total). A vehicle treated hRBM20R634QhRBM20+mice group and a humanized wildtype (hRBM20+ / hRBM20+) mice control group were included in the study.
[0371] Starting from 4-week post-injection time point, heart function of PE treated, vehicle treated, and wildtype control animals was measured by echocardiogram (Echo) once every 4 weeks. As shown in FIG. 8B, ejection fractions (EF), left ventricle internal diameters at diastolic stage (LVID;d), and left ventricle internal diameters at systolic stage (LVID;s) were measured at pretreatment baseline and post-injection time points. FIG. 8C shows ejection fractions (EF), left ventricle internal diameters at diastolic stage (LVID;d), and left ventricle internal diameters at systolic stage (LVID;s) measured before dosing and at 16-week post-injection time point.
[0372] Results showed PE-hRBM20-4.9, and PE-hRBM20-4.11, and PE-hRBM20-4.12 all improved heart function compared to vehicle treatment.
[0373] Heart samples were collected after the 16-week post-injection Echo timepoint for ex vivo analysis. FIG. 8D and FIG. 8E plot gene editing results at the humanized RBM20 locus from heart RNA samples. PE-hRBM20-4.9, PE-hRBM20-4.11, and PE-hRBM20-4.12 successfully edited the target region, reducing the percentage of mutant RBM20 R634Q encoding sequences detected in this assay. RBM20 is a gene splicing factor and the RBM20 R634Q mutation leads to abnormal splicing of RBM20 targets. FIG.8F and FIG. 8G show expression levels of mRNA isoforms of the Ttn gene and the Camk2d gene. The splicing of the Ttn gene and the Camk2d gene is also regulated by RBM20. PE-hRBM20-4.9, PE-hRBM20-4.11, and PE-hRBM20-4.12 treatments all reverted their splicing patterns towards wildtype.
[0374] These data demonstrate the potential of in vivo prime editing drugs in treating human RBM20 cardiomyopathy.Example 8. Self-inactivating prime editing system
[0375] A self-inactivating prime editing system was designed to introduce edit(s) at a target genomic locus while simultaneously editing its own expression cassette(s). This is achieved by incorporating a self-inactivation site into one or more of the expression cassettes in the prime editing system. In the split PE system described herein, the self-inactivation site can be added to either the PE-N cassette, the PE-C cassette, or both expression cassettes. The self-inactivation site in PE-N andAttorney Docket No.:TYA-072WO the self-inactivation site in PE-C can be the same or different. The presence of a self-inactivation site still allows the expression and function of the prime editor protein.
[0376] A self-inactivation site comprises a nucleic acid sequence that can be recognized and edited by the prime editing system. The editing event at the self-inactivation site leads to protein coding change(s), such as amino acid change(s) and / or premature stop codon(s), thereby reducing or abolishing the expression or function of the prime editor protein. Once this system is delivered to cells, it achieves both targeted genome editing and the inactivation of itself.
[0377] FIG. 10A depicts a generic example of such design. In this example, a self-inactivation site (SI site) is added to the PE-N cassette. FIG. 10B depicts a specific example in which a selfinactivation site is inserted after the start codon of the PE-N cassette. The self-inactivation site can be recognized and edited by the pegRNA included in the prime editing system, which targets the mouse Rbm20 locus. This pegRNA is part of the prime editing system. Without the editing event, the selfinactivation site adds a peptide to the prime editor protein. The editing events at the self-inactivation site introduce amino acid changes, frameshift mutations, and a premature stop codon, thereby abolishing the protein coding capability of the cassette and inactivating the expression of the prime editor protein.
[0378] FIGs. 10C and 10D present an in vivo study testing this example in mouse. This selfinactivating prime editing system example shows successful editing at both its target genomic locus and its self-inactivation site, demonstrating the feasibility of self-inactivating prime editing systems for therapeutic applications where temporal control of editor expression is desired.
[0379] Exemplary sequences used in the self-inactivation design are provided in Table 18.Table 18. Exemplary sequences for self-inactivationDescription Sequence SEQ ID NO.Full PE-N AAV ctgcgcgctcgctcgctcactgaggccgcccgggcaaagcccgggcgtcgggcg 208 cassette sequence acctttggtcgcccggcctcagtgagcgagcgagcgcgcagagagggagtggcc aactccatcactaggggttcctggtaccgttgccttctgccccccaaccctgctccca gctggccctcccaggcctgggttgctggcctctgctttatcaggattctcaagaggg acagctggtttatgttgcatgactgttccctgcatatctgctctggttttaaatagcttatc tgagcagctggaggaccacatgggcttatatggcgtggggtacatgttcctgtagcc ttgtccctggcacctgccaaaatagcagccaacaccccccacccccaccgccatccccctgccccacccgtcccctgtcgcacattcctccctccgcagggctggctcaccaggccccagcccacatgcctgcttaaagccctctccatcctctgcctcacccagtcccAttorney Docket No.:TYA-072WO Description Sequence SEQ ID NO.cgctgagactgagcagacgcctccaggccaccatgcctcccaggtatggtccaga gcggccacgctcgaagtccaatgagctaaaacggacagccgacggaagcgagtt cgagtcaccaaagaagaagcggaaagtcgacaagaagtacagcatcggcctgga catcggcaccaactctgtgggctgggccgtgatcaccgacgagtacaaggtgccc agcaagaaattcaaggtgctgggcaacaccgaccggcacagcatcaagaagaac ctgatcggagccctgctgttcgacagcggcgaaacagccgaggccacccggctg aagagaaccgccagaagaagatacaccagacggaagaaccggatctgctatctgc aagagatcttcagcaacgagatggccaaggtggacgacagcttcttccacagactg gaagagtccttcctggtggaagaggataagaagcacgagcggcaccccatcttcg gcaacatcgtggacgaggtggcctaccacgagaagtaccccaccatctaccacct gagaaagaaactggtggacagcaccgacaaggccgacctgcggctgatctatctg gccctggcccacatgatcaagttccggggccacttcctgatcgagggcgacctgaa ccccgacaacagcgacgtggacaagctgttcatccagctggtgcagacctacaac cagctgttcgaggaaaaccccatcaacgccagcggcgtggacgccaaggccatc ctgtctgccagactgagcaagagcagaaagctggaaaatctgatcgcccagctgcc cggcgagaagaagaatggcctgttcggaaacctgattgccctgagcctgggcctg acccccaacttcaagagcaacttcgacctggccgaggatgccaaactgcagctga gcaaggacacctacgacgacgacctggacaacctgctggcccagatcggcgacc agtacgccgacctgtttctggccgccaagaacctgtccgacgccatcctgctgagc gacatcctgagagtgaacaccgagatcaccaaggcccccctgagcgcctctatgat caagagatacgacgagcaccaccaggacctgaccctgctgaaagctctcgtgcgg cagcagctgcctgagaagtacaaagagattttcttcgaccagagcaagaacggcta cgccggctacattgacggcggagccagccaggaagagttctacaagttcatcaag cccatcctggaaaagatggacggcaccgaggaactgctcgtgaagctgaagaga gaggacctgctgcggaagcagcggaccttcgacaacggcagcatcccccaccag atccacctgggagagctgcacgccattctgcggcggcaggaagatttttacccattc ctgaaggacaaccgggaaaagatcgagaagatcctgaccttccgcatcccctacta cgtgggccctctggccaggggaaacagcagattcgcctggatgaccagaaagag cgaggaaaccatcaccccctggaacttcgaggaagtggtggacaagggcgcttcc gcccagagcttcatcgagcggatgaccaacttcgataagaacctgcccaacgaga aggtgctgcccaagcacagcctgctgtacgagtacttcaccgtgtataacgagctga ccaaagtgaaatacgtgaccgagggaatgagaaagcccgccttcctgagcggcga gcagaaaaaggccatcgtggacctgctgttcaagaccaaccggaaagtgaccgtg aagcagctgaaagaggactacttcaagaaaatcgagtgcttcgactccgtggaaat ctccggcgtggaagatcggttcaacgcctccctgggcacataccacgatctgctga aaattatcaaggacaaggacttcctggacaatgaggaaaacgaggacattctggaa gatatcgtgctgaccctgacactgtttgaggacagagagatgatcgaggaacggct gaaaacctatgcccacctgttcgacgacaaagtgatgaagcagctgaagcggcgg agatacaccggctggggcaggctgagccggaagctgatcaacggcatccgggac aagcagtccggcaagacaatcctggatttcctgaagtccgacggcttcgccaacag aaacttcatgcagctgatccacgacgacagcctgacctttaaagaggacatccaga aagcccaggtgtccggccagggcgatagcctgcacgagcacattgccaatctggc cggcagccccgccattaagaagggcatcctgcagacagtgaaggtggtggacgagctcgtgaaagtgatgggccggcacaagcccgagaacatcgtgatcgaaatggccagagagaaccagaccacccagaagggacagaagaacagccgcgagagaatgaAttorney Docket No.:TYA-072WO Description Sequence SEQ ID NO.agcggatcgaagagggcatcaaagagctgggcagccagatcctgaaagaacacc ccgtggaaaacacccagctgcagaacgagaagctgtacctgtactacctgcagaat gggcgggatatgtacgtggaccaggaactggacatcaaccggctgtccgactacg atgtggacgctatcgtgcctcagagctttctgaaggacgactccatcgacaacaagg tgctgaccagaagcgacaagaaccggggcaagagcgacaacgtgccctccgaa gaggtcgtgaagaagatgaagaactactggcggcagctgctgaacgccaagctga ttacccagagaaagttcgacaatctgaccaaggccgagagaggcggcctgagcga actggataaggccggcttcatcaagagacagctggtggaaacccggcagatcaca aagcacgtggcacagatcctggactcccggatgaacactaagtacgacgagaatg acaagctgatccgggaagtgaaagtgatcaccctgaagtccaagctggtgtccgatt tccggaaggatttccagttttacaaagtgcgcgagatcaacaactaccaccacgccc acgacgcctacctgaacgccgtcgtgggaaccgccctgatcaaaaagtaccctaa gctggaaagcgagttcgtgtacggcgactacaaggtgtacgacgtgcggaagatg atcgccaagtgcctgtcctacgagacagagatcctgacagtggagtatggcctgct gccaatcggcaagatcgtggagaagaggatcgagtgtaccgtgtactctgtggata acaatggcaacatctatacacagcccgtggcacagtggcacgataggggagagca ggaggtgttcgagtattgcctggaggacggcagcctgatcagggcaaccaaggac cacaagttcatgacagtggatggccagatgctgcccatcgacgagattttcgagcg ggagctggacctgatgagagtggataacctgcctaattctggcggctcaaaaagaa ccgccgacggcagcgaattcgagtctcccaagaagaagaggaaagtctaattgcc agccatctgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactccc actgtcctttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattct attctggggggtggggtggggcaggacagcaagggggaggattgggaagacaat agcaggcatgctggggactcgagcaaaaaagcaccgactcggtgccactttttcaa gttgataacggactagccttattttaacttgctatttctagctctaaaaccaagatcccat agtcccccacggtgtttcgtcctttccacaagatatataaagccaagaaatcgaaata ctttcaagttacggtaagcatatgatagtccattttaaaacataattttaaaactgcaaac tacccaagaaattattactttctacgtcacgtattttgtactaatatctttgtgtttacagtc aaattaattctaattatctctctaacagccttgtatcgtatatgcaaatatgaaggaatca tgggaaataggccctcgctagcaggaacccctagtgatggagttggccactccctct ctgcgcgctcgctcgctcactgaggccgggcgaccaaaggtcgcccgacgcccg ggctttgcccgggcggcctcagtgagcgagcgagcgcgcagFull PE-C AAV ctgcgcgctcgctcgctcactgaggccgcccgggcaaagcccgggcgtcgggcg 209 cassette sequence acctttggtcgcccggcctcagtgagcgagcgagcgcgcagagagggagtggcc aactccatcactaggggttcctggtaccgttgccttctgccccccaaccctgctccca gctggccctcccaggcctgggttgctggcctctgctttatcaggattctcaagaggg acagctggtttatgttgcatgactgttccctgcatatctgctctggttttaaatagcttatc tgagcagctggaggaccacatgggcttatatggcgtggggtacatgttcctgtagcc ttgtccctggcacctgccaaaatagcagccaacaccccccacccccaccgccatcc ccctgccccacccgtcccctgtcgcacattcctccctccgcagggctggctcacca ggccccagcccacatgcctgcttaaagccctctccatcctctgcctcacccagtccc cgctgagactgagcagacgcctccaggccaccatgaaacggacagccgacggaagcgagttcgagtcaccaaagaagaagcggaaagtcatcaagattgctacacggaaatacctgggaaagcagaacgtgtacgacatcggcgtggagcgggatcacaacttcAttorney Docket No.:TYA-072WO Description Sequence SEQ ID NO.gccctgaagaatggctttatcgccagcaattgtttcaacgaaatcggcaaggctacc gccaagtacttcttctacagcaacatcatgaactttttcaagaccgagattaccctggc caacggcgagatccggaagcggcctctgatcgagacaaacggcgaaaccgggg agatcgtgtgggataagggccgggattttgccaccgtgcggaaagtgctgagcatg ccccaagtgaatatcgtgaaaaagaccgaggtgcagacaggcggcttcagcaaag agtctatcctgcccaagaggaacagcgataagctgatcgccagaaagaaggactg ggaccctaagaagtacggcggcttcgacagccccaccgtggcctattctgtgctgg tggtggccaaagtggaaaagggcaagtccaagaaactgaagagtgtgaaagagct gctggggatcaccatcatggaaagaagcagcttcgagaagaatcccatcgactttct ggaagccaagggctacaaagaagtgaaaaaggacctgatcatcaagctgcctaag tactccctgttcgagctggaaaacggccggaagagaatgctggcctctgccggcga actgcagaagggaaacgaactggccctgccctccaaatatgtgaacttcctgtacct ggccagccactatgagaagctgaagggctcccccgaggataatgagcagaaaca gctgtttgtggaacagcacaagcactacctggacgagatcatcgagcagatcagcg agttctccaagagagtgatcctggccgacgctaatctggacaaagtgctgtccgcct acaacaagcaccgggataagcccatcagagagcaggccgagaatatcatccacct gtttaccctgaccaatctgggagcccctgccgccttcaagtactttgacaccaccatc gaccggaagaggtacaccagcaccaaagaggtgctggacgccaccctgatccac cagagcatcaccggcctgtacgagacacggatcgacctgtctcagctgggaggtg actccggcggaagctctggtggcagcaagcggaccgccgacggctctgaattcga gagccctaagaagaaaagaaaggtgagcggaggctctagcggcggaagcaccct gaacattgaagacgagtatagactgcatgaaacaagcaaggaacccgacgtgtcc ctgggctccacctggctgtccgactttccccaggcctgggccgagacaggaggaat gggcctggccgtgcggcaggcacccctgatcatccctctgaaggccacctctacac ccgtgagcatcaagcagtaccctatgtctcaggaggccagactgggcatcaagcct cacatccagaggctgctggaccagggcatcctggtgccatgccagagcccctgga acacaccactgctgcccgtgaagaagccaggcaccaatgactatagacccgtgca ggatctgagagaggtgaacaagagggtggaggatatccaccccaccgtgcccaa cccttacaatctgctgtccggcctgcccccttctcaccagtggtatacagtgctggac ctgaaggatgccttcttttgtctgagactgcaccctaccagccagccactgttcgcctt tgagtggagggaccctgagatgggcatctctggccagctgacctggacacgcctg cctcagggcttcaagaatagcccaacactgtttaacgaggccctgcaccgcgacct ggcagatttccggatccagcacccagatctgatcctgctgcagtacgtggacgatct gctgctggccgccaccagcgagctggattgccagcagggaacacgcgccctgct gcagaccctgggaaacctgggatatagggcatccgccaagaaggcccagatctgt cagaagcaggtgaagtacctgggctatctgctgaaggagggccagagatggctga cagaggccaggaaggagacagtgatgggccagccaacacccaagaccccaaga cagctgagggagttcctgggcaaagcaggattttgcaggctgttcatcccaggattc gcagagatggcagcacctctgtacccactgaccaagccgggcaccctgtttaattg gggccctgaccagcagaaggcctatcaggagatcaagcaggccctgctgacagc accagccctgggcctgccagacctgaccaagcctttcgagctgtttgtggatgagaa gcagggctacgccaagggcgtgctgacccagaagctgggaccatggagacggc ccgtggcctatctgtccaagaagctggacccagtggcagcaggatggccaccatgcctgaggatggtggcagcaatcgccgtgctgacaaaggatgccggcaagctgaccatgggacagccactggtcatcctggcaccacacgcagtggaggccctggtgaagcAttorney Docket No.:TYA-072WO Description Sequence SEQ ID NO.agcctccagatcgctggctgtctaacgcccggatgacacactaccaggccctgctg ctggacaccgatcgcgtgcagtttggccctgtggtggccctgaatccagccaccct gctgcctctgccagaggagggcctgcagcacaactgtctggactccggaggatcta gcggaggctcctctggctctgagacacctggcacaagcgagagcgcaacacctga aagcagcgggggcagcagcggggggtcaatggctgaaaatggtgataatgaaaa gatggctgccctggaggccaaaatctgtcatcaaattgagtattattttggcgacttca atttgccacgggacaagtttctaaaggaacagataaaactggatgaaggctgggtac ctttggagataatgataaaattcaacaggttgaaccgtctaacaacagactttaatgta attgtggaagcattgagcaaatccaaggcagaactcatggaaatcagtgaagataa aactaaaatcagaaggtctccaagcaaacccctacctgaagtgactgatgagtataa aaatgatgtaaaaaacagatctgtttatattaaaggcttcccaactgatgcaactcttga tgacataaaagaatggttagaagataaaggtcaagtactaaatattcagatgagaag aacattgcataaagcatttaagggatcaatttttgttgtgtttgatagcattgaatctgcta agaaatttgtagagacccctggccagaagtacaaagaaacagacctgctaatactttt caaggacgattactttgccaaaaaaaatgaatctggcggctcaaaaagaaccgccg acggcagcgaattcgagtctcccaagaagaagaggaaagtctaattgccagccatc tgttgtttgcccctcccccgtgccttccttgaccctggaaggtgccactcccactgtcc tttcctaataaaatgaggaaattgcatcgcattgtctgagtaggtgtcattctattctggg gggtggggtggggcaggacagcaagggggaggattgggaagacaatagcagg catgctggggactcgagcaaaaaattctagttggtttaacgcgtaactagatagaacc gcgtatagtgtggtatggtccagagcgaccacggtctcgaagtccaatgagcaccg actcggtgccactttttcaagttgataacggactagccttattttaacttgctatttctagc tctaaaacctctggaccatacctgggagcggtgtttcgtcctttccacaagatatataa agccaagaaatcgaaatactttcaagttacggtaagcatatgatagtccattttaaaac ataattttaaaactgcaaactacccaagaaattattactttctacgtcacgtattttgtact aatatctttgtgtttacagtcaaattaattctaattatctctctaacagccttgtatcgtatat gcaaatatgaaggaatcatgggaaataggccctcgctagcaggaacccctagtgat ggagttggccactccctctctgcgcgctcgctcgctcactgaggccgggcgaccaa aggtcgcccgacgcccgggctttgcccgggcggcctcagtgagcgagcgagcgcgcagSelf-inactivation cctcccaggtatggtccagagcggccacgctcgaagtccaatgagcta 210 sitepegRNA spacer ctcccaggtatggtccagag 211 pegRNA reverse tcattggacttcgagaccgtggtcgctctggaccatacc 212 transcriptiontemplate andprime binding site
Claims
Attorney Docket No.:TYA-072WO CLAIMS1. A genetically-modified, non-human animal comprising a humanized RNA binding motif protein 20 (RBM20) gene, wherein nucleic acids 250-264 of SEQ ID NO: 1 within the Rbm20 gene are replaced with a human RBM20 nucleic acid sequence, optionally wherein the non-human animal is a mouse.
2. The non-human animal of claim 1, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4.
3. The non-human animal of claim 2, wherein the humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) at an amino acid position corresponding to position 634 of human RBM20 protein.
4. The non-human animal of claim 1, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5.
5. The non-human animal of claim 4, wherein the humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q).
6. A genetically-modified, non-human animal comprising a humanized RNA binding motif protein 20 (RBM2O) gene, wherein at least 15 consecutive nucleic acids between 230-284 of SEQ ID NO: 1 are replaced with a human RBM20 nucleic acid sequence, optionally wherein the non-human animal is a mouse.
7. The non-human animal of claim 6, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 6.
8. The non-human animal of claim 6, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 7.
9. The non-human animal of claim 6, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 8.Attorney Docket No.:TYA-072WO 10. The non-human animal of claim 6, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 9.
11. The non-human animal of any one of claims 1-10, wherein the human RBM20 nucleic acid sequence replaces at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 120, at least 130, at least 140, at least 150, at least 160, at least 170, at least 180, at least 190, at least 200, at least 210, or at least 220 nucleotides upstream and / or downstream of nucleic acids 250-264 of SEQ ID NO: 1 within the murine Rbm20 gene.
12. The non-human animal of claim 11, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 4 and:(a) a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11-49 upstream of SEQ ID NO: 4; and / or(b) a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 4.
13. The non-human animal of claim 11, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 5 and:(a) a sequence comprising any one of “GGCCG” or SEQ ID NOs: 11 -49 upstream of SEQ ID NO: 5; and / or(b) a sequence comprising any one of “GGTGA” or SEQ ID NOs: 50-88 downstream of SEQ ID NO: 5.
14. The non-human animal of any one of claims 1-13, wherein the human RBM20 nucleic acid sequence replaces nucleic acids 31-294 of SEQ ID NO: 1 within the murine Rbm20 gene.
15. The non-human animal of claim 14, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 89.
16. The non-human animal of claim 14, wherein the human RBM20 nucleic acid sequence comprises the sequence of SEQ ID NO: 90.Attorney Docket No.:TYA-072WO 17. The non-human animal of any one of claims 1-16, wherein a first allele comprises a first humanized RBM20 gene and a second allele comprises a murine Rbm20 gene.
18. The non-human animal of any one of claim 1-16, wherein a first allele comprises a first humanized RBM20 gene and a second allele comprises a second humanized RBM20 gene.
19. The non-human animal of claim 17 or 18, wherein the first humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) at an amino acid position corresponding to position 634 of human RBM20 protein.
20. The non-human animal of claim 17 or 18, wherein the first humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q).
21. The non-human animal of claim 18, wherein the second humanized RBM20 gene encodes an RBM20 protein comprising an arginine (R) to glutamine (Q) mutation at an amino acid position corresponding to position 634 of human RBM20 protein (R634Q).
22. The non-human animal of any one of claims 1-21, wherein a human RBM20-targeting guide RNA specifically binds to the humanized RBM20 gene.
23. The non-human animal of any one of claims 17-22, wherein the humanized RBM20 gene allows targeting of a nucleic acid-guided nuclease to the RBM20 nucleotide sequence using a human RBM20-targeting guide RNA.
24. The non-human animal of any one of claims 1-23, wherein the endogenous and human RBM20 sequences encode a humanized RBM20 protein with substantially similar activity to a wild-type RBM20 protein.
25. The non-human animal of any one of claims 1-24, wherein the animal has substantially similar ejection fraction (EF), left ventricle internal diameter at systolic stage (LVID;s), and / or left ventricle internal diameters at diastolic stage (LVID;d) to a non-human animal comprising two wild-type RBM20 alleles.Attorney Docket No.:TYA-072WO 26. The non-human animal of any one of claims 4-18 or 20-23, wherein the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, and wherein the humanized RBM20 protein comprising an R634Q mutation has altered activity compared to a wild-type RBM20 protein.
27. The non-human animal of any one of claims 4-18, 20-23, and 26, wherein the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, and wherein the animal has reduced ejection fraction (EF), increased left ventricle internal diameters at systolic stage (LVID;s), and / or increased left ventricle internal diameter at diastolic stage (LVID;d) compared to a non-human animal comprising two wild-type RBM20 alleles.
28. The non-human animal of any one of claims 4-18, 20-23, 26, and 27, wherein the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation, and wherein the animal has increased levels of Nppa transcripts, Nppb transcripts, natriuretic peptide A, and / or natriuretic peptide B.
29. A method of determining the efficacy of a gene editing system targeting the human RBM20 gene, comprising:(a) contacting a cell of the non-human animal of any one of claims 1-28 expressing a humanized RNA binding motif protein 20 (RBM20) gene with the gene editing system, wherein the gene editing system comprises a guide RNA comprising a spacer sequence that binds to a target sequence of the humanized RBM20 gene and an RNA-guided nuclease.
30. The method of claim 29, further comprising:(b) analyzing the humanized RBM20 gene by NGS to determine the presence or absence of editing events in the humanized RBM20 gene, wherein the gene editing system is determined to be efficacious by the presence of editing events; and / or(c) monitoring the heart functions of the animal, wherein the gene editing system is determined to be efficacious by the improvement of heart functions.
31. The method of claim 29 or 30, wherein the cell is in vivo or ex vivo.
32. The method of claim 29 or 30, wherein the guide RNA comprises the sequence of any one of SEQ ID NOs: 95, 96, 98-103. 195, 199, and 200.Attorney Docket No.:TYA-072WO 33. The method of any one of claims 29-32, wherein the guide RNA comprises the sequence of SEQ ID NO: 96 and / or SEQ ID NO: 103.
34. The method of any one of claims 29-32, wherein the guide RNA comprises the sequence of SEQ ID NO: 96 and / or SEQ ID NO: 199.
35. The method of claim 33, wherein the guide RNA comprises the sequence of SEQ ID NO: 109.
36. The method of claim 34, wherein the guide RNA comprises the sequence of SEQ ID NO: 187.
37. A population of cardiomyocytes isolated from the non-human animal of any one of claims 1-28.
38. A guide RNA polynucleotide comprising (a) a spacer sequence; (b) a scaffold sequence; and (c) a template + primer binding site sequence, wherein the spacer sequence binds a target sequence within a human RBM20 gene or a humanized RBM20 gene.
39. The guide RNA of claim 38, wherein the target sequence is within a mutant human RBM20 gene or a mutant humanized RBM20 gene.
40. The guide RNA of claim 38 or 39, wherein the human RBM20 gene or the humanized RBM20 gene comprises the sequence of SEQ ID NO: 5.
41. The guide RNA of any one of claims 38-40, wherein the human RBM20 gene or the humanized RBM20 gene encodes an RBM20 protein comprising an R634Q mutation.
42. The guide RNA of any one of claim 38-41, wherein the human RBM20 gene or the humanized RBM20 gene comprises the sequence of SEQ ID NO: 7 or SEQ ID NO: 9.
43. The guide RNA of any one of claims 38-42, wherein the spacer sequence comprises the nucleotide sequence of any one of SEQ ID NOs: 95, 96, and 195.
44. The guide RNA of claim 43, wherein the spacer sequence comprises the sequence of SEQ ID NO: 96.Attorney Docket No.:TYA-072WO 45. The guide RNA of any one of claims 38-44, wherein the scaffold sequence comprises the nucleotide sequence of SEQ ID NO: 97.
46. The guide RNA of any one of claims 38-45, wherein the template + primer binding site sequence comprises the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199, and 200.
47. The guide RNA of claim 46, wherein template + primer binding site sequence comprises the nucleotide sequence of SEQ ID NO: 103.
48. The guide RNA of claim 46, wherein template + primer binding site sequence comprises the nucleotide sequence of SEQ ID NO: 199.
49. The guide RNA of any one of claims 38-45, wherein the guide RNA comprises:(a) a spacer comprising the nucleotide sequence of SEQ ID NO: 95 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 98;(b) a spacer comprising the nucleotide sequence of SEQ ID NO: 95 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 99;(c) a spacer comprising the nucleotide sequence of SEQ ID NO: 95 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 100;(d) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 101;(e) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 102;(f) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 103; or (g) a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 199.
50. The guide RNA of claim 49, wherein the guide RNA comprises a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 103.Attorney Docket No.:TYA-072WO 51. The guide RNA of claim 49, wherein the guide RNA comprises a spacer comprising the nucleotide sequence of SEQ ID NO: 96 and a template + PBS sequence comprising the nucleotide sequence of SEQ ID NO: 199.
52. The guide RNA of any one of claims 38-51, wherein the guide RNA further comprises a linker and a 3’ motif, wherein the linker and 3’ motif comprise the nucleotide sequence of any one of SEQ ID NOs: 126 and 202-204.
53. The guide RNA of any one of claims 38-52, wherein the guide RNA comprises the nucleotide sequence of any one of SEQ ID NOs: 104-109, 187, and 188.
54. The guide RNA of claim 53, wherein the guide RNA comprises the nucleotide sequence of SEQ ID NO: 109.
55. The guide RNA of claim 53, wherein the guide RNA comprises the nucleotide sequence of SEQ ID NO: 187.
56. A gene editing system comprising the guide RNA of any one of claims 38-55 and an effector protein.
57. The gene editing system of claim 56, wherein the effector protein is a nucleic acid-guided nuclease.
58. The gene editing system of claim 57, wherein the nucleic acid-guided nuclease is a Cas nuclease.
59. The gene editing system of claim 57, wherein the nucleic acid-guided nuclease is an RNA-guided nickase.
60. The gene editing system of claim 56, wherein the effector protein comprises an RNA-guided nickase and a reverse transcriptase.
61. The gene editing system of claim 60, wherein the effector protein is a fusion protein comprising an RNA-guided nickase and a reverse transcriptase.Attorney Docket No.:TYA-072WO 62. The gene editing system of claim 61, wherein the fusion protein comprises:(a) an N-terminal fragment comprising an N-terminal fragment of an RNA-guided nickase and a split intein domain; and(b) a C-terminal fragment comprising a C-terminal fragment of the RNA-guided nickase, a reverse transcriptase, and a split intein domain.
63. The gene editing system of any one of claims 56-62, wherein the system further comprises a nicking guide RNA (ngRNA).
64. The gene editing system of claim 63, wherein the ngRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 110-117 and 194.
65. The gene editing system of claim 64, wherein the ngRNA comprises the nucleotide sequence of SEQ ID NO: 113.
66. The gene editing system of claim 64, wherein the ngRNA comprises the nucleotide sequence of SEQ ID NO: 194.
67. The gene editing system of any one of claims 63-66, wherein the ngRNA comprises the nucleotide sequence of SEQ ID NO: 118-125 and 193.
68. The gene editing system of claim 67, wherein the ngRNA comprises the nucleotide sequence of SEQ ID NO: 121.
69. The gene editing system of claim 67, wherein the ngRNA comprises the nucleotide sequence of SEQ ID NO: 193.
70. A vector comprising a nucleotide sequence encoding (a) the gRNA of any one of claims 38-55; or (b) the gene editing system of any one of claims 56-69 or a portion thereof.
71. The vector of claim 70, wherein the vector is an AAV viral vector.
72. A method for expressing a polynucleotide in a target cell, comprising contacting the target cell with the gene editing system of any one of claims 56-69 or transducing the target cell with one or more vectors of claim 70 or 71.Attorney Docket No.:TYA-072WO 73. The method of claim 72, wherein the target cell is a cardiac cell, a muscle cell, an induced pluripotent stem cell-derived cardiomyocyte (iPSC-CM), a cardiomyocyte, or an iPSC.
74. The method of claim 72 or 73, wherein the target cell comprises at least one mutant human RBM20 allele or at least one mutant humanized RBM20 allele.
75. The method of claim 74, wherein the mutant human RBM20 allele or mutant humanized RBM20 allele encodes an RBM20 protein comprising an R634Q mutation.
76. The method of claim 74 or 75, wherein the method results in editing of the mutant human RBM20 allele or mutant humanized RBM20 allele.
77. The method of any one of claims 74-76, wherein the method does not, or substantially does not. edit the wild-type human RBM20 allele or wild-type humanized RBM20 allele.
78. A method for preventing and / or treating a disease or condition in a subject in need thereof, comprising administering to the subject the gene editing system of any one of claims 56-69 or one or more vectors of claim 70 or 71, wherein the disease or condition is a cardiac pathology.
79. The method of claim 78, wherein the subject comprises at least one mutant human RBM20 allele or at least one mutant humanized RBM20 allele.
80. The method of claim 79, wherein the mutant human RBM20 allele or mutant humanized RBM20 allele encodes an RBM20 protein comprising an R634Q mutation.
81. The method of claim 79 or 80, wherein the method results in editing of the mutant human RBM20 allele or mutant humanized RBM20 allele.
82. The method of any one of claims 79-81, wherein the method does not, or substantially does not, edit the wild-type human RBM20 allele or wild-type humanized RBM20 allele.
83. A method for identifying guide RNAs suitable for gene editing systems, wherein the guide RNAs allow efficient editing of a human RBM20 gene or a humanized RBM20 geneAttorney Docket No.:TYA-072WO 84. The method of claim 83, wherein the guide RNAs provide editing with an efficiency of at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%. or at least 70%.
85. The method of claim 83 or 84, wherein the guide RNAs are pegRNAs and / or ngRNAs.
86. A nicking guide RNA (ngRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 110-117 and 194.
87. The ngRNA of claim 86, comprising the nucleotide sequence of SEQ ID NO: 113.
88. The ngRNA of claim 86, comprising the nucleotide sequence of SEQ ID NO: 194.
89. The ngRNA of claim 86, wherein the ngRNA comprises the nucleotide sequence of any one of SEQ ID NOs: 118-125 and 193.
90. The ngRNA of claim 89, comprising the nucleotide sequence of SEQ ID NO: 121.
91. The ngRNA of claim 89, comprising the nucleotide sequence of SEQ ID NO: 193.
92. A gene editing system comprising:(a) a first expression cassette comprising:(i) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N- terminal fragment of a split-intein; and(ii) a second polynucleotide encoding a nicking guide RNA (ngRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 110-117 and 194; and (b) a second expression cassette comprising:(iii) a third polynucleotide encoding a C-terminal fragment of a fusion protein comprising a C-terminal fragment of an RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein; and(iv) a fourth polynucleotide encoding a prime editing guide RNA (pegRNA) comprising the nucleotide sequence of any one of SEQ ID NOs: 98-103, 199. and 200.Attorney Docket No.:TYA-072WO 93. The gene editing system of claim 92, wherein the ngRNA comprises the sequence of any one of SEQ ID NOs: 118-125 and 193, and the pegRNA comprises the sequence of any one of SEQ ID NOs: 104-109. 187, and 188.
94. The gene editing system of claim 92 or 93, wherein the N-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 129 and the C-terminal fragment of the fusion protein comprises the sequence of SEQ ID NO: 132.
95. The gene editing system of claim 94, wherein:(a) the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 177; or(b) the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 178.
96. The gene editing system of claim 94 or 95, wherein:(a) the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 180; or(b) the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 181.
97. The gene editing system of claim 94, wherein:(a) the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 189; or(b) the first expression cassette comprises the sequence of SEQ ID NO: 176 and the second expression cassette comprises the sequence of SEQ ID NO: 190.
98. The gene editing system of claim 94 or 97, wherein:(a) the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 191; or(b) the first expression cassette comprises the sequence of SEQ ID NO: 179 and the second expression cassette comprises the sequence of SEQ ID NO: 192.Attorney Docket No.:TYA-072WO 99. A gene editing system comprising a first expression cassette comprising the sequence of SEQ ID NO: 176 and a second expression cassette comprising the sequence of SEQ ID NO: 177 or SEQ ID NO: 178.
100. A gene editing system comprising a first expression cassette comprising the sequence of SEQ ID NO: 179 and a second expression cassette comprising the sequence of SEQ ID NO: 180 or SEQ ID NO: 181.
101. A gene editing system comprising a first expression cassette comprising the sequence of SEQ ID NO: 176 and a second expression cassette comprising the sequence of SEQ ID NO: 189 or SEQ ID NO: 190.
102. A gene editing system comprising a first expression cassette comprising the sequence of SEQ ID NO: 179 and a second expression cassette comprising the sequence of SEQ ID NO: 191 or SEQ ID NO: 192.
103. A system comprising:(a) a first vector comprising a first expression cassette comprising the sequence of SEQ ID NO: 176; and(b) a second vector comprising a second expression cassette comprising the sequence of any one of SEQ ID NOs: 177, 178, 189, and 190.
104. A system comprising:(a) a first vector comprising a first expression cassette comprising the sequence of SEQ ID NO: 179; and(b) a second vector comprising a second expression cassette comprising the sequence of any one of SEQ ID NOs: 180, 181, 191, and 192.
105. The system of claim 103 or 104, wherein the first vector and the second vector are viral vectors.
106. The system of claim 105, wherein the viral vectors are AAV vectors.Attorney Docket No.:TYA-072WO 107. A cell comprising the guide RNA of any one of claims 38-55, the gene editing system of any one of claims 56-69 and 92-102 or a portion thereof, the vector of claim 70 or 71, or the system of any one of claims 103-106 or a portion thereof.
108. The cell of claim 107, wherein the cell is a cardiac cell, a muscle cell, an iPSC-CM, a cardiomyocyte, or an iPSC.
109. A cell produced by the method of any one of claims 72-77.
110. The cell of claim 109, wherein the cell is a cardiac cell, a muscle cell, an iPSC-CM, a cardiomyocyte, or an iPSC.
111. A gene editing system capable of editing a target genomic locus and self-inactivation, the system comprising:(a) one or more polynucleotides encoding a prime editor protein or fragments thereof, wherein at least one polynucleotide comprises a self-inactivation site positioned upstream of or within a coding region; and(b) a polynucleotide encoding a pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site;wherein editing of the self-inactivation site by the pegRNA introduces one or more protein coding changes that reduce or eliminate expression or function of the prime editor protein.
112. The gene editing system of claim 111, wherein the one or more polynucleotides encoding the prime editor protein or fragments thereof is located in a single vector.
113. The gene editing system of claim 111 or claim 112, wherein the one or more polynucleotides encoding the prime editor protein or fragments thereof and the polynucleotide encoding the pegRNA that recognizes and edits both the target genomic locus and the self-inactivation site are located in the same vector.
114. The gene editing system of claim 111, wherein the one or more polynucleotides encoding the prime editor protein or fragments thereof and the polynucleotide encoding the pegRNA thatAttorney Docket No.:TYA-072WO recognizes and edits both the target genomic locus and the self-inactivation site are located in at least two separate vectors.
115. The gene editing system of any one of claims 111-114, wherein the self-inactivation site encodes a peptide sequence of 1 to 100 amino acids that does not substantially impair expression or function of the prime editor protein prior to editing of the self-inactivation site.
116. The gene editing system of any one of claims 111-115, wherein the one or more protein coding changes comprise a frameshift mutation, a premature stop codon, an amino acid change, or a combination thereof.
117. The gene editing system of any one of claims 111-116, wherein the self-inactivation site is positioned within the coding region within 500 nucleotides downstream of a start codon.
118. The gene editing system of any one of claims 111-117, wherein editing of the self-inactivation site reduces or eliminates expression or function of the prime editor protein by at least at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or at least 99%.
119. The gene editing system of any one of claims 114-118, wherein the one or more polynucleotides encoding the prime editor protein or fragments thereof comprise:(a) a first polynucleotide encoding an N-terminal fragment of a fusion protein comprising an N-terminal fragment of an RNA-guided nickase and an N-terminal fragment of a split-intein; and (b) a second polynucleotide encoding a C-terminal fragment of the fusion protein comprising a C-terminal fragment of the RNA-guided nickase, a polymerase, and a C-terminal fragment of a split-intein.
120. The gene editing system of claim 119, wherein the first polynucleotide and / or the second polynucleotide comprises a self-inactivation site.
121. The gene editing system of claim 119, wherein both the first polynucleotide and the second polynucleotide comprise self-inactivation sites.Attorney Docket No.:TYA-072WO 122. The gene editing system of any one of claims 119-121, wherein the first polynucleotide and the polynucleotide encoding the pegRNA are packaged in a first AAV vector, and the second polynucleotide is packaged in a second AAV vector.
123. The gene editing system of claim 122, wherein the first AAV vector and the second AAV vector are delivered to cardiac tissue.
124. The gene editing system of any one of claims 111-123, wherein the target genomic locus is the RBM20 locus.
125. The gene editing system of any one of claims 92-102, wherein at least one of the first polynucleotide or the third polynucleotide comprises a self-inactivation site that is recognized and edited by the pegRNA, and wherein editing of the self-inactivation site introduces one or more protein coding changes that reduce or eliminate expression or function of the gene editing system.
126. The gene editing system of claim 125, wherein the self-inactivation site is positioned within the coding region within 500 nucleotides downstream of a start codon and encodes a peptide of 1 to 100 amino acids.
127. The gene editing system of claim 125 or 126, wherein editing of the self-inactivation site introduces a frameshift mutation, one or more premature stop codons, one or more amino acid changes, or a combination thereof.
128. The system of any one of claims 103-106, wherein at least one of the first expression cassette or the second expression cassette comprises a self-inactivation site that is recognized and edited by the pegRNA, wherein editing of the self-inactivation site introduces one or more protein coding changes that reduce or eliminate expression or function of the gene editing system.
129. The gene editing system or system of any one of claims 111-128, for use in treating a genetic disorder in a subject.