Engineered constructs for improved RNA payload transcription
By providing an expression cassette containing a specific promoter, payload sequence and termination sequence, the problem of low transcription and expression efficiency of RNA payload in the prior art is solved, and the effect of improving the expression level of RNA payload is achieved, and it has potential application value for the treatment of gene mutation diseases.
Patent Information
- Application Number
- CN202380074917.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-05-15
- Filing Date
- 2023-08-23
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively improve the transcription and expression of RNA payload, especially in the treatment of diseases caused by gene mutations.
An expression cassette is provided that comprises a promoter, payload sequence and termination sequence of a specific sequence for improving transcription and expression of RNA payload. The expression cassette includes a specific promoter sequence, such as SEQ ID NO: 17, SEQ ID NO: 1250 or SEQ ID NO: 1262, a payload sequence controlled by transcription of these sequences, contains a small RNA payload, and a termination sequence includes a specific sequence, such as SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265.
By using this expression cassette, the expression level of RNA payload can be effectively increased, thereby potentially used to treat diseases caused by gene mutations.
Smart Images

Figure CN120112640A_ABST
Abstract
Description
[0001] Cross-references
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 400,583 filed on August 24, 2022 (invention title: “Engineered constructs for increasing transcription of RNA payloads”), U.S. Provisional Application No. 63 / 419,889 filed on October 27, 2022 (invention title: “Engineered constructs for increasing transcription of RNA payloads”), U.S. Provisional Application No. 63 / 453,584 filed on March 21, 2023 (invention title: “Engineered constructs for increasing transcription of RNA payloads”), and U.S. Provisional Application No. 63 / 466,625 filed on May 15, 2023 (invention title: “Engineered constructs for increasing transcription of RNA payloads”), all of which are incorporated herein by reference in their entirety.
[0003] Sequence Listing
[0004] This application contains a sequence listing, which has been submitted electronically in extensible markup language (XML) format and is hereby incorporated by reference in its entirety. The XML copy was created on August 21, 2023, is named "421688-712021_SL.xml", and is 1.29 megabytes in size. Background Art
[0005] Many diseases and conditions are caused by mutation, deletion, altered expression or splicing of genes. RNA can be used as a mechanism for gene therapy, such as by editing mutated RNA sequences associated with a disease. An expression cassette is required to increase or regulate the expression of the RNA payload. Summary of the invention
[0006] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1007 ... NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269.
[0007] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence that is transcriptionally controlled by the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence.
[0008] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any one of the following: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1007, SEQ ID NO: 1021, SEQ ID NO: 1242, SEQ ID NO: 1254, SEQ ID NO: 1255, SEQ ID NO: 1257 or SEQ ID NO: 1269.
[0009] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any of the following: SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of the following: SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: 1257, SEQ ID NO: 1264-SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: NO:1287–SEQ ID NO:1289.
[0010] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any one of: SEQ ID NO:13–SEQ ID NO:17, SEQ ID NO:167–SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248–SEQ ID NO:1253, or SEQ ID NO:1259–SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence.
[0011] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of the following: SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289.
[0012] In some aspects, the promoter sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263. In some aspects, the promoter sequence comprises a sequence having at least 95% sequence identity to any one of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263.
[0013] In some aspects, the termination sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289. In some aspects, the termination sequence comprises a sequence having at least 95% sequence identity to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289.
[0014] In some aspects, the promoter sequence comprises SEQ ID NO: 17. In some aspects, the promoter sequence comprises SEQ ID NO: 1262. In some aspects, the promoter sequence comprises SEQ ID NO: 1250. In some aspects, the promoter sequence comprises SEQ ID NO: 1251. In some aspects, the promoter sequence comprises SEQ ID NO: 1252. In some aspects, the promoter sequence comprises SEQ ID NO: 1253.
[0015] In some aspects, the termination sequence comprises SEQ ID NO: 1264. In some aspects, the termination sequence comprises SEQ ID NO: 1265. In some aspects, the termination sequence comprises SEQ ID NO: 1254. In some aspects, the termination sequence comprises SEQ ID NO: 1255. In some aspects, the termination sequence comprises SEQ ID NO: 1257. In some aspects, the termination sequence comprises SEQ ID NO: 60. In some aspects, the termination sequence comprises SEQ ID NO: 1242. In some aspects, the termination sequence comprises SEQ ID NO: 1269. In some aspects, the termination sequence comprises SEQ ID NO: 1017.
[0016] In some aspects, the small RNA payload comprises an engineered guide RNA capable of hybridizing with a target sequence. In some aspects, the engineered guide RNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% reverse complementarity with the target sequence. In some aspects, the engineered guide RNA comprises at least one base pair mismatch relative to the target sequence. In some aspects, the target sequence comprises an adenosine residue. In some aspects, the target sequence is an RNA sequence. In some aspects, the RNA sequence is an mRNA or pre-mRNA.
[0017] In some aspects, the target sequence comprises a G to A mutation relative to the wild-type sequence. In some aspects, the target sequence comprises a missense mutation or a nonsense mutation relative to the wild-type sequence. In some aspects, the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with Charcot-Marie-Tooth disease type 1A (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2).
[0018] In some aspects, the payload sequence has at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity with SEQ ID NO: 1273, SEQ ID NO: 1274 or SEQ ID NO: 61. In some aspects, the small RNA payload includes antisense oligonucleotides, siRNA, shRNA, miRNA or tracrRNA. In some aspects, the length of the small RNA payload is not less than 20 nucleotide residues and not more than 500 nucleotide residues. In some aspects, the length of the small RNA payload is not less than 60 residues and not more than 100 residues. In some aspects, the length of the small RNA payload is not less than 80 residues and not more than 120 residues. In some aspects, the length of the small RNA payload is not less than 100 residues and not more than 140 residues. In some aspects, the length of the small RNA payload is not less than 130 residues and not more than 170 residues. In some aspects, the payload sequence also includes an Sm binding sequence or a hairpin sequence. In some aspects, the hairpin sequence comprises a U7 hairpin. In some aspects, the hairpin sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:52 or SEQ ID NO:54, or the Sm binding sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:56 or SEQ ID NO:58.
[0019] In some aspects, the length of the expression cassette is not less than 1300 nucleotide residues and no more than 2160 nucleotide residues. In some aspects, the expression cassette has at least 80% sequence identity with a U1 sequence or a U7 sequence. In some aspects, the U1 sequence is a mouse U1 sequence or a human U1 sequence. In some aspects, the U7 sequence is a mouse U7 sequence or a human U7 sequence.
[0020] In some aspects, the promoter sequence comprises a zinc finger 143 motif capable of recruiting a ZNF143 transcription factor. In some aspects, the promoter sequence comprises an OCT-1 transcription factor binding sequence capable of recruiting an OCT-1 transcription factor. In some aspects, the promoter sequence comprises a proximal sequence element capable of recruiting SNAPc. In some aspects, the proximal sequence element is capable of integron-dependent recruitment of RNA polymerase II.
[0021] In some aspects, the small RNA payload is capable of forming a guide-target RNA scaffold comprising a structural feature after hybridization of the small RNA payload to the target sequence. In some aspects, the structural feature is a protrusion, a mismatch, an inner loop, a hairpin, or a combination thereof. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is a symmetrical protrusion. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is an asymmetric protrusion. In some aspects, the structural feature comprises an inner loop, and wherein the inner loop is a symmetrical inner loop. In some aspects, the structural feature comprises an inner loop, and wherein the inner loop is an asymmetric inner loop. In some aspects, the structural feature comprises a hairpin, and wherein the hairpin is a recruiting hairpin or a non-recruiting hairpin. In some aspects, the guide-target RNA scaffold comprises a wobble base pair.
[0022] In various aspects, the present disclosure provides recombinant polynucleotides encoding one or more expression cassettes as described herein.
[0023] In some aspects, the recombinant polynucleotide encodes two expression cassettes as described herein, comprising a first promoter, a second promoter, a first termination sequence, and a second termination sequence. In some aspects, the first promoter and the second promoter are identical. In some aspects, the first promoter and the second promoter are different. In some aspects, the first termination sequence and the second termination sequence are identical. In some aspects, the first termination sequence and the second termination sequence are different. In some aspects, the first promoter comprises SEQ ID NO: 17. In some aspects, the second promoter comprises SEQ ID NO: 1262. In some aspects, the first termination sequence comprises SEQ ID NO: 1264. In some aspects, the second termination sequence comprises SEQ ID NO: 1265. In some aspects, (a) the first promoter sequence includes SEQ ID NO:17, the first termination sequence includes SEQ ID NO:1264, the second promoter sequence includes SEQ ID NO:1262, and the second termination sequence includes SEQ ID NO:1265; or (b) the first promoter sequence includes SEQ ID NO:17, the first termination sequence includes SEQ ID NO:1265, the second promoter sequence includes SEQ ID NO:1262, and the second termination sequence includes SEQ ID NO:1264.
[0024] In various aspects, the present disclosure provides viral vectors that encapsidate an expression cassette as described herein or a recombinant polynucleotide as described herein.
[0025] In some aspects, the viral vector comprises two or more, three or more, or four or more expression cassettes as described herein. In some aspects, the viral vector is an adeno-associated viral vector. In some embodiments, the adeno-associated viral vector is selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV17, AAV18, AAV19, AAV20, AAV21, AAV32, AAV40, AAV53, AAV6, AAV7, AAV8, AAV19, AAV22, AAV33, AAV40, AAV54, AAV6, AAV7, AAV8 V7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-DJ, AAV-DJ / 8, AAV-DJ / 9, AAV1 / 2, AAV.rh8, AAV.rh10, AAV.rh20, AAV.rh39, AAV.Rh43, AAV.Rh74, AAV.v66, AAV.Oligo001, AAV.SCH9, AAV.r3.45, AAV.RHM4-1, AAV.hu37, AAV.Anc80, AAV.Anc80L65, AAV.7m8, AAV.PhP.eB, AAV.PhP.V1, AAV.PHP.B , AAV.PhB.C1, AAV.PhB.C2, AAV.PhB.C3, AAV.PhB.C6, AAV.cy5, AAV2.5, AAV2tYF, AAV3B, AAV.LK03, AAV.HSC1, AAV.HSC2, AAV.HSC3, AAV.HSC4, AAV.HSC5, AAV.HSC6, AAV.HSC7, AAV.HSC8, AAV.HSC9, AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAV.HSC16, AAV.HSC17, AAVhu68, chimeras thereof, and combinations thereof.
[0026] In various aspects, the present disclosure provides a pharmaceutical composition comprising an expression cassette as described herein, a recombinant polynucleotide as described herein, or a viral vector as described herein and a pharmaceutically acceptable excipient, carrier, diluent, or a combination thereof.
[0027] In various aspects, the present disclosure provides a method of expressing a small RNA payload in a cell, the method comprising delivering an expression cassette as described herein, a recombinant polynucleotide as described herein, a viral vector as described herein, or a pharmaceutical composition as described herein to the cell and expressing the small RNA payload encoded by the expression cassette in the cell.
[0028] In various aspects, the present disclosure provides a method for editing a target sequence, the method comprising: delivering an expression cassette as described herein, a recombinant polynucleotide as described herein, a viral vector as described herein, or a pharmaceutical composition as described herein to a cell encoding a target sequence; expressing a small RNA payload in the cell, wherein the small RNA payload comprises an engineered guide RNA capable of hybridizing with a target sequence; forming a guide-target RNA scaffold after hybridization of the small RNA payload to the target sequence; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0029] In various aspects, the present disclosure provides a method of editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any of the following: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% identity to any of the following: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 8 NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269; expressing a small RNA payload in the cell; forming a guide-target RNA scaffold after the small RNA payload hybridizes with the target sequence; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0030] In various aspects, the present disclosure provides a method for editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any one of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence; expressing the small RNA payload in the cell; the small RNA payload hybridizing to the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0031] In various aspects, the present disclosure provides a method of editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% identity to any of the following: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1007, SEQ ID NO: 1021, SEQ ID NO: 1242, SEQ ID NO: 1254, SEQ ID NO: 1255, SEQ ID NO: 1257, or SEQ ID NO:1269; expressing a small RNA payload in the cell; the small RNA payload hybridizes with a target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0032] In various aspects, the present disclosure provides a method of editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any one of: SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% identity to any one of: SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: 1257, SEQ ID NO: 1268, SEQ ID NO: 1270, SEQ ID NO: 1271, SEQ ID NO: 1272, SEQ ID NO: 1273, SEQ ID NO: 1274, SEQ ID NO: 1275 NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289; expressing a small RNA payload in the cell; forming a guide-target RNA scaffold after hybridization of the small RNA payload to a target sequence; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0033] In various aspects, the present disclosure provides a method for editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any one of the following: SEQ ID NO:13-SEQ ID NO:17, SEQ ID NO:167-SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248-SEQ ID NO:1253, or SEQ ID NO:1259-SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence; expressing the small RNA payload in the cell; the small RNA payload hybridizing to the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0034] In various aspects, the present disclosure provides a method for editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence, the termination sequence comprising a sequence having at least 80% identity to any one of the following: SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275 or SEQ ID NO:1287-SEQ ID NO:1289; expressing the small RNA payload in the cell; the small RNA payload hybridizing to the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0035] In some aspects, the promoter sequence comprises SEQ ID NO: 17. In some aspects, the promoter sequence comprises SEQ ID NO: 1262. In some aspects, the promoter sequence comprises SEQ ID NO: 1250. In some aspects, the promoter sequence comprises SEQ ID NO: 1251. In some aspects, the promoter sequence comprises SEQ ID NO: 1252. In some aspects, the promoter sequence comprises SEQ ID NO: 1253.
[0036] In some aspects, the termination sequence comprises SEQ ID NO: 1264. In some aspects, the termination sequence comprises SEQ ID NO: 1265. In some aspects, the termination sequence comprises SEQ ID NO: 1254. In some aspects, the termination sequence comprises SEQ ID NO: 1255. In some aspects, the termination sequence comprises SEQ ID NO: 1257. In some aspects, the termination sequence comprises SEQ ID NO: 60. In some aspects, the termination sequence comprises SEQ ID NO: 1242. In some aspects, the termination sequence comprises SEQ ID NO: 1269. In some aspects, the termination sequence comprises SEQ ID NO: 1017.
[0037] In some aspects, the target sequence comprises a mutation relative to the wild-type sequence. In some aspects, editing the target sequence corrects the mutation in the target sequence. In some aspects, the mutation is a missense mutation. In some aspects, the mutation is a nonsense mutation. In some aspects, the mutation is a G to A mutation. In some aspects, the mutation is associated with a disease. In some aspects, the disease is a synucleinopathy, Parkinson's disease, dementia with Lewy bodies, multiple system atrophy, Charcot-Marie-Tooth disease, hereditary compressive susceptibility neuropathy, Yuan-Harel-Lupski syndrome, Tauopathy, Alzheimer's disease, frontotemporal dementia, progressive supranuclear palsy, corticobasal degeneration, chronic traumatic encephalopathy, autism, traumatic brain injury, Dravet syndrome, Crohn's disease, muscular dystrophy, B-cell leukemia, Dejerine-Sottas disease, Stargardt disease, alpha-1 antitrypsin deficiency, Tay-Sachs disease, cystic fibrosis, liposomal acid lipase deficiency, or Gaucher disease.
[0038] In some aspects, the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with Charcot-Marie-Tooth disease type 1A (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2).
[0039] In some aspects, editing the target sequence comprises editing an untranslated region of the target. In some aspects, the untranslated region is a 5' untranslated region or a 3' untranslated region. In some aspects, the 3' untranslated region is a polyadenylation sequence. In some aspects, editing the target sequence comprises editing a translation start site.
[0040] In some aspects, editing the target sequence changes the expression of the target sequence. In some aspects, editing the target sequence increases the expression of the target sequence. In some aspects, editing the target sequence decreases the expression of the target sequence.
[0041] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette as described herein, a recombinant polynucleotide as described herein, a viral vector as described herein, or a pharmaceutical composition as described herein; delivering the expression cassette to cells of the subject; and expressing a small RNA payload in the cells, thereby treating the disease.
[0042] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 772, SEQ ID NO: 773, SEQ ID NO: 774, SEQ ID NO: 775, SEQ ID NO: 776, SEQ ID NO: 777, SEQ ID NO: 778, SEQ ID NO: 779, SEQ ID NO: 781, SEQ ID NO: 783, SEQ ID NO: 784, SEQ ID NO: 785, SEQ ID NO: 786, SEQ ID NO: 787, SEQ ID NO: 788, SEQ ID NO: 789, SEQ ID NO: 790, SEQ ID NO: 801, SEQ ID NO: 802, SEQ ID NO: 803, SEQ ID NO: 804, SEQ ID NO: 805 NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0043] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0044] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence, which comprises a sequence having at least 80% sequence identity to any of the following: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1007, SEQ ID NO: 1021, SEQ ID NO: 1242, SEQ ID NO: 1254, SEQ ID NO: 1255, SEQ ID NO: 1257, or SEQ ID NO: 1269; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0045] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any of the following: SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of the following: SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: 1257, SEQ ID NO: 1264-SEQ ID NO: ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287-SEQ ID NO: 1289; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0046] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising a sequence having at least 80% sequence identity to any one of: SEQ ID NO:13–SEQ ID NO:17, SEQ ID NO:167–SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248–SEQ ID NO:1253, or SEQ ID NO:1259–SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0047] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence, which comprises a sequence having at least 80% sequence identity to any of the following: SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0048] In some aspects, the promoter sequence comprises SEQ ID NO: 17. In some aspects, the promoter sequence comprises SEQ ID NO: 1262. In some aspects, the promoter sequence comprises SEQ ID NO: 1250. In some aspects, the promoter sequence comprises SEQ ID NO: 1251. In some aspects, the promoter sequence comprises SEQ ID NO: 1252. In some aspects, the promoter sequence comprises SEQ ID NO: 1253.
[0049] In some aspects, the termination sequence comprises SEQ ID NO: 1264. In some aspects, the termination sequence comprises SEQ ID NO: 1265. In some aspects, the termination sequence comprises SEQ ID NO: 1254. In some aspects, the termination sequence comprises SEQ ID NO: 1255. In some aspects, the termination sequence comprises SEQ ID NO: 1257. In some aspects, the termination sequence comprises SEQ ID NO: 60. In some aspects, the termination sequence comprises SEQ ID NO: 1242. In some aspects, the termination sequence comprises SEQ ID NO: 1269. In some aspects, the termination sequence comprises SEQ ID NO: 1017.
[0050] In some aspects, the disease is a synucleinopathy, Parkinson's disease, dementia with Lewy bodies, multiple system atrophy, Charcot-Marie-Tooth disease, hereditary compressive susceptibility neuropathy, Yuan-Harel-Lupski syndrome, Tauopathy, Alzheimer's disease, frontotemporal dementia, progressive supranuclear palsy, corticobasal degeneration, chronic traumatic encephalopathy, autism, traumatic brain injury, Dravet syndrome, Crohn's disease, muscular dystrophy, B-cell leukemia, Dejerine-Sottas disease, Stargardt disease, alpha-1 antitrypsin deficiency, Tay-Sachs disease, cystic fibrosis, liposomal acid lipase deficiency, or Gaucher disease.
[0051] In some aspects, the small RNA payload comprises an engineered guide RNA that hybridizes to a target sequence, and wherein the cell encodes the target sequence. In some aspects, the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with type 1A Charcot-Marie-Tooth disease (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1) or methyl CpG binding protein 2 (MECP2).
[0052] In some aspects, the method further comprises forming a guide-target RNA scaffold after the engineered guide RNA hybridizes with the target sequence, recruiting an editing enzyme to the target sequence, and editing the target sequence with the editing enzyme. In some aspects, the target sequence comprises a mutation relative to a wild-type sequence. In some aspects, editing the target sequence corrects the mutation in the target sequence. In some aspects, the mutation is a missense mutation. In some aspects, the mutation is a nonsense mutation. In some aspects, the mutation is a G to A mutation. In some aspects, the mutation is associated with a disease. In some aspects, editing the target sequence comprises editing an untranslated region of the target. In some aspects, the untranslated region is a 5' untranslated region or a 3' untranslated region. In some aspects, the 3' untranslated region is a polyadenylation sequence. In some aspects, editing the target sequence comprises editing a translation start site.
[0053] In some aspects, editing the target sequence changes the expression of the target sequence. In some aspects, editing the target sequence increases the expression of the target sequence. In some aspects, editing the target sequence decreases the expression of the target sequence.
[0054] In some aspects, the guide-target RNA scaffold comprises a structural feature. In some aspects, the structural feature is a protrusion, a mismatch, an internal loop, a hairpin, or a combination thereof. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is a symmetrical protrusion. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is an asymmetric protrusion. In some aspects, the structural feature comprises an internal loop, and wherein the internal loop is a symmetrical internal loop. In some aspects, the structural feature comprises an internal loop, and wherein the internal loop is an asymmetric internal loop. In some aspects, the structural feature comprises a hairpin, and wherein the hairpin is a recruiting hairpin or a non-recruiting hairpin. In some aspects, the guide-target RNA scaffold comprises a wobble base pair.
[0055] In some aspects, the editing enzyme includes ADAR, APOBEC or Cas nuclease. In some aspects, the ADAR includes ADAR1, ADAR2, ADAR3 or a combination thereof. In some aspects, the target sequence comprises RNA or DNA. In some aspects, the target sequence is mRNA or pre-mRNA. In some aspects, editing the target sequence includes deamidating the nucleotides of the target sequence. In some aspects, the target sequence is edited with an efficiency of at least 10%, at least 20% or at least 25%.
[0056] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising: a zinc finger 143 motif, an OCT-1 transcription factor binding sequence, a proximal sequence element; a payload sequence transcriptionally controlled by the promoter sequence, the payload sequence comprising a small RNA payload; and a transcription termination sequence; wherein the expression cassette includes one or more sequence elements selected from the group consisting of: a) a zinc finger 143 motif having at least 80% sequence identity with any one of SEQ ID NO:24–SEQ ID NO:26, b) an OCT-1 transcription factor binding sequence having at least 80% sequence identity with any one of SEQ ID NO:27–SEQ ID NO:30, c) a proximal sequence element having at least 80% sequence identity with any one of SEQ ID NO:31–SEQ ID NO:37, and d) a combination thereof.
[0057] In some aspects, the zinc finger 143 motif has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with any one of SEQ ID NO:24-SEQ ID NO:26. In some aspects, the zinc finger 143 motif has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with SEQ ID NO:20. In some aspects, the OCT-1 transcription factor binding sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with any one of SEQ ID NO:27-SEQ ID NO:30. In some aspects, the proximal sequence element has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity with any one of SEQ ID NO:31-SEQ ID NO:37.
[0058] In some aspects, the transcription termination sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NO:40-SEQ ID NO:42. In some aspects, the transcription termination sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:60, SEQ ID NO:1242-SEQ ID NO:1247, or SEQ ID NO:1254-SEQ ID NO:1257. In some aspects, the transcription termination sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:1242. In some aspects, the transcription termination sequence includes the sequence SEQ ID NO:1242. In some aspects, the transcription termination sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 60. In some aspects, the transcription termination sequence comprises the sequence SEQ ID NO: 60. In some aspects, the transcription termination sequence comprises the sequence SEQ ID NO: 38 or SEQ ID NO: 39.
[0059] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising a proximal sequence element, wherein the promoter sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, wherein the proximal sequence element of the promoter sequence is replaced by a sequence of any one of SEQ ID NO: 67-SEQ ID NO: 120; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a transcription termination sequence comprising a 3' cassette sequence element, wherein the transcription termination sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: A sequence having at least 75% sequence identity to any one of SEQ ID NO: 1257, wherein the 3' box sequence element of the termination sequence is replaced by a sequence of any one of SEQ ID NO: 121 - SEQ ID NO: 166.
[0060] In some aspects, the promoter sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, wherein the proximal sequence elements of the promoter sequence are replaced by the sequence of any one of SEQ ID NO: 67-SEQ ID NO: 120. In some aspects, the promoter sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, wherein the proximal sequence elements of the promoter sequence are replaced by the sequence of any one of SEQ ID NO: 67-SEQ ID NO: 120. In some aspects, the termination sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708 - SEQ ID NO: 1240, SEQ ID NO: 1242 - SEQ ID NO: 1247, SEQ ID NO: 1254 - SEQ ID NO: 1257, wherein the 3' box sequence element of the termination sequence is replaced by the sequence of any one of SEQ ID NO: 121 - SEQ ID NO: 166. In some aspects, the termination sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708 - SEQ ID NO: 1240, SEQ ID NO: 1242 - SEQ ID NO: 1247, SEQ ID NO: 1254 - SEQ ID NO: 1257, wherein the 3' box sequence element of the termination sequence is replaced by the sequence of any one of SEQ ID NO: 121 - SEQ ID NO: 166.
[0061] In various aspects, the present disclosure provides an expression cassette comprising: a promoter sequence comprising a sequence having at least 75% sequence identity to any one of the following: SEQ ID NO:16–SEQ ID NO:17, SEQ ID NO:167–SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248–SEQ ID NO:1253; a payload sequence transcriptionally controlled by the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 75% sequence identity to any one of the following: SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257.
[0062] In some aspects, the promoter sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NO: 16-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253. In some aspects, the promoter sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO: 16-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253. In some aspects, the termination sequence comprises a sequence having at least 80% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: 1257. In some aspects, the termination sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO:60, SEQ ID NO:708 - SEQ ID NO:1240, SEQ ID NO:1242 - SEQ ID NO:1247, SEQ ID NO:1254 - SEQ ID NO:1257.
[0063] In some aspects, the promoter sequence is SEQ ID NO:376. In some aspects, the promoter sequence is SEQ ID NO:1250. In some aspects, the transcription termination sequence is SEQ ID NO:917. In some aspects, the transcription termination sequence is SEQ ID NO:1254. In some aspects, the promoter sequence is SEQ ID NO:168. In some aspects, the promoter sequence is SEQ ID NO:1251. In some aspects, the transcription termination sequence is SEQ ID NO:709. In some aspects, the transcription termination sequence is SEQ ID NO:1255. In some aspects, the promoter sequence is SEQ ID NO:1241. In some aspects, the transcription termination sequence is SEQ ID NO:1242 or SEQ ID NO:60. In some aspects, the promoter sequence is SEQ ID NO:17. In some aspects, the transcription termination sequence is SEQ ID NO:1242 or SEQ ID NO:60.
[0064] In some aspects, the small RNA payload comprises an engineered guide RNA capable of hybridizing to a target sequence.
[0065] In some aspects, the engineered guide RNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% or 100% reverse complementarity with the target sequence. In some aspects, the engineered guide RNA comprises at least one base pair mismatch relative to the target sequence. In some aspects, the target sequence comprises an adenosine residue. In some aspects, the target sequence is an RNA sequence. In some aspects, the RNA sequence is an mRNA or pre-mRNA.
[0066] In some aspects, the target sequence comprises a G to A mutation relative to the wild-type sequence. In some aspects, the target sequence comprises a missense mutation or a nonsense mutation relative to the wild-type sequence. In some aspects, the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with type 1A Charcot-Marie-Tooth disease (CMT1A), ATP binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase β (GBA), PTEN-induced kinase 1 (PINK1) or methyl CpG binding protein 2 (MECP2). In some aspects, the payload sequence has at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity to SEQ ID NO:1273, SEQ ID NO:1274 or SEQ ID NO:61.
[0067] In some aspects, the small RNA payload includes antisense oligonucleotides, siRNA, shRNA, miRNA or tracrRNA. In some aspects, the length of the small RNA payload is not less than 20 nucleotide residues and no more than 500 nucleotide residues. In some aspects, the length of the small RNA payload is not less than 60 residues and no more than 100 residues. In some aspects, the length of the small RNA payload is not less than 80 residues and no more than 120 residues. In some aspects, the length of the small RNA payload is not less than 100 residues and no more than 140 residues. In some aspects, the length of the small RNA payload is not less than 130 residues and no more than 170 residues.
[0068] In some aspects, the payload sequence further comprises an Sm binding sequence or a hairpin sequence. In some aspects, the hairpin sequence comprises a U7 hairpin. In some aspects, the hairpin sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO:52, SEQ ID NO:54, SEQ ID NO:56, or SEQ ID NO:58.
[0069] In some aspects, the expression cassette comprises two or more sequence elements. In some aspects, the expression cassette comprises three or more sequence elements. In some aspects, the length of the expression cassette is not less than 1300 nucleotide residues and not more than 2160 nucleotide residues. In some aspects, the expression cassette has at least 80% sequence identity with a U1 sequence or a U7 sequence. In some aspects, the U1 sequence is a mouse U1 sequence or a human U1 sequence. In some aspects, the U7 sequence is a mouse U7 sequence or a human U7 sequence.
[0070] In some aspects, the promoter sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253. In some aspects, the promoter sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 1241. In some aspects, the promoter sequence comprises the sequence SEQ ID NO: 1241. In some aspects, the transcription termination sequence is SEQ ID NO: 1242 or SEQ ID NO: 60. In some aspects, the promoter sequence has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to SEQ ID NO: 17. In some aspects, the promoter sequence comprises the sequence SEQ ID NO: 17. In some aspects, the transcription termination sequence is SEQ ID NO: 1242 or SEQ ID NO: 60.
[0071] In some aspects, the expression cassette has at least 80%, at least 85%, at least 90%, at least 95%, or 100% sequence identity to any one of SEQ ID NO: 1 to SEQ ID NO: 12 or SEQ ID NO: 59. In some aspects, the zinc finger 143 motif is capable of recruiting ZNF143 transcription factor. In some aspects, the OCT-1 transcription factor binding sequence is capable of recruiting OCT-1 transcription factor. In some aspects, the proximal sequence element is capable of recruiting SNAPc. In some aspects, the proximal sequence element is capable of integron-dependent recruitment of RNA polymerase II.
[0072] In some aspects, the small RNA payload is capable of forming a guide-target RNA scaffold comprising a structural feature after hybridization of the small RNA payload to the target sequence. In some aspects, the structural feature is a protrusion, a mismatch, an inner loop, a hairpin, or a combination thereof. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is a symmetrical protrusion. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is an asymmetric protrusion. In some aspects, the structural feature comprises an inner loop, and wherein the inner loop is a symmetrical inner loop. In some aspects, the structural feature comprises an inner loop, and wherein the inner loop is an asymmetric inner loop. In some aspects, the structural feature comprises a hairpin, and wherein the hairpin is a recruiting hairpin or a non-recruiting hairpin. The guide-target RNA scaffold comprises a wobble base pair.
[0073] In various aspects, the present disclosure provides methods of expressing a small RNA payload in a cell, the method comprising delivering an expression cassette as described herein to a cell and expressing the small RNA payload encoded by the expression cassette in the cell.
[0074] In various aspects, the present disclosure provides a method for editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising: a zinc finger 143 motif, an OCT-1 transcription factor binding sequence, and a proximal sequence element; a payload sequence transcriptionally controlled by the promoter sequence, the payload sequence comprising a small RNA payload, wherein the small RNA payload comprises an engineered guide RNA sequence capable of hybridizing with the target sequence and a transcription termination sequence; expressing the small RNA payload in the cell; forming a guide-target RNA scaffold after the small RNA payload hybridizes with the target sequence; recruiting an editing enzyme to the target sequence; and editing the target sequence using the editing enzyme.
[0075] In various aspects, the present disclosure provides a method of editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising a proximal sequence element, wherein the promoter sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, wherein the proximal sequence element of the promoter sequence is replaced by a sequence of any one of SEQ ID NO: 67-SEQ ID NO: 120; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a transcription termination sequence comprising a 3' cassette sequence element, wherein the transcription termination sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: NO:1257 has at least 75% sequence identity, wherein the 3' box sequence element of the termination sequence is replaced by the sequence of any one of SEQ ID NO:121-SEQ ID NO:166; expressing a small RNA payload in the cell; the small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0076] In various aspects, the present disclosure provides a method of editing a target sequence, the method comprising: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: a promoter sequence comprising a proximal sequence element, wherein the promoter sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 16–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a transcription termination sequence comprising a 3' cassette sequence element, wherein the transcription termination sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708–SEQ ID NO: 1240, SEQ ID NO: 1242–SEQ ID NO: 1247, SEQ ID NO: 1254–SEQ ID NO: NO: any one of 1257 having at least 75% sequence identity; expressing a small RNA payload in the cell; the small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0077] In some aspects, the promoter sequence is SEQ ID NO:376. In some aspects, the promoter sequence is SEQ ID NO:1250. In some aspects, the transcription termination sequence is SEQ ID NO:917. In some aspects, the transcription termination sequence is SEQ ID NO:1254. In some aspects, the promoter sequence is SEQ ID NO:168. In some aspects, the promoter sequence is SEQ ID NO:1251. In some aspects, the transcription termination sequence is SEQ ID NO:709. In some aspects, the transcription termination sequence is SEQ ID NO:1255. In some aspects, the promoter sequence is SEQ ID NO:1241. In some aspects, the transcription termination sequence is SEQ ID NO:1242 or SEQ ID NO:60. In some aspects, the promoter sequence is SEQ ID NO:17. In some aspects, the transcription termination sequence is SEQ ID NO:1242 or SEQ ID NO:60.
[0078] In various aspects, the present disclosure provides a method for editing a target sequence, the method comprising: delivering an expression cassette as described herein to a cell encoding the target sequence; expressing a small RNA payload in the cell, wherein the small RNA payload comprises an engineered guide RNA capable of hybridizing to the target sequence; forming a guide-target RNA scaffold after hybridization of the small RNA payload to the target sequence; recruiting an editing enzyme to the target sequence; and editing the target sequence with the editing enzyme.
[0079] In some aspects, the target sequence comprises a mutation relative to the wild-type sequence. In some aspects, editing the target sequence corrects the mutation in the target sequence. In some aspects, the mutation is a missense mutation. In some aspects, the mutation is a nonsense mutation. In some aspects, the mutation is a G to A mutation. In some aspects, the mutation is associated with a disease.
[0080] In some aspects, the disease is a synucleinopathy, Parkinson's disease, dementia with Lewy bodies, multiple system atrophy, Charcot-Marie-Tooth disease, hereditary compressive susceptibility neuropathy, Yuan-Harel-Lupski syndrome, Tauopathy, Alzheimer's disease, frontotemporal dementia, progressive supranuclear palsy, corticobasal degeneration, chronic traumatic encephalopathy, autism, traumatic brain injury, Dravet syndrome, Crohn's disease, muscular dystrophy, B-cell leukemia, Dejerine-Sottas disease, Stargardt disease, alpha-1 antitrypsin deficiency, Tay-Sachs disease, cystic fibrosis, liposomal acid lipase deficiency, or Gaucher disease. In some aspects, the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with Charcot-Marie-Tooth disease type 1A (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2).
[0081] In some aspects, editing the target sequence comprises editing an untranslated region of the target. In some aspects, the untranslated region is a 5' untranslated region or a 3' untranslated region. In some aspects, the 3' untranslated region is a polyadenylation sequence. In some aspects, editing the target sequence comprises editing a translation start site. In some aspects, editing the target sequence changes the expression of the target sequence. In some aspects, editing the target sequence increases the expression of the target sequence. In some aspects, editing the target sequence reduces the expression of the target sequence.
[0082] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising: a zinc finger 143 motif, an OCT-1 transcription factor binding sequence, and a proximal sequence element; and a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0083] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising a proximal sequence element, wherein the promoter sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, wherein the proximal sequence element of the promoter sequence is replaced with a sequence of any one of SEQ ID NO: 67-SEQ ID NO: 120; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a transcription termination sequence comprising a 3' cassette sequence element, wherein the transcription termination sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: NO:1257, wherein the 3' box sequence element of the termination sequence is replaced by the sequence of any one of SEQ ID NO:121-SEQ ID NO:166; delivering the expression cassette to the cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0084] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette, the expression cassette comprising: a promoter sequence comprising a proximal sequence element, wherein the promoter sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 16-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a transcription termination sequence comprising a 3' cassette sequence element, wherein the transcription termination sequence comprises a sequence having at least 75% sequence identity to any one of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: NO: 1257 has at least 75% sequence identity; delivering the expression cassette to cells of the subject; and expressing the small RNA payload in the cells, thereby treating the disease.
[0085] In some aspects, the promoter sequence is SEQ ID NO:376. In some aspects, the promoter sequence is SEQ ID NO:1250. In some aspects, the transcription termination sequence is SEQ ID NO:917. In some aspects, the transcription termination sequence is SEQ ID NO:1254. In some aspects, the promoter sequence is SEQ ID NO:168. In some aspects, the promoter sequence is SEQ ID NO:1251. In some aspects, the transcription termination sequence is SEQ ID NO:709. In some aspects, the transcription termination sequence is SEQ ID NO:1255. In some aspects, the promoter sequence is SEQ ID NO:1241. In some aspects, the transcription termination sequence is SEQ ID NO:1242 or SEQ ID NO:60. In some aspects, the promoter sequence is SEQ ID NO:17. In some aspects, the transcription termination sequence is SEQ ID NO:1242 or SEQ ID NO:60.
[0086] In various aspects, the present disclosure provides a method of treating a disease in a subject, the method comprising: administering to the subject a composition comprising an expression cassette as described herein; delivering the expression cassette to cells of the subject; and expressing a small RNA payload in the cells, thereby treating the disease.
[0087] In some aspects, the disease is a synucleinopathy, Parkinson's disease, dementia with Lewy bodies, multiple system atrophy, Charcot-Marie-Tooth disease, hereditary compressive susceptibility neuropathy, Yuan-Harel-Lupski syndrome, Tauopathy, Alzheimer's disease, frontotemporal dementia, progressive supranuclear palsy, corticobasal degeneration, chronic traumatic encephalopathy, autism, traumatic brain injury, Dravet syndrome, Crohn's disease, muscular dystrophy, B-cell leukemia, Dejerine-Sottas disease, Stargardt disease, alpha-1 antitrypsin deficiency, Tay-Sachs disease, cystic fibrosis, liposomal acid lipase deficiency, or Gaucher disease. In some aspects, the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with type 1A Charcot-Marie-Tooth disease (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase β (GBA), PTEN-induced kinase 1 (PINK1) or methyl CpG binding protein 2 (MECP2). In some aspects, the small RNA payload comprises an engineered guide RNA that hybridizes to the target sequence, and wherein the cell encodes the target sequence.
[0088] In some aspects, the method further comprises forming a guide-target RNA scaffold after hybridization of the engineered guide RNA to the target sequence, recruiting an editing enzyme to the target sequence, and editing the target sequence with the editing enzyme. In some aspects, the target sequence comprises a mutation relative to the wild-type sequence. In some aspects, editing the target sequence corrects the mutation in the target sequence. In some aspects, the mutation is a missense mutation. In some aspects, the mutation is a nonsense mutation. In some aspects, the mutation is a G to A mutation. In some aspects, the mutation is associated with a disease.
[0089] In some aspects, editing the target sequence comprises editing an untranslated region of the target. In some aspects, the untranslated region is a 5' untranslated region or a 3' untranslated region. In some aspects, the 3' untranslated region is a polyadenylation sequence. In some aspects, editing the target sequence comprises editing a translation start site. In some aspects, editing the target sequence changes the expression of the target sequence. In some aspects, editing the target sequence increases the expression of the target sequence. In some aspects, editing the target sequence reduces the expression of the target sequence.
[0090] In some aspects, the guide-target RNA scaffold comprises a structural feature. In some aspects, the structural feature is a protrusion, a mismatch, an internal loop, a hairpin, or a combination thereof. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is a symmetrical protrusion. In some aspects, the structural feature comprises a protrusion, and wherein the protrusion is an asymmetric protrusion. In some aspects, the structural feature comprises an internal loop, and wherein the internal loop is a symmetrical internal loop. In some aspects, the structural feature comprises an internal loop, and wherein the internal loop is an asymmetric internal loop. In some aspects, the structural feature comprises a hairpin, and wherein the hairpin is a recruiting hairpin or a non-recruiting hairpin. In some aspects, the guide-target RNA scaffold comprises a wobble base pair.
[0091] In some aspects, the editing enzyme includes ADAR, APOBEC or Cas nuclease. In some aspects, the ADAR includes ADAR1, ADAR2, ADAR3 or a combination thereof. In some aspects, the target sequence comprises RNA or DNA. In some aspects, the target sequence is mRNA or pre-mRNA. In some aspects, editing the target sequence includes deamidating the nucleotides of the target sequence. In some aspects, the target sequence is edited with an efficiency of at least 10%, at least 20% or at least 25%.
[0092] In some aspects, the expression cassette is delivered to the cell via a viral vector. In some aspects, the viral vector is an adenoviral vector, an adeno-associated viral vector, or a lentiviral vector. In some embodiments, the adeno-associated virus vector is selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-DJ, AAV-DJ / 8, AAV-DJ / 9, AAV1 / 2, AAV.rh8, AAV.rh10, AAV.rh20, AAV.rh39, AAV.Rh43, AAV.Rh74, AAV.v66, AAV.Oligo001, AAV.SCH9, AAV.r3.45, AAV.RHM4-1, AAV.hu37, AAV.Anc80, AAV.Anc80L65, AAV.7m8, AAV. AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAV.HSC16, AAV.HSC17, AAVhu68, chimeras thereof, and combinations thereof.
[0093] In various aspects, the present disclosure provides viral vectors that encapsidate expression cassettes as described herein.
[0094] In some aspects, the viral vector is an adeno-associated viral vector. In some embodiments, the adeno-associated virus vector is selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV10, AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-DJ, AAV-DJ / 8, AAV-DJ / 9, AAV1 / 2, AAV.rh8, AAV.rh10, AAV.rh20, AAV.rh39, AAV.Rh43, AAV.Rh74, AAV.v66, AAV.Oligo001, AAV.SCH9, AAV.r3.45, AAV.RHM4-1, AAV.hu37, AAV.Anc80, AAV.Anc80L65, AAV.7m8, AAV. AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAV.HSC16, AAV.HSC17, AAVhu68, chimeras thereof, and combinations thereof.
[0095] In various aspects, the present disclosure provides a pharmaceutical composition comprising an expression cassette as described herein or a viral vector as described herein and a pharmaceutically acceptable excipient, carrier, diluent, or a combination thereof.
[0096] Incorporated by Reference
[0097] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The novel features of the present invention are particularly set forth in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description setting forth illustrative embodiments, in which the principles of the invention are utilized, and in the accompanying drawings:
[0099] Figure 1AAn example configuration of an engineered guide RNA expression cassette based on the mouse U7 (mU7) promoter is schematically illustrated. The expression cassette encodes a payload sequence that is transcriptionally controlled by the mU7 promoter. The mU7 promoter contains an SPH element (e.g., a zinc finger 143 motif), an OCT-1 transcription factor binding sequence, and a proximal sequence element (PSE). The payload sequence begins at the transcription start site and ends at a termination sequence, including an engineered guide RNA sequence ("guide") that is operably linked to a Sm binding sequence (smOPT).
[0100] Figure 1B An example configuration of an engineered guide RNA expression cassette based on the human U1 (hU7) promoter is schematically illustrated. The expression cassette encodes a payload sequence that is transcriptionally controlled by the hU1 promoter. The hU1 promoter contains an SPH element (e.g., a zinc finger 143 motif), an OCT-1 transcription factor binding sequence, and a proximal sequence element (PSE). The payload sequence begins at the transcription start site and ends at a termination sequence, including an engineered guide RNA sequence ("guide") operably linked to a Sm binding sequence (smOPT).
[0101] Figure 2A Schematic illustration of a reporter construct used to measure expression of an engineered guide RNA sequence and subsequent editing of a target RNA sequence. The reporter construct includes a target sequence (e.g., CDS1) containing an ATG start site that can be edited to ITG (pronounced GTG) via ADAR-catalyzed deamidation. Conversion of ATG to GTG results in increased expression of luciferase (NanoLuc).
[0102] Figure 2B Bar graphs showing luciferase assays show that reporter constructs are edited by engineered guide RNA constructs. Unedited (ATG) constructs express basal levels of luciferase, resulting in background luciferase activity levels. Edited (GTG) constructs express higher levels of luciferase, resulting in increased luciferase activity relative to unedited constructs.
[0103] Figure 3 Bar graphs showing luciferase activity in the presence of unedited (A) or edited (G) reporter genes SEQ ID NO: 48 ("fPMP22-cDNA (ATG)"), SEQ ID NO: 49 ("fSNCA-pre (ATG)"), and SEQ ID NO: 50 ("fSNCA-cDNA (ATG)"). For each reporter gene, the edited constructs expressed higher levels of luciferase compared to the unedited constructs, resulting in increased levels of luciferase activity.
[0104] Figure 4 The workflow for generating and evaluating the expression of expression cassette constructs is schematically illustrated. Cells are transfected with plasmids encoding engineered guide RNAs, and the expression of engineered guide RNAs is evaluated by luciferase activity. The expression of engineered guide RNAs can be further evaluated using mirVANA total RNA isolation, DNaseI treatment, ddPCR guide quantification assays, or Sanger editing.
[0105] Figure 5 Bar graphs of luciferase assays are shown to evaluate the expression of SNCA-targeted engineered guide RNAs (SEQ ID NO: 1274) controlled by mouse U7 promoters with various OCT-1 transcription factor binding sequences. The original OCT-1 transcription factor binding sequence (SEQ ID NO: 21) in the SNCA-targeted guide RNA expression cassette (SEQ ID NO: 6) was replaced with variant OCT-1 transcription factor binding sequences of each of SEQ ID NO: 27–SEQ ID NO: 30 or a random sequence of SEQ ID NO: 45 or a repeated random sequence of SEQ ID NO: 46. A construct encoding only a GFP cassette ("GFP control") was used as a negative control. Higher luciferase activity indicates increased expression of the engineered guide RNA.
[0106] Figure 6 Bar graphs of luciferase assays are shown to evaluate the expression of SNCA-targeted engineered guide RNAs (SEQ ID NO: 1274) controlled by mouse U7 promoters with various zinc finger 143 motifs. The original zinc finger 143 motif (SEQ ID NO: 20) in the SNCA-targeted guide RNA expression cassette (SEQ ID NO: 6) was replaced by variant zinc finger 143 motifs of each of SEQ ID NO: 24-SEQ ID NO: 26 or a random sequence of SEQ ID NO: 43. A construct encoding only a GFP cassette ("GFP control") was used as a negative control. Higher luciferase activity indicates increased expression of the engineered guide RNA.
[0107] Figure 7Bar graphs of luciferase assays are shown to evaluate the expression of SNCA-targeted engineered guide RNA (SEQ ID NO: 1274) controlled by the mouse U7 promoter with various proximal sequence elements (PSEs). The original PSE (SEQ ID NO: 22) in the SNCA-targeted guide RNA expression cassette (SEQ ID NO: 6) was replaced by variant PSEs of each of SEQ ID NO: 31-SEQ ID NO: 37 or a random sequence of SEQ ID NO: 44. A construct encoding only a GFP cassette ("GFP control") was used as a negative control. Higher luciferase activity indicates increased expression of the engineered guide RNA.
[0108] Figure 8 Bar graphs of luciferase assays are shown to evaluate the expression of SNCA-targeted engineered guide RNA (SEQ ID NO: 1274) controlled by mouse U7 promoters with various transcription termination sequences. The original termination sequence (SEQ ID NO: 23) in the SNCA-targeted guide RNA expression cassette (SEQ ID NO: 6) was replaced by variant termination sequences of each of SEQ ID NO: 40-SEQ ID NO: 42 or a random sequence of SEQ ID NO: 47. A construct encoding only a GFP cassette ("GFP control") was used as a negative control. Higher luciferase activity indicates increased expression of the engineered guide RNA.
[0109] Fig. 9A Bar graphs of luciferase assays are shown to evaluate the expression of PMP22-targeted engineered guide RNAs (SEQ ID NO: 1273) controlled by mouse U7 promoters with various combinations of engineered sequence elements. SEQ ID NO: 2 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1. SEQ ID NO: 3 comprises a variant termination sequence of SEQ ID NO: 41 relative to SEQ ID NO: 1. SEQ ID NO: 4 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1 and a variant termination sequence of SEQ ID NO: 41. SEQ ID NO: 5 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1, a variant termination sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. Expression is quantified relative to a construct encoding only a GFP cassette ("GFP"). Higher luciferase activity indicates increased guide RNA expression.
[0110] Fig. 9BBar graphs of luciferase assays are shown to evaluate the expression of SNCA-targeted engineered guide RNAs (SEQ ID NO: 1274) controlled by a mouse U7 promoter with various combinations of engineered sequence elements. Expression of SNCA-targeted guide RNAs was also tested under the control of a human U1 promoter (SEQ ID NO: 13) and a human U7 promoter (SEQ ID NO: 14). SEQ ID NO: 9 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6. SEQ ID NO: 10 comprises a variant stop sequence of SEQ ID NO: 41 relative to SEQ ID NO: 6. SEQ ID NO: 11 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6 and a variant stop sequence of SEQ ID NO: 41. SEQ ID NO: 12 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6, a variant stop sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. Expression is quantified relative to a construct encoding only the GFP cassette ("GFP"). Higher luciferase activity indicates increased guide RNA expression.
[0111] Fig. 10A Bar graphs of guide quantification assays are shown to evaluate the expression of PMP22-targeted engineered guide RNAs (SEQ ID NO: 1273) controlled by mouse U7 promoters with various combinations of engineered sequence elements. SEQ ID NO: 2 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1. SEQ ID NO: 3 comprises a variant termination sequence of SEQ ID NO: 41 relative to SEQ ID NO: 1. SEQ ID NO: 4 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1 and a variant termination sequence of SEQ ID NO: 41. SEQ ID NO: 5 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1, a variant termination sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. Expression is quantified relative to a construct encoding only a GFP cassette ("GFP"). A higher guide RNA to GAPDH ratio indicates increased guide RNA expression.
[0112] Fig. 10BBar graphs showing guide quantification assays to evaluate the expression of SNCA-targeted engineered guide RNAs (SEQ ID NO: 1274) controlled by a mouse U7 promoter with various combinations of engineered sequence elements. Expression of SNCA-targeted guide RNAs was also tested under the control of a human U1 promoter (SEQ ID NO: 13) and a human U7 promoter (SEQ ID NO: 14). SEQ ID NO: 9 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6. SEQ ID NO: 10 comprises a variant stop sequence of SEQ ID NO: 41 relative to SEQ ID NO: 6. SEQ ID NO: 11 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6 and a variant stop sequence of SEQ ID NO: 41. SEQ ID NO: 12 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6, a variant stop sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. Expression is quantified relative to a construct encoding only the GFP cassette ("GFP"). A higher guide RNA to GAPDH ratio indicates increased guide RNA expression.
[0113] Fig.11A Bar graph showing Sanger editing of ATG sequences to GTG to evaluate expression and editing activity of PMP22-targeted engineered guide RNA (SEQ ID NO: 1273) controlled by mouse U7 promoter with various combinations of engineered sequence elements. SEQ ID NO: 2 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1. SEQ ID NO: 3 comprises a variant termination sequence of SEQ ID NO: 41 relative to SEQ ID NO: 1. SEQ ID NO: 4 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1 and a variant termination sequence of SEQ ID NO: 41. SEQ ID NO: 5 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1, a variant termination sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. A construct encoding only a GFP cassette ("GFP") was used as a negative control. A higher editing percentage indicates increased guide RNA expression.
[0114] Fig. 11BBar graph showing Sanger editing of ATG sequences to GTG to evaluate expression and editing activity of SNCA-targeted engineered guide RNA (SEQ ID NO: 1274) controlled by mouse U7 promoter with various combinations of engineered sequence elements. SEQ ID NO: 9 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6. SEQ ID NO: 10 comprises a variant stop sequence of SEQ ID NO: 41 relative to SEQ ID NO: 6. SEQ ID NO: 11 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6 and a variant stop sequence of SEQ ID NO: 41. SEQ ID NO: 12 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6, a variant stop sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. A construct encoding only a GFP cassette ("GFP") was used as a negative control. A higher percentage of editing indicates increased guide RNA expression.
[0115] Fig. 12A Bar graphs showing Sanger editing of the -3 residue to assess expression and editing activity of a PMP22-targeted engineered guide RNA (SEQ ID NO: 1273) controlled by a mouse U7 promoter with various combinations of engineered sequence elements. SEQ ID NO: 2 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1. SEQ ID NO: 3 comprises a variant termination sequence of SEQ ID NO: 41 relative to SEQ ID NO: 1. SEQ ID NO: 4 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1 and a variant termination sequence of SEQ ID NO: 41. SEQ ID NO: 5 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 1, a variant termination sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. A construct encoding only a GFP cassette ("GFP") was used as a negative control. A higher percentage of editing indicates increased guide RNA expression.
[0116] Fig. 12BBar graph showing Sanger editing of the -5 residue to assess expression and editing activity of an SNCA-targeted engineered guide RNA (SEQ ID NO: 1274) controlled by a mouse U7 promoter with various combinations of engineered sequence elements. SEQ ID NO: 9 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6. SEQ ID NO: 10 comprises a variant stop sequence of SEQ ID NO: 41 relative to SEQ ID NO: 6. SEQ ID NO: 11 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6 and a variant stop sequence of SEQ ID NO: 41. SEQ ID NO: 12 comprises a variant PSE of SEQ ID NO: 31 relative to SEQ ID NO: 6, a variant stop sequence of SEQ ID NO: 41, and a variant OCT-1 transcription factor binding sequence of SEQ ID NO: 28. A construct encoding only a GFP cassette ("GFP") was used as a negative control. A higher percentage of editing indicates increased guide RNA expression.
[0117] Fig.13A A scatter plot using a linear fit is shown, showing Fig. 10B Guided quantification of assay results and Fig. 9B Correlation between the luciferase assay results.
[0118] Fig. 13B A scatter plot using a linear fit is shown, showing Fig. 11B Sanger editing assay results and Fig. 9B Correlation between the luciferase assay results.
[0119] Fig. 13C A scatter plot using a linear fit is shown, showing Fig. 10B Guided quantification of assay results and Fig. 11B Correlation between the Sanger editing assay results.
[0120] Fig.14A A scatter plot using a linear fit is shown, showing Fig. 10A Guided quantification of assay results and Fig. 9A Correlation between the luciferase assay results.
[0121] Fig. 14B A scatter plot using a linear fit is shown, showing Fig. 12A Sanger editing assay results and Fig. 9A Correlation between the luciferase assay results.
[0122] Fig. 14C A scatter plot using a linear fit is shown, showing Fig. 10A Guided quantification of assay results and Fig.11A Correlation between the Sanger editing assay results.
[0123] Fig.15 Shown are sequences with a single copy of the promoter variant integrated into the HEK293T cell genome (left) and an engineered guide RNA targeting RAB7A ( Fig.15 Figure (continued)), GAPDH ( Fig.15 Middle Figure (continued)) and SNCA ( Fig.15 Comparison of copy integration in the figure below (continued). Fig.15 SEQ ID NO: 1283 and SEQ ID NO: 1284 are disclosed in order of appearance, respectively.
[0124] Fig.16 A diagram showing various exemplary structural features present in the guide-target RNA scaffold formed after hybridization of the potential guide RNA of the present disclosure with the target RNA. The exemplary structural features shown include an 8 / 7 asymmetric loop (i. 8 nucleotides on the target RNA side and 7 nucleotides on the guide RNA side), a 2 / 2 symmetric protrusion (ii. 2 nucleotides on the target RNA side and 2 nucleotides on the guide RNA side), a 1 / 1 mismatch (iii. 1 nucleotide on the target RNA side and 1 nucleotide on the guide RNA side), a 5 / 5 symmetric inner loop (iv. 5 nucleotides on the target RNA side and 5 nucleotides on the guide RNA side), a 24 bp region (v. 24 nucleotides on the target RNA side are base paired with 24 nucleotides on the guide RNA side) and a 2 / 3 asymmetric protrusion (vi. 2 nucleotides on the target RNA side and 3 nucleotides on the guide RNA side). Fig.16 SEQ ID NO: 1285 and SEQ ID NO: 1286 are disclosed in order of appearance, respectively.
[0125] Fig.17ABar graphs quantifying the expression of SNCA targeting guide RNA (SEQ ID NO: 1274, left) or PMP22 targeting guide RNA (SEQ ID NO: 1273, right) in ARPE-19 cells are shown. The expression of SNCA targeting guide RNA in ARPE-19 cells was compared by an expression cassette controlled by a wild-type mouse U7 promoter (SEQ ID NO: 6) or an expression cassette controlled by an engineered mouse U7 promoter (SEQ ID NO: 12) (left panel). The expression of PMP22 targeting guide RNA in ARPE-19 cells was compared by an expression cassette controlled by a wild-type mouse U7 promoter (SEQ ID NO: 1) or an expression cassette controlled by an engineered mouse U7 promoter (SEQ ID NO: 5) (right panel). The engineered expression cassettes of SEQ ID NO: 12 and SEQ ID NO: 5 comprise an engineered promoter of SEQ ID NO: 17, which comprises an OCT-1 transcription factor binding sequence of SEQ ID NO: 28 and a PSE of SEQ ID NO: 31, and an engineered termination sequence of SEQ ID NO: 60, which comprises a termination sequence motif of SEQ ID NO: 41. Expression is quantified relative to a construct encoding only a GFP cassette ("GFP"). A higher guide RNA to GAPDH ratio indicates increased guide RNA expression.
[0126] Fig. 17B Bar graphs for quantifying the expression of SERPINA1 targeting guide RNA (SEQ ID NO:61) in HepG2 cells are shown. Expression cassettes controlled by wild-type mouse U7 promoter ("mU7-WT") or by engineered mouse U7 promoter (SEQ ID NO:59) were compared to the expression of SERPINA1 targeting guide RNA in HepG2 cells (right figure). The engineered expression cassette of SEQ ID NO:59 includes an engineered promoter (which includes a PSE of SEQ ID NO:31) of SEQ ID NO:16 and an engineered termination sequence (which includes a termination sequence motif of SEQ ID NO:41) of SEQ ID NO:60. Expression is quantified relative to a construct encoding only a GFP box ("GFP"). Higher guide RNA and GAPDH ratios indicate that guide RNA expression increases.
[0127] Fig.18 An exemplary novel promoter of the present disclosure tested on antisense oligonucleotides for clinically relevant Duchenne muscular dystrophy (DMD) exon splicing in differentiating muscle cells is shown. Engineered guide RNA expression constructs were randomly integrated into the genome and evaluated after 10 days of muscle cell differentiation.
[0128] Fig.19AExemplary combinations of promoters, promoter variants, 3' box stop sequences, and truncated 3' box stop sequences of the present disclosure for driving guide RNA expression are shown.
[0129] Fig.19B Exemplary combinations of promoters, promoter variants, 3' box stop sequences, and truncated 3' box stop sequences of the present disclosure for driving guide RNA expression are shown.
[0130] Fig. 20A A bar graph quantifying the expression of PMP22 targeting guide RNA containing a luciferase reporter gene (Reporter 1) in HEK293 cells is shown. The PMP22 targeting engineered guide RNA constructs with engineered promoter elements contained in SEQ ID NO: 2, SEQ ID NO: 3, and SEQ ID NO: 5 have increased expression folds on the expression of PMP22 targeting guide RNA in HEK293 cells relative to the control mU7 wild type guide RNA construct (SEQ ID NO: 1).
[0131] Fig. 20B Bar graph quantifying the expression of SNCA targeting guide RNA containing a luciferase reporter gene (Reporter 2) in HEK293 cells is shown. The expression of Reporter 2 targeting guide RNA in HEK293 cells by SNCA targeting engineered guide RNA constructs with engineered promoter elements contained in SEQ ID NO:9, SEQ ID NO:10 and SEQ ID NO:11 was increased by fold relative to the expression of the control mU7 wild-type guide RNA construct (SEQ ID NO:6).
[0132] Fig.21A Bar graphs are shown, wherein the left subgraph shows quantification of the expression of PMP22 targeting guide RNA containing a luciferase reporter gene (reporter gene 1) in HEK293T cells. Expression of reporter gene 1 guide RNA in HEK293T cells by the PMP22 targeting engineered guide RNA construct with an engineered promoter element contained in SEQ ID NO: 5, fold increase relative to the control mU7 wild-type guide RNA construct (SEQ ID NO: 1), and increased expression when compared to the control PMP22 targeting guide RNA controlled by the wild-type human U1 promoter (SEQ ID NO: 13). Negative control expression was also quantified by a construct encoding only a GFP box ("GFP control"). Fig.21AThe right panel of the figure shows a bar graph quantifying the expression of SNCA targeting guide RNA with a luciferase reporter gene (Reporter 2) in HEK293T cells. The expression of the reporter 2 targeting guide RNA in HEK293T cells by the SNCA targeting engineered guide RNA construct with an engineered promoter element contained in SEQ ID NO: 12 was increased fold relative to the expression of the control mU7 wild-type guide RNA construct (SEQ ID NO: 6). Negative control expression was also quantified by a construct encoding only a GFP cassette ("GFP control").
[0133] Fig. 21B Bar graphs are shown, wherein the left subgraph shows quantification of the expression of PMP22 targeting guide RNA containing a luciferase reporter gene (reporter 1) in HEK293T cells. The expression of reporter 1 guide RNA in HEK293T cells by the engineered PMP22 targeting guide RNA controlled by the engineered hU1 promoter (SEQ ID NO: 1241) has a greater expression fold relative to the control PMP22 targeting guide RNA controlled by the wild-type human U1 promoter (SEQ ID NO: 13). Negative control expression was also quantified by a construct encoding only a GFP box ("GFP"). Fig. 21B The right panel of the Figure shows a bar graph quantifying the expression of SNCA-targeted guide RNA with a luciferase reporter gene (Reporter 2) in HEK293T cells. The expression of the Reporter 2 guide RNA in HEK293T cells by the engineered SNCA-targeted guide RNA controlled by the engineered hU1 promoter (SEQ ID NO: 1241) has a greater expression fold relative to the control hU1 wild-type guide RNA construct (SEQ ID NO: 7). Negative control expression was also quantified by a construct encoding only the GFP cassette ("GFP").
[0134] Fig.22A Bar graphs are shown quantifying SNCA guide RNA expression of constructs containing promoter sequences comprising the full-length WT mU7 promoter sequence (SEQ ID NO: 15), a variant of the WT mU7 promoter sequence with a 100 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1248), an engineered mU7 promoter sequence (SEQ ID NO: 17), or a variant of the engineered mU7 promoter sequence with a 100 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1249). Guide RNA expression was quantified by ddPCR and normalized to a housekeeping gene (GAPDH). A higher guide RNA expression to GAPDH expression (gRNA / GAPDH) ratio indicates increased guide RNA expression.
[0135] Fig. 22B Bar graphs are shown quantifying PMP22 guide RNA expression of expression cassette constructs containing promoter sequences, including full-length WT mU7 promoter sequence (SEQ ID NO: 15), variants of WT mU7 promoter sequence with 100 base deletions between DSE and PSE promoter elements (SEQ ID NO: 1248), engineered mU7 promoter sequence (SEQ ID NO: 17), or variants of engineered mU7 promoter sequence with 100 base deletions between DSE and PSE promoter elements (SEQ ID NO: 1249). Guide RNA expression was quantified by ddPCR and normalized to a housekeeping gene (GAPDH). Higher guide RNA expression to GAPDH expression (gRNA / GAPDH) ratios indicate increased guide RNA expression.
[0136] Fig.23 Shown are bar graphs quantifying Rab7a editing in expression cassette constructs comprising promoter sequences, including the full-length WT mU7 promoter sequence (SEQ ID NO: 15), a variant of the WT mU7 promoter sequence having a 50 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1258), a variant of the WT mU7 promoter sequence having a 75 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1259), a variant of the WT mU7 promoter sequence having a 100 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1248), a variant of the WT mU7 promoter sequence having a 126 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1260), and a variant of the WT mU7 promoter sequence having a 135 base deletion between the DSE and PSE promoter elements (SEQ ID NO: 1261).
[0137] Fig.24Bar graphs quantifying GFP expression from expression constructs with a squirrel monkey herpesvirus U-RNA element (HSUR) are shown. The HSUR element was extracted from NCBI NC_001350 and bound downstream of a gRNA cassette with the RNU5B1 promoter (SEQ ID NO: 1250) and a GFP gRNA targeting a GFP-G67R reporter gene, where deamination of the AGA codon to GGA restored fluorescence in a correlated manner. The expression construct was introduced as a single copy by BxbI integrase and enriched for 14 days by puromycin. GFP expression was quantified by flow cytometry as the geometric mean fluorescence intensity (GFP gMFI), and cells upstream of mCherry fluorescence were gated so that only cells positive for the cassette were plotted. GFP expression from expression constructs containing the termination sequences SEQ ID NO: 1266–SEQ ID NO: 1272 was quantified and compared to GFP expression from an expression construct with the termination sequence SEQ ID NO: 1254.
[0138] Fig.25A Shown is a bar graph quantifying GFP guide RNA expression of expression cassette constructs comprising a promoter sequence of SEQ ID NO: 17 and a terminator sequence of SEQ ID NO: 60 (SEQ ID NO: 17 / SEQ ID NO: 60), a promoter sequence of SEQ ID NO: 15 and a terminator sequence of SEQ ID NO: 1243 (SEQ ID NO: 15 / SEQ ID NO: 1243), a promoter sequence of SEQ ID NO: 1250 and a terminator sequence of SEQ ID NO: 1254 (SEQ ID NO: 1250 / SEQ ID NO: 1254), a promoter sequence of SEQ ID NO: 1252 and a terminator sequence of SEQ ID NO: 1256 (SEQ ID NO: 1252 / SEQ ID NO: 1256), a promoter sequence of SEQ ID NO: 1251 and a terminator sequence of SEQ ID NO: 1255 (SEQ ID NO: 1251 / SEQ ID NO: 1255), or a promoter sequence of SEQ ID NO: 1253 and a terminator sequence of SEQ ID NO: 1257 (SEQ ID NO: 1258). NO:1253 / SEQ ID NO:1257). Guide RNA expression was quantified by ddPCR and normalized to a housekeeping gene (GAPDH). A higher guide RNA expression to GAPDH expression (gRNA / GAPDH) ratio indicates increased guide RNA expression.
[0139] Fig.25BBar graphs are shown quantifying SNCA guide RNA expression of expression cassette constructs comprising a promoter sequence of SEQ ID NO: 17 and a terminator sequence of SEQ ID NO: 60 (SEQ ID NO: 17 / SEQ ID NO: 60), a promoter sequence of SEQ ID NO: 15 and a terminator sequence of SEQ ID NO: 1243 (SEQ ID NO: 15 / SEQ ID NO: 1243), a promoter sequence of SEQ ID NO: 1250 and a terminator sequence of SEQ ID NO: 1254 (SEQ ID NO: 1250 / SEQ ID NO: 1254), a promoter sequence of SEQ ID NO: 1252 and a terminator sequence of SEQ ID NO: 1256 (SEQ ID NO: 1252 / SEQ ID NO: 1256), a promoter sequence of SEQ ID NO: 1251 and a terminator sequence of SEQ ID NO: 1255 (SEQ ID NO: 1251 / SEQ ID NO: 1255), or a promoter sequence of SEQ ID NO: 1253 and a terminator sequence of SEQ ID NO: 1257 (SEQ ID NO: 1258). NO:1253 / SEQ ID NO:1257). Guide RNA expression was quantified by ddPCR and normalized to a housekeeping gene (GAPDH). A higher guide RNA expression to GAPDH expression (gRNA / GAPDH) ratio indicates increased guide RNA expression.
[0140] Fig.26 Schematic diagram of the flow sequencing process for screening promoter or terminator sequences. The screening begins with a pool of HEK293 cells with a single attp1 sequence. The next intermediate generated contains two boxes, one with a GFP-G67R ORF, which is non-fluorescent but has BFP indicating enrichment. The second box contains blasticidin resistance and BxbI integrase. The promoter or terminator sequence library is cloned into a plasmid containing mCherry and puromycin resistance. The pooled promoter or terminator sequence plasmid preparation can be transfected into intermediate cells and integrated by puromycin resistance enrichment with mCherry as an enrichment marker.
[0141] Fig. 27 Show Fig.26 Results of flow sequencing analysis described in , where each point represents the normalized performance of each termination sequence pooled from each of the three promoter sequences. The data points indicated by arrows represent superior termination sequences that entered the single copy evaluation, including SEQ ID NO: 1254 and SEQ ID NO: 1255, which showed similar expression compared to the WT mU7 termination sequence (SEQ ID NO: 1243).
[0142] Fig.28 The quantification of the DNA sequence of the cells with the termination sequence identified in the flow cytometry screening (eg Fig. 27 Bar graph of quantified GFP expression from expression constructs of the invention (described above). GFP expression was quantified by flow cytometry as the geometric mean fluorescence intensity (GFP gMFI). GFP expression from expression cassettes containing the terminator sequences SEQ ID NO:712, SEQ ID NO:868, SEQ ID NO:1021, SEQ ID NO:930, SEQ ID NO:1017, SEQ ID NO:1254, SEQ ID NO:771, SEQ ID NO:906, SEQ ID NO:1007, and SEQ ID NO:1002 was quantified and compared to the engineered mU7 terminator sequence SEQ ID NO:60. DETAILED DESCRIPTION
[0143] The present disclosure provides an expression cassette for expressing an RNA payload. The expression cassette described herein can be engineered to increase the expression of an encoded RNA payload sequence. In some embodiments, certain elements of the expression cassette, such as a promoter sequence, a core promoter sequence, or a transcription termination sequence, can be engineered to enhance payload expression. These sequence elements can be engineered from various endogenous promoters such as U1, U6, or U7 promoters to increase payload expression. Each sequence element of the expression cassette can be engineered to enhance the expression of an encoded RNA payload.
[0144] Promoter and terminator sequences
[0145] The expression cassette of the present disclosure may include a promoter sequence, an RNA payload encoding sequence and a termination sequence. The promoter may recruit transcription factors, polymerases (e.g., RNA polymerase II or RNA polymerase III) or other transcription mechanisms to promote the transcription of RNA payloads. For example, the expression cassette may initiate the transcription of guide RNA for RNA editing, guide RNA for DNA editing, tracrRNA, siRNA, shRNA or miRNA or antisense oligonucleotides. In some embodiments, the promoter may be engineered to increase the expression of the RNA payload controlled by promoter transcription. The termination sequence may enhance the termination of transcription and initiate the turnover of transcription, thereby increasing the transcription of payload. In some embodiments, the termination sequence may be engineered to enhance the expression of RNA payloads. Sequence elements (e.g., transcription factor binding sequences, transcription initiation sequences, termination sequences or combinations thereof) in the promoter or termination sequence may be engineered to enhance payload expression. The sequence elements may be interchanged with sequence elements (such as U1, U6 or U7 promoters) of endogenous RNA promoters.
[0146] The expression cassette can be engineered from an endogenous sequence. For example, the expression cassette can be engineered from an endogenous U1, U2, U3, U4, U5, U6 or U7 sequence. The endogenous sequence can be from any organism, including humans, mice or other mammals. In some embodiments, the expression cassette can include a promoter engineered from an endogenous promoter such as an endogenous U1, U2, U3, U4, U5, U6 or U7 promoter. In some embodiments, the expression cassette can include a transcription termination sequence engineered from an endogenous transcription termination sequence such as an endogenous U1, U2, U3, U4, U5, U6 or U7 transcription termination sequence.
[0147] The present disclosure provides regulatory elements for enhancing the optimal expression of small RNA payloads such as engineered guide RNA. Regulatory elements may refer to many different regions in the natural human genome, but as disclosed herein, screening has been performed in large-scale assays to determine the regulatory element combination that provides enhanced guide RNA expression. The expression cassette of the present disclosure includes both regulatory elements and payloads. For example, the expression cassette may include a regulatory element, and the regulatory element includes a portion of a natural human genome or a natural mouse genome promoter region. In some embodiments, the expression cassette may include a regulatory element, and the regulatory element includes a squirrel monkey herpes virus U-RNA (HSUR) element. In some embodiments, the expression cassette may include a regulatory element, and the regulatory element includes a mutant form of a natural human genome promoter region or a mutant form of a natural mouse genome promoter region. In some embodiments, the carrier of the present disclosure provides two expression cassettes, wherein there is a natural promoter region and a mutant promoter region. The expression cassette of the present disclosure is engineered to locate the promoter region 5' or upstream of a therapeutic payload (e.g., a small RNA sequence, such as an engineered guide RNA).
[0148] In addition, the regulatory element may include a portion of a natural human genome termination region, a natural mouse genome termination sequence, or a squirrel monkey herpes virus U-RNA (HSUR) termination sequence. The regulatory element may also include a portion of a mutated human genome termination region or a mutated mouse genome termination sequence. In some embodiments, the vector of the present disclosure provides two expression cassettes, wherein a natural termination region and a mutated termination region are present. The expression cassette of the present disclosure is engineered to position the termination region 3' or downstream of the therapeutic payload.
[0149] The promoter region of the present disclosure can be decomposed into multiple elements, including (from 5' to 3') distal sequence elements (DSE) and proximal sequence elements (PSE). These different elements can play different roles in the rate and efficiency of transcription of downstream payloads. In some embodiments, the PSE is part of the core promoter region. The PSE may be bound by the snRNA activated protein complex (SNAPc). SNAPc is a transcription factor that is important for transcription initiation and may promote the binding or recruitment of other transcription factors (such as TBP, TFIIA, TFIIB, TFIIE and TFIIF). In some embodiments, the DSE is part of an enhancer region. The DSE may bind to factors that help stabilize the transcription factors and transcription machinery on the PSE. In some embodiments, the DSE comprises an SPH element that recruits STAF transcription factors (e.g., ZNF143 transcription factors). The STAF transcription factor (e.g., ZNF143 transcription factor) is a zinc finger protein and comprises an activation domain that can activate RNA polymerase promoters (e.g., mRNA type RNA polymerase II promoter, type 3 RNA polymerase III promoter, and RNA polymerase II snRNA promoter). The SPH element may also comprise a ZNF143 motif capable of recruiting a zinc finger 143 (ZNF143) transcription factor. In some embodiments, the DSE comprises an OCT-1 element, which comprises an octamer sequence that recruits the Oct-1 transcription factor. Modification of any one of the DSE and PSE regions or other portions of the promoter region, or combination selection of different DSE and PSE regions, can improve the rate and efficiency of downstream payload transcription. The distance between the DSE and PSE can vary. In some embodiments, the distance between the DSE and PSE is shortened compared to the native promoter sequence. In some embodiments, the distance between the DSE and PSE is extended compared to the native promoter sequence. In some embodiments, the present disclosure provides a promoter from a natural human genome that has been adapted for use in a heterologous system that requires transcription of a therapeutic payload. In some embodiments, the present disclosure provides a promoter having modifications in a DSE as compared to a natural human genome DSE or a natural mouse genome DSE, and these modifications are part of the enhancer region of the promoter. The regions in the DSE that are important for engineering include SPH elements (recruiting the transcription factor STAF) and OCT-1 transcription factor (TF) binding sequences. In some embodiments, the SPH element comprises a zinc finger 143 (ZNF143) motif (recruiting zinc fingers). In some embodiments, the SPH element is a ZNF143 element (e.g., a zinc finger 143 (ZNF143) motif (recruiting zinc fingers)).These SPH regions (e.g., ZNF143 motifs) and OCT-1TF binding regions may also be referred to as regulatory factors. As disclosed herein, a promoter sequence having optimal elements within the DSE may result in enhanced transcription of a downstream small RNA payload. In some embodiments, the promoter sequence disclosed herein has one or more regions within it that correspond to SPH elements (e.g., ZNF143 motifs) and OCT-1TF binding sequences.
[0150] Sequence element
[0151] Engineering the expression cassette may include integrating or replacing engineered sequence elements into the expression construct. In some embodiments, elements present in the DSE or PSE in the promoter may be integrated or replaced by engineered elements. In some embodiments, sequence elements present in the termination sequence may be integrated or replaced by engineered elements. For example, endogenous transcription factor binding sequences (e.g., endogenous SPH elements, such as ZNF143 binding sequences, endogenous OCT-1 binding sequences, or endogenous GABP binding sequences) present in the DSE may be replaced by engineered transcription factor binding sequences (e.g., engineered SPH elements, such as ZNF143 binding sequences, engineered OCT-1 binding sequences, or engineered GABP binding sequences). Alternatively or additionally, endogenous core promoter sequence elements (e.g., endogenous proximal sequence elements or endogenous TATA boxes) may be replaced by engineered core promoter sequences (e.g., engineered proximal sequence elements or engineered TATA boxes). Alternatively or additionally, an endogenous termination sequence element (eg, an endogenous 3'box sequence element) can be replaced by an engineered termination sequence element (eg, an engineered 3'box sequence element). Table 1 provides examples of engineered sequence elements that can be inserted or replaced into an expression cassette.
[0152] Table 1 - Exemplary engineered sequence elements
[0153]
[0154] In some embodiments, the expression cassette can comprise one or more engineered sequence elements provided in Table 1. For example, the expression cassette can comprise a DSE having an engineered SPH element (e.g., a ZNF143 element) comprising a zinc finger 143 motif of any one of SEQ ID NO: 24-SEQ ID NO: 26 that binds to a ZNF143 transcription factor, a DSE having an engineered OCT-1 transcription factor binding site of any one of SEQ ID NO: 27-SEQ ID NO: 30 that binds to an OCT-1 transcription factor, an engineered proximal sequence element (PSE) of any one of SEQ ID NO: 31-SEQ ID NO: 37 that recruits SNAPc and phosphorylates RNA polymerase II transcription machinery, an engineered transcription termination sequence element (e.g., a 3' box sequence element) of any one of SEQ ID NO: 38-SEQ ID NO: 42 that initiates transcription termination, or a combination thereof.
[0155] The engineered SPH element comprising the zinc finger 143 motif can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity to any one of SEQ ID NO: 24-SEQ ID NO: 26. In some embodiments, the SPH element comprising the engineered zinc finger 143 motif can replace an endogenous SPH element comprising the zinc finger 143 motif of SEQ ID NO: 20.
[0156] The engineered OCT-1 transcription factor binding site may have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity to any one of SEQ ID NO: 27-SEQ ID NO: 30. In some embodiments, the engineered OCT-1 transcription factor binding site may replace the endogenous OCT-1 transcription factor binding site of SEQ ID NO: 21 in a distal sequence element (DSE).
[0157] Table 2 provides additional exemplary PSE sequences of the present disclosure.
[0158] Table 2 - Additional exemplary PSE sequences
[0159]
[0160]
[0161] The PSE may have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity to any one of SEQ ID NO:31-SEQ ID NO:37 or SEQ ID NO:67-SEQ ID NO:120. In some embodiments, the PSE may replace the endogenous PSE of SEQ ID NO:22. In some embodiments, the PSE that may be included in the engineered promoter sequence may have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity to any one of SEQ ID NO:67-SEQ ID NO:120. In some embodiments, the promoter sequence comprises a PSE sequence of SEQ ID NO:31-SEQ ID NO:37 or SEQ ID NO:67-SEQ ID NO:120. In some embodiments, the PSE is selected from SEQ ID NO:31-SEQ ID NO:37 or SEQ ID NO:67-SEQ ID NO:120. The PSE can be selected or engineered from a PSE of an endogenous gene. For example, the PSE can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity with a PSE from a U1, U2, U4, U5, U6, U7, U3, SNORD13, SNORD118, RPPH1, TRNAU1, 7SK, RNY3, or RNY4 gene. In some embodiments, the engineered promoter can include a PSE (e.g., any one of SEQ ID NO:31-SEQ ID NO:37 or SEQ ID NO:67-SEQ ID NO:120). In some embodiments, the engineered promoter can include a PSE (e.g., any one of SEQ ID NO:31-SEQ ID NO:37 or SEQ ID NO:67-SEQ ID NO:120) substituted for the PSE of SEQ ID NO:22.
[0162] In some embodiments, the engineered promoter may comprise a repetitive sequence element (e.g., a repeated transcription factor binding site) to enhance payload expression. For example, the engineered promoter may comprise a DSE having two or more SPH elements comprising a zinc finger 143 motif (e.g., SEQ ID NO: 20 or two or more of SEQ ID NO: 24-SEQ ID NO: 26, or a combination thereof). In another example, the engineered promoter may comprise a DSE having two or more OCT-1 transcription factor binding sites (e.g., SEQ ID NO: 21 or two or more of SEQ ID NO: 27-SEQ ID NO: 30, or a combination thereof). In another example, the engineered promoter may comprise two or more proximal sequence elements (PSEs) (e.g., two or more of SEQ ID NO: 22, SEQ ID NO: 31-SEQ ID NO: 37, SEQ ID NO: 67-SEQ ID NO: 120, or a combination thereof). The repetitive sequences may be separated by spacer sequences.
[0163] In some embodiments, the engineered promoter may comprise a plurality of promoter elements (e.g., an SPH element comprising a zinc finger 143 motif, an OCT-1 transcription factor binding site, or a proximal sequence element). In some embodiments, the engineered promoter may include one or more SPH elements comprising an engineered zinc finger 143 motif of any one of SEQ ID NO:24-SEQ ID NO:26 (which binds to a ZNF143 transcription factor), one or more engineered OCT-1 transcription factor binding sites of any one of SEQ ID NO:27-SEQ ID NO:30 (which binds to an OCT-1 transcription factor), or one or more engineered proximal sequence elements (PSEs) of any one of SEQ ID NO:31-SEQ ID NO:37, SEQ ID NO:67-SEQ ID NO:120. The engineered promoter may also comprise an endogenous SPH element comprising the zinc finger 143 motif of SEQ ID NO:20, the endogenous OCT-1 transcription factor binding site of SEQ ID NO:21, or the endogenous proximal sequence element (PSE) of SEQ ID NO:22.
[0164] The engineered transcription termination sequence may comprise a 3' box sequence element. The 3' box element may have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity to any one of SEQ ID NO:40-SEQ ID NO:42. In some embodiments, the 3' box sequence element may comprise the sequence GTTYN 0-3 AARRYAGA (SEQ ID NO: 38), wherein each N is independently A, T, C, or G, each R is independently A or G, and each Y is independently C or T. In some embodiments, the 3' box sequence element may comprise the sequence GTTTN 1-4 AANARNAGA (SEQ ID NO: 39), wherein each N is independently A, T, C, or G, and each R is independently A or G. In some embodiments, the engineered transcription termination sequence can replace the endogenous 3' box sequence element of sequence SEQ ID NO:23.
[0165] Table 3 provides additional exemplary 3' box sequence elements that may be included in the engineered termination sequences of the present disclosure.
[0166] Table 3 - Additional exemplary 3' box sequence elements
[0167]
[0168]
[0169] The 3' box element may have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity with any one of SEQ ID NO: 40-SEQ ID NO: 42. In some embodiments, the engineered transcription termination sequence may replace the endogenous 3' box sequence element of SEQ ID NO: 23. In some embodiments, the 3' box sequence element that may be included in the engineered promoter sequence may have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity with any one of SEQ ID NO: 121-SEQ ID NO: 166. In some embodiments, the termination sequence comprises a 3' box sequence element SEQ ID NO: 40 - SEQ ID NO: 42 or SEQ ID NO: 121 - SEQ ID NO: 166. In some embodiments, the 3' box sequence element is selected from SEQ ID NO: 40 - SEQ ID NO: 42 or SEQ ID NO: 121 - SEQ ID NO: 166. The 3' box sequence element can be selected or engineered from a 3' box sequence element of an endogenous gene. For example, the 3'box sequence element can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 93%, at least about 95%, at least about 97%, or about 100% sequence identity with the 3'box sequence element from U1, U2, U4, U5, U6, U7, U3, SNORD13, SNORD118, RPPH1, TRNAU1, 7SK, RNY3 or RNY4 gene. In some embodiments, the engineered termination sequence can include a 3'box sequence element (e.g., any one of SEQ ID NO: 40-SEQ ID NO: 42 or SEQ ID NO: 121-SEQ ID NO: 166). In some embodiments, the engineered termination sequence can include a 3' box sequence element (eg, any one of SEQ ID NO:40 - SEQ ID NO:42 or SEQ ID NO:121 - SEQ ID NO:166) in place of the 3' box sequence element of SEQ ID NO:23.
[0170] Promoter
[0171] The expression cassette may include a promoter. The promoter may be an endogenous promoter. The promoter may be an engineered promoter that is engineered to increase the expression of an RNA payload sequence that is transcriptionally controlled by the promoter. Table 4 provides examples of endogenous promoters (e.g., SEQ ID NO: 13-SEQ ID NO: 15), engineered promoters (e.g., SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 1241, SEQ ID NO: 1248, SEQ ID NO: 1249, SEQ ID NO: 1252, SEQ ID NO: 1253, and SEQ ID NO: 1258-SEQ ID NO: 1261), and additional promoters (e.g., SEQ ID NO: 1250, SEQ ID NO: 1251, SEQ ID NO: 1262, and SEQ ID NO: 1263).
[0172] Table 4 - Exemplary promoter sequences
[0173]
[0174]
[0175]
[0176]
[0177] In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to any of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263. In some embodiments, the engineered promoter used to enhance expression of the RNA payload can be a variant of the promoter (e.g., a variant of any one of SEQ ID NO: 13 - SEQ ID NO: 15, SEQ ID NO: 1250, SEQ ID NO: 1251, SEQ ID NO: 1262, and SEQ ID NO: 1263). In some embodiments, the engineered promoter can include a variant of any of SEQ ID NO: 13 - SEQ ID NO: 15, SEQ ID NO: 1250, SEQ ID NO: 1251, SEQ ID NO: 1262, and SEQ ID NO: 1263, which sequences have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to any of SEQ ID NO: 13 - SEQ ID NO: 15, SEQ ID NO: 1250, SEQ ID NO: 1251, SEQ ID NO: 1262, and SEQ ID NO: 1263, and have a sequence identity of at least about 100% with respect to any of SEQ ID NO: 13 - SEQ ID NO: 15, SEQ ID NO: 1250, SEQ ID NO: 1251, SEQ ID NO: 1262, and SEQ ID NO: 1263. Any one of SEQ ID NO: 1250, SEQ ID NO: 1251, SEQ ID NO: 1262 and SEQ ID NO: 1263 has at least one nucleotide substitution.
[0178] In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:13. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:15. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:17. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1241. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:1250.In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1251. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1252. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1253. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1262. In some embodiments, the promoter used to enhance expression of the RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1263.
[0179] The engineered promoter can enhance expression of an RNA payload controlled by the engineered promoter compared to an endogenous promoter (eg, an endogenous U1 promoter, an endogenous U6 promoter, or an endogenous U7 promoter). In some embodiments, the engineered promoter (e.g., a promoter comprising any one of the sequences SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 1241, SEQ ID NO: 1248, SEQ ID NO: 1249, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1258-SEQ ID NO: 1261) can increase expression of the RNA payload by at least about 5%, at least about 10%, at least about 15%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, or at least about 50% relative to an endogenous promoter (e.g., an endogenous U1 promoter, an endogenous U6 promoter, or an endogenous U7 promoter). In some embodiments, the engineered promoter (e.g., a promoter comprising a sequence of any one of SEQ ID NO: 16, SEQ ID NO: 17, SEQ ID NO: 1241, SEQ ID NO: 1248, SEQ ID NO: 1249, SEQ ID NO: 1252, SEQ ID NO: 1253, and SEQ ID NO: 1258-SEQ ID NO: 1261) can increase expression of the RNA payload by from about 5% to about 50%, from about 10% to about 50%, from about 15% to about 50%, from about 20% to about 50%, from about 25% to about 50%, from about 30% to about 50%, from about 35% to about 50%, from about 40% to about 50%, relative to an endogenous promoter (e.g., an endogenous U1 promoter, an endogenous U6 promoter, or an endogenous U7 promoter). From about 45% to about 50%, from about 5% to about 40%, from about 10% to about 40%, from about 15% to about 40%, from about 20% to about 40%, from about 25% to about 40%, from about 30% to about 40%, from about 35% to about 40%, from about 5% to about 30%, from about 10% to about 30%, from about 15% to about 30%, from about 20% to about 30%, from about 5% to about 30%, from about 10% to about 20% or from about 15% to about 20%.
[0180] In some embodiments, the promoter sequence can enhance the transcription of the RNA payload. The promoter sequence can be located upstream of the payload sequence. Table 5 provides additional exemplary promoter sequences of the present disclosure.
[0181] Table 5 - Additional exemplary promoter sequences
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188]
[0189]
[0190]
[0191]
[0192]
[0193]
[0194]
[0195]
[0196]
[0197]
[0198]
[0199]
[0200]
[0201]
[0202]
[0203]
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210]
[0211]
[0212]
[0213]
[0214]
[0215]
[0216]
[0217]
[0218]
[0219]
[0220]
[0221]
[0222]
[0223]
[0224]
[0225]
[0226]
[0227]
[0228]
[0229]
[0230]
[0231]
[0232]
[0233]
[0234]
[0235]
[0236]
[0237]
[0238]
[0239]
[0240]
[0241]
[0242]
[0243]
[0244]
[0245]
[0246]
[0247]
[0248]
[0249]
[0250]
[0251]
[0252]
[0253]
[0254]
[0255]
[0256]
[0257]
[0258]
[0259]
[0260]
[0261]
[0262]
[0263]
[0264]
[0265]
[0266]
[0267]
[0268]
[0269]
[0270]
[0271]
[0272]
[0273]
[0274]
[0275]
[0276]
[0277]
[0278]
[0279]
[0280]
[0281]
[0282]
[0283] In some embodiments, the promoter sequence can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to any of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263. In some aspects, the promoter sequence includes the sequence SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263. In some aspects, the promoter sequence is selected from the sequence SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263. In some embodiments, the PSE of any one of the promoter sequences SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263 is replaced with a PSE of any one of SEQ ID NO: 31–SEQ ID NO: 37, or SEQ ID NO: 67–SEQ ID NO: 120. In some aspects, the PSE of any one of SEQ ID NO:31-SEQ ID NO:37 or SEQ ID NO:67-SEQ ID NO:120 is inserted or replaced into any one of the promoters SEQ ID NO:13-SEQ ID NO:17, SEQ ID NO:167-SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248-SEQ ID NO:1253, or SEQ ID NO:1259-SEQ ID NO:1263.In some embodiments, the PSE sequence is extracted from any one of SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263 and inserted into a different promoter (e.g., any one of SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263). In some embodiments, the PSE of any one of the promoters SEQ ID NO:13–SEQ ID NO:17, SEQ ID NO:167–SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248–SEQ ID NO:1253, or SEQ ID NO:1259–SEQ ID NO:1263 is replaced with a PSE extracted from a different promoter (e.g., any one of SEQ ID NO:13–SEQ ID NO:17, SEQ ID NO:167–SEQ ID NO:707, SEQ ID NO:1241, SEQ ID NO:1248–SEQ ID NO:1253, or SEQ ID NO:1259–SEQ ID NO:1263), or is replaced with a PSE of any one of SEQ ID NO:31–SEQ ID NO:37, or SEQ ID NO:67–SEQ ID NO:120.
[0284] The promoter of the present disclosure may have nucleotide insertions or deletions on either side of the promoter. Nucleotide bases may be inserted or deleted between the promoter and the 5'ITR or between the promoter and the payload. In some embodiments, the promoter sequence of the present disclosure (e.g., any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253 or SEQ ID NO: 1259-SEQ ID NO: 1263) may be shortened by 1 to 2, 1 to 3, 1 to 5, 1 to 10 or 1 to 20 nucleotide bases from the 5' end, the 3' end or both the 5' end and the 3' end. In some embodiments, the promoter (e.g., any of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263) can be truncated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the 5' end, the 3' end, or both the 5' end and the 3' end. In some embodiments, 1 to 2, 1 to 3, 1 to 5, 1 to 10, or 1 to 20 nucleotide bases may be added to the 5' end, the 3' end, or both the 5' end and the 3' end of a promoter sequence of the present disclosure (e.g., any one of SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263). In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides may be added to the 5' end, 3' end, or both the 5' end and 3' end of the promoter (e.g., any one of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263). The nucleotide added to the 5' end or 3' end of the promoter may be selected from any nucleotide (e.g., A, T, C or G).For example, SEQ ID NO: 1250 comprises an 18-nucleotide base truncation at the 5' end of SEQ ID NO: 376. In another example, SEQ ID NO: 1251 comprises a 2-nucleotide base truncation at the 5' end and a 2-nucleotide base addition at the 3' end of SEQ ID NO: 168.
[0285] The promoter (e.g., any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263) can have nucleotides added to the 5' end to extend the expression cassette. In some embodiments, the promoter (e.g., any one of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263) can have additional nucleotides added to the 5' end to extend the promoter to a total length of 200 nucleotides, 300 nucleotides, 400 nucleotides, or 500 nucleotides. For example, SEQ ID NO: 1262 is an extended version of SEQ ID NO: 1250, wherein an additional 118 nucleotides are added at the 5' end to extend to a total length of 400 nucleotides. In another example, SEQ ID NO: 1263 is an extended version of SEQ ID NO: 1251, wherein an additional 118 nucleotides are added at the 5' end to extend to a total length of 400 nucleotides.
[0286] Termination sequence
[0287] The expression cassette may include a terminator sequence (also referred to as a terminator). The terminator sequence may be an endogenous terminator sequence. The terminator sequence may be an engineered terminator sequence that is engineered to increase RNA payload expression. Table 6 provides examples of endogenous terminator sequences (e.g., SEQ ID NO: 1243) engineered terminator sequences (e.g., SEQ ID NO: 60, SEQ ID NO: 1242, SEQ ID NO: 1256, SEQ ID NO: 1257, SEQ ID NO: 1275, or SEQ ID NO: 1287–SEQ ID NO: 1289) and other terminator sequences (e.g., SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1002, SEQ ID NO: 1007, SEQ ID NO: 1017, SEQ ID NO: 1021, SEQ ID NO: 1244–SEQ ID NO: 1247, SEQ ID NO: 1254, SEQ ID NO: 1255, or SEQ ID NO: 1264–SEQ ID NO: 1272).
[0288] Table 6 - Exemplary Termination Sequences
[0289]
[0290]
[0291]
[0292] In some embodiments, the expression cassette comprises an engineered termination sequence (e.g., SEQ ID NO: 60, SEQ ID NO: 1242, SEQ ID NO: 1256, SEQ ID NO: 1257, SEQ ID NO: 1275, or SEQ ID NO: 1287-SEQ ID NO: 1289). The engineered termination sequence can enhance expression of a payload (e.g., a small RNA payload) encoded by the expression cassette. In some embodiments, the engineered promoter sequence can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:60, SEQ ID NO:1242, SEQ ID NO:1256, SEQ ID NO:1257, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289.
[0293] In some embodiments, the expression cassette comprises a termination sequence that can enhance expression of a payload (e.g., a small RNA payload) encoded by the expression cassette. In some embodiments, the termination sequence can be identical to SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1002, SEQ ID NO: 1007, SEQ ID NO: 1017, SEQ ID NO: 1021, SEQ ID NO: 1242, SEQ ID NO: 1243-SEQ ID NO: 1247, SEQ ID NO: 1254, SEQ ID NO: 1255, SEQ ID NO: 1256, SEQ ID NO: 1257, SEQ ID NO: 1264-SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287-SEQ ID NO: 1291. NO:1289 has at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity.
[0294] In some embodiments, the 3' box sequence element that can be included in the engineered termination sequence can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or about 100% sequence identity to any one of SEQ ID NO:40-SEQ ID NO:42 or SEQ ID NO:121-SEQ ID NO:166. In some embodiments, the termination sequence comprises the sequence SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1002, SEQ ID NO:1007, SEQ ID NO:1017, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1256, SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289. In some embodiments, the termination sequence is selected from SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1002, SEQ ID NO:1007, SEQ ID NO:1017, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1256, SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289.
[0295] In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:60. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:771. In some embodiments, the termination sequence for enhancing expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO:930. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1002. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1007.In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1017. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1021. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1242. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1254. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1255.In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1257. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1264. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1265. In some embodiments, the termination sequence used to enhance expression of an RNA payload can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1269.
[0296] In some embodiments, a termination sequence (also referred to as a terminator) can enhance transcription of an RNA payload. The termination sequence can be located downstream of the payload sequence. Table 7 provides additional exemplary termination sequences of the present disclosure.
[0297] Table 7 - Additional exemplary termination sequences
[0298]
[0299]
[0300]
[0301]
[0302]
[0303]
[0304]
[0305]
[0306]
[0307]
[0308]
[0309]
[0310]
[0311]
[0312]
[0313]
[0314]
[0315]
[0316]
[0317]
[0318]
[0319]
[0320]
[0321]
[0322]
[0323]
[0324]
[0325]
[0326]
[0327]
[0328]
[0329]
[0330]
[0331]
[0332] In some embodiments, the terminator sequence can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to any of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289. In some embodiments, the termination sequence comprises the sequence SEQ ID NO: 60, SEQ ID NO: 708 - SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243 - SEQ ID NO: 1247, SEQ ID NO: 1254 - SEQ ID NO: 1257, SEQ ID NO: 1264 - SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287 - SEQ ID NO: 1289. In some embodiments, the termination sequence is selected from the group consisting of SEQ ID NO: 60, SEQ ID NO: 708 - SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243 - SEQ ID NO: 1247, SEQ ID NO: 1254 - SEQ ID NO: 1257, SEQ ID NO: 1264 - SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287 - SEQ ID NO: 1289.In some embodiments, the 3' box sequence element of the termination sequence of any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289 is replaced with the 3' box sequence element of any one of SEQ ID NO:40–SEQ ID NO:42 or SEQ ID NO:121–SEQ ID NO:166. In some embodiments, the 3' box sequence element of any one of SEQ ID NO:40-SEQ ID NO:42 or SEQ ID NO:121-SEQ ID NO:166 is inserted or replaced into the termination sequence of any one of SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289. In some embodiments, the 3' box sequence element is extracted from any one of SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289 and is inserted into a different termination sequence (e.g., SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289). NO: any one of 1289).In some embodiments, the 3' box sequence element of the termination sequence of any one of SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289 is extracted from a different termination sequence (e.g., SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289). NO:1289), or replaced by the 3' box sequence element of any one of SEQ ID NO:40-SEQ ID NO:42 or SEQ ID NO:121-SEQ ID NO:166.
[0333] The termination sequence of the present disclosure can have the insertion or deletion of nucleotides on either side of the termination sequence. Nucleotide bases can be inserted into the 3' end of the termination sequence or lack at the 3' end of the termination sequence to extend the length of the box. In some embodiments, the termination sequence of the present disclosure (for example, SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275 or SEQ ID NO:1287-SEQ ID NO:1289) can be shortened by 1 to 2, 1 to 3, 1 to 5, 1 to 10 or 1 to 20 nucleotide bases from the 5' end, 3' end or both the 5' end and the 3' end. In some embodiments, the termination sequence (e.g., any of SEQ ID NO:60, SEQ ID NO:708 - SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243 - SEQ ID NO:1247, SEQ ID NO:1254 - SEQ ID NO:1257, SEQ ID NO:1264 - SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287 - SEQ ID NO:1289) can be truncated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides from the 5' end, the 3' end, or both the 5' and 3' ends. In some embodiments, 1 to 2, 1 to 3, 1 to 5, 1 to 10, or 1 to 20 nucleotide bases can be added to the 5' end, the 3' end, or both the 5' end and the 3' end of the termination sequence (e.g., any of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289).In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19 or 20 nucleotides can be added to the 5' end, 3' end, or both of the 5' end and 3' end of the termination sequence (e.g., any one of SEQ ID NO: 60, SEQ ID NO: 708 - SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243 - SEQ ID NO: 1247, SEQ ID NO: 1254 - SEQ ID NO: 1257, SEQ ID NO: 1264 - SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287 - SEQ ID NO: 1289). The nucleotide added to the 5' end or 3' end of the termination sequence can be selected from any nucleotide (e.g., A, T, C or G). For example, SEQ ID NO: 1254 comprises 1 nucleotide base deletion on the 5' end and 2 nucleotide deletions on the 3' end of SEQ ID NO: 917. In another example, SEQ ID NO: 1255 comprises 1 nucleotide base deletion on the 5' end and 1 nucleotide base addition on the 3' end of SEQ ID NO: 709. For example, SEQ ID NO: 1287 comprises 2 nucleotide base deletions on the 5' end of SEQ ID NO: 60. For example, SEQ ID NO: 1288 comprises 4 nucleotide base deletions on the 5' end of SEQ ID NO: 60. For example, SEQ ID NO: 1289 comprises 6 nucleotide base deletions on the 5' end of SEQ ID NO: 60.
[0334] A termination sequence (e.g., any one of SEQ ID NO:60, SEQ ID NO:708 - SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243 - SEQ ID NO:1247, SEQ ID NO:1254 - SEQ ID NO:1257, SEQ ID NO:1264 - SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287 - SEQ ID NO:1289) may be added to the 3' end to extend the length of the expression cassette. In some embodiments, the termination sequence (e.g., any one of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: 1257, SEQ ID NO: 1264-SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287-SEQ ID NO: 1289) may have additional nucleotides added to the 3' end to extend the termination sequence to a total length of 100 nucleotides, 150 nucleotides, 200 nucleotides, or 300 nucleotides. For example, SEQ ID NO: 1264 is an extended version of SEQ ID NO: 1002, in which an additional 100 nucleotides are added to the 3' end to extend to a total length of 200 nucleotides. In another example, SEQ ID NO: 1265 is an extended version of SEQ ID NO: 1017, wherein an additional 100 nucleotides are added to the 3' end to extend to a total length of 200 nucleotides.
[0335] Small noncoding RNA (snRNA) undergoes post-transcriptional cap switching, in which the monomethylguanosine (MMG) cap is converted to a trimethylguanosine (TMG) cap by the TGSI enzyme. Effective cap switching is essential for the formation of mature snRNA and subsequent transport to the nucleus by snurportin1. The dipurine (adenine or guanine) sequence on the 5' end of the guide RNA may contribute to effective cap switching. The present disclosure provides an expression cassette, wherein the expressed gRNA has an additional 2 bases at the 5' end, wherein the additional 2 bases are purines (adenine or guanine). Therefore, the present disclosure provides, in some embodiments, an expression cassette having a gRNA starting with AA, GG, GA or AG. For example, the SNCA guide RNA (SEQ ID NO: 1290) can have an additional G at the 5' end, thereby generating a SNCA guide RNA sequence SEQ ID NO: 1274 containing GA at the 5' end.
[0336] Promoter and terminator sequence pairing
[0337] The expression cassettes of the present disclosure may comprise a promoter sequence (e.g., any one of SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263), a payload sequence under the transcriptional control of the promoter sequence, and a termination sequence (e.g., any one of SEQ ID NO: 60, SEQ ID NO: 708–SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243–SEQ ID NO: 1247, SEQ ID NO: 1254–SEQ ID NO: 1257, SEQ ID NO: 1264–SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287–SEQ ID NO: 1289).
[0338] In some embodiments, the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1007, SEQ ID NO: 1021, SEQ ID NO: NO: 1242, SEQ ID NO: 1254, SEQ ID NO: 1255, SEQ ID NO: 1257, or SEQ ID NO: 1269. In some embodiments, the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO: 17, SEQ ID NO: 1250, or SEQ ID NO: 1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO: 1241, SEQ ID NO: 1251, SEQ ID NO: 1252, SEQ ID NO: 1253, or SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence.In some embodiments, the expression cassette comprises: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of the following: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264, or SEQ ID NO: 1265; or b) SEQ ID NO: 60, SEQ ID NO: 771, SEQ ID NO: 930, SEQ ID NO: 1007, SEQ ID NO: 1021, SEQ ID NO: 1242, SEQ ID NO: 1254, SEQ ID NO: 1255, SEQ ID NO: 1257, or SEQ ID NO: 1269.
[0339] In some embodiments, the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any of SEQ ID NO: 13-SEQ ID NO: 17, SEQ ID NO: 167-SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248-SEQ ID NO: 1253, or SEQ ID NO: 1259-SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of SEQ ID NO: 60, SEQ ID NO: 708-SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243-SEQ ID NO: 1247, SEQ ID NO: 1254-SEQ ID NO: 1257, SEQ ID NO: 1264-SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: NO: 1287 - SEQ ID NO: 1289. In some embodiments, the expression cassette comprises: a promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence. In some embodiments, the expression cassette comprises: a promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a termination sequence comprising a sequence having at least 80% sequence identity to any of the following: SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289.
[0340] In one embodiment, the expression cassette comprises:
[0341] (i) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1264;
[0342] (ii) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1265;
[0343] (iii) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1254;
[0344] (iv) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1255;
[0345] (v) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1257;
[0346] (vi) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 60;
[0347] (vii) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1242;
[0348] (viii) promoter SEQ ID NO: 1262 and terminator sequence SEQ ID NO: 1264;
[0349] (ix) promoter SEQ ID NO: 1262 and termination sequence SEQ ID NO: 1265;
[0350] (x) promoter SEQ ID NO: 1262 and terminator sequence SEQ ID NO: 1254;
[0351] (xi) promoter SEQ ID NO: 1262 and termination sequence SEQ ID NO: 1255; (xii) promoter SEQ ID NO: 1262 and termination sequence SEQ ID NO: 1257; (xiii) promoter SEQ ID NO: 1262 and termination sequence SEQ ID NO: 60;
[0352] (xiv) promoter SEQ ID NO:1262 and termination sequence SEQ ID NO:1242; (xv) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1264; (xvi) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1265; (xvii) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1254; (xviii) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1255; (xix) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1257; (xx) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:60;
[0353] (xxi) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1242; (xxii) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1264; (xxiii) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1265; (xxiv) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1254; (xxv) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1255; (xxvi) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1257; (xxvii) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:60; (xxviii) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1242; (xxix) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1264; (xxx) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1265; (xxxi) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1254; (xxxii) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1255; (xxxiii) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1257; (xxxiv) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:60; (xxxv) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1242; (xxxvi) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1264; (xxxvii) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1265; (xxxviii) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1254; (xxxix) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1255; (xl) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1257; (xli) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:60;
[0354] (xlii) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1242; (xliii) promoter SEQ ID NO:17 and termination sequence SEQ ID NO:1269; (xliv) promoter SEQ ID NO:1262 and termination sequence SEQ ID NO:1269; (xlv) promoter SEQ ID NO:1250 and termination sequence SEQ ID NO:1269; (xlvi) promoter SEQ ID NO:1251 and termination sequence SEQ ID NO:1269; (xlvii) promoter SEQ ID NO:1252 and termination sequence SEQ ID NO:1269; (xlviii) promoter SEQ ID NO:1253 and termination sequence SEQ ID NO:1269;
[0355] (xlix) promoter SEQ ID NO: 17 and terminator sequence SEQ ID NO: 1017;
[0356] (1) promoter SEQ ID NO: 1262 and terminator sequence SEQ ID NO: 1017;
[0357] (li) promoter SEQ ID NO: 1250 and terminator sequence SEQ ID NO: 1017;
[0358] (lii) promoter SEQ ID NO: 1251 and terminator sequence SEQ ID NO: 1017;
[0359] (liii) promoter SEQ ID NO: 1252 and terminator sequence SEQ ID NO: 1017; or
[0360] (liv) promoter SEQ ID NO: 1253 and terminator sequence SEQ ID NO: 1017;
[0361] Further additional promoter / terminator sequence pairs
[0362] In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 17 and a termination sequence of SEQ ID NO: 1264. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 1262 and a termination sequence of SEQ ID NO: 1265. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 1250 and a termination sequence of SEQ ID NO: 1254. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 1251 and a termination sequence of SEQ ID NO: 1255. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 1252 and a termination sequence of SEQ ID NO: 1255. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 1253 and a termination sequence of SEQ ID NO: 1255. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 17 and a termination sequence of SEQ ID NO: 60. In one embodiment, the expression cassette comprises a promoter of SEQ ID NO: 17 and a termination sequence of SEQ ID NO: 1242. In one embodiment, the expression cassette comprises the promoter SEQ ID NO: 1262 and the termination sequence SEQ ID NO: 1269. In one embodiment, the expression cassette comprises the promoter SEQ ID NO: 17 and the termination sequence SEQ ID NO: 1265. In one embodiment, the expression cassette comprises the promoter SEQ ID NO: 17 and the termination sequence SEQ ID NO: 1017.
[0363] Payload
[0364] The expression cassette of the present disclosure can encode an RNA payload that is transcriptionally controlled by a promoter (e.g., an engineered promoter). In some embodiments, the RNA payload can encode a small RNA payload, such as a guide sequence (e.g., for RNA or DNA editing), tracrRNA, siRNA, shRNA or miRNA, an antisense oligonucleotide (e.g., for expression knockout), a structural element (e.g., an RNA hairpin), or a combination thereof. Provided herein are polynucleotides for engineering RNA payloads and editing the RNA payloads; and compositions comprising the engineered RNA payloads or the polynucleotides. As used herein, the term "engineered" with respect to RNA payloads or polynucleotides encoding the RNA payloads refers to non-naturally occurring RNA or polynucleotides encoding the RNA. For example, the present disclosure provides engineered polynucleotides encoding engineered guide RNAs. In some embodiments, the engineered guide comprises RNA. In some embodiments, the engineered guide comprises DNA. In some examples, the engineered guide comprises modified DNA bases and unmodified RNA bases. In some embodiments, the engineered guide comprises modified DNA bases or unmodified DNA bases. In some examples, the engineered guide comprises both DNA bases and RNA bases.
[0365] Guide RNA Payloads for RNA Editing
[0366] The expression cassettes described herein can be used to enhance engineered guide RNAs and engineered polynucleotides encoding the engineered guide RNAs for site-specific, selective editing of target RNAs via RNA editing entities or biologically active fragments thereof. The engineered guide RNAs disclosed herein can comprise a latent structure such that when the engineered guide RNAs hybridize with target RNAs to form guide-target RNA scaffolds, at least a portion of the latent structure exhibits at least a portion of the structural features described herein.
[0367] As described herein, the engineered guide RNA comprises a targeting domain complementary to the target RNA described herein. In this way, the guide RNA can be engineered to site-specifically / selectively target a specific target RNA and hybridize with it, thereby facilitating editing of specific nucleotides in the target RNA by RNA editing entities or their biologically active fragments. The targeting domain can include nucleotides positioned so that when the guide RNA hybridizes with the target RNA, the nucleotides are relative to the bases to be edited by the RNA editing entity or its biologically active fragment, and are not base-paired or incompletely base-paired with the bases to be edited. This mispairing helps to locate the editing of the RNA editing entity to the desired base of the target RNA. However, in some cases, in addition to the desired editing, there may also be some off-target editing, and in some cases it is a significant off-target editing.
[0368] The hybridization of the target RNA and the targeting domain of the guide RNA may produce a specific secondary structure in the guide-target RNA scaffold, which appears after hybridization and is referred to as a "potential structure" in this article. The potential structure may become a structural feature described herein when it is expressed, including mismatches, protrusions, internal loops and hairpins. Without being bound by theory, the existence of the structural features described herein produced after the guide RNA hybridizes with the target RNA configures the guide RNA to promote specific or selective targeted editing of the target RNA by RNA editing entities or their biologically active fragments. In addition, compared with constructs containing only mismatches or constructs with perfect complementarity to the target RNA, the structural features combined with the above-mentioned mismatches generally promote an increase in the amount of editing of target residues (e.g., adenosine residues), less off-target editing or both. Therefore, the potential structure in the engineered guide RNA of the present disclosure is rationally designed to produce specific structural features in the guide-target RNA scaffold can be a powerful tool to promote target RNA editing with high specificity, selectivity and robust activity.
[0369] In some examples, the engineered guides provided herein include engineered guides that can be configured to at least partially form a guide-target RNA scaffold with at least a portion of a target RNA molecule upon hybridization with the target RNA molecule, wherein the guide-target RNA scaffold includes at least one structural feature, and wherein the guide-target RNA scaffold recruits RNA editing entities and facilitates chemical modification of nucleotide bases in the target RNA molecule by the RNA editing entities.
[0370] In some examples, the target RNA of the engineered guide RNA of the present disclosure can be pre-mRNA or mRNA. In some embodiments, the engineered guide RNA of the present disclosure hybridizes with the sequence of the target RNA. In some embodiments, a portion (e.g., a targeting domain) of the engineered guide RNA hybridizes with the sequence of the target RNA. The portion of the engineered guide RNA hybridized with the target RNA is sufficiently complementary to the sequence of the target RNA to hybridize.
[0371] Targeting domain. The engineered guide RNA disclosed herein can be engineered in any manner suitable for RNA editing. In some examples, the engineered guide RNA typically comprises at least a targeting sequence that allows it to hybridize with a region of the target RNA molecule. The targeting sequence may also be referred to as a "targeting domain" or "targeting region."
[0372] As used herein, the term "targeting sequence" is used interchangeably with "targeting domain" or "targeting region" and refers to a polynucleotide sequence within an engineered guide RNA sequence that is at least partially complementary to a target polynucleotide. A target polynucleotide (e.g., a target RNA or a target DNA) can be a region of a polynucleotide of interest, such as a gene or a messenger RNA. As used herein, a "complementary" sequence refers to a sequence that is the reverse complement relative to a second sequence.
[0373] The targeting sequence of the engineered guide RNA allows the engineered guide RNA to hybridize with a target polynucleotide (e.g., a target RNA) by base pairing (such as Watson Crick base pairing). The targeting sequence can be located at the N-terminus or C-terminus or both of the engineered guide RNA, or the targeting sequence can be located within the engineered guide RNA. The targeting sequence can have any length sufficient to hybridize with the target polynucleotide. In some cases, the length of the targeting sequence can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 6, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112 2, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157 , 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or up to about 200 nucleotides.In one embodiment, the engineered polynucleotide comprises a length of about 25 to 200, 50 to 150, 75 to 100, 80 to 110, 90 to 120, 95 to 115, 60 to 200, 60 to 180, 60 to 160, 60 to 140, 70 to 200, 70 to 180, 70 to 160, 70 to 140, 80 to 200, 80 to 190, 80 to 170, 80 to 1 60, 80 to 150, 80 to 140, 80 to 130, 80 to 120, 90 to 200, 90 to 190, 90 to 180, 90 to 170, 90 to 160, 90 to 150, 90 to 140, 90 to 130, 90 to 120, 100 to 200, 100 to 190, 100 to 180, 100 to 170, 100 to 160, 100 to 150, 100 to 140, 100 to 130, 100 to 120, 110 to 200, 110 to 190, 110 to 180, 110 to 170, 110 to 160, 110 to 150, 110 to 140, 110 to 120, 120 to 200, 120 to 190, 120 to 180, 120 to 170, 120 to 160, 120 to 150, 120 to 140, 130 from 150 to 200, 130 to 190, 130 to 180, 130 to 170, 130 to 160, 130 to 150, 140 to 200, 140 to 190, 140 to 180, 140 to 170, 140 to 160, 150 to 200, 150 to 190, 150 to 180, 150 to 170, 160 to 200, 160 to 190 or 160 to 180 nucleotides.
[0374] The targeting sequence has at least partial sequence complementarity with the target polynucleotide. The targeting sequence can have a sequence complementarity sufficient to hybridize with the target nucleotide with the target polynucleotide. In some cases, the targeting sequence has 95%, 96%, 97%, 98%, 99% or 100% sequence complementarity with the target RNA. In some cases, the targeting sequence has less than 100% complementarity with the target polynucleotide sequence. For example, the targeting sequence may have a single base mismatch relative to the target polynucleotide when combined with the target polynucleotide. In other cases, the targeting sequence comprises at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 20, 30, 40 or up to about 50 base mismatches relative to the polynucleotide when combined with the target polynucleotide. In some aspects, nucleotide mismatches may be associated with structural features provided herein. In some aspects, the targeting sequence comprises at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or up to about 15 nucleotides that differ in complementarity from a wild-type polynucleotide of the subject target polynucleotide.
[0375] The targeting sequence comprises nucleotide residues that are complementary to the target polynucleotide. The targeting sequence may have a plurality of residues that are complementary to the target polynucleotide and are sufficient to hybridize with the target polynucleotide. The complementary residues may be continuous or discontinuous. In some cases, the targeting sequence comprises at least 50 nucleotides that are complementary to the target polynucleotide. In some cases, the targeting sequence comprises from 50 to 150 nucleotides that are complementary to the target polynucleotide. In some cases, the targeting sequence comprises from 50 to 200 nucleotides that are complementary to the target polynucleotide. In some cases, the targeting sequence comprises from 50 to 250 nucleotides that are complementary to the target polynucleotide. In some cases, the targeting sequence comprises from 50 to 300 nucleotides that are complementary to the target polynucleotide.In some cases, the targeting sequence comprises 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117 7, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 1 76, 177, 178, 179, 180, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, or 300 nucleotides complementary to the target polynucleotide. In some cases, the targeting sequence comprises more than 50 nucleotides total and has at least 50 nucleotides that are complementary to the target polynucleotide.In some cases, the targeting sequence comprises a total of from 50 to 400 nucleotides and has from 50 to 150 nucleotides with complementarity to the target polynucleotide. In some cases, the targeting sequence comprises a total of from 50 to 400 nucleotides and has from 50 to 200 nucleotides with complementarity to the target polynucleotide. In some cases, the targeting sequence comprises a total of from 50 to 400 nucleotides and has from 50 to 250 nucleotides with complementarity to the target polynucleotide. In some cases, the targeting sequence comprises a total of from 50 to 400 nucleotides and has from 50 to 300 nucleotides with complementarity to the target polynucleotide. In some cases, the at least 50 nucleotides with complementarity to the target polynucleotide are separated by one or more mispairings, one or more protrusions or one or more loops, or any combination thereof. In some cases, the nucleotides with complementarity to the target polynucleotide from 50 to 150 are separated by one or more mispairings, one or more protrusions or one or more loops, or any combination thereof. In some cases, the nucleotides from 50 to 200 with the target polynucleotide having complementarity are separated by one or more mispairings, one or more protrusions or one or more loops or any combination thereof. In some cases, the nucleotides from 50 to 250 with the target polynucleotide having complementarity are separated by one or more mispairings, one or more protrusions or one or more loops or any combination thereof. In some cases, the nucleotides from 50 to 300 with the target polynucleotide having complementarity are separated by one or more mispairings, one or more protrusions or one or more loops or any combination thereof. For example, the targeting sequence comprises a total of 54 nucleotides, wherein 25 nucleotides are complementary to the target polynucleotides successively, 4 nucleotides form protrusions, and 25 nucleotides are complementary to the target polynucleotides. As another example, the targeting sequence comprises a total of 118 nucleotides, wherein 25 nucleotides are complementary to the target polynucleotides successively, 4 nucleotides form protrusions, 25 nucleotides are complementary to the target polynucleotides, 14 nucleotides form loops, and 50 nucleotides are complementary to the target polynucleotides.
[0376] In some cases, the targeting domain has 95%, 96%, 97%, 98%, 99% or 100% sequence complementarity with the target RNA. In some cases, the targeting sequence has less than 100% complementarity with the target RNA sequence. For example, the targeting sequence and the region of the target RNA that can be bound by the targeting sequence can have a single base mismatch.
[0377] The targeting sequence may have sufficient complementarity with the target RNA to allow the targeting sequence to hybridize with the target RNA. In some embodiments, the targeting sequence has a minimum antisense complementarity of about 50 nucleotides or more with the target RNA. In some embodiments, the targeting sequence has a minimum antisense complementarity of about 60 nucleotides or more with the target RNA. In some embodiments, the targeting sequence has a minimum antisense complementarity of about 70 nucleotides or more with the target RNA. In some embodiments, the targeting sequence has a minimum antisense complementarity of about 80 nucleotides or more with the target RNA. In some embodiments, the targeting sequence has a minimum antisense complementarity of about 90 nucleotides or more with the target RNA. In some embodiments, the targeting sequence has a minimum antisense complementarity of about 100 nucleotides or more with the target RNA. In some embodiments, antisense complementarity refers to a non-continuous segment of a sequence. In some embodiments, antisense complementarity refers to a continuous segment of a sequence.
[0378] In some embodiments, the targeting sequence may exhibit potential structural features when hybridized with the target RNA to form a guide-target RNA scaffold. For example, potential structural features may include symmetrical protrusions, asymmetrical protrusions, symmetrical inner loops, asymmetrical inner loops, or combinations thereof. In some embodiments, the potential structural features may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 unpaired nucleotides on the target RNA side. In some embodiments, the potential structural features may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 unpaired nucleotides on the guide RNA side.
[0379] In some embodiments, the engineered guide RNA for RNA editing can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to SEQ ID NO: 1273, SEQ ID NO: 1274, SEQ ID NO: 61, or SEQ ID NO: 1290. For example, the engineered guide RNA of SEQ ID NO: 1273 can be used to target PMP22. In another example, the engineered guide RNA of SEQ ID NO: 1274 can be used to target SNCA. In another example, the engineered guide RNA of SEQ ID NO: 1290 can be used to target SNCA. In another example, the engineered guide RNA of SEQ ID NO: 61 can be used to target SERPINA1. Table 8 provides examples of engineered guide RNAs.
[0380] Table 8 - Engineered guide RNA
[0381]
[0382] Engineered guide RNA with a recruitment domain. In some instances, the subject engineered guide RNA includes a recruitment domain for recruiting RNA editing entities (e.g., ADARs), wherein in some cases, the recruitment domain is formed and exists in the absence of binding to the target RNA." Recruitment domain" may be referred to as "recruitment sequence" or "recruitment region" here. In some instances, the subject engineered guide may promote the editing of bases of nucleotides in the target sequence of the target RNA, and the editing results in regulating the expression of a polypeptide encoded by the target RNA. In some cases, regulation may be an increase or decrease in the expression of the polypeptide. In some cases, the engineered guide may be configured to promote RNA editing entities (e.g., ADARs or APOBECs) to edit bases of nucleotides or polynucleotides in RNA regions. To promote editing, the engineered polynucleotides of the present disclosure may recruit RNA editing entities (e.g., ADARs or APOBECs). Various RNA editing entities may be used to recruit domains. In some examples, the recruitment domain comprises: glutamate ionotropic receptor AMPA-type subunit 2 (GluR2), an Alu sequence, or in the case of recruiting APOBEC, an APOBEC recruitment domain.
[0383] In some instances, more than one recruitment domain may be included in the engineering guide of the present disclosure. In instances where a recruitment domain may be present, after the target sequence hybridizes with the target sequence of the target RNA, the recruitment domain may be used to position the RNA editing entity to effectively react with the subject target RNA. In some cases, the recruitment domain may allow temporary binding of the RNA editing entity to the engineering guide. In some instances, the recruitment domain allows permanent binding of the RNA editing entity to the engineering guide. The recruitment domain may be of any length. 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, up to about 80 nucleotides in length. 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, or 80 nucleotides in length. In some cases, the length of the recruitment domain can be about 45 nucleotides. In some cases, at least a portion of the recruitment domain comprises at least 1 to about 75 nucleotides. In some cases, at least a portion of the recruitment domain comprises about 45 nucleotides to about 60 nucleotides.
[0384] In some embodiments, the recruitment domain comprises a GluR2 sequence or a functional fragment thereof. In some cases, the GluR2 sequence can be recognized by an RNA editing entity (such as an ADAR or a biologically active fragment thereof). In some embodiments, the GluR2 sequence can be a non-naturally occurring sequence. In some cases, the GluR2 sequence can be modified, for example, for enhancing recruitment. In some embodiments, the GluR2 sequence can comprise a portion of a naturally occurring GluR2 sequence and a synthetic sequence.
[0385] In some examples, the recruitment domain comprises a GluR2 sequence, or a sequence having at least about 70%, 80%, 85%, 90%, 95%, 98%, 99% or 100% identity and / or length to GUGGAAU AGUAUAACAAUAUGCUAAAUGUUGUUAUAGUAUCCCAC (SEQ ID NO: 51). In some cases, the recruitment domain may have at least about 80% sequence homology to at least about 10, 15, 20, 25 or 30 nucleotides of SEQ ID NO: 51. In some examples, the recruitment domain may have a sequence homology and / or length to at least about 90%, 95%, 96%, 97%, 98% or 99% to SEQ ID NO: 51.
[0386] In addition, RNA editing entity recruitment domains are also contemplated. In one embodiment, the recruitment domain comprises an apolipoprotein B mRNA editing enzyme, a catalytic polypeptide-like (APOBEC) domain. In some cases, the APOBEC domain may comprise a non-naturally occurring sequence or a naturally occurring sequence. In some embodiments, the APOBEC domain encoding sequence may comprise a modified portion. In some cases, the APOBEC domain encoding sequence may comprise a portion of a naturally occurring APOBEC domain encoding sequence. In another embodiment, the recruitment domain may be from an Alu domain.
[0387] Any number of recruitment domains may be present in the engineered guides of the present disclosure. In some instances, at least about 1, 2, 3, 4, 5, 6, 7, 8, 9 or up to about 10 recruitment domains may be included in the engineered guide. The recruitment domain may be located at any position of the engineered guide RNA. In some cases, the recruitment domain may be located at the N-terminus, middle or C-terminus of the engineered guide RNA. The recruitment domain may be upstream or downstream of the targeting sequence. In some cases, the recruitment domain flanks the targeting sequence of the subject guide. The recruitment sequence may contain all ribonucleotides or deoxyribonucleotides, although in some cases a recruitment domain containing ribonucleotides and deoxyribonucleotides may not be excluded.
[0388] Engineered guide RNA with potential structure. In some examples, the engineered guide disclosed herein for promoting RNA editing entity editing target RNA can be an engineered potential guide RNA. "Engineering potential guide RNA" refers to an engineered guide RNA comprising a potential structure. "Potential structure" refers to a structural feature that is formed substantially only after the guide RNA hybridizes with the target RNA. For example, the sequence of the guide RNA provides one or more structural features, but these structural features are substantially only formed after hybridization with the target RNA, so one or more potential structural features appear as structural features after hybridization with the target RNA. Structural features are formed after the guide RNA hybridizes with the target RNA, and the potential structure provided in the guide RNA is thus revealed. The formation and structure of the potential structural features after binding to the target RNA depends on the guide RNA sequence. For example, the formation and structure of the potential structural features may depend on the pattern of complementary and mismatched residues in the guide RNA sequence relative to the target RNA. The guide RNA sequence can be engineered to have potential structural features formed after binding to the target RNA.
[0389] A double-stranded RNA (dsRNA) substrate can be formed after hybridization of an engineered guide RNA of the present disclosure with a target RNA. The resulting dsRNA substrate is also referred to herein as a "guide-target RNA scaffold".
[0390] Fig.16 A diagram showing various exemplary structural features present in the guide-target RNA scaffold formed after hybridization of the potential guide RNA disclosed herein with the target RNA. The exemplary structural features shown include an 8 / 7 asymmetric loop (i. 8 nucleotides on the target RNA side and 7 nucleotides on the guide RNA side), a 2 / 2 symmetric protrusion (ii. 2 nucleotides on the target RNA side and 2 nucleotides on the guide RNA side), a 1 / 1 mismatch (iii. 1 nucleotide on the target RNA side and 1 nucleotide on the guide RNA side), a 5 / 5 symmetric inner loop (iv. 5 nucleotides on the target RNA side and 5 nucleotides on the guide RNA side), a 24 bp region (v. 24 nucleotides on the target RNA side are base paired with 24 nucleotides on the guide RNA side) and a 2 / 3 asymmetric protrusion (vi. 2 nucleotides on the target RNA side and 3 nucleotides on the guide RNA side).
[0391] Unless otherwise stated, the number of participating nucleotides in a given structural feature is expressed as the ratio of the number of nucleotides on the target RNA side to the number of nucleotides on the guide RNA side. The position annotation of each figure is also shown in this legend. For example, the target nucleotide to be edited is designated as position 0. Downstream (3') of the target nucleotide to be edited, each nucleotide is counted in increments of +1. Upstream (5') of the target nucleotide to be edited, each nucleotide is counted in increments of -1. Therefore, the example 2 / 2 symmetrical protrusions in this legend are located at positions +12 to +13 in the guide-target RNA scaffold. Similarly, the 2 / 3 asymmetric protrusions in this legend are located at positions -36 to -37 in the guide-target RNA scaffold. As used herein, position annotations are provided for the target RNA side of the target nucleotide to be edited and the guide-target RNA scaffold. As used herein, if a single position is annotated, the structural feature extends from the position away from position 0 (target nucleotide to be edited). For example, if a potential guide RNA is annotated herein as forming a 2 / 3 asymmetric bulge at position -36, then the 2 / 3 asymmetric bulge is formed on the target RNA side of the guide-target RNA scaffold relative to the target nucleotide to be edited (position 0) from -36 to -37. As another example, if a potential guide RNA is annotated herein as forming a 2 / 2 symmetric bulge at position +12, then the 2 / 2 symmetric bulge is formed on the target RNA side of the guide-target RNA scaffold relative to the target nucleotide to be edited (position 0) from +12 to +13.
[0392] In some examples, the engineered guides disclosed herein lack a recruitment region, and the recruitment of the RNA editing entity can be achieved by the structural features of the guide-target RNA scaffold formed by hybridization of the engineered guide RNA and the target RNA. In some examples, when present in an aqueous solution and not bound to a target RNA molecule, the engineered guide does not contain structural features that recruit RNA editing entities (e.g., ADARs or APOBECs). The engineered guide RNA forms one or more structural features of recruiting RNA editing entities (e.g., ADARs or APOBECs) together with the target RNA molecule after hybridization with the target RNA.
[0393] In the absence of a recruitment sequence, the engineered guide RNA is still able to associate with a subject RNA editing entity (e.g., ADAR or APOBEC) to promote editing of a target RNA and / or regulate expression of a polypeptide encoded by the subject target RNA. This can be achieved by structural features formed in the guide-target RNA scaffold formed after hybridization of the engineered guide RNA with the target RNA. The structural features may include any of the following: mismatches, symmetrical protrusions, asymmetrical protrusions, symmetrical internal loops, asymmetrical internal loops, hairpins, wobble base pairs, or any combination thereof.
[0394] Structural features that may be present in the guide-target RNA scaffolds of the present disclosure are described herein. Examples of features include mismatches, protrusions (symmetrical protrusions or asymmetrical protrusions), inner loops (symmetrical inner loops or asymmetrical inner loops) or hairpins (raising hairpins or non-raising hairpins). The engineered guide RNA of the present disclosure may have 1 to 50 features. The engineered guide RNA of the present disclosure may have from 1 to 5, from 5 to 10, from 10 to 15, from 15 to 20, from 20 to 25, from 25 to 30, from 30 to 35, from 35 to 40, from 40 to 45, from 45 to 50, from 5 to 20, from 1 to 3, from 4 to 5, from 2 to 10, from 20 to 40, from 10 to 40, from 20 to 50, from 30 to 50, from 4 to 7, or from 8 to 10 features. In some embodiments, structural features (e.g., mismatches, bulges, internal loops) can be formed from potential structures in an engineered potential guide RNA after the engineered potential guide RNA hybridizes with a target RNA and thereby forms a guide-target RNA scaffold. In some embodiments, structural features are not formed from potential structures, but rather preformed structures (e.g., GluR2 recruiting hairpins or hairpins from U7 snRNA).
[0395] The guide-target RNA scaffold can be formed after the engineered guide RNA of the present disclosure hybridizes with the target RNA. As disclosed herein, "mismatch" refers to the non-pairing of a single nucleotide in the guide RNA with a relative single nucleotide in the target RNA within the guide-target RNA scaffold. A mismatch can include any two single nucleotides that are not base paired. When the number of participating nucleotides on the guide RNA side and the target RNA side exceeds 1, the resulting structure is no longer considered a mismatch, but is considered a protrusion or an inner loop, depending on the size of the structural feature. In some embodiments, the mismatch is an A / C mismatch. An A / C mismatch can include a C in the engineered guide RNA of the present disclosure that is opposite to an A in the target RNA. An A / C mismatch can include an A in the engineered guide RNA of the present disclosure that is opposite to a C in the target RNA. A / G mismatch can include a G in the engineered guide RNA of the present disclosure that is opposite to a G in the target RNA.
[0396] In some embodiments, a mismatch located 5' of the editing site can promote base flipping of the target A to be edited. Mismatches also help to confer sequence specificity. Therefore, mismatches can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0397] In another aspect, the structural feature includes a wobble base. A wobble base pair refers to two bases that are weakly base paired. For example, a wobble base pair of the present disclosure may refer to a G paired with a U. Thus, a wobble base pair may be a structural feature formed by a potential structure provided by an engineered potential guide RNA.
[0398] In some cases, the structural feature can be a hairpin. As disclosed herein, a hairpin includes an RNA duplex, wherein a portion of a single RNA chain folds itself to form an RNA duplex. Due to the nucleotide sequences having base pairing with each other, a portion of a single RNA chain folds itself, wherein the nucleotide sequence is separated by an insertion sequence that is not base paired with itself, thereby forming a base pairing portion and a non-base paired insertion loop portion. The hairpin can have 10 to 500 nucleotides over the length of the entire duplex structure. The length of the loop portion of the hairpin can be 3 to 15 nucleotides. The hairpin can be present in any engineered guide RNA disclosed herein. The engineered guide RNA disclosed herein can have from 1 to 10 hairpins. In some embodiments, the engineered guide RNA disclosed herein has 1 hairpin. In some embodiments, the engineered guide RNA disclosed herein has 2 hairpins. As disclosed herein, the hairpin can include a raised hairpin or a non-raised hairpin. The hairpin can be located at any position in the engineered guide RNA disclosed herein. In some embodiments, one or more hairpins are near or present at the 3' end of an engineered guide RNA of the disclosure, near or present at the 5' end of an engineered guide RNA of the disclosure, near or within a targeting domain of an engineered guide RNA of the disclosure, or any combination thereof.
[0399] In some aspects, structural features include non-raising hairpins. As disclosed herein, non-raising hairpins do not have the primary function of raising RNA editing entities. In some cases, non-raising hairpins do not raise RNA editing entities. In some cases, non-raising hairpins have a dissociation constant for binding RNA editing entities under physiological conditions, and the dissociation constant is not enough to bind. For example, the dissociation constant of non-raising hairpins in binding RNA editing entities at 25 ° C is greater than about 1mM, 10mM, 100mM or 1M, as determined in an in vitro assay. Non-raising hairpins can show the function of improving the localization of the target RNA by the engineered guide RNA. In some embodiments, non-raising hairpins improve nuclear retention. In some embodiments, non-raising hairpins include hairpins from U7 snRNA. Therefore, non-raising hairpins (such as hairpins from U7 snRNA) are pre-formed structural features that can be present in a construct comprising an engineered guide RNA construct, rather than a structural feature formed by the potential structure provided in the engineered potential guide RNA.
[0400] The hairpins of the present disclosure can be of any length. In one aspect, the hairpin can be from about 10-500 or more nucleotides. 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148 8, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 201, 202, 203, 204, 205, 206, 207, 208 08, 209, 210, 211, 212, 213, 214, 215, 216, 217, 218, 219, 220, 221, 222, 223, 224, 225, 226, 227, 228, 229, 230, 231, 232, 233, 234, 235, 236, 237, 238, 239, 240, 241, 242, 243, 244, 245, 246, 247, 248, 249, 250, 251, 252, 253, 254, 255, 256, 257, 258, 259, 260, 261, 262, 263, 264, 265, 266, 267,268, 269, 270, 271, 272, 273, 274, 275, 276, 277, 278, 279, 280, 281, 282, 283, 284, 285, 286, 287, 288, 289, 290, 291, 292, 293, 294, 295, 296, 297, 298, 299, 300, 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, 320, 321, 322, 323, 324, 325, 326 ,327,328,329,330,331,332,333,334,335,336,337,338,339,340,341,342,343,344,345,346,347,348,349,350,351,352,353,354,355,356,357,358,359,360,361,362,363,364,365,366,367,368,369,370,371,372,373,374,375,376,377,378,379,380,381,382,383,384,385 5, 386, 387, 388, 389, 390, 391, 392, 393, 394, 395, 396, 397, 398, 399, 400, 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, 420, 421, 422, 423, 424, 425, 426, 427, 428, 429, 430, 431, 432, 433, 434, 435, 436, 437, 438, 439, 440, 441, 442, 443, 444 44, 445, 446, 447, 448, 449, 450, 451, 452, 453, 454, 455, 456, 457, 458, 459, 460, 461, 462, 463, 464, 465, 466, 467, 468, 469, 470, 471, 472, 473, 474, 475, 476, 477, 478, 479, 480, 481, 482, 483, 484, 485, 486, 487, 488, 489, 490, 491, 492, 493, 494, 495, 496, 497, 498, 499, 500 or more nucleotides. In other cases, the hairpin may also include 10 to 20, 10 to 30, 10 to 40, 10 to 50, 10 to 60, 10 to 70,10 to 80, 10 to 90, 10 to 100, 10 to 110, 10 to 120, 10 to 130, 10 to 140, 10 to 150, 10 to 160, 10 to 170, 10 to 180, 10 to 190, 10 to 200, 10 to 210, 10 to 220, 10 to 230, 10 to 240, 10 to 250, 10 to 260, 10 to 270, 10 to 280, 10 to 290 400, 10-450, 10-460, 10-470, 10-480, 10-490, or 10-500 nucleotides.
[0401] The guide-target RNA scaffold can be formed after the engineered guide RNA of the present disclosure hybridizes with the target RNA. As disclosed herein, a protrusion refers to a structure that is substantially formed only after the guide-target RNA scaffold is formed, wherein the continuous nucleotides in the engineered guide RNA or target RNA are not complementary to their positional counterparts on the opposite strand. The protrusion can change the secondary or tertiary structure of the guide-target RNA scaffold. The protrusion can independently have 0 to 4 continuous nucleotides on the guide RNA side of the guide-target RNA scaffold, and 1 to 4 continuous nucleotides on the target RNA side of the guide-target RNA scaffold, or the protrusion can independently have 0 to 4 nucleotides on the target RNA side of the guide-target RNA scaffold, and 1 to 4 continuous nucleotides on the guide RNA side of the guide-target RNA scaffold. However, as used herein, the protrusion does not refer to a structure in which a single participating nucleotide of the engineered guide RNA and a single participating nucleotide of the target RNA are not base-paired, and a single participating nucleotide of the engineered guide RNA that is not base-paired and a single participating nucleotide of the target RNA are referred to herein as mismatches. In addition, when the number of participating nucleotides on the guide RNA side or the target RNA side exceeds 4, the resulting structure is no longer considered a protrusion, but rather an internal loop. In some embodiments, the guide-target RNA scaffold of the present disclosure has 2 protrusions. In some embodiments, the guide-target RNA scaffold of the present disclosure has 3 protrusions. In some embodiments, the guide-target RNA scaffold of the present disclosure has 4 protrusions. Therefore, the protrusion can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0402] In some embodiments, the presence of a protrusion in a guide-target RNA scaffold can position or can help position an ADAR to selectively edit a target A in a target RNA and reduce off-target editing of non-target A in a target RNA. In some embodiments, the presence of a protrusion in a guide-target RNA scaffold can recruit or help recruit an additional amount of ADAR. The protrusions in the guide-target RNA scaffold disclosed herein can recruit other proteins, such as other RNA editing entities. In some embodiments, the protrusion located at the 5' of the editing site can promote base flipping of the target A to be edited. Relative to other A present in the target RNA, the protrusion can also help confer A sequence specificity to the target RNA to be edited. For example, the protrusion can help guide ADAR editing by constraining it in the direction of selective editing that produces target A.
[0403] The guide-target RNA scaffold can be formed after the engineered guide RNA of the present disclosure hybridizes with the target RNA. The protrusion can be a symmetrical protrusion or an asymmetrical protrusion. When there are the same number of nucleotides on each side of the protrusion, a "symmetrical protrusion" is formed. For example, the symmetrical protrusion in the guide-target RNA scaffold of the present disclosure can have the same number of nucleotides on the engineered guide RNA side and the target RNA side of the guide-target RNA scaffold. The symmetrical protrusion of the present disclosure can be formed by 2 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 2 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetrical protrusion of the present disclosure can be formed by 3 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 3 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetrical protrusion of the present disclosure can be formed by 4 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 4 nucleotides on the target RNA side of the guide-target RNA scaffold. Therefore, the symmetrical protrusion can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0404] The guide-target RNA scaffold can be formed after the engineered guide RNA of the present disclosure is hybridized with the target RNA. The protrusion can be a symmetrical protrusion or an asymmetrical protrusion. When there are different numbers of nucleotides on each side of the protrusion, an "asymmetric protrusion" is formed. For example, the asymmetric protrusion in the guide-target RNA scaffold of the present disclosure can have different numbers of nucleotides on the engineered guide RNA side and the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present disclosure can be formed by 0 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 1 nucleotide on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present disclosure can be formed by 0 nucleotides on the target RNA side of the guide-target RNA scaffold and 1 nucleotide on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present disclosure can be formed by 0 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 2 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 0 nucleotides on the target RNA side of the guide-target RNA scaffold and 2 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 0 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 3 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 0 nucleotides on the target RNA side of the guide-target RNA scaffold and 3 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 0 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 4 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 0 nucleotides on the target RNA side of the guide-target RNA scaffold and 4 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 1 nucleotide on the engineered guide RNA side of the guide-target RNA scaffold and 2 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 1 nucleotide on the target RNA side of the guide-target RNA scaffold and 2 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 1 nucleotide on the engineered guide RNA side of the guide-target RNA scaffold and 3 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 1 nucleotide on the target RNA side of the guide-target RNA scaffold and 3 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 1 nucleotide on the engineered guide RNA side of the guide-target RNA scaffold and 4 nucleotides on the target RNA side of the guide-target RNA scaffold.The asymmetric protrusion of the present invention can be formed by 1 nucleotide on the target RNA side of the guide-target RNA scaffold and 4 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 2 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 3 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 2 nucleotides on the target RNA side of the guide-target RNA scaffold and 3 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 2 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 4 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present invention can be formed by 2 nucleotides on the target RNA side of the guide-target RNA scaffold and 4 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present disclosure can be formed by 3 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 4 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric protrusion of the present disclosure can be formed by 3 nucleotides on the target RNA side of the guide-target RNA scaffold and 4 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. Therefore, the asymmetric protrusion can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0405] In some cases, the structural feature may be an inner loop. As disclosed herein, an inner loop refers to a structure that is substantially formed only after the guide-target RNA scaffold is formed, wherein the nucleotides in the engineered guide RNA or target RNA are not complementary to their positional counterparts on the relative chain, and wherein one side of the inner loop (the target RNA side of the guide-target RNA scaffold or the engineered guide RNA side) has 5 nucleotides or more. When the number of participating nucleotides on the guide RNA side and the target RNA side drops to less than 5, the resulting structure is no longer considered to be an inner loop, but is considered to be a protrusion or mismatch, depending on the size of the structural feature. The inner loop may be a symmetrical inner loop or an asymmetrical inner loop. The inner loop present near the editing site can help the base flipping of the target A in the target RNA to be edited.
[0406] One side of the inner loop (whether on the target RNA side of the guide-target RNA scaffold or the engineered guide RNA side) can be formed by 5 to 150 nucleotides. One side of the inner loop can be formed by 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 120, 135, 140, 145, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900 or 1000 nucleotides, or any number of nucleotides therebetween. One side of the inner loop can be formed by 5 nucleotides. One side of the inner loop may be formed by 10 nucleotides. One side of the inner loop may be formed by 15 nucleotides. One side of the inner loop may be formed by 20 nucleotides. One side of the inner loop may be formed by 25 nucleotides. One side of the inner loop may be formed by 30 nucleotides. One side of the inner loop may be formed by 35 nucleotides. One side of the inner loop may be formed by 40 nucleotides. One side of the inner loop may be formed by 45 nucleotides. One side of the inner loop may be formed by 50 nucleotides. One side of the inner loop may be formed by 55 nucleotides. One side of the inner loop may be formed by 60 nucleotides. One side of the inner loop may be formed by 65 nucleotides. One side of the inner loop may be formed by 70 nucleotides. One side of the inner loop may be formed by 75 nucleotides. One side of the inner loop may be formed by 80 nucleotides. One side of the inner loop may be formed by 85 nucleotides. One side of the inner loop may be formed by 90 nucleotides. One side of the inner loop may be formed by 95 nucleotides. One side of the inner loop may be formed by 100 nucleotides. One side of the inner loop may be formed by 110 nucleotides. One side of the inner loop may be formed by 120 nucleotides. One side of the inner loop may be formed by 130 nucleotides. One side of the inner loop may be formed by 140 nucleotides. One side of the inner loop may be formed by 150 nucleotides. One side of the inner loop may be formed by 200 nucleotides. One side of the inner loop may be formed by 250 nucleotides. One side of the inner loop may be formed by 300 nucleotides. One side of the inner loop may be formed by 350 nucleotides. One side of the inner loop may be formed by 400 nucleotides. One side of the inner loop may be formed by 450 nucleotides. One side of the inner loop may be formed by 500 nucleotides. One side of the inner loop may be formed by 600 nucleotides. One side of the inner loop may be formed by 700 nucleotides. One side of the inner loop may be formed by 800 nucleotides. One side of the inner loop may be formed by 900 nucleotides. One side of the internal loop can be formed by 1000 nucleotides. Thus, the internal loop can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0407] The inner loop may be a symmetric inner loop or an asymmetric inner loop. A "symmetric inner loop" is formed when the same number of nucleotides are present on each side of the inner loop. For example, a symmetric inner loop in the guide-target RNA scaffold of the present disclosure may have the same number of nucleotides on the engineered guide RNA side and the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 5 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 6 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 6 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 7 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 7 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 8 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 8 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 9 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 9 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 10 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 10 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 15 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold target and 15 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 20 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 20 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 30 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 30 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 40 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 40 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 50 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 60 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 60 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 70 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 70 nucleotides on the target RNA side of the guide-target RNA scaffold.The symmetric inner loop of the present disclosure may be formed by 80 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 80 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 90 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 90 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 100 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 110 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 110 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 120 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 120 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 130 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 130 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 140 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 140 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 150 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 200 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 250 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 250 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 300 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 350 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 350 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 400 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 450 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 450 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 500 nucleotides on the target RNA side of the guide-target RNA scaffold.The symmetric inner loop of the present disclosure may be formed by 600 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 600 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 700 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 700 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 800 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 800 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 900 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 900 nucleotides on the target RNA side of the guide-target RNA scaffold. The symmetric inner loop of the present disclosure may be formed by 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 1000 nucleotides on the target RNA side of the guide-target RNA scaffold. Therefore, the symmetric internal loop can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0408] An asymmetric internal loop is formed when there are different numbers of nucleotides on each side of the internal loop. For example, an asymmetric internal loop in a guide-target RNA scaffold of the present disclosure can have different numbers of nucleotides on the engineered guide RNA side and the target RNA side of the guide-target RNA scaffold.
[0409] The asymmetric inner loop of the present disclosure can be formed by 5 to 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 5 to 150 nucleotides on the target RNA side of the guide-target RNA scaffold, wherein the number of nucleotides on the engineered side of the guide-target RNA scaffold is different from the number of nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 to 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 5 to 1000 nucleotides on the target RNA side of the guide-target RNA scaffold, wherein the number of nucleotides on the engineered side of the guide-target RNA scaffold is different from the number of nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 6 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 6 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 7 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 7 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 8 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 8 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and a 9 nucleotide inner loop on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 9 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 10 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 10 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 6 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and a 7 nucleotide inner loop on the target RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure can be formed by 6 nucleotides on the target RNA side of the guide-target RNA scaffold and 7 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 6 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 8 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 6 nucleotides on the target RNA side of the guide-target RNA scaffold and 8 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 6 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and a 9 nucleotide inner loop on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 6 nucleotides on the target RNA side of the guide-target RNA scaffold and 9 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 6 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 10 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 6 nucleotides on the target RNA side of the guide-target RNA scaffold and 10 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 7 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 8 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 7 nucleotides on the target RNA side of the guide-target RNA scaffold and 8 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 7 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 9 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 7 nucleotides on the target RNA side of the guide-target RNA scaffold and 9 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 7 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 10 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 7 nucleotides on the target RNA side of the guide-target RNA scaffold and 10 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 8 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 9 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 8 nucleotides on the target RNA side of the guide-target RNA scaffold and 9 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure can be formed by 8 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 10 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 8 nucleotides on the target RNA side of the guide-target RNA scaffold and 10 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 9 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold and 10 nucleotides on the target RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 9 nucleotides on the target RNA side of the guide-target RNA scaffold and 10 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 5 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 50 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure may be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present invention can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 50 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present invention can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present invention can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present invention can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present invention can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 100 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure can be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 100 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 150 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 5 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 150 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 200 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 200 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 300 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 300 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold.The asymmetric inner loop of the present disclosure may be formed by 400 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 400 nucleotides on the engineered polynucleotide side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 500 nucleotides on the target RNA side of the guide-target RNA scaffold and 1000 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. The asymmetric inner loop of the present disclosure may be formed by 1000 nucleotides on the target RNA side of the guide-target RNA scaffold and 500 nucleotides on the engineered guide RNA side of the guide-target RNA scaffold. Therefore, the asymmetric internal loop can be a structural feature formed by the potential structure provided by the engineered potential guide RNA.
[0410] As described herein, a "micro-footprint" sequence refers to a sequence with a potential structure, which, when it is manifested, promotes the editing of adenosine of a target RNA by adenosine deaminase. Macro footprints can be used to guide or concentrate RNA editing entities (e.g., ADARs) and direct their activity to micro-footprints. In some embodiments, nucleotides are included in the micro-footprint sequence, and the position of the nucleotides is such that when the guide RNA is hybridized with the target RNA, the nucleotides are opposite to the adenosine to be edited by the ADAR enzyme, and are not paired with the adenosine base to be edited. This nucleotide is referred to as "mismatch position" or "mismatch" in this article, and can be cytosine. The micro-footprint sequence as described herein has at least one structural feature selected from the group consisting of the following after the engineered guide RNA is hybridized with the target RNA: protrusions, inner loops, mismatches, hairpins, and any combination thereof. Engineered guide RNAs with high-quality micro-footprint sequences can be selected based on their ability to promote the editing of specific target RNAs. The engineered guide RNA selected for its ability to promote the editing of specific targets can adopt various micro-footprint potential structures, and the structure can vary according to the target.
[0411] The guide RNA of the present disclosure may also include a macro footprint. In some embodiments, the macro footprint includes a barbell macro footprint. Micro footprints can be used to guide or concentrate RNA editing enzymes and direct their activity to the target adenosine to be edited. "Barbell" as described herein refers to a pair of inner loop potential structures that appear after the guide RNA hybridizes with the target RNA. In some embodiments, each inner loop is located at the 5' end or 3' end of the guide-target RNA scaffold formed after the guide RNA hybridizes with the target RNA. In some embodiments, each inner loop is flanked on the opposite side of the micro footprint sequence. When the guide RNA hybridizes with the target RNA, the insertion of the barbell macro footprint sequence flanked on the opposite side of the micro footprint sequence results in the formation of a barbell inner loop on the opposite side of the micro footprint, which in turn contains at least one structural feature that promotes the editing of a specific target RNA.
[0412] In some embodiments, the presence of a barbell flanking a microfootprint can improve one or more aspects of editing. For example, relative to an otherwise comparable guide RNA lacking a barbell, the presence of a barbell macrofootprint in addition to a microfootprint can result in a greater amount of on-target adenosine editing. Additionally and or alternatively, relative to an otherwise comparable guide RNA lacking a barbell, the presence of a barbell macrofootprint in addition to a microfootprint can result in a smaller amount of local off-target adenosine editing. In addition, while the effects of various microfootprint structural features can vary from target to target based on selection in high-throughput screening, the increase in one or more editing aspects provided by the barbell macrofootprint structure may be unrelated to a specific target RNA. Therefore, the inclusion of a barbell structure can provide a simple method for improving the editing of a guide RNA previously selected to promote the editing of a target RNA of interest. For example, relative to an otherwise comparable guide RNA lacking a barbell, macrofootprints (e.g., barbell macrofootprints) and microfootprints can provide an increased amount of on-target adenosine editing. In other embodiments, the presence of a barbell macrofootprint in addition to a microfootprint can result in a lower amount of localized off-target adenosine editing relative to an otherwise comparable guide RNA that forms a guide-target RNA scaffold lacking a barbell when the guide RNA hybridizes to the target RNA.
[0413] As disclosed herein, a "macrofootprint" sequence can be positioned so that it is flanked by a microfootprint sequence. In addition, although a macrofootprint sequence can be flanked by a microfootprint sequence, additional potential structures flanking either end of the macrofootprint can also be incorporated. In some embodiments, such additional potential structures are included as part of the macrofootprint. In some embodiments, such additional potential structures are separate, different, or both separate and different from the macrofootprint. In some embodiments, a macrofootprint sequence may include a barbell macrofootprint sequence, the barbell macrofootprint sequence including a potential structure that, when manifested, produces a first inner loop and a second inner loop.
[0414] In some embodiments, the first inner loop of the barbell or the second inner loop of the barbell is located at least about 5 bases (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 bases) away from the A / C mismatch relative to the base of the first inner loop or the second inner loop that is closest to the A / C mismatch. In some embodiments, the first inner loop of the barbell or the second inner loop of the barbell is located at most about 50 bases (e.g., 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, or 5) away from the A / C mismatch relative to the base of the first inner loop or the second inner loop that is closest to the A / C mismatch.
[0415] In some embodiments, the first internal loop or the second internal loop independently comprises at least about 5 bases or more (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150); about 150 bases or less (e.g., 145, 135, 125, 115, 95, 85, 75, 65, 55, 45, 35, 25, 5, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5); or at least about 5 bases to at least about 150 bases (e.g., 5-150, 6-145, 7-140, 8-135, 9-130, 10-125, 11-120, 12-115, 13-110, 14-105, 15-100, 16-95, 17-90, 18-85, 19-80, 20-75, 21-70, 22-65, 23-60, 24-55 , 25-50) base number and at least about 5 bases or more (e.g., 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150) of the target RNA; about 150 bases or less (e.g., 145, 135, 125, 115, 95, 85, 75, 65, 55, 45, 35, 25, 19, 18, 17, 16, 1 or the number of bases from at least about 5 bases to at least about 150 bases (e.g., 5-150, 6-145, 7-140, 8-135, 9-130, 10-125, 11-120, 12-115, 13-110, 14-105, 15-100, 16-95, 17-90, 18-85, 19-80, 20-75, 21-70, 22-65, 23-60, 24-55, 25-50).
[0416] As disclosed herein, a "base pairing (bp) region" refers to a region of a guide-target RNA scaffold where a base in the guide RNA pairs with the opposite base in the target RNA. The base pairing region may extend from one end or the proximal end of one end of the guide-target RNA scaffold to or the proximal end of the other end of the guide-target RNA scaffold. The base pairing region may extend between two structural features. The base pairing region may extend from one end or the proximal end of one end of the guide-target RNA scaffold to or the proximal end of the structural feature. The base pairing region may extend from the structural feature to the other end of the guide-target RNA scaffold. In some embodiments, the base pairing region has 1 bp to 100 bp, 1 bp to 90 bp, 1 bp to 80 bp, 1 bp to 70 bp, 1 bp to 60 bp, 1 bp to 50 bp, 1 bp to 45 bp, 1 bp to 40 bp, 1 bp to 35 bp, 1 bp to 30 bp, 1 bp to 25 bp, 1 bp to 20 bp, 1 bp to 15 bp, 1 bp to 10 bp, 1 bp to 5 bp, 5 bp to 10 bp, 5 bp to 20 bp, 10 bp to 20 bp, 10 bp to 50 bp, 5bp to 50bp, at least 1bp, at least 2bp, at least 3bp, at least 4bp, at least 5bp, at least 6bp, at least 7bp, at least 8bp, at least 9bp, at least 10bp, at least 12bp, at least 14bp, at least 16bp, at least 18bp, at least 20bp, at least 25bp, at least 30bp, at least 35bp, at least 40bp, at least 45bp, at least 50bp, at least 60bp, at least 70bp, at least 80bp, at least 90bp, at least 100bp.
[0417] Guide RNA expression cassette. The guide RNA expression cassette can include a promoter (e.g., any one of SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263), a guide RNA sequence, a structural element, and a termination sequence (e.g., any one of SEQ ID NO: 60, SEQ ID NO: 708–SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243–SEQ ID NO: 1247, SEQ ID NO: 1254–SEQ ID NO: 1257, SEQ ID NO: 1264–SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287–SEQ ID NO: 1289). The guide RNA sequence can target a target RNA. In some embodiments, the target RNA encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with type 1A Charcot-Marie-Tooth disease (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2). Examples of engineered guide RNA expression cassettes comprising promoters, guide RNA sequences, structural elements, and termination sequences are provided in Table 9.
[0418] Table 9 - Exemplary engineered guide RNA expression cassettes
[0419]
[0420]
[0421]
[0422]
[0423]
[0424] In some embodiments, the engineered guide RNA expression cassette can have at least about 70%, at least about 75%, at least about 80%, at least about 83%, at least about 85%, at least about 87%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence identity to any one of SEQ ID NO:1-SEQ ID NO:12 or SEQ ID NO:59.
[0425] The engineered guide RNA expression cassette can comprise a promoter (e.g., any one of SEQ ID NO: 13–SEQ ID NO: 17, SEQ ID NO: 167–SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248–SEQ ID NO: 1253, or SEQ ID NO: 1259–SEQ ID NO: 1263), a guide RNA sequence, structural elements, and a termination sequence (e.g., any one of SEQ ID NO: 60, SEQ ID NO: 708–SEQ ID NO: 1240, SEQ ID NO: 1242, SEQ ID NO: 1243–SEQ ID NO: 1247, SEQ ID NO: 1254–SEQ ID NO: 1257, SEQ ID NO: 1264–SEQ ID NO: 1272, SEQ ID NO: 1275, or SEQ ID NO: 1287–SEQ ID NO: 1289).
[0426] For example, the engineered guide RNA expression cassette of SEQ ID NO: 1 comprises a promoter of SEQ ID NO: 15, a PMP22 guide RNA sequence of SEQ ID NO: 1273, and a termination sequence of SEQ ID NO: 1243. For example, the engineered guide RNA expression cassette of SEQ ID NO: 2 comprises a promoter of SEQ ID NO: 16, a PMP22 guide RNA sequence of SEQ ID NO: 1273, and a termination sequence of SEQ ID NO: 1243. For example, the engineered guide RNA expression cassette of SEQ ID NO: 3 comprises a promoter of SEQ ID NO: 15, a PMP22 guide RNA sequence of SEQ ID NO: 1273, and a termination sequence of SEQ ID NO: 1275. For example, the engineered guide RNA expression cassette of SEQ ID NO: 4 comprises a promoter of SEQ ID NO: 16, a PMP22 guide RNA sequence of SEQ ID NO: 1273, and a termination sequence of SEQ ID NO: 60. For example, the engineered guide RNA expression cassette of SEQ ID NO:5 comprises a promoter of SEQ ID NO:17, a PMP22 guide RNA sequence of SEQ ID NO:1273, and a termination sequence of SEQ ID NO:60.
[0427] For example, the engineered guide RNA expression cassette of SEQ ID NO:6 comprises a promoter of SEQ ID NO:15, a SNCA guide RNA sequence of SEQ ID NO:1274, and a termination sequence of SEQ ID NO:1243. For example, the engineered guide RNA expression cassette of SEQ ID NO:7 comprises a promoter of SEQ ID NO:13, a SNCA guide RNA sequence of SEQ ID NO:1290, and a termination sequence of SEQ ID NO:1243. For example, the engineered guide RNA expression cassette of SEQ ID NO:8 comprises a promoter of SEQ ID NO:14, a SNCA guide RNA sequence of SEQ ID NO:1274, and a termination sequence of SEQ ID NO:1243. For example, the engineered guide RNA expression cassette of SEQ ID NO:9 comprises a promoter of SEQ ID NO:16, a SNCA guide RNA sequence of SEQ ID NO:1274, and a termination sequence of SEQ ID NO:1243. For example, the engineered guide RNA expression cassette of SEQ ID NO:10 comprises a promoter of SEQ ID NO:15, a SNCA guide RNA sequence of SEQ ID NO:1274, and a termination sequence of SEQ ID NO:1275. For example, the engineered guide RNA expression cassette of SEQ ID NO: 11 comprises a promoter of SEQ ID NO: 16, a SNCA guide RNA sequence of SEQ ID NO: 1274, and a termination sequence of SEQ ID NO: 60. For example, the engineered guide RNA expression cassette of SEQ ID NO: 12 comprises a promoter of SEQ ID NO: 17, a SNCA guide RNA sequence of SEQ ID NO: 1274, and a termination sequence of SEQ ID NO: 60.
[0428] For example, the engineered guide RNA expression cassette of SEQ ID NO:59 comprises a promoter of SEQ ID NO:16, a SERPINA1 guide RNA sequence of SEQ ID NO:61, and a termination sequence of SEQ ID NO:60.
[0429] Additional Engineered Guide RNA Components
[0430] The present disclosure provides engineered guide RNAs with additional structural features and components. For example, the engineered guide RNAs described herein can be circular. In another example, the engineered guide RNAs described herein can include U7, SmOPT sequences, or a combination of the two sequences.
[0431] In some cases, the engineered guide RNA can be circularized. In some cases, the engineered guide RNA provided herein can be circular or in a circular configuration. In some aspects, at least a portion of the circular guide RNA lacks a 5' hydroxyl or a 3' hydroxyl.
[0432] In some instances, the engineered guide RNA can include a backbone comprising a plurality of sugar and phosphate moieties covalently linked together. In some instances, the backbone of the engineered guide RNA can include a phosphodiester bond (between the first hydroxyl in the phosphate group on the 5' carbon of deoxyribose in DNA or ribose in RNA and the second hydroxyl on the 3' carbon of deoxyribose in DNA or ribose in RNA.
[0433] In some embodiments, the backbone of the engineered guide RNA may lack a 5' reduced hydroxyl group, a 3' reduced hydroxyl group, or both that can be exposed to a solvent. In some embodiments, the backbone of the engineered guide may lack a 5' reduced hydroxyl group, a 3' reduced hydroxyl group, or both that can be exposed to a nuclease. In some embodiments, the backbone of the engineered guide may lack a 5' reduced hydroxyl group, a 3' reduced hydroxyl group, or both that can be exposed to a hydrolase. In some cases, the backbone of the engineered guide may be represented as a polynucleotide sequence in a circular 2-dimensional format of one nucleotide by one nucleotide. In some cases, the backbone of the engineered guide may be represented as a polynucleotide sequence in a circular 2-dimensional format of one nucleotide by one nucleotide. In some cases, the 5' hydroxyl group, the 3' hydroxyl group, or both may be connected by a phosphorus-oxygen bond. In some cases, the 5' hydroxyl group, the 3' hydroxyl group, or both may be modified to a phosphate with a phosphorus-containing moiety.
[0434] As described herein, the engineering guide can include a circular structure. The engineered polynucleotide can be cyclized from a precursor engineered polynucleotide. Such precursor engineered polynucleotides can be precursor engineered linear polynucleotides. In some cases, the precursor engineered linear polynucleotide can be a precursor of a circular engineered guide RNA. For example, a precursor engineered linear polynucleotide can be a linear mRNA transcribed from a plasmid, which can be configured to be cyclized in a cell using the technology described herein. The precursor engineered linear polynucleotide can be constructed to have a domain that allows cyclization when inserted into a cell, such as a ribozyme domain and a connecting domain. The ribozyme domain can include a domain (e.g., adjacent to a connecting domain) that can cleave a linear precursor RNA at a specific site. The precursor engineered linear polynucleotide can include from 5' to 3': a 5' ribozyme domain, a 5' connecting domain, a cyclization region, a 3' connecting domain, and a 3' ribozyme domain. In some cases, the cyclization region can include a guide RNA as described herein. In some cases, the precursor polynucleotide can be specifically processed by 5' ribozymes and 3' ribozymes at two sites to release the exposed ends on the 5' linking domain and the 3' linking domain. The free exposed ends can be link competent so that the ends can be connected to form a mature cyclized structure. For example, the free ends may include 5'-OH and 2', 3'-cyclic phosphates, which are connected via RNA junctions in cells. Linear polynucleotides with connection and ribozyme domains can be transfected into cells, where they can be cyclized by endogenous cellular enzymes. In some cases, polynucleotides can encode engineered guide RNAs comprising ribozymes and linking domains as described herein, which can be cyclized in cells. For example, PCT / US2021 / 034301 provides descriptions of circular guide RNAs and structures thereof, sequences of circular guide RNAs, and methods for engineering cyclized polynucleotide domains, and each of these descriptions in PCT / US2021 / 034301 is incorporated herein by reference.
[0435] As described herein, engineered polynucleotides (e.g., circularized guide RNA) may include a spacer domain. As described herein, a spacer domain may refer to a domain that provides space between other domains. The spacer domain may be used between the region to be circularized and the flanking linking sequence to increase the overall size of the mature circularized guide RNA. When the region to be circularized includes a targeting domain configured to associate with a target sequence as described herein, relative to a comparable engineered polynucleotide lacking a spacer domain, the addition of a spacer can provide an improvement (e.g., increased specificity, enhanced editing efficiency, etc.) for the engineered polynucleotide for the target polynucleotide. In some cases, the spacer domain is configured not to hybridize with the target RNA. In some embodiments, the precursor engineered polynucleotide or the circular engineered guide may include in the order of 5' to 3': a first ribozyme domain; a first linking domain; a first spacer domain; a targeting domain, a second spacer domain, a second linking domain, and a second ribozyme domain that may be at least partially complementary to the target RNA. In some cases, when the targeting domain binds to the target RNA, the first spacer domain, the second spacer domain, or both are configured not to bind to the target RNA.
[0436] The compositions and methods disclosed herein provide engineered polynucleotides encoding guide RNAs that are operably connected to a portion of a small nuclear ribonucleic acid (snRNA) sequence. The engineered polynucleotides may include at least a portion of a small nuclear ribonucleic acid (snRNA) sequence. U7 and U1 small nuclear RNAs (whose natural role is spliceosome processing of pre-mRNA) have been redesigned for decades to change the splicing at the desired disease target. A portion (e.g., the first 18 nucleotides of U7 snRNA) that hybridizes naturally with the spacer element of histone pre-mRNA in U7 snRNA is replaced with a short targeting (or antisense) sequence of a disease gene, redirecting the splicing mechanism to change the splicing around the target site. In addition, the wild-type U7 Sm domain binding site is converted to an optimized consensus Sm binding sequence (SmOPT) that can increase the expression level, activity, and subcellular localization of artificial antisense engineered U7 snRNA. Many subsequent groups have adapted this modified U7SmOPT snRNA chassis to the antisense sequences of other genes to recruit spliceosome elements and modify RNA splicing for other disease targets.
[0437] snRNA is a class of small RNA molecules found in the nucleus of eukaryotic cells. They participate in a variety of important processes, such as RNA splicing (removing introns from pre-mRNA), regulation of transcription factors (7SK RNA) or RNA polymerase II (B2 RNA) and maintenance of telomeres. They are always associated with specific proteins, and the resulting RNA-protein complex is referred to as small nuclear ribonucleoprotein (snRNP) or sometimes referred to as snurps. There are many snRNAs, which are named U1, U2, U3, U4, U5, U6, U7, U8, U9 and U10.
[0438] U7-type snRNAs are generally involved in the maturation of histone mRNAs. Such snRNAs have been identified in a large number of eukaryotic species (56 to date), and the U7 snRNAs of each of these species should be considered equally convenient for the present disclosure.
[0439] The wild-type U7 snRNA includes a stem-loop structure, a U7-specific Sm sequence, and a sequence that is antisense to the 3' end of the histone pre-mRNA.
[0440] In addition to the SmOPT domain, U7 contains a sequence that is antisense to the 3' end of the histone pre-mRNA. When this sequence is replaced by a targeting sequence that is antisense to another target pre-mRNA, U7 is redirected to the new target pre-mRNA. Therefore, the stable expression of the modified U7 snRNA containing the SmOPT domain and the targeted antisense sequence leads to specific changes in mRNA splicing. Although the AAV-2 / 1-based vector expressing the appropriately modified mouse U7 gene and its natural promoter and 3' elements can efficiently transfer the gene to skeletal muscle and complete the rescue of anti-dystrophin by covering and skipping mouse Dmd exon 23, the engineered polynucleotides as described herein (whether directly administered or administered via, for example, AAV vectors) can promote the editing of the target RNA by deaminase.
[0441] The engineered polynucleotide may at least partially comprise a snRNA sequence. The snRNA sequence may be a U1, U2, U3, U4, U5, U6, U7, U8, U9 or U10 snRNA sequence.
[0442] In some cases, relative to comparable polynucleotides lacking these features, engineered polynucleotides comprising at least a portion of snRNA sequences (e.g., snRNA promoters, snRNA hairpins, etc.) may have excellent properties for treating or preventing diseases or disorders. For example, as described herein, engineered polynucleotides comprising at least a portion of snRNA sequences can promote exon skipping of exons with higher efficiency than comparable polynucleotides lacking such features. In addition, as described herein, engineered polynucleotides comprising at least a portion of snRNA sequences can more effectively promote the editing of nucleotide bases in target RNA (e.g., pre-mRNA or mature RNA) than comparable polynucleotides lacking these features. Promoter and snRNA components are described in PCT / US2021 / 028618 and PCT / US2022 / 078801, and each of these descriptions in PCT / US2021 / 028618 and PCT / US2022 / 078801 is incorporated herein by reference.
[0443] Disclosed herein is an engineered RNA comprising (a) an engineered guide RNA as described herein, and (b) a U7 snRNA hairpin sequence, a SmOPT sequence, or a combination thereof. In some embodiments, the U7 hairpin comprises a human U7 hairpin sequence or a mouse U7 hairpin sequence. In some cases, the human U7 hairpin sequence comprises TAGGCTTTCTGGCTTTTTACCG GAAAGCCCCT (SEQ ID NO: 52) or RNA: UAGGCUUUCUGGCU UUUUACCGGAAAGCCCCU (SEQ ID NO: 53). In some cases, the mouse U7 hairpin sequence comprises CAGGTTTTCTGACTTCGGTCGGAAAACCCC T (SEQ ID NO: 54) or RNA: CAGGUUUUCUGACUUCGGUCGGA AAACCCCU (SEQ ID NO: 55). In some embodiments, the SmOPT sequence has the sequence AATTTTTTGGAG (SEQ ID NO: 56) or RNA: AAUUUUU GGAG (SEQ ID NO: 57). In some embodiments, the RNA payload can include a guide RNA, a U7 hairpin sequence (e.g., a human or mouse U7 hairpin sequence), a SmOPT sequence, or a combination thereof. For example, the RNA payload can include the sequence AATTTTTTGGAGC AGGTTTTCTGACTTCGGTCGGAAAACCCCTCCCAATTTCACTGGTCTACAATGAAAGCAAAACAGTTCTCTTCCCCGCTCCCCGGTGTGTGAGAGGGGCTTTGATCCTTCTCTGGTTTCCTAGGAAACGCGTATGTG (SEQ ID NO: 58). In some cases, the combination of a U7 hairpin sequence and a SmOPT sequence can include a SmOPT U7 hairpin sequence, wherein the SmOPT sequence is connected to the U7 sequence. In some cases, the U7 hairpin sequence, the SmOPT sequence, or a combination thereof is located downstream (e.g., 3') of an engineered guide RNA disclosed herein.
[0444] Guide RNA payloads for DNA editing
[0445] The expression cassette described herein can be used to enhance the expression of RNA components to perform site-specific, selective editing of target DNA by DNA editing entities or their biologically active fragments. The RNA components for site-specific DNA editing may include guide RNA, trans-activating CRISPR RNA (tracrRNA), single guide RNA, or engineered polynucleotides encoding the same DNA. The engineered guide RNA as described herein may include a sequence complementary to the target DNA described herein. In this way, the guide RNA can be engineered to site-specifically / selectively target a specific target DNA and hybridize with it, thereby facilitating editing of specific nucleotides in the target DNA by DNA editing entities or their biologically active fragments. DNA editing can be promoted by nucleases such as Cas nucleases. In some embodiments, the Cas nuclease may be Cas9, Cas12, or Cas14.
[0446] In some embodiments, the engineered guide RNA hybridizes with the sequence of the target DNA. In some embodiments, a portion of the engineered guide RNA hybridizes with the sequence of the target DNA. The portion of the engineered guide RNA hybridized with the target DNA has sufficient complementarity with the sequence of the target DNA to hybridize. In some embodiments, the guide RNA may include a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or about 100% sequence complementarity with the target DNA. The guide RNA encoded by the expression cassette of the present disclosure may include a length of from about 15 to about 70 nucleotides, from about 40 to about 70 nucleotides, or from about 70 to about 100 nucleotides. In some embodiments, the region of the guide RNA hybridized with the target may include a length of from about 18 to about 44 nucleotides.
[0447] In some instances, the engineering guide RNA can promote the editing of the bases of nucleotides in the target sequence of the target DNA, resulting in the regulation of the expression of the gene encoded by the target DNA. In some cases, the regulation can be an increase or decrease in the expression of the gene. In some cases, the engineering guide can be configured to promote the editing of the bases of nucleotides or polynucleotides in the DNA region by a DNA editing entity (e.g., Cas nuclease).
[0448] In some embodiments, the expression cassette described herein can be used to enhance the expression of trans-activated crRNA (tracrRNA) and the engineered polynucleotide encoding the same tracrRNA, to edit the target DNA via a DNA editing entity or its biologically active fragment. TracrRNA can bind to and activate DNA editing enzymes (e.g., Cas nucleases). The tracrRNA encoded by the expression cassette of the present disclosure can include a length of from about 75 to about 100 nucleotides.
[0449] In some embodiments, the expression cassettes described herein can be used to enhance the expression of single guide RNA and engineered polynucleotides encoding the same single guide RNA to edit the target DNA via a DNA editing entity or a biologically active fragment thereof. The single guide RNA may include a region that binds and activates a DNA editing enzyme (e.g., Cas nuclease) and a region that hybridizes with a target DNA sequence. The portion of the single guide RNA hybridized with the target DNA has sufficient complementarity with the sequence of the target DNA to hybridize. In some embodiments, the single guide RNA may include a sequence having at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or about 100% sequence complementarity with the target DNA. The single guide RNA encoded by the expression cassette of the present disclosure may include a length of from about 80 to about 120 nucleotides. In some embodiments, the region of the single guide RNA hybridized with the target may include a length of from about 18 to about 44 nucleotides.
[0450] Other RNA targeting oligonucleotides
[0451] Expression cassette described herein can be used to enhance the expression of engineered polynucleotides hybridized with target RNA (for example, target mRNA or target pre-mRNA) of other engineered RNA targeting oligonucleotides (including antisense oligonucleotides, siRNA, shRNA and miRNA) and encoding the same content. Engineered oligonucleotides as described herein can include a targeting domain with complementarity with target RNA as described herein. Therefore, oligonucleotides can be engineered to hybridize with and hybridize with a specific target RNA of targeting, thereby changing the expression of a polypeptide encoded by the target RNA.
[0452] In some embodiments, the engineered oligonucleotides (e.g., antisense oligonucleotides, siRNA, shRNA or miRNA) of the present disclosure are hybridized with the sequence of the target RNA. In some embodiments, the part of the engineered oligonucleotide (e.g., targeting domain) is hybridized with the sequence of the target RNA. The part of the engineered oligonucleotide hybridized with the target RNA has enough complementarity to hybridize with the sequence of the target RNA. The targeting sequence may also be referred to as "targeting domain" or "targeting region". In some embodiments, the combination of the engineered oligonucleotide and the target RNA can recruit additional components, such as RISC components.
[0453] Therapeutic applications
[0454] The expression cassettes of the disclosed RNA payloads encoded by the engineered promoter transcriptional control can have a variety of therapeutic applications. The engineered promoters described herein can promote therapeutic uses by increasing payload expression and enhancing the therapeutic effects produced by the payload. For example, increasing guide RNA payload expression can enhance the editing efficiency of the target DNA or RNA. In another example, increasing antisense oligonucleotide expression can enhance the target knockdown efficiency.
[0455] RNA editing
[0456] RNA editing may refer to the process in which RNA can be enzymatically modified after synthesis at a specific nucleoside. RNA editing may include any of the insertion, deletion or substitution of nucleotides. Examples of RNA editing include chemical modifications such as pseudouridylation (isomerization of uridine residues) and deamination (removal of amine groups from cytidine to produce uridine, or C to U editing; or from adenosine to inosine, or A to I editing). RNA editing can be used to correct mutations (e.g., correct missense mutations) to restore protein expression, or introduce mutations or edit RNA coding or noncoding regions to inhibit RNA translation and achieve protein knockdown. The expression cassette disclosed herein can be used to express engineered guide RNA to promote RNA editing performed by RNA entities (e.g., adenosine deaminase (ADAR) acting on RNA) or their biologically active fragments.
[0457] Described herein is an engineered guide RNA that facilitates RNA editing by an RNA editing entity (e.g., an adenosine deaminase (ADAR) acting on RNA) or a biologically active fragment thereof. In some cases, ADAR can be an enzyme that catalyzes the chemical conversion of adenosine in RNA to inosine. Because the properties of inosine are similar to those of guanosine (e.g., inosine will form two hydrogen bonds with cytosine), inosine can be recognized as guanosine by the translation cell machinery. Therefore, "adenosine to inosine (A to I) RNA editing" effectively changes the primary sequence of the RNA target. Typically, ADAR enzymes share a common domain structure, including a variable number of amino-terminal dsRNA binding domains (dsRBDs) and a carboxyl-terminal catalytic deaminase domain. Human ADARs have two or three dsRBDs. There is evidence that ADARs can form homodimers and heterodimers with other ADARs when bound to double-stranded RNA, but it is not yet certain whether dimerization is necessary for editing to occur. The engineered guide RNA disclosed herein can facilitate RNA editing by any one or any combination of the three human ADAR genes (ADAR 1-3) that have been identified. ADARs have a typical modular domain organization containing at least two copies of dsRNA binding domains (dsRBDs; ADAR1 has three dsRBDs; ADAR2 and ADAR3 have two dsRBDs each) in the N-terminal region of the ADAR, followed by a C-terminal deaminase domain.
[0458] The engineered guide RNA disclosed herein promotes RNA editing performed by endogenous ADAR enzymes. In some embodiments, exogenous ADARs can be delivered together with the engineered guide RNA disclosed herein to promote RNA editing. In some embodiments, ADAR is human ADAR1. In some embodiments, ADAR is human ADAR2. In some embodiments, ADAR is human ADAR3. In some embodiments, ADAR is human ADAR1, human ADAR2, human ADAR2, or any combination thereof.
[0459] In some embodiments, the disclosure provides an engineered guide RNA that facilitates editing at a specific region of a target RNA (e.g., mRNA or pre-mRNA). For example, the engineered guide RNA disclosed herein can target a coding sequence or a non-coding sequence of an RNA. For example, the target region in an RNA coding sequence can be a translation initiation site (TIS). In some embodiments, the target region in a non-coding sequence of an RNA can be a polyadenylation (polyA) signal sequence.
[0460] Missense mutation. In some embodiments, the engineered guide RNA of the present disclosure can target missense mutations in target RNA sequences. The engineered guide RNA can promote the conversion of target adenosine (A) RNA editing mediated by ADAR to inosine (I), which can be read as guanosine (G). A to I conversion via ADAR-mediated RNA editing can correct G to A missense mutations. For example, ADAR-mediated editing can correct valine to isoleucine or valine to methionine mutations by converting isoleucine codons (AUU, AUC or AUA) or methionine codons (AUG) to valine codons (AUA, GUC, GUU or GUG). In another example, ADAR-mediated editing can correct cysteine to tyrosine or mutations by converting tyrosine codons (AUA or UAC) to cysteine codons (UGU or UGC). Alternatively, or in addition, the engineered guide RNA can promote APOBEC-mediated RNA editing of target cytosine (C) to convert to uracil (U). Conversion of C to U via APOBEC-mediated RNA editing can correct U to C missense mutations. The engineered guide RNAs disclosed herein can target one missense mutation or any combination of missense mutations of a target sequence (e.g., SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1, or MECP2).
[0461] Nonsense mutation. In some embodiments, the engineered guide RNA of the present disclosure can target nonsense mutations in target RNA sequences. The engineered guide RNA can promote ADAR-mediated RNA editing of target adenosine (A) to inosine (I), which can be read as guanosine (G). A to I conversion via ADAR-mediated RNA editing can correct G to A nonsense mutations. For example, ADAR-mediated editing can correct tryptophan to terminate nonsense mutations by converting the UAG stop codon to a tryptophan codon (UGG). In another example, ADAR-mediated editing can correct tryptophan to terminate nonsense mutations by converting the UGA stop codon to a tryptophan codon (UGG). Correcting nonsense mutations by ADAR-mediated editing can increase the expression of the target sequence. The engineered guide RNAs disclosed herein can target one missense mutation or any combination of missense mutations of a target sequence (e.g., SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1, or MECP2).
[0462] Translation initiation site (TIS). In some embodiments, the engineered guide RNA of the present disclosure targets adenosine at the translation initiation site (TIS). The engineered guide RNA promotes ADAR-mediated RNA editing of TIS (AUG) to GUG. This results in inhibition of RNA translation, thereby causing protein knockdown. Protein knockdown can also refer to reduced expression of wild-type protein. The engineered guide RNA of the present disclosure can target a missense mutation or any missense mutation combination of the target sequence (e.g., SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1 or MECP2).
[0463] 3'UTR. In some embodiments, the engineered guide RNA of the present disclosure targets one or more adenosines in the 3' untranslated region (3'UTR). In some embodiments, the engineered guide RNA promotes ADAR-mediated RNA editing of one or more adenosines in the 3'UTR, thereby reducing mRNA output from the nucleus and inhibiting translation, resulting in protein knockdown. In some embodiments, the target sequence can be SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1 or MECP2.
[0464] Poly (A) signal sequence. In some embodiments, the engineered guide RNA of the present disclosure targets one or more adenosines in a poly (A) signal sequence. In some embodiments, the engineered guide RNA promotes the RNA editing mediated by ADAR of one or more adenosines in a poly (A) signal sequence, thereby causing the interruption of RNA processing and the degradation of target mRNA, and thus causing protein to be knocked down. In some embodiments, the target may have one or more poly (A) signal sequences. In these cases, one or more engineered guide RNAs of the present disclosure that are different in their respective sequences may be multiplexed to target adenosines in one or more poly (A) signal sequences. In both cases, the engineered guide RNA of the present disclosure promotes the RNA editing mediated by ADAR of adenosine to inosine (interpreted as guanosine by cell mechanisms) in a poly (A) signal sequence, thereby causing protein to be knocked down. In some embodiments, the target sequence may be SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1 or MECP2.
[0465] DNA editing
[0466] DNA editing can refer to the process in which DNA is enzymatically (e.g., by an RNA-guided endonuclease). DNA editing can include any of the insertion, deletion, or substitution of one or more nucleotides. DNA editing can be used to correct mutations (e.g., correct missense mutations) to restore protein expression, or introduce mutations or edit coding or non-coding regions of DNA to inhibit DNA transcription and achieve protein knockdown. The expression cassette disclosed herein can be used to express engineered guide RNA to promote DNA editing of a DNA entity (e.g., CRISPR / Cas endonuclease) or its biologically active fragment. Described herein is an engineered guide RNA that promotes DNA editing of a DNA editing entity (e.g., CRISPR / Cas endonuclease) or its biologically active fragment.
[0467] The engineered guide RNA disclosed herein can promote DNA editing of endogenous Cas enzymes. In some embodiments, exogenous Cas enzymes can be delivered together with the engineered guide RNA disclosed herein to promote DNA editing. In some embodiments, the Cas nuclease is Cas9. In some embodiments, the Cas nuclease is Cas12. In some embodiments, the Cas nuclease is Cas14.
[0468] In some embodiments, the present disclosure provides engineered guide RNAs that facilitate editing at specific regions of target DNA (e.g., mRNA or pre-mRNA). For example, the engineered guide RNAs disclosed herein can target coding or non-coding sequences of DNA.
[0469] The engineered guide RNA of the present disclosure can recruit CRISPR / Cas endonucleases (e.g., Cas9 nucleases) to form ribonucleoprotein (RNP) complexes, which target specific sites in target polynucleotides (e.g., target DNA) via base pairing between the guide RNA and the target region in the target polynucleotide. The engineered guide RNA may include a targeting sequence complementary to the target site of the target polynucleotide. Therefore, the engineered guide RNA forms a complex with the Cas nuclease, and the guide RNA provides sequence specificity for the RNP complex via the targeting sequence. When recruited to the target polynucleotide, the Cas nuclease can site-specifically edit the target polynucleotide (e.g., target DNA). In some embodiments, the target polynucleotide may be SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1 or MECP2.
[0470] Expression knockdown
[0471] The expression cassette disclosed herein can be used to express engineered RNA targeting oligonucleotides (e.g., antisense oligonucleotides, siRNA, shRNA or miRNA) to promote the knockdown expression of target RNA. In some embodiments, the combination of the RNA targeting oligonucleotide and the target RNA can recruit other components (e.g., RISC complex components) to the target RNA, thereby reducing the expression of the peptide encoded by the target RNA. For example, the combination of siRNA can recruit RISC and promote the cutting of the target RNA. In another example, the combination of miRNA or shRNA can recruit RISC and inhibit the translation of the target RNA. In some embodiments, the target sequence can encode SNCA, PMP22, DUX4, LRRK2, MAPT, GRN, ABCA4, APP, SERPINA1, HEXA, CFTR, LIPA, GBA, PINK1 or MECP2.
[0472] Therapeutic targets and approaches
[0473] The small RNA payloads disclosed herein, such as engineered guide RNAs, can be used in methods for treating a condition in a subject in need. The condition can be a disease, illness, genotype, phenotype, or any state associated with side effects. In some embodiments, treating a condition can include preventing the condition, slowing the development of the condition, reversing or alleviating the symptoms of the condition. The method for treating a condition can include delivering an engineered polynucleotide encoding an engineered guide RNA to a cell of a subject in need, and expressing the engineered guide RNA in the cell. In some embodiments, the engineered guide RNA disclosed herein can be used to treat genetic conditions (e.g., synuclein diseases, such as Alzheimer's disease (AD), frontotemporal dementia (FTD), Parkinson's disease). In some embodiments, the engineered guide RNA disclosed herein can be used to treat conditions associated with one or more mutations.
[0474] The present disclosure provides compositions and methods of use of expression cassettes encoding engineered guide payloads (e.g., engineered guide RNAs), such as methods of treatment. In some embodiments, the expression cassettes of the present disclosure encode guide RNAs targeting the coding sequence of RNA (e.g., RNAs encoding α-synuclein, PMP22, DUX4, LRRK2, tau, progranulin, ABCA4, amyloid precursor protein, or α-1 antitrypsin). In some embodiments, the engineered polynucleotides of the present disclosure encode guide RNAs targeting the non-coding sequence (e.g., poly(A) sequence) of RNA. In some embodiments, the present disclosure provides a composition of one or more than one engineered polynucleotides, the engineered polynucleotides encoding more than one engineered guide RNA targeting TIS, poly(A) sequence, or any other part of the coding sequence or non-coding sequence. The engineered guide RNA disclosed herein promotes ADAR-mediated RNA editing of adenosine in TIS, poly(A) sequence, any part of the coding sequence of RNA, any part of the non-coding sequence of RNA, or any combination thereof.
[0475] Table 10 provides examples of target genes that can be targeted by engineered RNA payloads encoded by the expression cassettes of the present disclosure. The target gene can be a wild-type gene, or the target gene can be a mutant gene. Targeting the gene using an engineered RNA payload can treat a disease associated with the target gene.
[0476] Table 10 - Exemplary gene targets and associated disorders
[0477]
[0478]
[0479] The expression cassettes disclosed herein can express payloads to target, modify and / or express any sequence of interest. Selected targets of interest that may be targeted by the payloads described herein for the treatment of relevant disorders are discussed below by way of example.
[0480] MAPT
[0481] The disclosure provides an expression cassette encoding an engineered guide RNA, which promotes RNA editing of MAPT to knock down the expression of Tau protein. Tau pathology may be the main driving factor of a variety of neurodegenerative diseases (collectively referred to as Tau disease). For example, the disease in which Tau can play a major role includes but is not limited to Alzheimer's disease (AD), frontotemporal dementia (FTD), Parkinson's disease, progressive supranuclear palsy (PSP), corticobasal degeneration (CBD) and chronic traumatic encephalopathy. The feature of Tau disease is the intracellular accumulation of neurofibrillary tangles (NFT) composed of aggregated, misfolded Tau (MAPT gene). Therefore, the disclosed engineered guide RNA targeting MAPT RNA is used for ADAR-mediated editing to knock down Tau protein, which can prevent or improve the progress of a variety of diseases, including but not limited to AD, FTD, autism, traumatic brain injury, Parkinson's disease and Dravet syndrome.
[0482] Therefore, the engineered guide RNA of the present disclosure can target MAPT for RNA editing, thereby driving the reduction of Tau protein expression. In some embodiments, Tau protein expression in human neurons is reduced. In some embodiments, the present disclosure provides a composition of an engineered guide RNA, the engineered guide RNA targets MAPT and promotes the ADAR-mediated RNA editing of MAPT, to reduce the pathogenic level of Tau by targeting the key adenosine for deamination present in the translation start site (TIS). In some embodiments, the engineered guide RNA of the present disclosure targets the coding sequence in MAPT. For example, the coding sequence can be the translation start site (TIS) (AUG) of MAPT, and the engineered guide RNA can promote the ADAR-mediated RNA editing of AUG to GUG. The engineered guide RNA of the present disclosure can target one or more TIS in MAPT to reduce or completely inhibit Tau protein expression.
[0483] For example, in some embodiments, the engineered guide RNA targets the AUG at the 18th nucleotide in exon 1 (c.1, Nm_005910.5; GRCh37 / Hg19; also referred to as "c.1" for coding nucleotide 1), referred to as a conventional TIS. In some embodiments, the engineered guide RNA targets the AUG at the 48th nucleotide (c.31) in exon 1. In some embodiments, the engineered guide RNA targets the AUG at the 6th nucleotide (c.379) in exon 5. With reference to the 2N4R Tau isoform containing 441 amino acids (Np_005901; GRCh37 / Hg19), these three TIS correspond to methionine (Met) 1, 11, and 127, respectively. In some embodiments, the engineered guide RNA targets the AUG at the 108th nucleotide (c.91) in exon 1. In some embodiments, one or more than one engineered guide RNA of the present disclosure targets any one or any combination of the four TISs. For example, a single engineered guide RNA disclosed herein can be designed to target more than one of the four TISs mentioned above. In some embodiments, more than one engineered guide RNA is designed to each independently target more than one of the four TISs mentioned above. In some embodiments, the engineered guide RNA disclosed herein can target any one or any combination of TISs in exon 1 (c.1, c.31, and c.91). Targeting these sites in MAPT facilitates editing, thereby inhibiting translation and reducing the expression of Tau protein. In some embodiments, the ratio of 3R to 4R subtypes of Tau can be measured by protein analysis (e.g., using ELISA or flow cytometry) to evaluate the effect of RNA editing, wherein a ratio of 1: 1 represents the ratio in a healthy adult brain. In some embodiments, any engineered guide RNA disclosed herein is packaged in an AAV vector and delivered by virus.
[0484] In some embodiments, the engineered guide RNA targets a non-coding sequence in MAPT. The non-coding sequence can be a poly(A) signal sequence, and the engineered guide RNA can promote the ADAR-mediated RNA editing of one or more adenosines in the poly(A) signal sequence of MAPT. In some embodiments, the engineered guide RNA of the present disclosure can be multiplexed to target more than one poly(A) signal sequence in MAPT. In some embodiments, the engineered guide RNA of the present disclosure can be multiplexed to target TIS and one or more poly(A) signal sequences in MAPT. In some embodiments, the engineered guide RNA can be multiplexed to target non-coding sequences and coding sequences in MAPT. The engineered guide RNA of the present disclosure promotes the ADAR-mediated RNA editing of MAPT, thereby achieving protein knockdown.
[0485] In some embodiments, the engineered guide RNAs of the present disclosure promote ADAR-mediated RNA editing of target adenosines from 1% to 100%. The engineered guide RNAs of the present disclosure can promote editing of target adenosines from 40% to 90%. In some embodiments, the engineered guide RNAs of the present disclosure can promote at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, 100%, from 5% to 20%, from 20% to 40%, from 40% to 60%, from 60% to 80%, from 80% to 100%, from 60% to 80%, from 70% to 90%, or up to 90% or more of target adenosine RNA editing. Optionally, in addition, the engineered guide RNAs of the present disclosure can promote these levels of on-target RNA editing while maintaining less than 10% off-target adenosine editing. Optionally, in addition, the engineered guide RNAs disclosed herein can promote these levels of on-target RNA editing while maintaining less than 30%, less than 25%, less than 20%, less than 15%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1%, or 0% off-target adenosine editing.
[0486] In some embodiments, the engineered guide RNA of the present disclosure promotes ADAR-mediated RNA editing of MAPT, which results in protein level knockdown. The protein level knockdown is quantified as a reduction in Tau protein expression. The engineered guide RNA of the present disclosure can promote Tau protein knockdown from 1% to 100%. The engineered guide RNA of the present disclosure can promote from 1% to 10%, from 10% to 20%, from 20% to 30%, from 30% to 40%, from 40% to 50%, from 50% to 60%, from 60% to 70%, from 70% to 80%, from 80% to 90%, from 90% to 100%, from 20% to 40%, from 30% to 50%, from 40% to 60%, from 50% to 70%, From 60% to 80%, from 20% to 50%, from 30% to 60%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, or at least 90% Tau protein knockdown. In some embodiments, the engineered guide RNA of the present disclosure promotes from 30% to 60% Tau protein knockdown. Tau protein knockdown can be measured by an assay comparing a sample or subject treated with an engineered guide RNA to a control sample or subject not treated with the engineered guide RNA.
[0487] α-Synuclein
[0488] The α-synuclein gene consists of 5 exons and encodes a protein of 140 amino acids with a predicted molecular weight of about 14.5 kDa. The encoded product is a protein of unknown function that is essentially disordered. Typically, α-synuclein is a monomer. Under certain stress conditions or other unknown reasons, α-synuclein will self-aggregate into oligomers. Lewy-related pathology (LRP), mainly composed of α-synuclein present in the brains of more than 50% of Alzheimer's patients confirmed by autopsy. Although the molecular mechanism of how α-synuclein affects the development of Alzheimer's disease is still unclear, experimental evidence shows that α-synuclein interacts with Tau-p and may trigger the intracellu...
Claims
1. An expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO:17, SEQ ID NO:1250 or SEQ ID NO:1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO:1241, SEQ ID NO:1251, SEQ ID NO:1252, SEQ ID NO:1253 or SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A termination sequence comprising a sequence having at least 80% identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269.
2. An expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO:17, SEQ ID NO:1250 or SEQ ID NO:1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO:1241, SEQ ID NO:1251, SEQ ID NO:1252, SEQ ID NO:1253 or SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and Terminates a sequence.
3. An expression cassette comprising: Promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A termination sequence comprising a sequence having at least 80% identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269.
4. An expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A stop sequence comprising a sequence at least 80% identical to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289.
5. An expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and Terminates a sequence.
6. An expression cassette, include: Promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A stop sequence comprising a sequence at least 80% identical to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289.
7. The expression cassette of any one of claims 4-6, wherein the promoter sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263.
8. The expression cassette of any one of claims 4-6, wherein the promoter sequence comprises a sequence having at least 95% sequence identity to any one of SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263.
9. The expression cassette of any one of claims 4-8, wherein the termination sequence comprises a sequence having at least 90% sequence identity to any one of SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289.
10. The expression cassette of any one of claims 4-8, wherein the termination sequence comprises a sequence having at least 95% sequence identity to any one of SEQ ID NO:60, SEQ ID NO:708-SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243-SEQ ID NO:1247, SEQ ID NO:1254-SEQ ID NO:1257, SEQ ID NO:1264-SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287-SEQ ID NO:1289.
11. The expression cassette of any one of claims 1-10, wherein the promoter sequence comprises SEQ ID NO:
17.
12. The expression cassette of any one of claims 1-10, wherein the promoter sequence comprises SEQ ID NO: 1262.
13. The expression cassette of any one of claims 1-10, wherein the promoter sequence comprises SEQ ID NO: 1250.
14. The expression cassette of any one of claims 1-10, wherein the promoter sequence comprises SEQ ID NO: 1251.
15. The expression cassette of any one of claims 1-10, wherein the promoter sequence comprises SEQ ID NO: 1252.
16. The expression cassette of any one of claims 1-10, wherein the promoter sequence comprises SEQ ID NO: 1253.
17. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1264.
18. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1265.
19. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1254.
20. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1255.
21. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1257.
22. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO:
60.
23. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1242.
24. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1269.
25. The expression cassette of any one of claims 1-16, wherein the termination sequence comprises SEQ ID NO: 1017.
26. The expression cassette of any one of claims 1-25, wherein the small RNA payload comprises an engineered guide RNA capable of hybridizing to a target sequence.
27. The expression cassette of claim 26, wherein the engineered guide RNA has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% reverse complementarity to the target sequence.
28. The expression cassette of claim 26 or claim 27, wherein the engineered guide RNA comprises at least one base pair mismatch relative to the target sequence.
29. The expression cassette of any one of claims 26-28, wherein the target sequence comprises an adenosine residue.
30. The expression cassette of any one of claims 26-29, wherein the target sequence is an RNA sequence.
31. The expression cassette of claim 30, wherein the RNA sequence is mRNA or pre-mRNA.
32. The expression cassette of any one of claims 26-31, wherein the target sequence comprises a G to A mutation relative to the wild-type sequence.
33. The expression cassette of any one of claims 26-32, wherein the target sequence comprises a missense mutation or a nonsense mutation relative to the wild-type sequence.
34. The expression cassette of any one of claims 26-33, wherein the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with Charcot-Marie-Tooth disease type 1A (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2).
35. The expression cassette of any one of claims 26-34, wherein the payload sequence has at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity to SEQ ID NO: 1273, SEQ ID NO: 1274 or SEQ ID NO:
61.
36. The expression cassette of any one of claims 1-35, wherein the small RNA payload comprises an antisense oligonucleotide, siRNA, shRNA, miRNA, or tracrRNA.
37. The expression cassette of any one of claims 1-36, wherein the small RNA payload has a length of no less than 20 nucleotide residues and no more than 500 nucleotide residues.
38. The expression cassette of any one of claims 1-37, wherein the small RNA payload has a length of no less than 60 residues and no more than 100 residues.
39. The expression cassette of any one of claims 1-37, wherein the small RNA payload is no less than 80 residues and no more than 120 residues in length.
40. The expression cassette of any one of claims 1-37, wherein the small RNA payload is no less than 100 residues and no more than 140 residues in length.
41. The expression cassette of any one of claims 1-37, wherein the small RNA payload is no less than 130 residues and no more than 170 residues in length.
42. The expression cassette of any one of claims 1-41, wherein the payload sequence further comprises an Sm binding sequence or a hairpin sequence.
43. The expression cassette of claim 42, wherein the hairpin sequence comprises a U7 hairpin.
44. An expression cassette as claimed in claim 42 or claim 43, wherein the hairpin sequence has at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity with SEQ ID NO:52 or SEQ ID NO:54, or the Sm binding sequence has at least 80%, at least 85%, at least 90%, at least 95% or 100% sequence identity with SEQ ID NO:56 or SEQ ID NO:
58.
45. The expression cassette of any one of claims 1-44, wherein the length of the expression cassette is no less than 1300 nucleotide residues and no more than 2160 nucleotide residues.
46. The expression cassette of any one of claims 1-45, wherein the expression cassette has at least 80% sequence identity to a U1 sequence or a U7 sequence.
47. The expression cassette of claim 46, wherein the U1 sequence is a mouse U1 sequence or a human U1 sequence.
48. The expression cassette of claim 46, wherein the U7 sequence is a mouse U7 sequence or a human U7 sequence.
49. The expression cassette of any one of claims 1-48, wherein the promoter sequence comprises a zinc finger 143 motif capable of recruiting the ZNF143 transcription factor.
50. The expression cassette of any one of claims 1-49, wherein the promoter sequence comprises an OCT-1 transcription factor binding sequence capable of recruiting an OCT-1 transcription factor.
51. The expression cassette of any one of claims 1-50, wherein the promoter sequence comprises a proximal sequence element capable of recruiting SNAPc.
52. The expression cassette of claim 51, wherein the proximal sequence element is capable of integron-dependent recruitment of RNA polymerase II.
53. The expression cassette of any one of claims 1-52, wherein the small RNA payload is capable of forming a guide-target RNA scaffold comprising structural features upon hybridization of the small RNA payload to a target sequence.
54. The expression cassette of claim 53, wherein the structural feature is a bulge, a mismatch, an internal loop, a hairpin, or a combination thereof.
55. The expression cassette of claim 54, wherein the structural feature comprises the protrusion, and wherein the protrusion is a symmetrical protrusion.
56. The expression cassette of claim 54, wherein the structural feature comprises the protrusion, and wherein the protrusion is an asymmetric protrusion.
57. The expression cassette of claim 54, wherein the structural feature comprises the internal loop, and wherein the internal loop is a symmetrical internal loop.
58. The expression cassette of claim 54, wherein the structural feature comprises the internal loop, and wherein the internal loop is an asymmetric internal loop.
59. The expression cassette of claim 54, wherein the structural feature comprises the hairpin, and wherein the hairpin is a recruiting hairpin or a non-recruiting hairpin.
60. The expression cassette of any one of claims 43-59, wherein the guide-target RNA scaffold comprises a wobble base pair.
61. A recombinant polynucleotide encoding one or more expression cassettes according to any one of claims 1-60.
62. The recombinant polynucleotide of claim 61, encoding two expression cassettes of any one of claims 1-60, wherein the expression cassettes comprise a first promoter, a second promoter, a first termination sequence, and a second termination sequence.
63. The recombinant polynucleotide of claim 62, wherein the first promoter and the second promoter are the same.
64. The recombinant polynucleotide of claim 62, wherein the first promoter and the second promoter are different.
65. The recombinant polynucleotide of any one of claims 62-64, wherein the first termination sequence and the second termination sequence are identical.
66. The recombinant polynucleotide of any one of claims 62-64, wherein the first termination sequence and the second termination sequence are different.
67. The recombinant polynucleotide of any one of claims 62-66, wherein the first promoter comprises SEQ ID NO:
17.
68. The recombinant polynucleotide of any one of claim 62 or claims 64-67, wherein the second promoter comprises SEQ ID NO: 1262.
69. The recombinant polynucleotide of any one of claims 62-68, wherein the first termination sequence comprises SEQ ID NO: 1264.
70. The recombinant polynucleotide of any one of claims 62-64 or claims 66-69, wherein the second termination sequence comprises SEQ ID NO: 1265.
71. A recombinant polynucleotide as described in claim 62, wherein (a) the first promoter sequence includes SEQ ID NO:17, the first termination sequence includes SEQ ID NO:1264, the second promoter sequence includes SEQ ID NO:1262, and the second termination sequence includes SEQ ID NO:1265; or (b) the first promoter sequence includes SEQ ID NO:17, the first termination sequence includes SEQ ID NO:1265, the second promoter sequence includes SEQ ID NO:1262, and the second termination sequence includes SEQ ID NO:1264.
72. A viral vector encapsulating the expression cassette of any one of claims 1-60 or the recombinant polynucleotide of any one of claims 61-71.
73. The viral vector of claim 72, wherein the viral vector comprises two or more, three or more, or four or more expression cassettes of any one of claims 1-60.
74. The viral vector of claim 72 or claim 73, wherein the viral vector is an adeno-associated viral vector.
75. The viral vector of claim 74, wherein the adeno-associated viral vector is selected from the group consisting of: AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, AAV 10. AAV11, AAV12, AAV13, AAV14, AAV15, AAV16, AAV-DJ, AAV-DJ / 8, AAV-DJ / 9, AAV1 / 2, AAV.rh8, AAV.rh10, AAV.rh20, AAV.rh39, AAV.Rh43, AAV.Rh74, AAV .v66, AAV.Oligo001, AAV.SCH9, AAV.r3.45, AAV.RHM4-1, AAV.hu37, AAV.Anc80, AAV.Anc80L65, AAV.7m8, AAV.PhP.eB, AAV.PhP.V1, AAV.PHP.B, AAV.PhB. C1, AAV.PhB.C2, AAV.PhB.C3, AAV.PhB.C6, AAV.cy5, AAV2.5, AAV2tYF, AAV3B, AAV.LK03, AAV.HSC1, AAV.HSC2, AAV.HSC3, AAV.HSC4, AAV.HSC5, AAV.HSC6, AAV.HSC7, AAV.HSC8, AAV.HSC9, AAV.HSC10, AAV.HSC11, AAV.HSC12, AAV.HSC13, AAV.HSC14, AAV.HSC15, AAV.HSC16, AAV.HSC17, AAVhu68, their chimeras, and combinations thereof.
76. A pharmaceutical composition comprising an expression cassette as described in any one of claims 1-60, a recombinant polynucleotide as described in any one of claims 61-71, or a viral vector as described in any one of claims 72-75 and a pharmaceutically acceptable excipient, carrier, diluent, or a combination thereof.
77. A method for expressing a small RNA payload in a cell, the method comprising delivering an expression cassette as described in any one of claims 1-60, a recombinant polynucleotide as described in any one of claims 61-71, a viral vector as described in any one of claims 72-75, or a pharmaceutical composition as described in claim 76 into a cell and expressing the small RNA payload encoded by the expression cassette in the cell.
78. A method for editing a target sequence, the method include: delivering the expression cassette of any one of claims 1-60, the recombinant polynucleotide of any one of claims 61-71, the viral vector of any one of claims 72-75, or the pharmaceutical composition of claim 76 to a cell encoding the target sequence; expressing the small RNA payload in the cell, wherein the small RNA payload comprises an engineered guide RNA capable of hybridizing to a target sequence; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and The target sequence is edited using the editing enzyme.
79. A method for editing a target sequence, the method include: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO:17, SEQ ID NO:1250 or SEQ ID NO:1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO:1241, SEQ ID NO:1251, SEQ ID NO:1252, SEQ ID NO:1253 or SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A termination sequence comprising a sequence having at least 80% identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269; expressing the small RNA payload in the cell; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and The target sequence is edited using the editing enzyme.
80. A method for editing a target sequence, the method include: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO:17, SEQ ID NO:1250 or SEQ ID NO:1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO:1241, SEQ ID NO:1251, SEQ ID NO:1252, SEQ ID NO:1253 or SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and Termination sequence; expressing the small RNA payload in the cell; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and The target sequence is edited using the editing enzyme.
81. A method for editing a target sequence, the method include: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: Promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A termination sequence comprising a sequence having at least 80% identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269; expressing the small RNA payload in the cell; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and The target sequence is edited using the editing enzyme.
82. A method for editing a target sequence, the method include: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a stop sequence comprising a sequence at least 80% identical to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289; expressing the small RNA payload in the cell; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and The target sequence is edited using the editing enzyme.
83. A method for editing a target sequence, the method include: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and Termination sequence; expressing the small RNA payload in the cell; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and The target sequence is edited using the editing enzyme.
84. A method for editing a target sequence, the method include: delivering an expression cassette to a cell encoding the target sequence, wherein the expression cassette comprises: Promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a stop sequence comprising a sequence at least 80% identical to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289; expressing the small RNA payload in the cell; The small RNA payload hybridizes with the target sequence to form a guide-target RNA scaffold; recruiting an editing enzyme to the target sequence; and editing the target sequence using the editing enzyme.
85. The method of any one of claims 77-84, wherein the promoter sequence comprises SEQ ID NO:
17.
86. The method of any one of claims 77-84, wherein the promoter sequence comprises SEQ ID NO: 1262.
87. The method of any one of claims 77-84, wherein the promoter sequence comprises SEQ ID NO: 1250.
88. The method of any one of claims 77-84, wherein the promoter sequence comprises SEQ ID NO: 1251.
89. The method of any one of claims 77-84, wherein the promoter sequence comprises SEQ ID NO: 1252.
90. The method of any one of claims 77-84, wherein the promoter sequence comprises SEQ ID NO: 1253.
91. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO: 1264.
92. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO: 1265.
93. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO: 1254.
94. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO:1255.
95. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO: 1257.
96. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO:
60.
97. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO: 1242.
98. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO:1269.
99. The method of any one of claims 77-90, wherein the termination sequence comprises SEQ ID NO:1017.
100. The method of any one of claims 78-99, wherein the target sequence comprises a mutation relative to a wild-type sequence.
101. The method of embodiment 100, wherein editing the target sequence corrects the mutation in the target sequence.
102. The method of claim 100 or claim 101, wherein the mutation is a missense mutation.
103. The method of claim 100 or claim 101, wherein the mutation is a nonsense mutation.
104. The method of any one of claims 100-103, wherein the mutation is a G to A mutation.
105. The method of any one of claims 100-104, wherein the mutation is associated with a disease.
106. The method of claim 105, wherein the disease is a synucleinopathy, Parkinson's disease, dementia with Lewy bodies, multiple system atrophy, Charcot-Marie-Tooth disease, hereditary compressive susceptibility neuropathy, Yuan-Harel-Lupski syndrome, Tauopathies, Alzheimer's disease, frontotemporal dementia, progressive supranuclear palsy, corticobasal degeneration, chronic traumatic encephalopathy, autism, traumatic brain injury, Dravet syndrome, Crohn's disease, muscular dystrophy, B-cell leukemia, Dejerine-Sottas disease, Stargardt disease, alpha-1 antitrypsin deficiency, Tay-Sachs disease, cystic fibrosis, liposomal acid lipase deficiency, or Gaucher disease.
107. The method of any one of claims 78-106, wherein the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with Charcot-Marie-Tooth disease type 1A (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2).
108. The method of claims 78-107, wherein editing the target sequence comprises editing an untranslated region of the target sequence.
109. The method of claim 108, wherein the untranslated region is a 5' untranslated region or a 3' untranslated region.
110. The method of claim 109, wherein the 3' untranslated region is a polyadenylation sequence.
111. The method of any one of claims 78-110, wherein editing the target sequence comprises editing a translation start site.
112. The method of any one of claims 78-111, wherein editing the target sequence alters expression of the target sequence.
113. The method of claim 112, wherein editing the target sequence increases expression of the target sequence.
114. The method of claim 112, wherein editing the target sequence reduces expression of the target sequence.
115. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette as described in any one of claims 1-60, a recombinant polynucleotide as described in any one of claims 61-71, a viral vector as described in any one of claims 72-75, or a pharmaceutical composition as described in claim 76; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
116. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO:17, SEQ ID NO:1250 or SEQ ID NO:1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO:1241, SEQ ID NO:1251, SEQ ID NO:1252, SEQ ID NO:1253 or SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A termination sequence comprising a sequence having at least 80% identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
117. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: a) SEQ ID NO:17, SEQ ID NO:1250 or SEQ ID NO:1262; b) SEQ ID NO: 13 or SEQ ID NO: 15; or c) SEQ ID NO:1241, SEQ ID NO:1251, SEQ ID NO:1252, SEQ ID NO:1253 or SEQ ID NO:1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and Termination sequence; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
118. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette comprising: Promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and A termination sequence comprising a sequence having at least 80% identity to any of: a) SEQ ID NO: 1002, SEQ ID NO: 1017, SEQ ID NO: 1264 or SEQ ID NO: 1265; or b) SEQ ID NO:60, SEQ ID NO:771, SEQ ID NO:930, SEQ ID NO:1007, SEQ ID NO:1021, SEQ ID NO:1242, SEQ ID NO:1254, SEQ ID NO:1255, SEQ ID NO:1257 or SEQ ID NO:1269; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
119. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a stop sequence comprising a sequence at least 80% identical to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
120. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette comprising: A promoter sequence comprising a sequence having at least 80% sequence identity to any of: SEQ ID NO: 13 - SEQ ID NO: 17, SEQ ID NO: 167 - SEQ ID NO: 707, SEQ ID NO: 1241, SEQ ID NO: 1248 - SEQ ID NO: 1253, or SEQ ID NO: 1259 - SEQ ID NO: 1263; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and Termination sequence; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
121. A method of treating a disease in a subject, the method include: administering to the subject a composition comprising an expression cassette comprising: Promoter sequence; a payload sequence under the transcriptional control of the promoter sequence, the payload sequence comprising a small RNA payload; and a stop sequence comprising a sequence at least 80% identical to any one of SEQ ID NO:60, SEQ ID NO:708–SEQ ID NO:1240, SEQ ID NO:1242, SEQ ID NO:1243–SEQ ID NO:1247, SEQ ID NO:1254–SEQ ID NO:1257, SEQ ID NO:1264–SEQ ID NO:1272, SEQ ID NO:1275, or SEQ ID NO:1287–SEQ ID NO:1289; delivering the expression cassette to cells of the subject; and The small RNA payload is expressed in the cell, thereby treating the disease.
122. The method of any one of claims 115-121, wherein the promoter sequence comprises SEQ ID NO:
17.
123. The method of any one of claims 115-121, wherein the promoter sequence comprises SEQ ID NO:1262.
124. The method of any one of claims 115-121, wherein the promoter sequence comprises SEQ ID NO:1250.
125. The method of any one of claims 115-121, wherein the promoter sequence comprises SEQ ID NO:1251.
126. The method of any one of claims 115-121, wherein the promoter sequence comprises SEQ ID NO:1252.
127. The method of any one of claims 115-121, wherein the promoter sequence comprises SEQ ID NO:1253.
128. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1264.
129. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1265.
130. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1254.
131. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1255.
132. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1257.
133. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:
60.
134. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1242.
135. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1269.
136. The method of any one of claims 115-127, wherein the termination sequence comprises SEQ ID NO:1017.
137. The method of any one of claims 115-136, wherein the disease is a synucleinopathy, Parkinson's disease, dementia with Lewy bodies, multiple system atrophy, Charcot-Marie-Tooth disease, hereditary compressive susceptibility neuropathy, Yuan-Harel-Lupski syndrome, Tauopathies, Alzheimer's disease, frontotemporal dementia, progressive supranuclear palsy, corticobasal degeneration, chronic traumatic encephalopathy, autism, traumatic brain injury, Dravet syndrome, Crohn's disease, muscular dystrophy, B-cell leukemia, Dejerine-Sottas disease, Stargardt disease, alpha-1 antitrypsin deficiency, Tay-Sachs disease, cystic fibrosis, liposomal acid lipase deficiency, or Gaucher disease.
138. The method of any one of claims 115-137, wherein the small RNA payload comprises an engineered guide RNA that hybridizes to a target sequence, and wherein the cell encodes the target sequence.
139. The method of claim 138, wherein the target sequence encodes alpha-synuclein (SNCA), peripheral myelin protein 22 (PMP22), double homeobox 4 (DUX4), leucine-rich repeat kinase 2 (LRRK2), Tau (MAPT), progranulin (GRN), repeats of PMP22 associated with Charcot-Marie-Tooth disease type 1A (CMT1A), ATP-binding cassette subfamily A member 4 (ABCA4), amyloid precursor protein (APP), alpha-1 antitrypsin (SERPINA1), hexosaminidase A (HEXA), cystic fibrosis transmembrane conductance regulator (CFTR), lipase A (LIPA), glucosylceramidase beta (GBA), PTEN-induced kinase 1 (PINK1), or methyl CpG binding protein 2 (MECP2).
140. The method of claim 138 or claim 139, further comprising forming a guide-target RNA scaffold after hybridization of the engineered guide RNA to the target sequence, recruiting an editing enzyme to the target sequence, and editing the target sequence with the editing enzyme.
141. The method of any one of claims 138-140, wherein the target sequence comprises a mutation relative to a wild-type sequence.
142. The method of claim 141, wherein editing the target sequence corrects the mutation in the target sequence.
143. The method of claim 141 or claim 142, wherein the mutation is a missense mutation.
144. The method of claim 141 or claim 142, wherein the mutation is a nonsense mutation.
145. The method of any one of claims 141-144, wherein the mutation is a G to A mutation.
146. The method of any one of claims 141-145, wherein the mutation is associated with the disease.
147. The method of any one of claims 140-146, wherein editing the target sequence comprises editing an untranslated region of the target sequence.
148. The method of claim 147, wherein the untranslated region is a 5' untranslated region or a 3' untranslated region.
149. The method of claim 148, wherein the 3' untranslated region is a polyadenylation sequence.
150. The method of any one of claims 140-149, wherein editing the target sequence comprises editing a translation start site.
151. The method of any one of claims 140-150, wherein editing the target sequence alters expression of the target sequence.
152. The method of claim 151, wherein editing the target sequence increases expression of the target sequence.
153. The method of claim 151, wherein editing the target sequence reduces expression of the target sequence.
154. The method of any one of claims 78-114 or 140-153, wherein the guide-target RNA scaffold comprises a structural feature.
155. The method of claim 154, wherein the structural feature is a protrusion, a mismatch, an internal loop, a hairpin, or a combination thereof.
156. The method of claim 155, wherein the structural feature comprises the protrusion, and wherein the protrusion is a symmetrical protrusion.
157. The method of claim 155, wherein the structural feature comprises the protrusion, and wherein the protrusion is an asymmetric protrusion.
158. The method of any one of claims 155 to 157, wherein the structural feature comprises the inner ring, and wherein the inner ring is a symmetrical inner ring.
159. The method of any one of claims 155 to 157, wherein the structural feature comprises the inner ring, and wherein the inner ring is an asymmetric inner ring.
160. The method of any one of claims 155 to 159, wherein the structural feature comprises the hairpin, and wherein the hairpin is a recruiting hairpin or a non-recruiting hairpin.
161. The method of any one of claims 78-114 or 140-160, wherein the guide-target RNA scaffold comprises a wobble base pair.
162. The method of any one of claims 78-114 or 140-161, wherein the editing enzyme comprises an ADAR, APOBEC, or Cas nuclease.
163. The method of claim 162, wherein the ADAR comprises ADAR1, ADAR2, ADAR3, or a combination thereof.
164. The method of any one of claims 78-114 or 140-163, wherein the target sequence comprises RNA or DNA.
165. The method of any one of claims 78-114 or 140-164, wherein the target sequence is mRNA or pre-mRNA.
166. The method of any one of claims 78-114 or 140-165, wherein editing the target sequence comprises deamidating nucleotides of the target sequence.
167. The method of any one of claims 78-114 or 140-166, wherein the target sequence is edited with an efficiency of at least 10%, at least 20%, or at least 25%.