Compositions, systems, and methods for RNA editing using dkc1
By using engineered gsnoRNA and DKC1 protein in host cells, highly efficient RNA editing was achieved, solving the problem of low target RNA editing efficiency in existing technologies, improving the functional editing effect of RNA, enhancing the expression of full-length proteins, and reducing the translation of meaningless stop codons.
Patent Information
- Application Number
- CN202280037916.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-05-26
- Filing Date
- 2022-05-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-05-26
AI Technical Summary
Existing RNA editing methods result in low editing efficiency of target RNAs, making it difficult to effectively alter RNA functionality through pseudouridineization, particularly by converting uridine residues in coding regions into pseudouridine residues.
Engineered gsnoRNA and DKC1 protein are used to edit target RNA in host cells. The gsnoRNA hybridizes with the target uridine residues of the target RNA and recruits the DKC1 protein to modify the target uridine residues into pseudouridine residues. The guide sequence of gsnoRNA and the catalytic action of DKC1 protein are used to achieve efficient editing.
It significantly improved the editing efficiency of target RNA and could effectively alter RNA functionality, such as converting uridine residues in the coding region into pseudouridine residues, thereby enhancing the functional editing effect of RNA, increasing the expression level of full-length protein, and reducing the translation of meaningless stop codons.
Smart Images

Figure BDA0004569025240000251 
Figure BDA0004569025240000261 
Figure BDA0004569025240000262
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to International Patent Application No. PCT / CN2021 / 096122, filed on May 26, 2021, the contents of which are incorporated herein by reference in their entirety.
[0003] Sequence list reference merging
[0004] The contents of the following submitted ASCII text file are incorporated herein by reference in their entirety: Computer-readable format of sequence lists (CRF) (filename: 165392000442SEQLIST.TXT, record date: May 24, 2022, size: 75,959 bytes). Technical Field
[0005] This application relates to compositions, systems, and methods for editing RNA by targeting pseudouridine hydratylation using DKC1. Background Technology
[0006] Pseudoruridides (Ψ) are the most abundant post-transcriptional modified nucleotides in stable RNAs (including tRNA, rRNA, snRNA, and mRNA), accounting for approximately 5% of total ribonucleotides. The conversion of uridine to Ψ (pseudoruridination) requires two distinct chemical reactions: the breaking of the C1'-N1 glycosidic bond and the formation of a new C1'-C5 bond that reattaches the base to the sugar. Pseudoruridination is a true isomerization reaction that generates additional hydrogen bond donors and affects a wide range of functional aspects, such as protein synthesis and increased terminator-codon readthrough, depending on the type of RNA carrying Ψ and its location within the RNA sequence (Yu and Meier, 2014, RNABiology 11:1483-1494). Many mRNAs with Ψ are located in coding regions, and most of them are responsive to environmental stresses, suggesting functional significance (Carlile et al., 2014, Nature 515:143).
[0007] In eukaryotes and archaea, pseudouridine conversion is introduced via box-H / ACA ribonucleoproteins (RNPs), each containing a unique small RNA (box-H / ACA RNA, which is one of two major classes of small nucleolar RNA or 'snoRNA') and four core proteins (dyskeratin (DKC1), NHP2, NOP10, and GAR1). Dyskeratin (DKC1; also known as NAP57 / CBF5) is a highly conserved multifunctional protein that acts as an RNA-guided pseudouridine synthase, directing the specific uridine enzyme to convert to pseudouridine. Dyskeratin is concentrated in the nucleolus and Cajalbody (CB), where it binds to three other highly conserved proteins (Nop10, Nhp2, and Gar1) to form a tetramer, which can be incorporated into the composition of different nuclear RNPs that play key biological roles. Within the nucleolus, the tetramer binds to H / ACA small nucleolar RNA (snoRNA) to form an H / ACA snoRNP, which regulates rRNA treatment and pseudouridine RNA targets via snoRNA-guided base complementarity. Within the CB, the tetramer binds to CB-specific small RNA (scaRNA) to form a scaRNP, which guides pseudouridine nucleotide ... Based on this guide-acceptor base pairing scheme, Karijolich and Yu (2011, Nature 474:395-398) designed an artificial box H / ACARNA to introduce Ψ into mRNA at the premature stop codon (PTC) in Saccharomyces cerevisiae. They confirmed that Ψ does indeed incorporate into TRM4 mRNA at this PTC. Pseudouridine-pTC promotes meaningless repression by altering ribosomal decoding (Fernandez et al., 2013, Nature 500:107-110; Wu et al., 2015, Methods in Enzymology 560:187-217; US 8,603,457). Using a similar strategy, other studies have shown that artificial H / ACARNAs can be site-specifically pseudouridine-modified pre-mRNAs after microinjection into Xenopus laevis oocytes (Chen et al., 2010, Mol Cell Biol 30:4108-4119). In both cases, the artificial H / ACARNAs were modified to alter the loop used as the guide sequence, but otherwise, these snoRNAs remained unchanged.
[0008] While site-specific pseudouridine cleavage or target RNA is a potentially powerful technique, available methods to date have resulted in low editing efficiency of target RNA. Therefore, there is a need in the art to optimize gsnoRNA, gsnoRNA-based gene editing systems, and methods for editing target RNA via pseudouridine cleavage. Summary of the Invention
[0009] This application provides a method for editing target RNA in host cells using gsnoRNA and DKC1 protein. An embodiment of this method, also referred to herein as the “RESTART” method, can be used to allow readthrough of RNA transcripts containing early stop codons (PTCs).
[0010] In some aspects, this application provides a method for editing target RNA in a host cell, comprising introducing a guide small nucleolar RNA (gsnoRNA) and a nucleic acid molecule encoding a DKC1 protein into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence in the target RNA containing a target uridine residue, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the gsnoRNA contains a scaffold sequence derived from wild-type H / ACA-snoRNAs selected from: ACA19, ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17.
[0011] In some aspects, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a scaffold sequence derived from wild-type ACA2b, ACA36, ACA44, ACA27, E2, ACA3, or ACA17, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the DKC1 protein is an endogenous DKC1 protein of the host cell. In some embodiments, the method further comprises introducing a nucleic acid encoding the DKC1 protein into the host cell.
[0012] In some aspects, this article provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence in the target RNA containing a target uridine residue, wherein the gsnoRNA contains a nucleotide sequence selected from the group consisting of nucleotide sequences provided in Tables 2, 3 or 4, and wherein the gsnoRNA recruits the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue.
[0013] In some aspects, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a nucleotide sequence selected from the following: SEQ ID NO: 4-6, 9-12, 15-19, 22-36 and 177-179, and wherein the gsnoRNA recruits the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue.
[0014] In some aspects, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a nucleotide sequence selected from SEQ ID NO: 20-21 and 145-150, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the DKC1 protein is an endogenous DKC1 protein of the host cell. In some embodiments, the method further comprises introducing a nucleic acid encoding the DKC1 protein into the host cell.
[0015] In some embodiments of any of the methods described above, the DKC1 protein has cytoplasmic localization in the host cell.
[0016] In some embodiments of any of the above methods, the DKC1 protein comprises a DKC1 protein fragment of amino acid residues 41 to 420 corresponding to human DKC1 isotype 3 protein, wherein the amino acid number is according to SEQ ID NO:2.
[0017] In some embodiments of the methods described above, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater) identity with SEQ ID NO:88. In some embodiments, the DKC1 protein comprises the amino acid sequence of SEQ ID NO:88. In some embodiments of the methods described above, the DKC1 protein comprises a naturally occurring DKC1 isotype having cytoplasmic localization in the host cell.
[0018] In some aspects, this article provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the host cell expresses a cytoplasm-localized DKC1 isotype, and wherein the gsnoRNA recruits the DKC1 isotype to modify the target uridine residue in the target RNA into a pseudouridine residue.
[0019] In some aspects, this document provides a method for editing target RNA in a host cell, comprising introducing (a) an engineered gsnoRNA and (b) a splice-switching antisense oligonucleotide (ASO) into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the ASO enhances the expression of a DKC1 protein, the DKC1 protein being an endogenous DKC1 isotype having cytoplasmic localization in the host cell, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue. In some embodiments, the gsnoRNA contains a scaffold sequence derived from wild-type H / ACA-snoRNAs selected from the following: ACA19, ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17.
[0020] In some embodiments of any of the methods described above, the DKC1 isotype corresponds to isotype 3 of the human DKC1 protein.
[0021] In some embodiments according to any of the methods described above, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater) identity with SEQ ID NO:2. In some embodiments, the DKC1 protein comprises the amino acid sequence of SEQ ID NO:2.
[0022] In some embodiments of any of the above methods, the DKC1 protein comprises a DKC1 protein fragment of amino acid residues 1 to 419 corresponding to the full-length human DKC1 isotype 1 protein, wherein the amino acid numbering is according to SEQ ID NO: 1.
[0023] In some embodiments of the methods described above, the gsnoRNA comprises a scaffold sequence derived from ACA2b. In some embodiments of the methods described above, the gsnoRNA comprises a scaffold sequence derived from ACA36. In some embodiments, the gsnoRNA comprises a mutation in the 3' hairpin of the ACA36 scaffold.
[0024] In some embodiments of any of the methods described above, the gsnoRNA contains a scaffold sequence derived from ACA19.
[0025] In some embodiments of any of the methods described above, the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA. In some embodiments, at least one of the one or more guide sequences is located in a hairpin structure (also referred to herein as a "3' hairpin structure") at the 3' end of the wild-type H / ACA-snoRNA. In some embodiments, at least one of the one or more guide sequences is located in a hairpin structure (also referred to herein as a "5' hairpin structure") at the 5' end of the wild-type H / ACA-snoRNA. In some embodiments, the gsnoRNA comprises a single guide sequence. In some embodiments, the gsnoRNA comprises two or more (e.g., 2, 3, 4, 5, 6 or more) guide sequences.
[0026] In some embodiments of any of the methods described above, the gsnoRNA contains one or more mutations (e.g., substitution, insertion, and / or deletion) in one or more hairpin structures (e.g., 3' and / or 5' hairpin structures) of wild-type ACA19.
[0027] In some embodiments of any of the above methods, the engineered gsnoRNA contains one or more substitution mutations in the nucleotides of the poly-U sequence of wild-type H / ACA-snoRNA, wherein the poly-U sequence contains at least four consecutive U residues.
[0028] In some embodiments of any of the above methods, the engineered gsnoRNA contains one or more insertion or deletion mutations between nucleotide residues in the guide region that hybridizes with the target uridine and the H / ACA box of the wild-type H / ACAsnoRNA, such that the engineered gsnoRNA contains 14 or 15 nucleotides between the nucleotide residues in the guide region that hybridizes with the target uridine and the H / ACA box.
[0029] In some embodiments, one or more mutations are selected from: replacing residues 26-29 with UUCU, replacing residues 26-29 with UGUU, adding G to the 3' hairpin structure after residue 115, and adding a dinucleotide sequence (XX, e.g., CU) to the 5' hairpin after residue 8, wherein X is a nucleotide selected from A, U, C, and G, and wherein the numbering is according to SEQ ID NO:37. In some embodiments, the dinucleotide sequence is a portion of a guide RNA designed to hybridize with the target RNA.
[0030] In some embodiments according to any of the above methods, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO:3-12, 15-19, 22-36, and 177-179. In some embodiments, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO:15-19.
[0031] In some embodiments of any of the methods described above, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO: 20-21 and 145-150.
[0032] In some embodiments of the methods described above, the method includes introducing a nucleic acid molecule encoding gsnoRNA into a host cell. In some embodiments, the nucleic acid molecule encoding gsnoRNA is under a small RNA promoter. In some embodiments, the nucleic acid molecule encoding gsnoRNA is under a promoter selected from the group consisting of U6 (e.g., transcribed by polymerase III) and U1 (e.g., transcribed by polymerase II) promoters. In some embodiments, the nucleic acid molecule encoding gsnoRNA is embedded in an intron sequence located between a first exon sequence and a second exon sequence, wherein the first exon sequence, the intron sequence, and the second exon sequence are derived from a naturally occurring gene. In some embodiments, the intron sequence is an intron of an endogenous gene from the host cell, wherein the gene is selected from EIF3A, SNHG12, RPL21, and RPSA. In some embodiments, the intron sequence is an intron of an exogenous gene (such as HBB). In some implementations, the nucleic acid encoding gsnoRNA is not embedded in an intron sequence.
[0033] In some embodiments of any of the methods described above, the nucleic acid molecule encoding the DKC1 protein is present in the viral vector. In some embodiments, the nucleic acid molecule encoding gsnoRNA is present in the viral vector.
[0034] In some embodiments of the methods described above, the method includes introducing a vector containing a first nucleic acid sequence encoding a DKC1 protein and a second nucleic acid sequence encoding gsnoRNA into a host cell. In some embodiments, the nucleic acid molecule encoding the DKC1 protein and the nucleic acid molecule encoding the gsnoRNA are present in different vectors. In some embodiments, the vector is a viral vector. In some embodiments, the vector is an adeno-associated virus (AAV) vector.
[0035] In some embodiments according to any of the above methods, the gsnoRNA comprises one or more chemically modified nucleosides and / or nucleoside linkages. In some embodiments, the gsnoRNA comprises one or more nucleosides modified with 2'-OMe or 2'-MOE. In some embodiments, the gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides. In some embodiments, the gsnoRNA comprises one or more phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA comprises a 5' cap modification. In some embodiments, the 5' cap modification is 7-methylguanosine (m 7G) Cap. In some embodiments, the gsnoRNA does not contain one or more chemically modified nucleosides or nucleoside-to-nucleoside linkages.
[0036] In some embodiments of any of the methods described above, the efficiency of editing the target RNA is at least 10% (e.g., at least about 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 60% or higher).
[0037] In some embodiments of any of the above methods, wherein the sequence in the target RNA containing the target uridine is an early stop codon in the sequence encoding the protein, the method results in the expression of the full-length protein in the host cell at at least 4% (e.g., at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%) of the expression level of the full-length protein without the early stop codon.
[0038] In some embodiments of the methods described above, wherein the sequence in the target RNA containing the target uridine is an early stop codon in the sequence encoding a protein, the method results in the expression of the full-length protein, wherein the expression of the protein can be detected without enrichment (e.g., without enrichment by immunoprecipitation). In some embodiments, the protein is detected by a tag (e.g., by a fluorescent tag). In some embodiments, the protein is detected by immunostaining according to methods known in the art.
[0039] In some embodiments of any of the methods described above, wherein the sequence in the target RNA containing the target uridine is an early stop codon in the sequence encoding the protein, the method results in the expression of the full-length protein in at least 20% of the host cells (e.g., at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of the host cells).
[0040] In some embodiments of any of the above methods, the target RNA is not ribosomal RNA (rRNA), such as endogenous rRNA of the host cell.
[0041] In some embodiments of any of the methods described above, the target RNA is messenger RNA (mRNA). In some embodiments, the sequence in the target RNA containing the target uridine is a stop codon, and modifying the target uridine to pseudouridine causes the stop codon to be translated into a coding codon. In some embodiments, the stop codon is a premature stop codon (PTC). In some embodiments, the PTC is associated with a genetic disease or condition. In some embodiments, the sequence in the target RNA containing the target uridine is a stop codon, and modifying the target uridine to pseudouridine reduces or prevents meaningless mediated decay (NMD).
[0042] In some embodiments according to any of the above methods, the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to gsnoRNA. In some embodiments, the RNP complex comprises NOP10, GAR1, and NHP2.
[0043] In some embodiments of the methods described above, the host cell is an archaea cell. In some embodiments, the host cell is a eukaryotic cell. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a human cell.
[0044] In some embodiments of any of the methods described above, the method is performed in vivo. In some embodiments, the method is performed ex vivo.
[0045] In some aspects, this article provides a method for treating a disease or condition in a subject that is associated with a PTC in a target RNA, comprising editing the target RNA in the subject's cells using any of the methods described above, wherein the gsnoRNA contains a guide sequence that hybridizes with the PTC in the target RNA, and wherein modifying a uridine residue in the PTC to a pseudouridine residue results in read-through translation of the PTC in the target RNA, thereby treating the subject's disease or condition.
[0046] In some implementations, the disease or condition is selected from the following: cystic fibrosis, Hurler syndrome, alpha-1-antitrypsin (A1AT) deficiency, Parkinson's disease, Alzheimer's disease, albinism, amyotrophic lateral sclerosis (ALS), asthma, thalassemia, Cadasil syndrome, Charcot-Marie-Tooth disease, chronic obstructive pulmonary disease (COPD), distal spinal muscular atrophy (DSMA), Duchenne / Becker muscular dystrophy, dystrophic epidermolysis bullosa, epidermolysis bullosa, Fabry disease, and Factor V Leiden associated disease. Disorders, familial adenomatous polyposis, galactosemia, Gaucher's disease, glucose-6-phosphate dehydrogenase, hemophilia, hereditary hemochromatosis, Hunter syndrome, Huntington's disease, inflammatory bowel disease (IBD), hereditary polyagglutination syndrome, Leber congenital amaurosis, Lesch-Nyhan syndrome, Lynch syndrome, Marfan syndrome, mucopolysaccharidosis, muscular dystrophy, myotonic dystrophy type I and II, neurofibromatosis, Niemann-Pick disease type A, B, and C. Diseases, NY-esol-related cancers, Peutz-Jeghers syndrome, phenylketonuria, Pompe's disease, primary ciliary body disease, prothrombin mutation-related diseases (such as prothrombin G20210A mutation), pulmonary hypertension, (autosomal dominant) retinitis pigmentosa, Sandhoff disease, severe combined immunodeficiency syndrome (SCID), sickle cell anemia, spinal muscular atrophy, Stargardt's disease, Tay-Sachs disease, Usher syndromeSyndrome, X-linked immune deficiency, Sturge-Weber syndrome, and cancer.
[0047] In some aspects, this document provides an engineered gsnoRNA comprising a guide sequence for hybridization with a sequence containing a target uridine residue in a target RNA of a host cell. In some embodiments, the gsnoRNA comprises a nucleotide sequence selected from SEQ ID NO: 4-6, 9-12, 15-19, 22-36, and 177-179, and the gsnoRNA is capable of recruiting the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the gsnoRNA comprises a nucleotide sequence selected from SEQ ID NO: 20-21 and 145-150, and the gsnoRNA is capable of recruiting the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue.
[0048] In some respects, this article provides an engineered gsnoRNA comprising a guide sequence that hybridizes to a target uridine residue in a target RNA of a host cell, wherein the gsnoRNA comprises a scaffold sequence derived from wild-type H / ACA-snoRNA selected from the following: ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17, and wherein the gsnoRNA can recruit the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue.
[0049] In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA contains a 5' cap modification. In some embodiments, the 5' cap modification is 7-methylguanosine (m... 7 G) Cap. In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA comprises one or more chemically modified nucleosides and / or nucleoside linkages. In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA comprises one or more nucleosides modified with 2'-OMe or 2'-MOE. In some embodiments, the gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides. In some embodiments, the gsnoRNA comprises one or more phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphate-thioester nucleoside linkages.
[0050] In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA comprises a scaffold sequence derived from ACA2b. In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA comprises a scaffold sequence derived from ACA36. In some embodiments, the gsnoRNA comprises a mutation in the 3' hairpin of the ACA36 scaffold.
[0051] In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA contains a scaffold sequence derived from ACA19.
[0052] In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA. In some embodiments, at least one of the one or more guide sequences is located in a hairpin structure at the 3' end of the wild-type H / ACA-snoRNA. In some embodiments, at least one of the one or more guide sequences is located in a hairpin structure at the 5' end of the wild-type H / ACA-snoRNA. In some embodiments, the gsnoRNA comprises a single guide sequence. In some embodiments, the gsnoRNA comprises two or more (e.g., 2, 3, 4, 5, 6 or more) guide sequences.
[0053] In some embodiments of any of the engineered gsnoRNAs described above, the gsnoRNA contains one or more mutations (e.g., substitution, insertion, and / or deletion) in one or more hairpin structures (e.g., 3' and / or 5' hairpin structures) of wild-type ACA19.
[0054] In some embodiments of any of the engineered gsnoRNAs described above, the engineered gsnoRNA contains one or more substitution mutations in the nucleotides of a polyU sequence of wild-type H / ACA-snoRNA, wherein the polyU sequence contains at least four consecutive U residues.
[0055] In some embodiments of any of the engineered gsnoRNAs described above, the engineered gsnoRNA contains one or more insertion or deletion mutations between nucleotide residues in the guide region that hybridizes with the target uridine and the H / ACA box of the wild-type H / ACAsnoRNA, such that the engineered gsnoRNA contains 14 or 15 nucleotides between the nucleotide residues in the guide region that hybridizes with the target uridine and the H / ACA box.
[0056] In some embodiments, one or more mutations are selected from the following: substitution of residues 26-29 with UUCU, substitution of residues 26-29 with UGUU, addition of G to the 3' hairpin structure after residue 115, and addition of a dinucleotide sequence (XX, e.g., CU) to the 5' hairpin after residue 8, wherein X is a nucleotide selected from A, U, C, and G, and wherein the numbering is according to SEQ ID NO:37. In some embodiments, the engineered gsnoRNA comprises a nucleotide sequence selected from the following: SEQ ID NO:15-19.
[0057] In some aspects, this document provides an engineered gsnoRNA comprising a guide sequence that hybridizes to a target uridine residue in a target RNA of a host cell, wherein the gsnoRNA comprises a scaffold sequence derived from wild-type ACA19, wherein the engineered gsnoRNA comprises one or more insertion or deletion mutations between nucleotide residues in the guide region for hybridization with the target uridine and the H / ACA box of the wild-type ACA19, wherein the engineered gsnoRNA comprises 14 or 15 nucleotides between nucleotide residues in the guide region for hybridization with the target uridine and the H / ACA box, and wherein the gsnoRNA can recruit the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the one or more mutations are selected from the following: substitution of residues 26-29 with UUCU, substitution of residues 26-29 with UGUU, addition of G to the 3' hairpin structure after residue 115, and addition of a dinucleotide sequence (XX, e.g., CU) to the 5' hairpin after residue 8, wherein X is a nucleotide selected from A, U, C, and G, and wherein the numbering is according to SEQ ID NO:37. In some embodiments, the gsnoRNA comprises a nucleotide sequence selected from the following: SEQ ID NO:15-19.
[0058] In some aspects, this document provides an isolated nucleic acid molecule comprising a sequence encoding an engineered gsnoRNA of any of the foregoing embodiments. In some embodiments, this document provides a vector (e.g., a viral vector) comprising the nucleic acid molecule.
[0059] In some aspects, this article provides an engineered RNA editing system comprising: (a) a gsnoRNA containing a guide sequence that hybridizes to a target RNA of a host cell containing a target uridine residue, or a nucleic acid molecule encoding the gsnoRNA; and (b) a DKC1 protein, or a nucleic acid molecule encoding the DKC1 protein, wherein the gsnoRNA can recruit the DKC1 protein to modify the target uridine residue in the target RNA into a pseudouridine residue.
[0060] In some embodiments, the DKC1 protein has cytoplasmic localization in host cells. In some embodiments according to any of the above methods, the DKC1 protein comprises a DKC1 protein fragment of amino acid residues 41 to 420 corresponding to human DKC1 isotype 3 protein, wherein the amino acid numbering is according to SEQ ID NO:2.
[0061] In some embodiments of the methods described above, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater) identity with SEQ ID NO:88. In some embodiments, the DKC1 protein comprises the amino acid sequence of SEQ ID NO:88. In some embodiments of the methods described above, the DKC1 protein comprises a naturally occurring DKC1 isotype having cytoplasmic localization in the host cell.
[0062] In some embodiments of any of the above-described engineered RNA editing systems, the DKC1 isotype corresponds to isotype 3 of the human DKC1 protein.
[0063] In some embodiments of any of the engineered RNA editing systems described above, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater) identity with SEQ ID NO:2. In some embodiments, the DKC1 protein comprises the amino acid sequence of SEQ ID NO:2.
[0064] In some embodiments of any of the engineered RNA editing systems described above, the DKC1 protein comprises a DKC1 protein fragment of amino acid residues 1 to 419 corresponding to the full-length human DKC1 isotype 1 protein, wherein the amino acid numbering is according to SEQ ID NO:1.
[0065] In some embodiments of any of the engineered RNA editing systems described above, the gsnoRNA comprises a scaffold sequence derived from ACA2b. In some embodiments of any of the engineered RNA editing systems described above, the gsnoRNA comprises a scaffold sequence derived from ACA36. In some embodiments, the gsnoRNA comprises a mutation in the 3' hairpin of the ACA36 scaffold.
[0066] In some embodiments of any of the above-described engineered RNA editing systems, the gsnoRNA contains a scaffold sequence derived from ACA19.
[0067] In some embodiments of any of the engineered RNA editing systems described above, the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA. In some embodiments, at least one of the one or more guide sequences is located in a hairpin structure at the 3' end of the wild-type H / ACA-snoRNA. In some embodiments, at least one of the one or more guide sequences is located in a hairpin structure at the 5' end of the wild-type H / ACA-snoRNA. In some embodiments, the gsnoRNA comprises a single guide sequence. In some embodiments, the gsnoRNA comprises two or more (e.g., 2, 3, 4, 5, 6 or more) guide sequences.
[0068] In some embodiments of any of the above-described engineered RNA editing systems, the gsnoRNA contains one or more mutations (e.g., substitution, insertion, and / or deletion) in one or more hairpin structures (e.g., 3' and / or 5' hairpin structures) of wild-type ACA19.
[0069] In some embodiments of any of the above-described engineered RNA editing systems, the engineered gsnoRNA contains one or more substitution mutations in the nucleotides of a poly-U sequence of wild-type H / ACA-snoRNA, wherein the poly-U sequence contains at least four consecutive U residues.
[0070] In some embodiments of any of the engineered RNA editing systems described above, the engineered gsnoRNA contains one or more insertion or deletion mutations between nucleotide residues in the guide region that hybridizes with the target uridine and the H / ACA box of the wild-type H / ACAsnoRNA, such that the engineered gsnoRNA contains 14 or 15 nucleotides between the nucleotide residues in the guide region that hybridizes with the target uridine and the H / ACA box.
[0071] In some embodiments of engineered RNA editing systems, the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to gsnoRNA. In some embodiments, the ribonucleoprotein complex comprises NOP10, GAR1, and / or NHP2.
[0072] In some embodiments, one or more mutations are selected from the following: replacing residues 26-29 with UUCU, replacing residues 26-29 with UGUU, adding G to the 3' hairpin structure after residue 115, and adding a dinucleotide sequence (XX, e.g., CU) to the 5' hairpin after residue 8, wherein X is a nucleotide selected from A, U, C, and G, and wherein the number is according to SEQ ID NO:37.
[0073] In some embodiments of any of the engineered RNA editing systems described above, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO: 3-12, 15-19, 22-36, and 177-179. In some embodiments, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO: 15-19.
[0074] In some embodiments of any of the engineered RNA editing systems described above, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO: 20-21 and 145-150.
[0075] In some respects, this article provides a pharmaceutical composition comprising any of the above-described gsnoRNA, nucleic acid molecules, or engineered RNA editing systems, and a pharmaceutically acceptable carrier.
[0076] In some respects, this article provides a host cell that contains any of the above-mentioned gsnoRNA, nucleic acid molecules, or engineered RNA editing systems.
[0077] In some respects, this article provides a kit for editing target RNA in host cells, comprising any of the above-described gsnoRNA, nucleic acid molecules, or engineered RNA editing systems.
[0078] This document also provides compositions, kits, and articles for use in any of the above methods. Attached Figure Description
[0079] Figure 1A-1F This demonstrates the readthrough of early stop codons mediated by engineered snoRNA. Figure 1A A schematic diagram of the “RESTART” method design is provided. The snoRNP complex is shown with a dashed box. Figure 1B A schematic diagram showing the structure of reporter gene-1 and the guide snoRNA construct is provided. In this reporter gene-1, a 15-base sequence is inserted between the codons for amino acids 154 and 155. The 15-base DNA sequence is shown, along with the premature stop codon (PTC) site (TAG). A positive control (Venus-GGT) is included. Figure 1C-1D HEK293T cells were co-transfected with reporter gene-1 and a guide snoRNA construct. Venus expression was detected using a high-content imaging system. Figure 1C A representative fluorescence image of the cells shows the expression level of Venus. Scale bar, 200 μm. Figure 1D The dot plot shows the relative fraction of Venus-positive cells. Figure 1E Western blot analysis showed the expression level of DKC1 protein after stable knockdown. Figure 1F The bar chart shows the relative fraction of Venus-positive cells in cells stably knocked down by shControl and DKC1 co-transfected with reporter gene-1 and gsnoRNA constructs.
[0080] Figure 2A-2C This illustrates the PTC readthrough effect mediated by gsnoRNA with different constructs. The dot plot shows the effect of co-transfection of reporter gene-1 and gsnoRNA into host introns. Figure 2A gsnoRNA within HBB introns ( Figure 2B ) or gsnoRNA transcribed from the small RNA promoter ( Figure 2C The relative fraction of Venus-positive cells. The structure of the gsnoRNA construct is shown at the bottom of each figure.
[0081] Figures 3A-3F Display for Figure 1A-1F The predicted secondary structures of the gsnoRNA scaffolds are shown above. The secondary structures and base pair probabilities mentioned above were predicted using the RNAfold server, as described by Gruber et al. (The Vienna RNAwebsuite. NucleicAcids Res 36, W70-4 (2008)), which is incorporated herein by reference in its entirety.
[0082] Figures 4A-4F The optimization of the gsnoRNA scaffold showed that it improved the efficiency of PTC readthrough. Figure 4A Predicted secondary structures of the gACA19, gACA2b, and gACA36 scaffolds. The secondary structures and base pair probabilities were predicted using the RNAfold server. Seven mutations were observed in the structure of the gACA19 scaffold. Figure 4B The structure of the gsnoRNA construct. Figure 4C The dot plot shows the path ( Figure 4B The relative fraction of Venus-positive cells co-transfected with the reporter gene-1 and gsnoRNA of the construct. Figure 4D ,through( Figure 4BRepresentative fluorescence images of cells co-transfected with the reporter gene-1 and gsnoRNA of the construct. Scale bar, 200 μm. Figure 4E The dot plot shows the relative fractions of Venus-positive cells co-transfected with reporter gene-1 and engineered gACA19 scaffolds with different mutations. The engineered location of gACA19 is annotated in ( Figure 4A )middle. Figure 4F , ( Figure 4E Representative fluorescence images of ). Scale bar, 200 μm.
[0083] Figures 5A-5D Display for Figures 4A-4F The predicted secondary structures of the gsnoRNA scaffold are shown. The aforementioned secondary structures and base pair probabilities were predicted using the RNAfold server.
[0084] Figures 6A-6B This demonstrates the engineering of the gACA36 support. Figure 6A The predicted secondary structure of the engineered gACA36 scaffold. The aforementioned secondary structure and base pair probabilities were predicted using the RNAfold server. Figure 6B The dot plot shows the relative fractions of Venus-positive cells co-transfected with reporter gene-1 and different gsnoRNA constructs.
[0085] Figure 7A-7I The results showed that exogenous DKC1-isotype 3 protein improved the efficiency of PTC-RT. Figure 7A The structures of two isotypes of the human DKC1 transcript. Exons are numbered at the top, coding regions are represented by solid boxes, and UTRs are represented by white boxes. NLS, nuclear localization signal. Figure 7B The diagram shows the structure of the reporter gene-3 construct, in which the gsnoRNA is arranged in tandem with the reporter gene. The sequence surrounding the PTC site and the aforementioned PTC site (TAA / TAG / TGA) are also shown. Figure 7C Western blot analysis revealed the expression level of DKC1 protein in HEK293T DKC1 stably overexpressing cells. Santa Cruz and Abcam anti-DKC1 antibodies targeted the C-terminal and N-terminal regions of the DKC1 protein, respectively. Figure 7D-7F The reporter gene-3 construct will be transfected into control HEK293T, DKC1-isotype 1 stably overexpressing, and DKC1-isotype 3 stably overexpressing cells, respectively. Figure 7D Representative fluorescence images of cells. Scale bar, 200 μm. Figure 7E The bar chart shows the relative percentage of EGFP-positive cells. Figure 7F The bar chart shows the relative fraction of EGFP intensity. Figure 7GThe bar chart shows the relative percentage of EGFP-positive cells in HEK293T cells transfected with different reporter gene-3 constructs and co-transfected with different reporter gene-3 and DKC1-isotype 3 (200 ng) constructs. Figure 7H The bar chart shows the relative fractions of EGFP intensity in HEK293T cells transfected with different reporter gene-3 constructs and co-transfected with different reporter gene-3 and DKC1-isotype 3 (200 ng) constructs. Figure 7I Locus-specific Ψ modifications in the reporter gene-3 transcript were detected using a label-free qPCR-based method. The curves were obtained through high-resolution melting analysis. The Ψ site is specifically labeled with a CMC chemical substance; upon reverse transcription, the Ψ-CMC adduct induces mutations / deletions at or around the Ψ site in the cDNA, thus producing a transition at melting temperature. HEK293T cells were co-transfected with different reporter gene-3 and DKC1-isotype 3 (200 ng) constructs.
[0086] Figures 8A-8D The results showed that exogenous DKC1-isotype 3 protein improved the efficiency of PTC-RT. Figures 8A-8C The reporter gene-1 and gsnoRNA constructs were co-transfected into control HEK293T, DKC1-isotype 1 stably overexpressing, and DKC1-isotype 3 stably overexpressing cells, respectively. Figure 8A Representative fluorescence images of cells. Scale bar, 200 μm. Figure 8B The bar chart shows the relative fraction of Venus-positive cells. Figure 8C The bar chart shows the relative fraction of Venus intensity. Figure 8D The bar chart shows the relative percentages of EGFP-positive cells co-transfected with reporter gene-3 along with empty vector (Vec), DKC1-isotype 1 (DKC1 iso1), or DKC1-isotype 3 (DKC1iso3) constructs. Statistical analysis of the bar chart was performed using an unpaired Student's t-test.
[0087] Figure 9 The reading efficiency of different truncated DKC1-isotype 3 constructs is shown. The dot plot shows the relative fraction of EGFP-positive cells co-transfected with reporter gene-3 and different truncated DKC1-isotype 3 constructs.
[0088] Figures 10A-10E This shows a comparison of readability for different stop codons. Figure 10A Representative fluorescence images of cells transfected with different reporter gene-3 constructs. Scale bar, 200 μm. Figure 10BRepresentative fluorescence images of cells co-transfected with reporter gene-3 and DKC1-isotype 3 (200 ng) constructs. Scale bar, 200 μm. Figure 10C-10E Bar graphs, they show the effects of reporter gene-3-TAA ( Figure 10C ), reporter gene-3-TAG ( Figure 10D ) or reporter gene-3-TGA ( Figure 10E The relative fraction of EGFP-positive cells co-transfected with a decreasing amount of DKC1-isotype 3 construct.
[0089] Figure 11A-11C Detection of locus-specific Ψ modifications. HEK293T cells were co-transfected with different reporter gene-3 and DKC1-isotype 3 (200 ng) constructs. Locus-specific Ψ modifications in reporter gene-3 transcripts were detected. Figure 11A ) and the Ψ1045 site in 18S rRNA ( Figure 11B-11C The results were detected using a non-radioactive label-based qPCR method. The curves were obtained through high-resolution denaturation analysis.
[0090] Figure 12. Guiding snoRNA to target hereditary diseases caused by nonsense mutations. Schematic diagram of the PTC disease reporter gene and gsnoRNA construct. Complementary region between gsnoRNA (top) and the target site in the PTC disease gene (bottom).
[0091] Figure 13 The complementary region between gsnoRNA (top) and the target site (bottom) in the PTC disease gene.
[0092] Figure 14A -C indicates nonsense mutations that can cause hereditary diseases through RESTART correction. Figure 14A -B, dot plots, these show the displayed gsnoRNA and RESTART v1 PTC disease reporter gene construct ( Figure 14A ) and RESTART v2 DKC1-isotype 3 construct ( Figure 14B The relative percentage of co-transfected EGFP-positive cells. Figure 14C The bar graph shows the relative fraction of EGFP-positive cells co-transfected with gsnoRNA and PTC disease reporter gene with or without the DKC1-isotype 3 construct.
[0093] Figures 15A-15E RESTART is shown to be delivered by RNA oligonucleotides. Figure 15A -C, the structure of gsnoRNA prepared by in vitro transcription. Figure 15DThe bar charts show the relative percentages of EGFP-positive cells transfected with the displayed gsnoRNA construct, in vitro transcribed gsnoRNA oligonucleotide, or chemically synthesized gsnoRNA oligonucleotide. Figure 15E The structure of chemically synthesized gsnoRNA oligonucleotides. Detailed Implementation
[0094] This application provides methods and compositions for editing target RNA in host cells, comprising introducing an engineered guide small nucleolar RNA (gsnoRNA) into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, and wherein the gsnoRNA recruits a DKC1 protein to modify the target uridine residue in the target RNA into a pseudouridine residue. In some aspects, the gsnoRNA is an engineered gsnoRNA containing one or more mutations compared to a wild-type H / ACA scaffold. In some embodiments, the one or more mutations increase the editing efficiency of the gsnoRNA. In some aspects, the method further includes increasing the amount of DKC1 protein with cytoplasmic localization, thereby increasing the editing efficiency of the gsnoRNA / DKC1 protein complex. In some aspects, the methods and compositions provided herein can be used to edit premature stop codons (PTCs) in target gene mRNA, thereby suppressing meaningless-mediated decay of the mRNA and promoting the translation of full-length proteins. In some implementations, the methods disclosed herein can be used to treat diseases associated with PTC in target genes.
[0095] In some aspects, the present invention provides engineered gsnoRNAs and gsnoRNA scaffolds or nucleic acid molecules encoding the aforementioned gsnoRNAs. In some embodiments, the engineered gsnoRNA scaffolds are based on wild-type H / ACA snoRNA scaffolds identified by the inventors as having higher editing efficiency compared to other scaffolds. In some embodiments, the engineered gsnoRNA scaffolds contain mutations that increase their editing efficiency.
[0096] The methods and compositions described in this application are based, at least in part, on the unexpected discovery that expression of cytoplasm-localized DKC1 isotypes (e.g., human DKC1 isotype 3) significantly increases the editing efficiency of target RNAs using the gsnoRNA / DKC1 system. In one aspect, the inventors recognized that gsnoRNA editing efficiency can be increased by introducing exogenous DKC1 isotypes with cytoplasm localization. In another aspect, the inventors identified truncated and deletion variants of the DKC1 protein that can be used to increase gsnoRNA editing efficiency.
[0097] In some aspects, this document provides nucleic acid constructs encoding gsnoRNAs used according to the methods described herein. In some embodiments, the inventors identify promoters and construct conformations for gsnoRNA expression that provide increased editing efficiency of the gsnoRNA.
[0098] I. Definition
[0099] Unless otherwise defined below, the terms are used as they are generally in the art.
[0100] The terms "polynucleotide," "nucleic acid," "nucleotide sequence," and "nucleic acid sequence" are used interchangeably. They refer to polymeric forms of nucleotides of any length, deoxyribonucleotides or ribonucleotides, or similar substances.
[0101] As used herein, “complementarity” refers to the ability of a nucleic acid to form hydrogen bonds with another nucleic acid via typical Watson-Crick and wobbling base pairing. The complementarity percentage shows the percentage of residues in a nucleic acid molecule that can form hydrogen bonds with a second nucleic acid (i.e., Watson-Crick and wobbling base pairing) (e.g., approximately 50%, 60%, 70%, 80%, 90%, and 100% complementarity out of 10, respectively). “Complete complementarity” means that all consecutive residues in the nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in the second nucleic acid sequence. “Substantially complementary” as used herein refers to complementarity at least approximately 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% over a region of approximately 40, 50, 60, 70, 80, 100, 150, 200, 250 or more nucleotides, or to two nucleic acids hybridizing under stringent conditions.
[0102] The term "hybridization" generally refers to specific hybridization, excluding non-specific hybridization. Specific hybridization occurs under selected experimental conditions, using techniques well-known in the art, to ensure that most stable interactions between the probe and the target are such that the probe and the target have at least 70%, preferably at least 80%, more preferably at least 90% sequence identity.
[0103] This article uses the term "mismatch" to refer to opposite nucleotides in a double-stranded RNA complex that do not form a complete base pair according to the Watson-Crick and wobble base pairing rules. Mismatched nucleotides are GA, CA, UC, AA, GG, CC, and UU pairs. Wobble base pairs are GU, IU, IA, and IC base pairs.
[0104] This disclosure provides several types of compositions based on polynucleotides or peptides (including variants and derivatives). These include, for example, substitutions, insertions, deletions, and covalent variants and derivatives. The term "derivative" is synonymous with the term "variant" and generally refers to a molecule that has been modified and / or altered in any way relative to a reference molecule or starting molecule.
[0105] Therefore, polynucleotides encoding peptides or polypeptides containing substitutions, insertions and / or additions, deletions, and covalent modifications relative to a reference sequence (specifically, the polypeptide sequences disclosed herein) are included within the scope of this invention. For example, sequence tags or amino acids (such as one or more lysines) may be added to a peptide sequence (e.g., at the N-terminus or C-terminus). Sequence tags can be used for peptide detection, purification, or localization. Lysines can be used to increase peptide solubility or allow biotinylation. Alternatively, amino acid residues at the carboxyl and amino terms of the amino acid sequence of a peptide or protein may be optionally deleted to provide a truncated sequence. Depending on the intended use of the sequence, certain amino acids (e.g., C-terminal or N-terminal residues) may be alternatively deleted, for example, to express the sequence as soluble or as a portion of a larger sequence linked to a solid carrier.
[0106] The term "identity" refers to the overall correlation between polymeric molecules, such as between polynucleotide molecules (e.g., DNA and / or RNA molecules) and / or between polypeptide molecules. The percentage of identity between two polynucleotide sequences can be calculated, for example, by aligning the two sequences for optimal comparison purposes (e.g., gaps can be introduced in one or both of the first and second nucleic acid sequences for optimal alignment, and inconsistencies can be ignored for comparison purposes). In some embodiments, the length of the sequence aligned for comparison purposes is at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% of the length of the reference sequence. Nucleotides at corresponding nucleotide positions are then compared. When a position in the first sequence is occupied by the same nucleotide as the corresponding position in the second sequence, the molecules are consistent at that position. The percentage of identity between two sequences is a function of the number of consistent positions shared by the sequences, taking into account the number and length of gaps introduced for optimal alignment of the two sequences. Mathematical algorithms can be used to compare sequences and determine the percentage of identity between two sequences. For example, the percentage of identity between two nucleic acid sequences can be determined using methods such as those described in: Computational Molecular Biology, Lesk, AM, ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, DW, ed., Academic Press, New York, 1993; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; Computer Analysis of Sequence Data, Part I, Griffin, AM and Griffin, HG, eds., Humana Press, New Jersey, 1994; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., Stockton Press, New York, 1991; each of which is incorporated herein by reference. For example, the percentage of identity between two nucleic acid sequences can be determined using the algorithm of Meyers and Miller (CABIOS, 1989, 4:11-17) (which has been incorporated into the ALIGN program (version 2.0)), using a PAM 120 weighted residual table, a gap length penalty of 12, and a gap penalty of 4.Alternatively, the percentage of identity between two nucleic acid sequences can be determined using the GAP procedure in the GCG software package, using the NWSgapdna.CMP matrix. Commonly used methods for determining the percentage of identity between sequences include (but are not limited to) those disclosed in Carillo, H. and Lipman, D., SIAM J Applied Math., 48:1073 (1988), which are incorporated herein by reference. Techniques for determining identity are incorporated into publicly available computer programs. Exemplary computer software for determining homology between two sequences includes (but is not limited to) the GCG software package (Devereux, J. et al., Nucleic Acids Research, 12(1), 387 (1984)), BLASTP, BLASTN, and FASTA (Altschul, SF et al., J. Molec. Biol., 215, 403 (1990)).
[0107] The "percentage of amino acid sequence identity (%)" for the polypeptide sequences identified herein is defined as the percentage of amino acid residues in the candidate sequence that are identical to amino acid residues in the compared polypeptide after sequence alignment, taking into account any conserved substitutions as part of sequence identity. Alignments for the purpose of determining the percentage of amino acid sequence identity can be performed in various ways within the scope of the art, for example, using publicly available computer software such as BLAST, BLAST-2, ALIGN, Megalign (DNASTAR), or MUSCLE software. Those skilled in the art can determine the appropriate parameters for determining alignments, including any algorithm required to achieve maximum alignment across the full length of the sequences being compared. However, for the purposes of this document, the amino acid sequence identity value % is generated using the sequence comparison computer program MUSCLE (Edgar, RC, Nucleic Acids Research 32(5):1792-1797, 2004; Edgar, RC, BMC Bioinformatics 5(1):113, 2004, each of which is incorporated herein by reference in its entirety for all purposes).
[0108] The terms “non-naturally occurring” and “engineered” are used interchangeably and indicate human involvement. When referring to a nucleic acid molecule or polypeptide, the above terms mean that the nucleic acid molecule or polypeptide contains at least one modification (e.g., at least one mutation, such as substitution, insertion, or deletion, or at least one non-naturally occurring chemical modification) compared to a naturally occurring nucleic acid molecule or polypeptide, or at least substantially does not contain at least one of the components that are naturally associated with them in nature and as found in nature.
[0109] As used in this article regarding ACA scaffold sequences, the term "wild-type" refers to the sequence of naturally occurring box H / ACA small nucleolar RNA.
[0110] As used herein, “expression” refers to the process of transcription of a polynucleotide from a DNA template (such as transcription into mRNA or other RNA transcripts) and / or the subsequent translation of the transcribed mRNA into peptides, polypeptides, or proteins. Transcripts and encoded polypeptides are collectively referred to as “gene products.” If the polynucleotide originates from genomic DNA, expression may include splicing of the mRNA in eukaryotic cells.
[0111] This article uses the term "polypeptide" or "peptide" to encompass all kinds of naturally occurring and synthetic proteins, including protein fragments of all lengths, fusion proteins, and modified proteins, including but not limited to glycoproteins, and all other types of modified proteins (e.g., proteins produced by phosphorylation, acetylation, myristylation, palmitoylation, glycosylation, oxidation, formylation, amidation, polyglutamylation, ADP-ribosylation, PEGylation, biotinylation, etc.).
[0112] The term "pharmaceutical composition" refers to a formulation which is in a form in which the biological activity of the active ingredient contained therein is permitted and which does not contain any other components that would have unacceptable toxicity to a subject to whom the formulation will be administered.
[0113] "Pharmaceutically acceptable carriers" refer to one or more components in a pharmaceutical preparation that are non-toxic to the subjects, excluding the active ingredient. Pharmaceutically acceptable carriers include (but are not limited to) buffers, excipients, stabilizers, cryoprotectants, tension agents, preservatives, and combinations thereof. Pharmaceutically acceptable carriers or excipients preferably meet toxicological and manufacturing testing requirements and / or are included in guidelines for inactive ingredients established by the U.S. Food and Drug Administration or other state / federal governments, or are listed in the United States Pharmacopeia or other pharmacopoeias recognized for use in mammals and more specifically in humans.
[0114] The term "packaging insert" refers to the instruction leaflet typically included in the commercial packaging of a therapeutic product, which contains information about the indications, usage, dosage, administration, combination therapy, contraindications, and / or warnings for using these therapeutic products.
[0115] "Article" is any manufactured article (e.g., package or container) or kit that contains at least one reagent, such as a pharmaceutical agent for treating a disease or condition (e.g., coronavirus infection), or a probe for specifically detecting a biomarker described herein. In some embodiments, the manufactured article or kit is promoted, distributed, or sold as a unit for performing the methods described herein.
[0116] It should be understood that the implementation schemes described herein include those "consisting of the implementation scheme" and / or "substantially consisting of the implementation scheme".
[0117] In this document, references to “about” values or parameters include (and descriptions) changes with respect to that value or parameter itself. For example, a description of “about X” includes a description of “X”.
[0118] As used herein, references to "NOT" values or parameters generally mean and describe "other than" a value or parameter. For example, "This method is not used to treat type X disease" means that the method is used to treat types of diseases other than X.
[0119] The term “about XY” used in this article has the same meaning as “about X to about Y”.
[0120] As used herein and in the claims of the appended application, the singular forms “a,” “an,” or “the” include the plural references unless the context clearly requires otherwise.
[0121] As used herein, the term "and / or" and phrases such as "A and / or B" are intended to include both A and B; A or B; A (alone); and B (alone). Similarly, as used herein, the term "and / or" and phrases such as "A, B, and / or C" are intended to include each of the following embodiments: A, B, and C; A, B, or C; A or C; A or B; B or C; A and C; A and B; B and C; A (alone); B (alone); and C (alone).
[0122] II. Compositions and Systems
[0123] In some aspects, this document provides an engineered gsnoRNA comprising a guide sequence for hybridization with a sequence containing a target uridine residue in a target RNA of a host cell, wherein the gsnoRNA comprises a nucleotide sequence selected from SEQ ID NO: 4-6, 9-12, 15-19, 22-36, and 177-179, and wherein the gsnoRNA is capable of recruiting the DKC1 protein in the host cell to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the gsnoRNA comprises one or more nucleosides modified with 2'-OMe or 2'-MOE. In some embodiments, the engineered gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides. In some embodiments, the engineered gsnoRNA comprises one or more phosphate thioester nucleoside linkages. In some embodiments, the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphate thioester nucleoside linkages. In some implementations, the gsnoRNA contains a 5' cap modification (e.g., 7-methylguanosine (m... 7G) cap modification). In some embodiments, the 5' cap modification is achieved by using m 7 In vitro transcriptional introduction of G(5′)ppp(5′)G cap analogues.
[0124] In some embodiments, the engineered gsnoRNA is generated via in vitro transcription. In some embodiments, the engineered gsnoRNA generated via in vitro transcription is a full-length gsnoRNA (e.g., containing a 3' hairpin, a 5' hairpin, an H-box, and an ACA box). In some embodiments, the engineered gsnoRNA generated by in vitro transcription contains a 5' cap modification (e.g., 7-methylguanosine (m... 7 G) Hat decoration).
[0125] In some embodiments, the engineered gsnoRNA comprises a single hairpin and an H-box, but not an ACA box. In some embodiments, the engineered gsnoRNA comprises the sequence of SEQ ID NO:179. In some embodiments, the engineered gsnoRNA comprises a single hairpin and an ACA box, but not an H-box. In some embodiments, the engineered gsnoRNA comprises the sequence of SEQ ID NO:180. In some embodiments, the gsnoRNA comprises one or more nucleosides modified with 2'-OMe or 2'-MOE. In some embodiments, the engineered gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides. In some embodiments, the engineered gsnoRNA comprises one or more phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphate-thioester nucleoside linkages.
[0126] In some aspects, this document provides an engineered gsnoRNA comprising a guide sequence that hybridizes to a target uridine residue in a target RNA of a host cell, wherein the gsnoRNA comprises a scaffold sequence derived from wild-type ACA2b or ACA36, and wherein the gsnoRNA is capable of recruiting the DKC1 protein in the host cell to modify the target uridine residue in the target RNA to a pseudouridine residue. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from the sequence of SEQ ID NO: 11 or 12. In some embodiments, the gsnoRNA comprises one, two, three, or four substitutions, deletions, and / or insertions compared to SEQ ID NO: 11 or 12. In some embodiments, the gsnoRNA comprises one or more nucleosides modified with 2'-OMe or 2'-MOE. In some embodiments, the engineered gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides. In some embodiments, the engineered gsnoRNA comprises one or more phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA contains no more than 10, 9, 8, or 6 phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA contains a 5' cap modification (e.g., 7-methylguanosine (m... 7 G) cap modification). In some embodiments, the 5' cap modification is achieved by using m 7 In vitro transcriptional introduction of G(5′)ppp(5′)G cap analogues.
[0127] In some aspects, this document provides an isolated nucleic acid molecule comprising a sequence encoding the gsnoRNA provided herein. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from wild-type H / ACA-snoRNAs selected from: ACA19, ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from: ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA36. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA2b. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA19. In some embodiments, the nucleic acid molecule further comprises a sequence encoding an agent that promotes the expression of isotype 3 of the DKC1 protein (e.g., a splice-conversion antisense oligonucleotide (ASO), wherein the ASO enhances the expression of the DKC1 protein, which is an endogenous DKC1 isotype having cytoplasmic localization in the host cell). In some embodiments, the nucleic acid molecule further comprises a sequence encoding a DKC1 isotype or a variant of the DKC1 protein, wherein the isotype or variant has cytoplasmic localization. An illustrative DKC1 protein is described in Section IIA below.
[0128] In some aspects, this document provides an engineered RNA editing system comprising: (a) a gsnoRNA (such as any of the gsnoRNAs described below in Section II B) containing a guide sequence that hybridizes to a target RNA of a host cell containing a target uridine residue, or a nucleic acid molecule encoding the gsnoRNA; and (b) a DKC1 protein (such as any of the DKC1 proteins described below in Section II A), or a nucleic acid molecule encoding the DKC1 protein, wherein the gsnoRNA is capable of recruiting the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype.
[0129] In some respects, this article provides a host cell that contains any of the gsnoRNA, nucleic acid constructs / molecules, or engineered RNA editing systems described herein.
[0130] In some respects, this article provides a kit for editing target RNA in host cells, comprising any of the gsnoRNA, nucleic acid constructs / molecules, or engineered RNA editing systems described herein.
[0131] A.DKC1 protein
[0132] This application provides, in some embodiments, an engineered DKC1 protein or a nucleic acid construct encoding the DKC1 protein.
[0133] Dyskeratin (DKC1) is a highly conserved multifunctional protein that acts as an RNA-guided pseudouridine synthase, directing the enzymatic conversion of specific uridines into pseudouridines. DKC1 is concentrated in the nucleolus and Cahal body (CB), where it binds to three other highly conserved proteins (Nop10, Nhp2, and Gar1) to form a tetramer. This tetramer can be incorporated into the composition of different nuclear RNPs that perform key biological functions. Within the nucleolus, the tetramer binds to H / ACA small nucleolar RNA (snoRNA) to form H / ACA snoRNPs, which regulate rRNA processing and pseudouridine RNA targets through snoRNA-guided base complementarity. Within the CB, the tetramer binds to CB-specific small RNA (scaRNA) to form scaRNPs, which guide the pseudouridine snoRNA sp- ...
[0134] Two DKC1 isoforms exist in human cells: DKC1 isoform 1 is the typical DKC1 form containing both bisected N and C-terminal nuclear localization signals (NLS); DKC1 isoform 3 is another splicing variant, which is generated through the retention of intron 12 and lacks C-terminal NLS. Figure 9 A). The endogenous mRNA expression level of isotype 1 is approximately 20 times higher than that of isotype 3. 5 Surprisingly, the inventors found that increasing the amount of DKC1 isotype 3 enhances the efficiency of gsnoRNA-guided target pseudouridine editing (e.g., the efficiency of target mRNA editing).
[0135] In some aspects, the compositions of the present invention comprise a nucleic acid construct for expressing the DKC1 protein. In some aspects, the compositions of the present invention comprise the DKC1 protein (e.g., a DKC1 protein complexed with gsnoRNA). In some embodiments, the DKC1 protein is isotype 3 of the mammalian DKC1 protein. In some embodiments, the DKC1 protein is homologous to isotype 3 of the human DKC1 protein. In some embodiments, the DKC1 protein has at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity with isotype 3 of the human DKC1 protein. In some embodiments, the DKC1 protein is isotype 3 of the human DKC1 protein. In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with SEQ ID NO:2. The sequences of full-length DKC1 (isotype 1) and isotype 3DKC1 are shown in Table 1 below.
[0136] In some embodiments, the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to gsnoRNA. In some embodiments, the ribonucleoprotein complex comprises NOP10, GAR1, and / or NHP2.
[0137] In some aspects, this document provides truncated variants of the DKC1 protein and nucleic acid constructs encoding them. In some embodiments, the DKC1 protein comprises a deletion of the nuclear localization signal (NLS) relative to the wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein comprises a deletion of amino acid residues 9-21 of DKC1 isotype 3, wherein the amino acid numbering is based on SEQ ID NO:2. In some embodiments, the DKC1 protein comprises amino acid residues 22-420 of DKC1 isotype 3, wherein the amino acid numbering is based on SEQ ID NO:2. In some embodiments, the DKC1 protein comprises amino acid residues 35-420 of DKC1 isotype 3, wherein the amino acid numbering is based on SEQ ID NO:2. In some embodiments, the DKC1 protein comprises amino acid residues 41-420 of DKC1 isotype 3, wherein the amino acid numbering is based on SEQ ID NO:2. Although the DKC1 sequence in SEQ ID NO:2 is homotype 3 of human DKC1, those skilled in the art will understand how to generate corresponding truncated and deleted variants of the homologous DKC1 protein (e.g., corresponding deleted / truncated variants of DKC1 protein in other mammalian species) based on sequence alignment.
[0138] In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with SEQ ID NO:85. In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with SEQ ID NO:86. In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with SEQ ID NO:87. In some embodiments, the DKC1 protein comprises an amino acid sequence that has at least 85% (e.g., at least 90%, 95%, 96%, 97%, 98%, 99%, or 100%) sequence identity with SEQ ID NO:88.
[0139] Table 1. DKC1 protein sequence
[0140]
[0141]
[0142] In some embodiments, amino acid sequence variants of the DKC1 protein provided herein are considered. For example, it may be desirable to improve the stability and / or other biological properties of DKC1 (e.g., the catalytic domain of DKC1) or its interactions with other proteins in the ribonucleoprotein complex. The structures of DKC1 and other proteins in the ribonucleoprotein complex have been described, for example, in Rashid et al. (Molecular Cell (2006) 21(2):249-260) and Czekay et al. (Front. Microbiol. (2021) 12:654370), the contents of which are incorporated herein by reference in their entirety. Amino acid sequence variants of the DKC1 protein may be prepared by introducing appropriate modifications into the nucleotide sequence encoding the target-binding moiety, or by peptide synthesis. These modifications include, for example, deletion and / or insertion and / or substitution of residues within the amino acid sequence of the target-binding moiety. Any combination of deletions, insertions, and substitutions may be made to achieve the final construct, provided that the final construct possesses the desired properties.
[0143] In some embodiments, DKC1 protein variants with one or more amino acid substitutions are provided. Amino acid substitutions can be introduced into the DKC1 protein, and products can be screened for desired activity.
[0144] Conservative substitutions are shown in Table A below.
[0145] Table A: Conservative Substitution
[0146]
[0147]
[0148] Amino acids can be classified into different categories based on their common side chain characteristics:
[0149] a. Hydrophobic: Leucine, Met, Ala, Val, Leu, Ile;
[0150] b. Neutral hydrophilic; Cys, Ser, Thr, Asn, Gln;
[0151] c. Acids: Asp, Glu;
[0152] d. Alkaline: His, Lys, Arg;
[0153] e. Residues that affect chain direction: Gly, Pro;
[0154] f. Aromatic tribes: Trp, Tyr, Phe.
[0155] Non-conservative replacement requires replacing members of one of these categories with members of another category.
[0156] Fusion proteins are also considered, which comprise fragments of the naturally occurring DKC1 protein or its functional variants and heterologous amino acid sequences (e.g., at the N-terminus, C-terminus, or internal location of the DKC1 fragment).
[0157] B. Nucleic acid constructs and engineered gsnoRNA
[0158] In some aspects, this document provides engineered gsnoRNAs based on H / ACA snoRNAs. In some embodiments, the gsnoRNA contains a single guide sequence. In some embodiments, the gsnoRNA contains two guide sequences. In some embodiments, the engineered gsnoRNA contains more than two (e.g., 3, 4, 5, 6 or more) guide sequences. For example, H / ACA snoRNA contains two hairpins followed by H and ACA box motifs. In some embodiments, both hairpins of the engineered gsnoRNA provided herein contain guide sequences that can target the target pseudouridine esterification site. In other embodiments, only one hairpin of the engineered gsnoRNA contains a guide sequence that can target the target pseudouridine esterification site. Illustrative engineered gsnoRNA sequences are provided in Tables 2 and 3 below.
[0159] In some aspects, the gsnoRNA disclosed herein is a synthetic oligonucleotide that can be synthesized according to methods known in the art. In some embodiments, the gsnoRNA according to the invention is an oligonucleotide (complete RNA). However, in some embodiments, the gsnoRNA of the invention may comprise DNA. In some embodiments, particularly when consisting only of nucleotides or linkages that can be expressed in a biological system, the gsnoRNA may be expressed in situ, for example, from a plasmid or viral vector.
[0160] In some aspects, the gsnoRNA comprises a scaffold sequence derived from wild-type H / ACA-snoRNA selected from ACA2b and ACA36. In some aspects, the editing efficiency of the gsnoRNA derived from the wild-type H / ACA scaffold is at least 5% in mammalian cells (e.g., in human cells such as HEK293T cells) (e.g., between or between about 5% and 15% or 5% and 10%). In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA2b. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA36. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA19.
[0161] In some aspects, this document discloses engineered gsnoRNAs and engineered gsnoRNA scaffolds derived from wild-type H / ACA-snoRNAs (e.g., derived from ACA2b, ACA36, or ACA19), wherein the gsnoRNA is a PTC in the RNA encoding a protein, wherein such modification results in the expression of the full-length protein. In some embodiments, the engineered gsnoRNA is capable of causing the expression of the full-length protein in host cells to be at least 4% (e.g., at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%) of the expression level of the full-length protein without premature stop codons. In some embodiments, the engineered gsnoRNA can induce the expression of the full-length protein, wherein the expression of the protein can be detected without enrichment (e.g., without enrichment by immunoprecipitation). In some embodiments, the protein is detected by tagging (e.g., by fluorescent tagging). In some embodiments, the protein is detected by immunostaining according to methods known in the art. In some implementations, the engineered gsnoRNA can induce expression of the full-length protein in at least 20% of host cells (e.g., at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of host cells).
[0162] In some embodiments, the gsnoRNA includes one or more guide sequences located in regions corresponding to the hairpin structure of the wild-type H / ACA-snoRNA. In some embodiments, the gsnoRNA includes one or more guide sequences located in the hairpin structure of the 3' end portion of the wild-type H / ACA-snoRNA. In some embodiments, the gsnoRNA includes one or more guide sequences located in the hairpin structure of the 5' end portion of the wild-type H / ACA-snoRNA.
[0163] In some embodiments, the gsnoRNA contains one or more mutations (e.g., substitutions, insertions, and / or deletions) in one or more hairpin structures (e.g., 3' and / or 5' hairpin structures) of wild-type ACA19. In some embodiments, the gsnoRNA contains one or more mutations relative to the wild-type scaffold that alter the distance between nucleotide residues in the guide region of the target uridine hybridization and the H / ACA box. In some embodiments, the one or more mutations contain the insertion or deletion of one or more nucleotide residues. In some embodiments, the engineered gsnoRNA contains 14 nucleotides between nucleotide residues in the guide region of the target uridine hybridization and the H / ACA box. In some embodiments, the engineered gsnoRNA contains 15 nucleotides between nucleotide residues in the guide region of the target uridine hybridization and the H / ACA box. In some embodiments, the above mutations increase the efficiency of pseudouridineization (e.g., PTC readthrough efficiency) by at least 1.2, 1.3, 1.4, 1.5, or 1.6 times compared to the wild-type scaffold.
[0164] In some embodiments, one or more mutations comprise substitutions in a small U-series (e.g., a sequence of four or more, or five or more consecutive uridine (U) residues). In some embodiments, the one or more mutations comprise altering the small U-series such that it contains no more than two consecutive U residues. In some embodiments, the one or more mutations comprise a single-base mutation in the “UUUU” sequence. In some embodiments, the mutation is a “UUCU” mutation or a “UGUU” mutation. In some embodiments, the mutated U-series is located in the loop region of the gsnoRNA scaffold. In some embodiments, the engineered gsnoRNA comprises the sequence of SEQ ID NO. 49 or 50. In some embodiments, the engineered gsnoRNA comprises the sequence of SEQ ID NO. 15 or 16. In some embodiments, the above mutations increase the efficiency of pseudouridineization (e.g., the efficiency of PTC readthrough) by at least 1.2, 1.3, 1.4, 1.5, or 1.6 times compared to the wild-type scaffold.
[0165] In some embodiments, one or more mutations comprise mutations that increase the openness of the guide region compared to the wild-type scaffold. In some embodiments, the one or more mutations reduce the base pairing probability of one or more residues within the guide region of the gsnoRNA scaffold (e.g., the 5' guide region of the gACA19 scaffold). In some embodiments, the one or more mutations comprise the insertion of one or more nucleotides. In some embodiments, the one or more mutations comprise the addition of CU after residue 8, wherein the number is according to SEQ ID NO:37. In some embodiments, the engineered gsnoRNA is the gsnoRNA of SEQ ID NO:53. The predicted secondary structure of gACA19-5addCU (SEQ ID NO:53) is shown in... Figure 5D In some implementations, the above mutation increases the efficiency of pseudouridineization (e.g., the efficiency of PTC readthrough) by at least 1.2, 1.3, 1.4, 1.5, or 1.6 times compared to the wild-type scaffold.
[0166] In some embodiments, one or more mutations are selected from the following: replacing residues 26-29 with UUCU, replacing residues 26-29 with UGUU, adding G to the 3' hairpin structure after residue 115, and adding a dinucleotide sequence (XX, e.g., CU) to the 5' hairpin after residue 8, wherein X is a nucleotide selected from A, U, C, and G, and wherein the number is according to SEQ ID NO:37.
[0167] In some implementations, gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO:17-19 and 22-29.
[0168] In some embodiments, the gsnoRNA comprises nucleotide sequences selected from the following: SEQ ID NO:4-6, 9-12, 15-19, 22-36, and 177-179.
[0169] In some implementations, the gsnoRNA contains nucleotide sequences selected from the following: SEQ ID NO:15-19.
[0170] In some embodiments, the gsoRNA comprises nucleotide sequences selected from the following: SEQ ID NO:20-21 and 145-150.
[0171] In some implementations, the gsnoRNA is a disease-targeting gsnoRNA (e.g., any of the gsnoRNA sequences provided in Table 4).
[0172] In some embodiments, the gsnoRNA comprises one or more chemically modified nucleosides and / or nucleoside linkages. In some embodiments, the gsnoRNA comprises one or more nucleosides modified with 2'O-methyl (2'-OMe) or 2'-O-methoxyethyl (2'-MOE). In some embodiments, the gsnoRNA according to the invention may be almost entirely chemically modified, for example by providing nucleotides having a 2'-O-methylated sugar moiety (2'-OMe) and / or having a 2'-O-methoxyethyl sugar moiety (2'-MOE). In some embodiments, the gsnoRNA comprises no more than 20, no more than 15, no more than 10, no more than 8, no more than 6, or no more than 4 2'-OMe or 2'-MOE modifications. In some embodiments, the gsnoRNA comprises about 2 to about 6 2'-OMe or 2'-MOE modifications. In some embodiments, the gsnoRNA comprises about 4 2'-OMe or 2'-MOE modifications. In some embodiments, the gsnoRNA contains no more than five modified sugars. In some embodiments, the gsnoRNA contains two nucleotides containing modified sugar moieties (e.g., 2'-OMe) at its 5' end and two nucleotides containing modified sugar moieties (e.g., 2'-OMe) at its 3' end. In some embodiments, the gsnoRNA contains no more than four, three, or two nucleotides containing modified sugar moieties (e.g., 2'-OMe) at its 5' end and no more than four, three, or two nucleotides containing modified sugar moieties (e.g., 2'-OMe) at its 3' end. In some embodiments, the gsnoRNA contains one or more phosphate-thioester nucleotide linkages. In some embodiments, the gsnoRNA contains no more than 20, no more than 15, no more than 10, no more than 8, or no more than 6 phosphate-thioester linkages. In some embodiments, the gsnoRNA contains about 2 to about 10 phosphate-thioester linkages. In some embodiments, the gsnoRNA contains about six phosphate-thioester linkages. In some embodiments, the gsnoRNA contains about three phosphate-thioester linkages at its 5' end and about three phosphate-thioester linkages at its 3' end. In some embodiments, the gsnoRNA contains no more than five, four, or three phosphate-thioester linkages at its 5' end and no more than five, four, or three phosphate-thioester linkages at its 3' end. The results provided in Example 7 demonstrate that a limited number of modifications are sufficient to ensure the stability and function of the gsnoRNA oligonucleotide.
[0173] In some embodiments, the gsnoRNA comprises one or more chemically modified nucleosides and / or nucleoside linkages. In some embodiments, the gsnoRNA comprises one or more nucleosides modified with 2'O-methyl (2'-OMe) or 2'-O-methoxyethyl (2'-MOE). In some embodiments, the gsnoRNA according to the invention may be almost entirely chemically modified, for example by providing nucleotides having a 2'-O-methylated sugar moiety (2'-OMe) and / or having a 2'-O-methoxyethyl sugar moiety (2'-MOE). In some embodiments, the gsnoRNA comprises a 5' hairpin, an H-box (consistent sequence ANANNA), a 3' hairpin, and an ACA-box (consistent sequence ANA). In some embodiments, the gsnoRNA comprises a single hairpin and an H-box (referred herein to as gH5 or rH5, corresponding to a 5' half gsnoRNA coding sequence or a gsnoRNA oligonucleotide, respectively), and lacks an ACA-box. In some embodiments, the gsnoRNA comprises a single hairpin and an ACA box (referred to herein as gH3 or rH3, corresponding to the 3' half gsnoRNA coding sequence or the gsnoRNA oligonucleotide, respectively), and lacks an H box. In some embodiments, the length of the gsnoRNA containing the single hairpin is between 60 and 70 nucleotides. In some embodiments, the length of the gsnoRNA containing the single hairpin is approximately 65 nucleotides.
[0174] In some embodiments, the gsnoRNA is prepared by in vitro transcription. In some embodiments, the gsnoRNA prepared by in vitro transcription contains the sequence of any one of SEQ ID NO: 4-6, 9-12, 15-19, 22-36. In some embodiments, the gsnoRNA prepared by in vitro transcription contains a 5' cap modification or a 5' hairpin (e.g., of the U6+U27 expression cassette). In some embodiments, the gsnoRNA prepared by in vitro transcription contains a 5' cap modification. In some embodiments, the 5' cap modification is m 7 G modifiers (e.g., cap 0, cap 1, or cap 2 modifiers) or m 6 Am modification. Methods suitable for adding a 5' cap to RNA oligonucleotides have been described (e.g.) in U.S. Patent No. 10,494,399, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the gsnoRNA further comprises a 3' hairpin (e.g., the gsnoRNA comprises the sequence of any one of SEQ ID NO: 4-6, 9-12, 15-19, and 22-36 and a 3' hairpin). In some embodiments, the gsnoRNA comprises a 5' cap modification and does not contain a 3' hairpin (e.g., as shown in the image). Figure 15A (as shown in the image). In some implementations, the 5' cap embellishment uses m 7G(5')ppp(5′)G cap analogues are introduced via in vitro transcription.
[0175] Various chemical compositions and modifications known in the field of oligonucleotides can be readily used according to the present invention. Conventional internucleotide linking bonds between nucleotides can be altered by monothiolation or dithiolation of phosphodiester bonds to produce thiophosphates or dithiophosphates, respectively. Other modifications to the aforementioned internucleotide linking bonds are possible, including amidation and peptide linkers. In a preferred aspect, the gsnoRNA of the present invention has one, two, three, four, five, six, or more thiophosphate linking bonds between most of the terminal nucleotides of the gsnoRNA (and thus preferably both at the 5' and 3' ends), meaning that in the case of three thiophosphate linking bonds, four nucleotides are ultimately linked accordingly. Those skilled in the art will understand that the number of these linking bonds can vary at each end, depending on the target sequence, or based on other aspects such as toxicity. However, some embodiments of this disclosure do not contain one or more PS linking bonds between any positions at the seven nucleotides at their ends of the gsnoRNA.
[0176] The ribose can be modified by replacing the 2'-O portion with a low-carbon alkyl group (C1-4, such as 2'-OMe), alkenyl group (C2-4), alkynyl group (C2-4), methoxyethyl group (2'-methoxyethoxy; or 2'-O-methoxyethyl; or 2'-MOE), or other substituents. In some embodiments, the substituent for the 2'OH group is methyl, methoxyethyl, or 3,3'-dimethylallyl. The latter is known to have the property of inhibiting nuclease sensitivity due to its large size, while improving hybridization efficiency. Alternatively, locked nucleic acid sequences (LNAs) can be applied, which contain 2'-4' intramolecular bridges (typically methylene bridges between the 2' oxygen and 4' carbon) within the ribose ring. Purine and / or pyrimidine nucleobases can be modified to alter their properties, for example, through amination or deamination of the heterocycle. Other modifications that may be present in the gsnoRNA of the present invention are 2'-F modified sugars, BNA, and cEt. The precise chemistry and form can vary depending on the different oligonucleotide constructs and applications, and can be tailored to the wishes and preferences of those skilled in the art.
[0177] Examples of chemical modifications in the gsnoRNA disclosed herein are modifications of the sugar moiety, including cross-linking of substituents within the sugar (ribose) moiety (e.g., in LNA or locked nucleic acids, BNA, cEt, etc.), and substitution of the 2′-O atom with alkyl (e.g., 2'-O-methyl), alkynyl (2'-O-alkynyl), alkenyl (2'-O-alkenyl), or alkoxyalkyl (e.g., 2'-O-methoxyethyl, 2'-MOE) groups having the lengths specified above. In the context of this invention, sugar “modification” also includes 2'-deoxyribose (e.g., in DNA). Additionally, the phosphodiester groups of the backbone can be modified by thiolation, dithiolation, amidation, etc., to generate nucleoside linkages such as thiophosphates, dithiophosphates, and aminophosphates. These nucleoside linkages can be completely or partially replaced by peptide linkages to generate peptide nucleic acid sequences, etc. Alternatively or additionally, nucleobases can be modified by (de)amination to generate inosine or 2'6'-diaminopurine, etc. Another modification could be the methylation of C5 in the cytidine moiety of the nucleotide to reduce the potential immunogenic properties known to be associated with the CpG sequence.
[0178] In some embodiments, the gsnoRNA does not contain one or more chemically modified nucleosides and / or inter-nucleoside linkages. In some embodiments, the gsnoRNA does not contain any non-natural inter-nucleoside linkages.
[0179] Mammalian H / ACAs noRNAs are typically embedded in the intron region of the pre-mRNA of protein-coding genes. During transcriptional elongation, several proteins that function in pseudouridineization (such as NOP10, dyskeratin (DKC1), or NHP2) bind to the nascent H / ACAs noRNA sequence. After splicing, the guide RNA is processed by debranching and exonuclease to obtain an RNA-protein complex (snRNP or snRNP complex) called a "small nucleoribonucleoprotein." Box H / ACA snoRNAs have no localization preference relative to the 5' or 3' end of the intron and can be present in small or very large introns, unlike box C / D snoRNAs, which are generally located 60 to 90 nucleotides upstream of the 3' splice site and encode in relatively small introns. Kiss and Filipowicz (1995, Genes Dev 9(11):1411-1424) proposed that a given snoRNA sequence could be fully processed by excising the intron region of any given actively spliced mRNA. To demonstrate the feasibility of this snoRNA processing independent of the host intron environment, Kiss and Filipowicz artificially inserted several snoRNAs (III 7a, U17b, and U19) into the second intron of the human β-globulin gene and expressed the resulting vector in fibroblast-like cells. After transfection, they found that the artificial, intron-derived snoRNA was accurately spliced from the human β-globulin intron with appropriate processing and β-globulin pre-mRNA. Darzacq et al. (2002, EMBO J21(11); 2746-2756) demonstrated that other guide RNAs can be inserted into the second intron of the human β-globulin gene using an expression vector under the control of the cytomegalovirus (CMV) promoter and delivered to mammalian cells via transfection.
[0180] The inventors of this application unexpectedly identified different host intron environment-dependent effects on the pseudouridine editing efficiency of different gsnoRNAs (as discussed in Example 1). For example, the inventors tested host genes based on wild-type ACA19 (embedded in the host intron of EIF3A), ACA-44 (embedded in the host intron of SNHG12), ACA27 (embedded in the host intron of RPL21), and E2 (embedded in the host intron of RPSA), and non-host introns embedded in the HBB gene (…). Figure 2A and 2BThe inventors investigated the PTC readthrough efficiency of gsnoRNAs. Surprisingly, they found that gsnoRNAs based on the E2 scaffold had lower editing efficiency when embedded in the HBB intron compared to the host RPSA intron, while gACA19 had similar editing efficiency when embedded in the HBB intron compared to the host EIF3A intron. Based on this observation that host gene sequences have different effects on different gsnoRNAs, the inventors hypothesized that direct expression of the gsnoRNA without host gene influence could further increase PTC readthrough efficiency. Therefore, the inventors designed a series of gsnoRNA expression constructs in which the nucleic acid molecule encoding the gsnoRNA was not embedded in an intron. As discussed in Example 1, the inventors demonstrated enhanced pseudouridine esterification activity of gsnoRNAs not embedded in introns, where the nucleic acid molecule encoding the gsnoRNA was driven by the hU6 (type III RNA polymerase III promoter) and hU1 (snRNA-type RNA polymerase II promoter) promoters. Therefore, in one aspect, this document provides a nucleic acid molecule encoding gsnoRNA, wherein the nucleic acid molecule is under the control of a small RNA promoter (e.g., a U6 or U1 promoter). In some embodiments, the nucleic acid encoding the gsnoRNA is not embedded in an intron sequence.
[0181] In some aspects, this document provides a nucleic acid construct encoding gsnoRNA. In some embodiments of the methods described herein, the method includes introducing a nucleic acid molecule encoding the gsnoRNA into a host cell. In some embodiments, the nucleic acid molecule encoding the gsnoRNA is under the control of a small RNA promoter. In some embodiments, the small RNA promoter is a U6 (transcribed by polymerase III) or U1 (transcribed by polymerase II) promoter. In some embodiments, expression of the gsnoRNA from the small RNA promoter according to the methods disclosed herein provides increased pseudouridineization efficiency (e.g., increased PTC readthrough efficiency) compared to the same gsnoRNA embedded in a host intron sequence or other intron sequences. In some embodiments, the pseudouridineization efficiency of gsnoRNA expressed under the control of a nucleic acid small RNA promoter is 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 times higher than that of the same gsnoRNA embedded in a host intron. Example 1 ( Figure 1A-1E and Figure 2A-2C The results provided confirm that gsnoRNA expressed from nucleic acids under the control of small RNA promoters enhances PTC readthrough compared to gsnoRNA embedded in introns.
[0182] In some embodiments, the nucleic acid molecule encoding gsnoRNA is embedded in an intron sequence located between the first exon sequence and the second exon sequence. In some embodiments, the first exon sequence, the intron sequence, and the second exon sequence are derived from a naturally occurring gene. In some embodiments, the intron may contain additional nucleotides (in addition to the nucleic acid molecule of the present invention, including a guide region). Since the guide region is expressed from the intron sequence, these additional nucleotides can be selected to most efficiently express from the intron. In some embodiments, the exon A / intron / exon B sequence is present in a vector (such as a plasmid or viral vector). This vector can be used to deliver the exon-intron-exon sequence to cells. Additional introns and exons may be present in this vector. In some embodiments, the exon A sequence (upstream of the intron carrying the nucleic acid encoding the gsnoRNA (which is expressed post-transcriptionally)) contains or is composed of exon 1 of the human β-globulin gene, and the exon B sequence (downstream of the intron carrying the nucleic acid encoding the gsnoRNA (which is expressed post-transcriptionally)) contains or is composed of exon 2 of the human β-globulin gene. In some embodiments, the exon A sequence (upstream of the intron carrying the nucleic acid encoding the gsnoRNA (which is expressed post-transcriptionally)) contains or is composed of exon 2 of the human hemoglobin subunit β (HBB) gene, and the exon B sequence (downstream of the intron carrying the nucleic acid encoding the gsnoRNA (which is expressed post-transcriptionally)) contains or is composed of exon 3 of the human hemoglobin subunit β (HBB) gene. In some embodiments, the nucleic acid molecule encoding the gsnoRNA is embedded in an intron sequence between a first exon sequence and a second exon sequence, wherein the intron sequence, the first exon sequence, and the second exon sequence correspond to the sequence of a host gene carrying the naturally occurring snoRNA. In some embodiments, the construct containing the gsnoRNA coding sequence embedded in the intron is under the control of a CMV promoter.
[0183] In some aspects, this document provides engineered gsnoRNAs targeting disease-associated PTC. In some embodiments, the aforementioned engineered gsnoRNAs targeting disease-associated PTC contain one or more mutations to enhance gsnoRNA editing efficiency and / or expression. In some embodiments, the engineered gsnoRNAs targeting disease-associated PTC are selected from SEQ ID NO:71-84 (shown in Figures 14-15). The sequences of exemplary engineered gsnoRNAs targeting disease-associated PTC are shown in Table 4 below.
[0184] In some embodiments, the gsnoRNA can be administered in a free form (or “naked,” carrier-free environment), or by other means (such as liposomes or nanoparticles), or delivered to cells using iontophoresis. In some embodiments, the gsnoRNA can be administered in the form of a ribonucleoprotein complex (e.g., a complex comprising DKC1, HNP2, NOP10, and / or GAR1). In some embodiments, the free gsnoRNA contains one or more chemically modified nucleosides and / or inter-nucleoside linkages as described above.
[0185] In some aspects, this document provides a nucleic acid construct encoding DKC1 (e.g., any of the DKC1 proteins described in Section IIA above). In some embodiments of the methods described herein, the method includes introducing a nucleic acid molecule encoding the DKC1 protein into a host cell. In some embodiments, the nucleic acid molecule comprises a promoter operatively linked to a nucleotide sequence encoding the DKC1. In some embodiments, the promoter is a Pol II promoter. In some embodiments, the promoter is a CMV promoter.
[0186] As disclosed herein, the vector may carry DNA or RNA and is generally used to express the gsnoRNA and / or DKC1 protein constructs of the present invention after the vector has been treated in a cell in which it has been introduced. This is generally achieved through transcription of the DNA or RNA present in the vector. In some embodiments, the vector is a viral vector (which can be used to infect the target cells to be treated) or a plasmid, which may be introduced into the cell in a variety of ways known to those skilled in the art.
[0187] In some embodiments, the nucleic acid molecule encoding the DKC1 protein and / or the nucleic acid molecule encoding gsnoRNA are present in a viral vector. In some embodiments, the method includes introducing a vector (e.g., a plasmid or viral vector) containing a first nucleic acid sequence encoding the DKC1 protein and a second nucleic acid sequence encoding the gsnoRNA into a host cell. In some embodiments, the vector is an adeno-associated virus (AAV) vector.
[0188] Exemplary engineered ACA scaffold sequences are shown in Table 2 below. The guide sequence is shown as (Xn) and underlined, where Xn is a sequence of X nucleotides of length n, where X is any of A, U, G, or C, and n is 4, 5, 6, 7, 8, 9, 10, 11, or 12. As those skilled in the art will appreciate, this guide sequence (Xn) can be modified to target gsnoRNA to a desired target site. In some embodiments, n is an integer representing the appropriate length of the guide region. In some embodiments, n is 4, 5, 6, 7, 8, or 9.
[0189] Exemplary engineered gsnoRNA sequences (including exemplary guide sequences) are shown in Table 3 below.
[0190] The exemplary engineered gsnoRNA sequences targeting PTC associated with the exemplary disease are shown in Table 4 below.
[0191] Table 2. ACA stent sequence.
[0192]
[0193]
[0194]
[0195]
[0196] Table 3. Illustrative gsnoRNA constructs (guide sequences are underlined)
[0197]
[0198]
[0199]
[0200]
[0201]
[0202] Table 4: gsnoRNAs targeting PTC associated with the disease (guide sequences are underlined)
[0203]
[0204]
[0205]
[0206] In one aspect, the inventors discovered that the editing efficiency of gsnoRNA was unexpectedly higher when it was tandemly encoded with its target RNA. The results provided in Example 3 confirm that using a reporter gene construct that tandemly encodes both gsnoRNA and target RNA increases editing efficiency.
[0207] Therefore, in some aspects, this document provides nucleic acid molecules comprising a nucleotide sequence encoding a guide small nucleolar RNA (gsnoRNA) tandem with a nucleotide sequence encoding a target RNA. In some embodiments, the nucleotide sequence encoding the gsnoRNA is driven by a U6 or U1 promoter. In some embodiments, the nucleotide sequence encoding the target RNA is driven by the same or different promoters. In some embodiments, the same gsnoRNA tandemly encoding the nucleotide sequence encoding the target RNA provides at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 times greater target RNA editing efficiency than gsnoRNA encoded in a nucleic acid molecule different from the target RNA.
[0208] III. Methods
[0209] In some embodiments, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered guide small nucleolar RNA (gsnoRNA) and a nucleic acid molecule encoding a DKC1 protein into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence in the target RNA containing a target uridine residue, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue. In some embodiments, the DKC1 protein is a DKC1 isotype with cytoplasmic localization (e.g., isotype 3). In some embodiments, the DKC1 protein is a portion of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA. In some embodiments, the DKC1 protein contains a deletion of the nuclear localization signal (NLS) relative to a wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein is a truncated DKC1 variant or a DKC1 variant containing deletions, such as any of the truncated or deleted variants described in Section IIA above. In some implementations, the gsnoRNA recruits NOP10, GAR1, and NHP2 in the host cell.
[0210] In some embodiments, the method includes introducing a nucleic acid (e.g., a nucleic acid vector) encoding gsnoRNA into a cell. In other embodiments, the method includes introducing a gsnoRNA oligonucleotide into the cell. In some embodiments, the gsnoRNA comprises a first hairpin and an H-box and a second hairpin and an ACA-box. In some embodiments, the gsnoRNA is prepared by in vitro transcription. In some embodiments, the gsnoRNA prepared by in vitro transcription comprises a sequence of any one of SEQ ID NO: 4-6, 9-12, 15-19, and 22-36. In some embodiments, the gsnoRNA prepared by in vitro transcription comprises a 5' cap modification (e.g., of a U6+U27 expression cassette) or a 5' hairpin. In some embodiments, the gsnoRNA prepared by in vitro transcription comprises a 5' cap modification. In some embodiments, the 5' cap modification is m 7 G modifiers (e.g., cap 0, cap 1, or cap 2 modifiers) or m 6 Am modification. Methods suitable for adding a 5' cap to RNA oligonucleotides have been described (e.g.) in U.S. Patent No. 10,494,399, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the gsnoRNA further comprises a 3' hairpin (e.g., the gsnoRNA comprises the sequence of any one of SEQ ID NO: 4-6, 9-12, 15-19, and 22-36 and a 3' hairpin). In some embodiments, the gsnoRNA comprises a 5' cap modification and does not contain a 3' hairpin (e.g., as shown in the image). Figure 15A (as shown in the image). In some embodiments, in vitro transcribed gsnoRNA can guide targeted pseudouridine esterification in this cell. In some embodiments, this 5' cap modification uses m 7 G(5′)ppp(5′)G cap analogues are introduced via in vitro transcription.
[0211] In some embodiments, the method includes introducing a nucleic acid (e.g., a nucleic acid vector) encoding a hemigne gsnoRNA (e.g., containing a single hairpin and H-box, or containing a single hairpin and ACA-box) into a cell. In other embodiments, the method includes introducing a gsnoRNA that is a hemigne gsnoRNA (e.g., containing a single hairpin and H-box, or containing a single hairpin and ACA-box) into a cell. In some embodiments, the gsnoRNA contains no more than 20, no more than 15, no more than 10, no more than 8, no more than 6, or no more than 4 2'-OMe or 2'-MOE modifications. In some embodiments, the gsnoRNA contains about 2 to about 6 2'-OMe or 2'-MOE modifications. In some embodiments, the gsnoRNA contains about 4 2'-OMe or 2'-MOE modifications. In some embodiments, the gsnoRNA contains no more than 5 modified sugars. In some embodiments, the gsnoRNA comprises two nucleotides containing a modified sugar moiety (e.g., 2'-OMe) at its 5' end and two nucleotides containing a modified sugar moiety (e.g., 2'-OMe) at its 3' end. In some embodiments, the gsnoRNA comprises no more than four, three, or two nucleotides containing a modified sugar moiety (e.g., 2'-OMe) at its 5' end and no more than four, three, or two nucleotides containing a modified sugar moiety (e.g., 2'-OMe) at its 3' end. In some embodiments, the gsnoRNA comprises one or more phosphate-thioester nucleoside linkages. In some embodiments, the gsnoRNA comprises no more than 20, no more than 15, no more than 10, no more than 8, or no more than 6 phosphate-thioester linkages. In some embodiments, the gsnoRNA comprises about 2 to about 10 phosphate-thioester linkages. In some embodiments, the gsnoRNA comprises about 6 phosphate-thioester linkages. In some embodiments, the gsnoRNA contains about three phosphate-thioester linkages at its 5' end and about three phosphate-thioester linkages at its 3' end. In some embodiments, the gsnoRNA contains no more than five, four, or three phosphate-thioester linkages at its 5' end and no more than five, four, or three phosphate-thioester linkages at its 3' end. In some embodiments, the gsnoRNA contains a 5' hairpin, an H-box (consistent sequence ANANNA), a 3' hairpin, and an ACA-box (consistent sequence ANA). In some embodiments, the gsnoRNA contains a single hairpin and an H-box (referred to herein as gH5 or rH5, corresponding to the 5' half gsnoRNA coding sequence or the gsnoRNA oligonucleotide, respectively), and lacks the ACA-box.In some embodiments, the gsnoRNA comprises a single hairpin and an ACA box (referred to herein as gH3 or rH3, corresponding to the 3' half gsnoRNA coding sequence or the gsnoRNA oligonucleotide, respectively), and lacks an H box. In some embodiments, the length of the gsnoRNA containing the single hairpin is between 60 and 70 nucleotides. In some embodiments, the length of the gsnoRNA containing the single hairpin is approximately 65 nucleotides.
[0212] This disclosure exemplifies the effects of reversing (but not limited to) meaningless termination mutations that typically lead to translation termination and mRNA degradation (through meaningless-mediated decay, see below). In another aspect, targeting pseudouridine can be used as a way to re-encode codons containing uridine, as a way to regulate protein function through amino acid substitution, for example, in key protein regions such as the active site of protein kinases.
[0213] One consequence of mutations leading to PTCs in the coding sequence of a gene is a decrease in mRNA content. This is due to a mechanism called nonsense-mediated decay (NMD), a cellular surveillance mechanism that degrades abnormal mRNA transcripts and prevents the translation of improperly processed transcripts. It is estimated that one-third of inherited diseases result from mutations leading to PTCs, such as those causing CF, retinitis pigmentosa (RP), and β-thalassemia. Under normal circumstances, exon jointer complexes (EJCs) are formed during splicing. These EJCs are then replaced by ribosomes during the first round of translation. On the other hand, when a PTC is located more than 50-54 nucleotides upstream of the last EJC, the NMD pathway is triggered by the formation of a termination complex composed of EJC-associated NMD factors. This situation This occurs during the precursor phase of the first round of translation, when ribosomes coexist with at least one EJC downstream of their location. This triggers decapping and 5'-to-3' exonuclease activity, as well as tail deadenylation and 3'-to-5' exonuclease-mediated transcriptolysis. To address the aforementioned genetic conditions, or any conditions due to similar mutations, it is crucial to inhibit this pathway in a gene-specific and sequence-specific manner.
[0214] In some aspects, this document provides methods for recoding PTCs, which result in an increase in mRNA content and translation readthrough of the recoded mRNA into a full-length protein. In some embodiments, the methods and compositions provided herein allow PTC readthroughs of more than 4%, more than 5%, more than 10%, more than 12%, more than 15%, more than 20%, or more than 30%. In some embodiments, the methods and compositions provided herein allow suppression of nonsense-mediated decay (NMD) of more than 10%, more than 12%, more than 15%, more than 20%, or more than 30%. PTC readthroughs can be analyzed by assessing protein content, by directly quantifying protein expression, or by analyzing the activity of the expressed protein. Methods for assessing NMD suppression are also known in the art. For example, to assess NMD suppression, known NMD suppression reporter gene analysis (Zhang et al. 1998, RNA4(7):80l-8l5) and translation readthroughs of genes carrying PTCs can be used. As illustrated herein, fluorescent reporter genes carrying nonsense mutations are used as target sequences. Without correction, this nonsense mutation results in reduced mRNA abundance (as a consequence of NMD) and truncated protein, leading to a lack of fluorescent signal. As shown herein, correction of this mutation by targeting pseudouridine esterification allows for the translation of the full-length protein from this mRNA. Those skilled in the art will understand that the PTC region of the fluorescent reporter construct described herein can be exchanged by any other model or targeted therapy-related target RNA.
[0215] In some embodiments, this document provides a method for recoding a PTC in RNA encoding a protein, wherein the method results in the expression of the full-length protein in host cells at at least 4% (e.g., at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, or at least 10%) of the expression level of the full-length protein without premature stop codons. In some embodiments, the method results in the expression of the full-length protein, wherein the expression of the protein can be detected without enrichment (e.g., without enrichment by immunoprecipitation). In some embodiments, the protein is detected by tagging (e.g., by a fluorescent tag). In some embodiments, the protein is detected by immunostaining according to methods known in the art. In some embodiments, the method results in the expression of the full-length protein in at least 20% of host cells (e.g., at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, or at least 50% of host cells).
[0216] In some aspects, this document provides methods for treating, preventing, and / or blocking meaningless-mediated RNA decay of target mRNA, the methods comprising introducing a guide small nucleolar RNA (gsnoRNA) and a nucleic acid molecule encoding a DKC1 protein into a host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a premature stop codon (PTC) sequence containing a target uridine residue in the target mRNA, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue, thereby pseudouridineization of the target uridine facilitating readthrough of the PTC. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype. In some embodiments, the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA. In some embodiments, the DKC1 protein contains a nuclear localization signal (NLS) absent compared to a wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein is a truncated DKC1 variant or a DKC1 variant containing deletions, such as either the truncated or deleted variants described in Section IIA above.
[0217] In some embodiments, the DKC1 protein is an endogenous protein of the host cell. In some embodiments, the DKC1 protein is an endogenous, naturally expressed DKC1 isotype of the host cell, wherein the DKC1 isotype is cytoplasmally localized within the host cell. In some embodiments, the DKC1 protein corresponds to isotype 2 of the human DKC1 protein.
[0218] In some embodiments, DKC1 and snoRNA can be delivered intracellularly together (e.g., as part of a ribonucleoprotein (RNP) complex). In some embodiments, the snoRNP comprises gsnoRNA and DKC1, NHP2, GAR1, and / or NOP10.
[0219] In some embodiments, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA (e.g., mRNA), wherein the host cell expresses a cytoplasm-localized DKC1 isotype, and wherein the gsnoRNA recruits the DKC1 isotype to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the method includes introducing a splice-conversion antisense oligonucleotide (ASO) into the host cell, wherein the ASO enhances the expression of a DKC1 protein, the DKC1 protein being an endogenous DKC1 isotype with cytoplasm localization in the host cell.
[0220] Splice-converting antisense oligonucleotides (ASOs) alter splicing by guiding splice site selection. Splice-converting ASOs can regulate pre-mRNA splicing by binding to target pre-mRNA and preventing the splicing machinery from entering specific splice sites, and can be used to generate novel splice variants, correct aberrant splicing, or manipulate alternative splicing. Methods for designing and delivering splice-converting antisense oligonucleotides to cells have been described, for example, in U.S. Patent Publications US20180334677 and US20120040917, U.S. Patent Nos. 10,190,117, and Disterer et al., Hum Gene Ther. 2014 Jul; 25(7):587-98, the contents of which are incorporated herein by reference in their entirety.
[0221] In some embodiments, the splice-converting ASO binds to the pre-mRNA of the DKC1 gene and directs the splicing of DKC1 isotype 3. In some embodiments, the introduction of this splice-converting ASO increases the expression of DKC1 protein (which is the endogenous DKC1 isotype with cytoplasmic localization in the host cell) by at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, or 10-fold compared to the expression of the same isotype in the absence of ASO in the host cell. In some embodiments, the application of this splice-converting ASO increases the expression of DKC1 isotype 3 by at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, or 10-fold compared to the expression of DKC1 isotype 3 in the absence of ASO in the host cell.
[0222] In some implementations, splice-switching ASOs can be delivered via aptamers, reverse molecular sentinel nanoprobes, ASO-encapsulated liposome-DNA-polycations, or ASO-encapsulated liposome-protamine-hyaluronic acid nanoparticles. For methods suitable for aptamer delivery, see Kotula, JW et al., Aptamer-mediated delivery of splice-switching oligonucleotides to the nuclei of cancer cells. Nucleic Acid Ther, 2012, 22(3): 187–95, the contents of which are incorporated herein by reference in their entirety.
[0223] In some embodiments, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a scaffold sequence derived from wild-type ACA19, ACA44, ACA27, E2, ACA3, ACA17, ACA2b, or ACA36, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA (e.g., mRNA) to a pseudouridine residue. In some embodiments, the gsnoRNA contains a scaffold sequence derived from wild-type ACA2b or ACA36. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype. In some embodiments, the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA.
[0224] In some embodiments, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a scaffold sequence derived from wild-type ACA19, ACA44, ACA27, E2, ACA3, ACA17, ACA2b, or ACA36, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA (e.g., mRNA) to a pseudouridine residue. The engineered gsnoRNA may be any of the engineered gsnoRNAs described in Section II B. In some embodiments, the gsnoRNA contains a nucleotide sequence selected from SEQ ID NO: 4-6, 9-12, 15-19, 22-36, and 177-179. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype. In some embodiments, the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA. In some embodiments, the DKC1 protein contains a nuclear localization signal (NLS) deletion relative to the wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein is a truncated DKC1 variant or a DKC1 variant containing deletions, such as any of the truncated or deleted variants described in Section IIA above.
[0225] In some embodiments, this document provides a method for editing target RNA in a host cell, comprising introducing an engineered gsnoRNA into the host cell, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a nucleotide sequence selected from the group consisting of SEQ ID NO: 4-6, 9-12, 15-19, 22-36, and 177-179, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA to a pseudouridine residue. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype. In some embodiments, the DKC1 protein comprises a DKC1 protein fragment corresponding to amino acid residues 1 to 419 of the full-length human DKC1 protein, wherein the amino acid numbering is according to SEQ ID NO: 1. In some embodiments, the DKC1 protein is a portion of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA. In some embodiments, the DKC1 protein contains a missing nuclear localization signal (NLS) relative to the wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein is a truncated DKC1 variant or a DKC1 variant containing a deletion, such as either the truncated or deleted variants described in Section IIA above.
[0226] In some embodiments, the methods provided herein include introducing a nucleic acid molecule containing a nucleotide sequence encoding a guide small nucleolar RNA (gsnoRNA) tandem with a nucleotide sequence encoding a target RNA into a host cell. In some embodiments, the nucleotide sequence encoding the gsnoRNA is driven by a U6 or U1 promoter. In some embodiments, the nucleotide sequence encoding the target RNA is driven by the same or different promoters. In some embodiments, the same gsnoRNA tandemly encoding the nucleotide sequence encoding the target RNA provides at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 times greater target RNA editing efficiency than gsnoRNA encoded in a nucleic acid molecule different from the target RNA.
[0227] In some embodiments, the methods provided herein include introducing a guide small nucleolar RNA (gsnoRNA) into an endogenous nucleic acid molecule of a host cell, wherein the endogenous nucleic acid molecule contains a nucleotide sequence encoding a target RNA. In some embodiments, the introduction includes inserting the nucleotide sequence encoding the gsnoRNA into a region of the endogenous nucleic acid molecule that is directly or indirectly adjacent to the region encoding the target RNA. In some embodiments, the nucleotide sequence encoding the gsnoRNA is driven by a U6 or U1 promoter. Methods known in the art for inserting nucleotide sequences into endogenous nucleic acid molecules, such as guided nuclease (e.g., CRISPR / Cas) editing and homology-directed repair, are employed. In some embodiments, inserting the same gsnoRNA into a region of an endogenous nucleic acid molecule directly or indirectly adjacent to the nucleotide sequence encoding the target RNA provides at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, or 2 times greater target RNA editing efficiency than inserting the same gsnoRNA into a nucleic acid molecule encoded in a different nucleic acid molecule than inserting the same gsnoRNA into a region directly or indirectly adjacent to the nucleotide sequence encoding the target RNA.
[0228] In some implementations, the extent to which pseudouridine entities residing in cells are recruited and redirected can be modulated by the administration and dosing regimen of gsnoRNA. This is determined by the experimenter (e.g., in vitro) or clinician, typically in phase I and / or phase II clinical trials.
[0229] In some embodiments, the methods provided herein include modifying target RNA (e.g., mRNA) sequences in eukaryotic cells (e.g., metazoan or mammalian cells, such as human cells). In some aspects, the methods and compositions provided herein can be used with cells from any organ (e.g., skin, lung, heart, kidney, liver, pancreas, intestine, muscle, gland, eye, brain, blood, etc.). The cells can be located in vitro or in vivo. An advantage of the methods, compositions, systems, kits, and articles of the present invention is that they can be used with cells in situ in a living organism, but also with cells in culture. In some embodiments, cells are treated in vitro and then introduced into a living organism (e.g., reintroduced into the organism of their original source). The methods, compositions, systems, kits, and articles of the present invention can also be used to edit target RNA sequences in cells within so-called organoids. Organoids can be considered as three-dimensional in vitro-derived tissues, but driven by specific conditions to produce individual, isolated tissues (see, for example, Lancaster and Knoblich. 2014, Science 345(6194):1247125). In a therapeutic setting, these organoids are useful because they can be derived in vitro from the patient's cells and then reintroduced as autologous material, with a lower likelihood of rejection compared to normal grafts. The cells to be treated will generally have a genetic mutation. This mutation can be heterozygous or homozygous. In some embodiments, the methods and compositions provided herein can be used to modify point mutations. In some embodiments, the methods and compositions provided herein are suitable for modifying sequences in cells, tissues, or organs associated with a disease state in a subject (e.g., a human subject), such as when the human subject suffers from a disease associated with PTC.
[0230] This disclosure provides a method for making alterations (pseudouridineization) in a target RNA sequence in eukaryotic cells using an oligonucleotide (e.g., any of the gsnoRNAs described in Section II B above, or any gsnoRNA based on an engineered scaffold described in Section II B above), the oligonucleotide targeting the site to be edited and recruiting an RNA editing protein (e.g., DKC1) to elicit an editing response (one or more). In some embodiments, the DKC1 is endogenous. In some embodiments, the DKC1 is delivered exogenously. In some embodiments, the method includes increasing the relative proportion of DKC1 isotype 3 or a cytoplasm-localized DKC1 protein. The target RNA sequence may contain mutations that a person skilled in the art would wish to correct or alter, such as point mutations (transitions or translocations). The target RNA may be any cellular or viral RNA sequence, but is more typically pre-mRNA or protein-coding mRNA. In some embodiments, the target sequence is endogenous to eukaryotic (e.g., mammalian, e.g., human) cells.
[0231] In some embodiments, the methods provided herein are suitable for facilitating the readthrough of a PTC, wherein the PTC is an opal codon (UGA), an amber codon (UAG), or an ochre codon (UAA). In some embodiments, the PTC is an opal codon, and the method results in a readthrough efficiency of at least 10%, at least 15%, at least 20%, or at least 25%, wherein the readthrough efficiency is analyzed as a percentage of protein expression or activity (e.g., fluorescence intensity) compared to a control lacking the PTC. In some embodiments, the PTC is an amber codon (UAG), and the method results in a readthrough efficiency of at least 2%, at least 5%, at least 10%, at least 12%, or at least 14%, wherein the readthrough efficiency is analyzed as a percentage of protein expression or activity (e.g., fluorescence intensity) compared to a control lacking the PTC. In some embodiments, the method results in cellular expression of at least 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, or 70% of the full-length protein encoded by the target gene containing the PTC.
[0232] In some implementations, the target uridine in the target RNA is an early stop codon in the sequence encoding the protein, wherein the method results in the expression of the full-length protein in the host cell at at least 4% (e.g., at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10% or higher) of the full-length protein without the early stop codon.
[0233] In some embodiments, the target uridine in the target RNA is an early stop codon in the sequence encoding a protein, wherein the method results in the expression of the full-length protein, and wherein the expression of the protein can be detected without enrichment (e.g., without enrichment by immunoprecipitation). In some embodiments, the protein is detected by tagging (e.g., by a fluorescent tag). In some embodiments, the protein is detected by immunostaining according to methods known in the art.
[0234] In some implementations, the target uridine in the target RNA is an early stop codon in the sequence encoding the protein, wherein the method results in the expression of the full-length protein in at least 20% of host cells, for example, at least 25%, at least 30%, at least 35%, at least 40%, at least 45%, at least 50% or higher percentage of host cells.
[0235] This invention also provides engineered gsnoRNA compositions or engineered RNA editing systems described herein for use in any of the methods described herein (such as methods for editing target RNA or therapeutic methods). Any of the engineered gsnoRNA compositions or engineered RNA editing systems described herein can be used to prepare medicaments for treating diseases or conditions.
[0236] A. Treatment methods
[0237] In some aspects, the methods provided herein include modifying a target RNA with gsnoRNA that recruits the DKC1 protein to modify the target RNA. In some embodiments, the gsnoRNA hybridizes with a target sequence containing target uridine residues, and the modification of the RNA involves modifying the target uridine into a pseudouridine.
[0238] In some embodiments, the target RNA is endogenous RNA of the cell (e.g., eukaryotic cells, such as mammalian or human cells). In some embodiments, the target RNA is endogenously transcribed RNA of the cell (e.g., transcribed from an endogenous nucleic acid sequence of the cell). In some embodiments, the target RNA is transcribed from a nucleic acid sequence introduced into the cell (e.g., RNA transcribed from an exogenously added nucleic acid molecule). In some embodiments, the target RNA is ribosomal RNA. In some embodiments, the target RNA is messenger RNA (mRNA).
[0239] In some embodiments, the sequence containing the target uridine in the target RNA is a stop codon, and modifying the target uridine to pseudouridine causes the stop codon to be translated into a coding codon. In some embodiments, the stop codon is a premature stop codon (PTC). In some embodiments, the PTC is associated with a genetic disease or condition. By using the manner and method of this disclosure, the target uridine in this PTC is converted to pseudouridine, which then leads to proper reading through the reading frame during translation, thereby providing a (partial or complete) functional full-length protein.
[0240] In some embodiments, this document provides a method for treating a disease or condition in a subject that is associated with a PTC in a target RNA, comprising editing the target RNA in the subject's cells using any of the RNA editing methods described herein, wherein the gsnoRNA contains a guide sequence that hybridizes to a PTC in the target RNA, and wherein modifying a uridine residue in the PTC to a pseudouridine residue causes a read-through translation of the PTC in the target RNA, thereby treating the subject's disease or condition.
[0241] In some embodiments, a method of treating a disease or condition associated with a PTC in a target RNA in a subject includes introducing an engineered gsnoRNA into the subject's host cells, wherein the gsnoRNA contains a guide sequence for hybridizing with a PTC containing a uridine residue in the target RNA, wherein the gsnoRNA contains a scaffold sequence derived from wild-type ACA2b, ACA36, ACA44, ACA27, E2, ACA3, or ACA17, and wherein the gsnoRNA recruits a DKC1 protein in the host cells to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the DKC1 protein is an endogenous DKC1 protein of the host cells. In some embodiments, the method further includes introducing a nucleic acid encoding the DKC1 protein into the host cells. In some embodiments, the DKC1 protein has cytoplasmic localization in the host cells. In some embodiments, the DKC1 protein comprises a DKC1 protein fragment of amino acid residues 41 to 420 corresponding to human DKC1 isotype 3 protein, wherein the amino acid numbering is according to SEQ ID NO:2. In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater) identity with SEQ ID NO:88. In some embodiments, the DKC1 protein comprises the amino acid sequence of SEQ ID NO:88.
[0242] In some embodiments, a method of treating a disease or condition associated with a PTC of a target RNA in a subject includes introducing an engineered gsnoRNA into the subject's host cells, wherein the gsnoRNA contains a guide sequence for hybridizing with a PTC containing a target uridine residue in the target RNA, wherein the gsnoRNA contains a nucleotide sequence selected from SEQ ID NO: 4-6, 9-12, 15-19, 22-36, and 177-179, and wherein the gsnoRNA recruits a DKC1 protein in the host cells to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the DKC1 protein is an endogenous DKC1 protein of the host cells. In some embodiments, the method further includes introducing a nucleic acid encoding the DKC1 protein into the host cells. In some embodiments, the DKC1 protein has cytoplasmic localization in the host cells. In some embodiments, the DKC1 protein comprises a DKC1 protein fragment of amino acid residues 41 to 420 corresponding to human DKC1 isotype 3 protein, wherein the amino acid numbering is according to SEQ ID NO:2. In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85% (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or greater) identity with SEQ ID NO:88. In some embodiments, the DKC1 protein comprises the amino acid sequence of SEQ ID NO:88.
[0243] In some embodiments, the gsnoRNA is a hemi-gsnoRNA, for example, comprising a single hairpin and H box, or a single hairpin and ACA box. In some embodiments, as shown in Figure 12B and... Figure 13 The data shows that the gsnoRNA contains or consists of any of the sequences described in SEQ ID NO:89-100 and 113-128.
[0244] In some implementations, a method of treating a disease or condition associated with PTC in a target RNA in a subject includes introducing an engineered gsnoRNA into the subject's host cells, wherein the gsnoRNA contains a sequence selected from SEQ ID NO:71-84. Sequences of exemplary engineered gsnoRNAs targeting uridine residues of PTC associated with illustrative diseases are shown in Table 4.
[0245] In some embodiments, a method of treating a disease or condition associated with a PTC in a target RNA in a subject includes introducing (a) an engineered gsnoRNA and (b) a splice-converting antisense oligonucleotide (ASO) into the subject's host cells, wherein the gsnoRNA contains a guide sequence that hybridizes to a sequence containing a target uridine residue in the target RNA, wherein the ASO enhances the expression of a DKC1 protein, which is an endogenous DKC1 isoform with cytoplasmic localization in the host cells, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA into a pseudouridine residue. In some embodiments, the splice-converting ASO binds to the pre-mRNA of the DKC1 gene and guides the splicing of DKC1 isoform 3. In some embodiments, the introduction of the splice-conversion ASO increases the expression of DKC1 protein (which is the endogenous DKC1 isotype with cytoplasmic localization in the host cell) by at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, or 10-fold compared to the expression of DKC1 isotype 3 in the host cell in the absence of the ASO. In some embodiments, the application of the splice-conversion ASO increases the expression of DKC1 isotype 3 by at least 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 3, 4, 5, or 10-fold compared to the expression of DKC1 isotype 3 in the host cell in the absence of the ASO. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from wild-type H / ACA-snoRNAs selected from: ACA19, ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA2b. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA36. In some embodiments, the gsnoRNA comprises a scaffold sequence derived from ACA19. In some embodiments, the gsnoRNA comprises nucleotide sequences selected from: SEQ ID NO:3-12, 15-19, 22-36, and 177-179. In some embodiments, the gsnoRNA comprises nucleotide sequences selected from: SEQ ID NO:15-19.
[0246] In some implementations, the disease or condition is selected from the following: cystic fibrosis, Heller syndrome, alpha-1-antitrypsin (A1AT) deficiency, Parkinson's disease, Alzheimer's disease, albinism, amyotrophic lateral sclerosis (ALS), asthma, thalassemia 8, Cadasil syndrome, Shaco-Malley-Duss disease, chronic obstructive pulmonary disease (COPD), distal spinal muscular atrophy (DSMA), Duchenne / Becker muscular dystrophy, dystrophic epidermolysis bullosa, epidermolysis bullosa, Fabry disease, Leiden factor V-related disease, familial adenomatous polyposis, galactosemia, Gaucher's disease, glucose-6-phosphate dehydrogenase, hemophilia, hereditary hemochromatosis, Huntington's disease, and inflammatory bowel disease (IBD). Hereditary polyagglutination syndromes, Leber congenital amaurosis, Leonor-Née II syndrome, Lynch syndrome, Marfan syndrome, mucopolysaccharidosis, muscular dystrophy, myotonic dystrophy of types I and II, neurofibromatosis, Niemann-Pick II disease of types A, B, and C, NY-esol-related cancers, Boytz-Yage syndrome, phenylketonuria, Pompe disease, primary ciliary body disorders, prothrombin mutation-related diseases (such as prothrombin G20210A mutation), pulmonary hypertension, (autosomal dominant) retinitis pigmentosa, Sandhoff's disease, severe combined immunodeficiency syndrome (SCID), sickle cell anemia, spinal muscular atrophy, Sturgeon's disease, Ty Sachs disease, Usher syndrome, X-linked immunodeficiency, Sturgeon-Weber syndrome, and cancer. Exemplary diseases or conditions associated with PTCs in target RNAs are listed in the Human Genetic Mutation Database (HCVM). Available at hgmd.cf.ac.uk and the ClinVar database (see Landrum et al., ClinVar: improvements to accessing data. Nucleic Acids Res. 2020; 48(D1):D835-D844; available at ncbi.nlm.nih.gov / clinvar / intro). In some implementations, threonine or serine is incorporated at the ΨAA and ΨAG codons, and phenylalanine or tyrosine is incorporated at the ΨGA codon.
[0247] In some embodiments, the present invention provides the use of nucleic acid molecules (encoding engineered gsnoRNAs as described herein) in the manufacture of medicaments for treating one or more of the diseases listed herein. In some embodiments, this invention provides engineered gsnoRNAs for treating cystic fibrosis (CF). Exemplary PTCs associated with CF are known in the art, for example as described in International Patent Publication WO2019191232, the contents of which are incorporated herein by reference in their entirety. PTC mutations associated with exemplary cystic fibrosis include (but are not limited to) G542X (UGA), W1282X (UGA), R553X (UGA), R1162X (UGA), Y122X (UAA), W1089X, W846X, and W401X mutations, which can be modified by pseudouridineization to codons encoding amino acids, thereby allowing translation into full-length proteins. For example, it is well established in the art that the codons ΨAA and ΨAG are translated as serine or threonine, while ΨGA is translated as tyrosine or phenylalanine, rather than being considered a stop codon (Karijolich and Yu, 2011). In some embodiments, the host cell is an archaea or eukaryotic cell. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a human cell. In some embodiments, the method is performed in vivo. In other embodiments, the method is performed in vitro.
[0248] The methods of this invention can be used to suppress NMD and / or promote PTC readthrough in disease-associated PTCs for a wide range of known disease-associated PTCs. Numerous human diseases arise from nonsense mutations in individual disease genes. For example, Usher syndrome is a hereditary retinal dystrophy (IRD) and a leading cause of combined deafness and blindness. Nonsense mutations occur in 12% of Usher syndrome patients and have been described in various genes, such as the USH2A gene. Some patients with Usher syndrome, who suffer from skeletal abnormalities and cognitive impairment, carry nonsense mutations in the IDUA gene, preventing the production of the functional full-length IDUA protein in these patients. A significant proportion of cystic fibrosis (CF) cases (a chronic disease affecting the lungs and digestive system) are due to nonsense mutations in the CFTR gene. PTCs resulting from these nonsense mutations have been identified at several different sites in the coding region, each leading to a complete lack of the functional full-length CFTR protein. Nonsense mutations have also been found in several related oncogenes in many cancer patients, resulting in a complete lack of the full-length protein product. Given the detrimental effects of nonsense mutations on gene expression and disease, nonsense suppression has become an attractive strategy and ultimate goal in combating these diseases.
[0249] C. Delivery to target cells
[0250] In some aspects, the methods described herein involve delivering (e.g., administering) gsnoRNA and / or DKC1 protein, or nucleic acids encoding such gsnoRNA and / or DKC1 protein, to host cells containing the target RNA. The amount, dosage, and administration regimen of the nucleic acid encoding gsnoRNA and / or DKC1 protein to be administered may vary depending on cell type, disease to be treated, target population, mode of administration (e.g., systemic vs. local), disease severity, and acceptable side effects, but these can and should be evaluated through trial and error during in vitro studies, preclinical studies, and clinical trials. The above-described assays are particularly straightforward when the modified sequence results in easily detectable phenotypic changes.
[0251] In some embodiments, the method includes delivering one or more nucleic acids (e.g., gsnoRNA or nucleic acids encoding the gsnoRNA and / or DKC1 protein) and / or a pre-formed gsnoRNA-protein complex (which may contain the gsnoRNA, DKC1 protein, NOP10 protein, GAR1 protein, and / or NHP2 protein) into cells (e.g., mammalian or human cells). Exemplary intracellular delivery methods include (but are not limited to): viruses or virus-like agents; chemical-based transfection methods, such as those using calcium phosphate, dendritic polymers, liposomes, or cationic polymers (e.g., DEAE-glucose or polyethyleneimine); non-chemical methods, such as microinjection, electroporation, cell extrusion, acoustic perforation, optical transfection, puncture infection, protoplast fusion, bacterial binding, plasmid or transposon delivery; particle-based methods, such as using a gene gun, magnetic transfection or magnet-assisted transfection, particle bombardment; and hybridization methods, such as nuclear transfection. In some embodiments, this application further provides cells produced by these methods, and organisms (e.g., non-human mammals) containing or produced by these cells.
[0252] Non-viral methods for nucleic acid delivery include lipid transfection, nuclear transfection, microinjection, gene gun method, virosomes, liposomes, immunoliposomes, polycationic or lipid:nucleic acid conjugates, naked DNA, artificial virions, and drug-enhanced DNA uptake. Lipid transfection is described (for example) in U.S. Patent Nos. 5,049,386, 4,946,787, and 4,897,355, and the lipid transfection reagents are commercially available (e.g., TRANSFECTAMINE). TM and In some implementations, the following is used: 2000 to transfect nucleic acids encoding gsnoRNA and / or DKC1 protein (e.g., nucleic acid vectors encoding the gsnoRNA and / or DKC1 protein).
[0253] A suitable assay technique involves delivering the nucleic acid molecule according to the invention to a cell extract, cell line, or test organism, followed by subsequent biopsy samples taken at different time points. The sequence of the target RNA can be evaluated in the biopsy sample, and the proportion of cells with the modification can be easily tracked subsequently. After one assay, knowledge can then be retained, and future deliveries can be made without taking biopsy samples. Therefore, the method of the invention may include the step of identifying the presence of desired changes in the target RNA sequence of cells, thereby verifying that the target RNA sequence has been modified. For example, when the protein encoded by the target RNA sequence is an ion channel, the change can be assessed based on the level of the protein (length, glycosylation, function, etc.) or by some functional readout (such as (induced) current). In the case of CFTR function, using a using chamber assay or NPD test in mammals (including humans) to assess the recovery or acquisition of function is well known to those skilled in the art.
[0254] After pseudouridineization has occurred in the cell, the modified RNA can become diluted over time, for example due to cell division, the limited half-life of the edited RNA, etc. Therefore, in practical therapeutic applications, the method of the present invention may involve repeated delivery of oligonucleotides until sufficient target RNA has been modified to provide tangible benefit to the patient and / or maintain that benefit over time.
[0255] In some implementations, gsnoRNA can be delivered to cells in naked nucleic acid form. Another way to deliver these constructs (gsnoRNA and / or DKC1 protein, or nucleic acid encoding the gsnoRNA and / or DKC1 protein) to cells (in vitro, ex vivo, or in vivo) is by using a delivery vector (such as a viral vector).
[0256] Traditional virus-based nucleic acid delivery systems include retroviruses, lentiviruses, adenoviruses, adeno-associated viruses (AAVs), and herpes simplex virus vectors. Retroviral, lentiviral, and AAV methods allow integration into the host genome, typically leading to long-term expression of the inserted transgene. Furthermore, high transduction efficiency has been observed in many different cell types. Retroviral tropism can be altered by incorporating exogenous envelope proteins, expanding the potential target population of target cells. Lentiviral vectors are retroviral vectors that can transduce or infect non-dividing cells and typically produce high viral titers. Retroviral vectors contain cis-acting long terminal repeats with packaging capacities up to 6–10 kb of exogenous sequences. Minimal cis-acting LTRs are sufficient to replicate and package the vector, which is then used to integrate nucleic acids into target cells to provide permanent transgene expression. Widely used retroviral vectors include those based on murine leukemia virus (MuLV), gibberish leukemia virus (GaLV), simian immunodeficiency virus (SIV), human immunodeficiency virus (HIV), and combinations thereof. In applications where transient expression is preferred, adenovirus-based systems can be used. Adenovirus-based vectors exhibit extremely high transduction efficiency in many cell types without requiring cell division. High titers and expression levels have been obtained using these vectors. These vectors can be mass-produced in relatively simple systems.
[0257] Packaging cells are typically used to form viral particles that infect host cells. These cells include 293T cells for packaging adenoviruses and ψ2 or PA317 cells for packaging retroviruses. Viral vectors are typically produced by generating cell lines that package nucleic acid vectors within viral particles. These vectors typically contain the minimum viral sequence required for packaging and subsequent integration into the host, along with other viral sequences replaced by expression cassettes of the polynucleotide(s) to be expressed. Missing viral functions are often supplied trans-form by the packaging cell lines. For example, AAV vectors used in gene therapy typically only have the ITR sequence from the AAV genome required for packaging and integration into the host genome. Viral DNA is packaged into a cell line containing helper plasmids encoding other AAV genes (i.e., rep and cap) but lacking the ITR sequence. This cell line can also be used as a helper to infect adenoviruses. The helper virus promotes the replication of the AAV vector and the expression of AAV genes from the helper plasmid. Due to the lack of the ITR sequence, the helper plasmid is not extensively packaged. Adenovirus contamination can be reduced by heat treatment, for example, which is more sensitive to adenovirus than to AAV.
[0258] In some embodiments, the viral vector is based on adeno-associated virus (AAV). In some embodiments, the viral vector is, for example, a retroviral vector (such as a lentiviral vector). Similarly, plasmids, artificial chromosomes, and plasmids that can be used in the human genome for targeted homologous recombination and integration into cells can be suitably used to deliver gsnoRNA as described herein. In some embodiments, when the gsnoRNA is delivered by a viral vector, the gsnoRNA is present in the form of an RNA transcript, a portion of which contains the sequence of oligonucleotides according to the invention. In some embodiments, the AAV vector according to the present disclosure is a recombinant AAV vector and refers to an AAV vector containing a portion of an AAV genome comprising an exon-intron-exon sequence according to the present disclosure encased in a protein shell of a capsid protein derived from an AAV serotype. The portion of the AAV genome may contain inverted terminal repeats (ITRs) derived from adeno-associated virus serotypes (such as AAV1, AAV2, AAV3, AAV4, AAV5, AAV6, AAV7, AAV8, AAV9, etc.). The protein shell containing the capsid protein can be derived from AAV serotypes such as AAV1, 2, 3, 4, 5, 6, 7, 8, and 9. The protein shell can also be referred to as the capsid protein shell. The AAV vector may lack one or all of the wild-type AAV genes, but may still contain a functional ITR nucleic acid sequence. The functional ITR sequence is essential for the replication, recovery, and packaging of AAV virions. The aforementioned ITR sequence can be a wild-type sequence or may have at least 80%, 85%, 90%, 95, or 100% sequence identity with the wild-type sequence, or may be altered (e.g., by nucleotide insertion, mutation, deletion, or substitution), as long as they maintain function. In this context, functionality refers to the ability to directly package the genome into the capsid shell and then allow expression in the host cell or target cell to be infected. In the context of this disclosure, the capsid protein shell may have a serotype different from the ITR of the AAV vector genome. Therefore, the AAV vector according to this disclosure may be composed of a capsid protein shell (i.e., an icosahedral capsid) containing capsid proteins (VP1, VP2, and / or VP3) of an AAV serotype (e.g., AAV serotype 2), and the ITR sequence contained in the AAV2 vector may be any of the aforementioned AAV serotypes, including the AAV2 vector. Thus, an "AAV2 vector" contains a capsid protein shell of AAV serotype 2, while, for example, an "AAV5 vector" contains a capsid protein shell of AAV serotype 5, thereby either one can encapsulate the genomic ITR of any AAV vector according to the present invention.In some embodiments, the recombinant AAV vector according to the present invention comprises a capsid protein shell of AAV serotype 2, 5, 8 or AAV serotype 9, wherein the AAV genome or ITR present in the AAV vector is derived from AAV serotype 2, 5, 8 or AAV serotype 9; this AAV vector is referred to as AAV2 / 2, AAV 2 / 5, AAV2 / 8, AAV2 / 9, AAV5 / 2, AAV5 / 5, AAV5 / 8, AAV 5 / 9, AAV8 / 2, AAV 8 / 5, AAV8 / 8, AAV8 / 9, AAV9 / 2, AAV9 / 5, AAV9 / 8 or AAV9 / 9 vector.
[0259] In some embodiments, the recombinant AAV vector according to the present invention comprises a capsid protein shell of AAV serotype 2 and the AAV genome or ITR present in the vector is derived from AAV serotype 5; this vector is referred to as an AAV 2 / 5 vector. In some embodiments, the recombinant AAV vector according to the present invention comprises a capsid protein shell of AAV serotype 2 and the AAV genome or ITR present in the vector is derived from AAV serotype 8; this vector is referred to as an AAV 2 / 8 vector. In some embodiments, the recombinant AAV vector according to the present invention comprises a capsid protein shell of AAV serotype 2 and the AAV genome or ITR present in the vector is derived from AAV serotype 9; this vector is referred to as an AAV 2 / 9 vector. In some embodiments, the recombinant AAV vector according to the present invention comprises a capsid protein shell of AAV serotype 2 and the AAV genome or ITR present in the vector is derived from AAV serotype 2; this vector is referred to as an AAV 2 / 2 vector. In some embodiments, a nucleic acid molecule having an exon-intron-guide RNA-intron-exon sequence according to the invention, represented by a selected nucleic acid sequence, is inserted between AAV genomes or ITR sequences as identified above, for example, an expression construct containing an expression regulatory element operatively linked to a coding sequence and a 3' terminator sequence. "AAV helper functions" generally refer to the corresponding AAV functions required for AAV replication and packaging, which are supplied trans-to the AAV vector. AAV helper functions supplement missing AAV functions in the AAV vector, but they lack AAV ITRs (which are provided by the AAV vector genome). AAV helper functions include the two main ORFs of AAV, namely the rep coding region and the cap coding region, or sequences with substantially equivalent functions. The Rep and Cap regions are well known in the art. AAV helper functions can be supplied to AAV helper constructs (which may be plasmids).
[0260] The introduction of helper constructs into host cells can occur, for example, before or simultaneously with the introduction of the AAV genome present in an AAV vector as identified herein, via transformation, transfection, or transduction. Therefore, the AAV helper constructs of the present invention can be selected such that they produce a desired combination, on the one hand, a capsid protein shell for the AAV vector, and on the other hand, a serotype of the AAV genome present in the AAV vector for replication and packaging. The “AAV helper virus” provides additional functionality required for AAV replication and packaging.
[0261] Suitable AAV helper viruses include adenoviruses, herpes simplex viruses (such as HSV types 1 and 2), and vaccinia viruses. Additional functions provided by this helper virus can also be introduced into host cells via a vector, as described in US 6,531,456. In some embodiments, the AAV genome present in the recombinant AAV vector according to the invention does not contain any nucleotide sequences encoding viral proteins, such as the AAV rep (replication) or cap (capsid) genes. The AAV genome may further contain marker or reporter genes, such as genes encoding antibiotic resistance genes (e.g., fluorescent proteins (e.g., gfp)) or genes encoding chemically, enzymatically, or otherwise detectable and / or selectable products known in the art (e.g., lacZ, aph, etc.). In some embodiments, the AAV vector according to the invention is an AAV2 / 5, AAV2 / 8, AAV2 / 9, or AAV2 / 2 vector.
[0262] In some embodiments, gsnoRNA and DKC1 are delivered to cells as a ribonucleoprotein complex (e.g., a complex comprising gsnoRNA, DKC1, NOP10, GAR1, and / or NHP2). Methods for intracellular delivery of proteins or protein complexes (such as pre-formed gsnoRNA-DKC1 / NOP10 / GAR1 / NHP2 complexes) include (but are not limited to) mechanical methods, such as microinjection into cells using microfluidic devices, electroporation, and mechanical denaturation of cells; and carrier-based methods, such as cell-penetrating peptides (CPPs), virus-like particles, supercharged proteins, nanocarriers, supramolecular carrier-based delivery systems, and nanoparticle-stabilized nanocapsules. See, for example, Fu et al., Bioconjugate Chem. 2014, 25, 1602-1608. Some mechanical methods (such as microinjection and electroporation) can be invasive and low-throughput. In some embodiments, the ribonucleoprotein complex is inserted into the complex via transmembrane insertion, while simultaneously allowing the cell to pass through a microfluidic system (such as a cell-derived fluid dynamics system). ) is delivered into the cell (see, for example, U.S. Patent Application Publication No. 20140287509).
[0263] As described above, the introduction of the nucleic acid molecule according to the invention into cells is performed using methods known to those skilled in the art. Following pseudouridineization, the readout of the effect (change in the target RNA sequence) can be monitored in various ways during an optional identification step. Therefore, the identification step to determine whether the desired pseudouridineization of the target uridine has indeed occurred generally depends on the position of the target uridine in the target RNA sequence and the effect induced by the presence of the uridine (point mutation, PTC). Thus, in some embodiments, depending on the final effect of the U-to-Ψ conversion, this identification step includes: assessing the presence of functional, elongated, full-length, and / or wild-type proteins; assessing whether pre-mRNA splicing is altered by the pseudouridineization; or using functional readout, wherein the target RNA encodes functional, full-length, elongated, and / or wild-type proteins after pseudouridineization. Functionality in any of the diseases mentioned herein is generally assessed according to methods known to those skilled in the art.
[0264] Nucleic acid molecules (such as gsnoRNA expression constructs or vectors) according to this disclosure are suitable for administration in aqueous solutions (e.g., saline) or in suspensions optionally containing additives, excipients, and other components compatible with pharmaceutical use. Administration can be by inhalation (e.g., via nebulization), intranasal, oral, by injection or infusion, intravenous, subcutaneous, intradermal, intracranial, intravitreal, intramuscular, intratracheal, intraperitoneal, rectal, etc. Administration can be in solid form, as powder, pills, or any other form compatible with pharmaceutical use in humans. This invention is particularly suitable for treating hereditary diseases such as CF.
[0265] In some embodiments, nucleic acid molecules (such as gsnoRNA, expression constructs, or vectors) can be delivered systemically. In some embodiments, nucleic acid molecules (such as gsnoRNA, expression constructs, or vectors) can be delivered to cells or locally to tissues where the target sequence is expressed. For example, mutations in CFTR cause CF, which is mainly found in lung epithelial tissue; therefore, in some embodiments, CFTR target sequences are used to specifically and directly deliver oligonucleotide constructs to the lungs. This can be achieved by inhalation (e.g.) of powders or sprays, typically conveniently using a nebulizer. In some embodiments, the nebulizer is a nebulizer using a so-called vibrating mesh, including PARI eFlow (Rapid) or i-neb from Respironics. Inhalation delivery of oligonucleotide constructs according to the invention is also expected to efficiently target these cells, which, in the case of CFTR gene targeting, could lead to improvement in gastrointestinal syndromes also associated with CF. In some diseases, the mucus layer shows increased thickness, resulting in reduced drug absorption through the lungs. One such disease is chronic bronchitis, and another example is CF. Various mucus normalizing agents are available, such as deoxyribonuclease, hypertonic saline, or mannitol, which can be purchased under the name Bronchitol. When mucus normalizing agents are used in combination with pseudouridine oligonucleotide constructs (such as the gsnoRNA construct according to the invention), they can increase the efficacy of those drugs. Therefore, administration of the oligonucleotide construct according to the invention to subjects (such as human subjects) can be combined with a mucus normalizing agent. Additionally, administration of the oligonucleotide construct according to the invention can be combined with administration of small molecules (such as potentiator compounds, e.g., Kalydeco (ivacaftor; VX-770), or corrector compounds, e.g., VX-809 (lumacaftor and / or VX-661) for the treatment of CF). Alternatively, in combination with the aforementioned mucus normalizing agents, the delivery of mucus-penetrating particles or nanoparticles can be used to efficiently deliver pseudouridine molecules to epithelial cells, such as those of the lungs and intestines. In some embodiments, administration of the oligonucleotide construct according to the invention to subjects (such as human subjects) is in combination with antibiotic treatment to… Reduce bacterial infections and symptoms such as thickened mucus and / or biofilm formation. The aforementioned antibiotics can be administered systemically, locally, or both. For use in patients with CF, the oligonucleotide constructs according to the invention, or the packaged or compound oligonucleotide constructs according to the invention, can be combined with any mucus normalizing agent (such as deoxyribonuclease, mannitol, hypertonic saline, and / or antibiotics and / or small molecules used to treat CF (such as potentiator compounds (e.g., ivacato) or corrective compounds (e.g., rumacapto and / or VX-661)). To increase access to target cells, bronchoalveolar lavage fluid (BAF) can be used to cleanse the lungs prior to the administration of the oligonucleotides according to the invention.
[0266] IV. Pharmaceutical compositions, reagent kits, and products
[0267] In some respects, this article provides a pharmaceutical composition comprising any of the gsnoRNA, nucleic acid constructs / molecules, or engineered RNA editing systems described herein, and a pharmaceutically acceptable carrier.
[0268] Pharmaceutical compositions can be prepared by mixing the therapeutic agent described herein with the desired purity with optional pharmaceutically acceptable carriers, excipients, or stabilizers (Remington's Pharmaceutical Sciences, 16th edition, Osol, A. ed. (1980)) into a lyophilized or aqueous formulation. Acceptable carriers, excipients, or stabilizers are non-toxic to the recipient at the doses and concentrations used and include buffers, antioxidants including ascorbic acid, methionine, vitamin E, and sodium metabisulfite; preservatives, isotonic agents (e.g., sodium chloride), stabilizers, metal complexes (e.g., Zn-protein complexes); chelating agents such as EDTA; and / or nonionic surfactants.
[0269] In some embodiments, the pharmaceutical composition is contained in a single-use vial (such as a single-use sealed vial). In some embodiments, the pharmaceutical composition is contained in a reusable vial. In some embodiments, the pharmaceutical composition is contained in bulk in a container. In some embodiments, the pharmaceutical composition is frozen.
[0270] In some embodiments, the pharmaceutical composition comprises gsnoRNA. In other embodiments, the pharmaceutical composition comprises a nucleic acid construct (e.g., a vector, such as a plasmid or viral vector) encoding the gsnoRNA. In some embodiments, the pharmaceutical composition comprises free gsnoRNA (“naked” gsnoRNA), or gsnoRNA bound to other components (such as ligands for targeting, uptake, and / or intracellular transport). The gsnoRNA may be used in aqueous solutions (generally pharmaceutically acceptable vectors and / or solvents) or formulated using transfection agents, liposomes, or nanoparticles (e.g., SNALP, LNP, etc.). These formulations may contain functional ligands to enhance bioavailability, etc.
[0271] This application further provides kits and articles of manufacture for any embodiment of the treatment methods described herein. The kits and articles of manufacture may comprise any of the formulations and pharmaceutical compositions described herein.
[0272] In some aspects, this document provides a kit for editing target RNA in host cells, comprising any of the gsnoRNA or nucleic acid molecules described in Section II B. In some embodiments, the kit further comprises an agent for enhancing the expression of endogenous DKC1 isotype 3 in the host cells. In some embodiments, the kit comprises a splice-conversion antisense oligonucleotide (ASO) that enhances the expression of a DKC1 protein, which is an endogenous DKC1 isotype with cytoplasmic localization in the host cells. In some embodiments, the kit further comprises the DKC1 protein or a nucleic acid encoding the DKC1 protein. In some embodiments, the DKC1 protein is a DKC1 isotype with cytoplasmic localization (e.g., isotype 3). In some embodiments, the DKC1 protein is a portion of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA. In some embodiments, the DKC1 protein contains a nuclear localization signal (NLS) absent relative to a wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein is a truncated DKC1 variant or a DKC1 variant containing deletions, such as any of the truncated or deleted variants described in Section IIA above. In some embodiments, the kit further includes instructions for editing the target RNA according to any of the methods described herein.
[0273] In some aspects, this document provides a kit for editing target RNA in host cells, comprising an engineered RNA editing system, wherein the engineered RNA editing system comprises: (a) a gsnoRNA containing a guide sequence that hybridizes to a target uridine residue sequence in the target RNA of the host cell, or a nucleic acid molecule encoding the gsnoRNA; and (b) a DKC1 protein, or a nucleic acid molecule encoding the DKC1 protein, wherein the gsnoRNA is capable of recruiting the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype. In some embodiments, the DKC1 protein is a cytoplasm-localized DKC1 isotype (e.g., isotype 3). In some embodiments, the DKC1 protein is a portion of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA. In some embodiments, the DKC1 protein contains a nuclear localization signal (NLS) absent, relative to a wild-type DKC1 protein of the same species. In some embodiments, the DKC1 protein is a truncated DKC1 variant or a DKC1 variant containing deletions, such as any of the truncated or deleted variants described in Section IIA above. In some embodiments, the kit further includes instructions for editing the target RNA according to any of the methods described herein.
[0274] The kit of the present invention is packaged in suitable packaging. Suitable packaging includes (but is not limited to) vials, bottles, wide-mouth flasks, flexible packaging (e.g., mylar or plastic bags), etc. The kit may optionally provide additional components, such as buffers and explanatory information. Therefore, this application also provides articles of manufacture including vials (such as mylars), bottles, wide-mouth flasks, flexible packaging, etc.
[0275] Instructions for use of the composition generally include information such as the dosage for the intended treatment, the dosing schedule, and the route of administration. Containers may be unit doses, bulk packaging (e.g., multi-dose packaging), or aliquots. For example, kits may be provided containing a sufficient dose of gsnoRNA and / or DKC1 protein, or nucleic acid molecules encoding gsnoRNA and / or DKC1 protein as disclosed herein, to provide effective treatment to a subject or a number of subjects. Alternatively, kits may be provided containing a sufficient dose of gsnoRNA and / or DKC1 protein, or nucleic acid molecules encoding such gsnoRNA and / or DKC1 protein, to allow for multiple administrations to a subject. Kits may also comprise multiple unit doses of the pharmaceutical composition and instructions for use, packaged in quantities sufficient for storage or use in pharmacies (e.g., hospital pharmacies and combination pharmacies).
[0276] In some embodiments, the kit includes a delivery system. This delivery system may be a unit-dose delivery system. Delivery systems for these various dosage forms may be syringes, dropper bottles, plastic squeeze units, nebulizers, atomizers, or unit-dose or multi-dose drug aerosols. In some embodiments, the present invention provides a delivery system for any of the gsnoRNA and / or DKC1 proteins, or nucleic acid molecules encoding the gsnoRNA and / or DKC1 proteins described herein, comprising the gsnoRNA and / or DKC1 protein, or nucleic acid molecules encoding the gsnoRNA and / or DKC1 protein, and means for delivering the gsnoRNA and / or DKC1 protein, or nucleic acid molecules encoding the gsnoRNA and / or DKC1 protein.
[0277] All features disclosed in this specification can be combined in any combination. Each feature disclosed in this specification may be replaced by an alternative feature for the same, equivalent, or similar purpose. Therefore, unless otherwise expressly stated, the features disclosed in this invention are merely examples of a general series of equivalent or similar features.
[0278] Example
[0279] This disclosure will be more fully understood by referring to the following examples. However, the examples above should not be construed as limiting the scope of this disclosure. It should be understood that the examples and embodiments described herein are for illustrative purposes only and will suggest various modifications or alterations thereto to those skilled in the art, and will be included within the spirit and scope of this application and the appended claims.
[0280] Example 1. Using engineered guide snoRNA to read out early stop codons
[0281] This embodiment demonstrates the efficiency of engineered guide snoRNAs in pseudouridinedation of target RNAs (e.g., pseudouridinedation of mRNA confirmed by pseudouridinedation-dependent PTC readout) and provides different expression systems for guide snoRNAs.
[0282] To achieve site-specific pseudouridineation in mRNA in vivo, artificially guided snoRNA (gsnoRNA) is engineered to target specific mRNAs for modification. Figure 1A The H / ACAsnoRNA contains two hairpins followed by H and ACA box motifs. The engineered snoRNA provided in this paper contains guide sequences targeting PTC sites in both hairpins. To assess PTC readthrough efficiency, a Venus reporter gene (reporter gene-1) was designed, expressing the Venus fluorescent reporter gene and inserting an amber codon (TAG) between amino acid codons 154 and 155 to prematurely terminate Venus translation. This reporter gene allows for the measurement of PTC readthrough efficiency by monitoring Venus expression levels. A positive control containing a glycine codon (GGT) at the same position was included. Figure 1B The reporter gene-1 (Venus-TAG) or control (Venus-GGT) was co-transfected with gsnoRNA expression constructs (gCtrl (SEQ ID NO:14), gACA19 (SEQ ID NO:37), gACA44 (SEQ ID NO:38), gACA27 (SEQ ID NO:39), gE2 (SEQ ID NO:40), gACA19-S (SEQ ID NO:41), or gACA19-L (SEQ ID NO:42)) into HEK293T cells to analyze PTC readthrough caused by snoRNA-guided pseudouridineization modification of corresponding stop codons. The effect of PTC readthrough was measured using a high-content imaging system and compared with the positive control group. Figure 1C , 1DThe effect of PTC readout was quantified. Relative Venus expression was reported as the percentage (%) of Venus detected compared to the control (Venus-GGT). These gsnoRNAs were used as the first-generation RESTART (RESTART v1).
[0283] The sequence of the control gsnoRNA (gCtrl) is provided below, with the guide region underlined (SEQ ID NO:14):
[0284]
[0285] In the human genome, over 90% of snoRNA genes are encoded in pre-mRNA introns. 1 The inventors first evaluated the effects of PTC readthrough mediated by several gsnoRNAs (RESTART v1.0) located in introns of the host gene. The inventors initially selected four endogenous snoRNAs with high expression levels in humans. 2 This includes ACA19, ACA44, ACA27, and E2 (located in the host genes of EIF3A, SNHG12, RPL21, and RPSA, respectively). Figure 2A-2C and Figures 3A-3F The reporter gene-1 was engineered with gsnoRNAs based on these scaffolds to target the Venus reporter gene PTC. Host gene fragments containing the snoRNAs were cloned into constructs driven by the CMV promoter. Following co-transfection of this reporter gene-1 with these gsnoRNA-expressing constructs (host-gCtrl; host-gACA19, host-gACA-44, host-gACA27, and host-gE2), evidence of PTC readthrough by Venus expression was observed: 5.2% and 5.0% of Venus-positive cells were detected from cells transfected with host-gACA19 and host-gE2, respectively (compared to the control Venus-GGT reporter gene), while others showed negligible signals. Figure 2A Venus expression was clearly sequence-dependent, as the control gsnoRNA (gCtrl) failed to activate Venus expression. The inventors recognized that the gACA19 and gE2 scaffolds, which exhibited higher activity than other scaffolds, were expected to have a more stable secondary structure than gACA44 and gACA27. Figures 3A-3DThis indicates that the higher efficiency of target modification confirmed by PTC readthrough may be associated with the stability of secondary structures. To test the effect of host gene sequences carrying gsnoRNAs on PTC readthrough efficiency, the inventors further compared the results by cloning different gsnoRNAs into introns between exon 2 and exon 3 of the hemoglobin subunit β (HBB) gene. gACA19 again showed the highest efficiency in mediating PTC readthrough of reporter gene-1 (relative to Venus-positive cells: 7.3%), and gE2 showed the second highest efficiency (relative to Venus-positive cells: 1.8%). Figure 2B ).
[0286] Based on the inventors' observations that host gene sequences have different effects on different gsnoRNAs (e.g. Figure 2A-2B As shown in the diagram, the inventors envisioned that directly expressing the aforementioned gsnoRNA without host gene intervention could further increase the efficiency of PTC readthrough. Therefore, the inventors designed a series of gsnoRNA expression constructs (RESTART v1.1) driven by hU6 (type III RNA polymerase III promoter) and hU1 (snRNA-type RNA polymerase II promoter) promoters. Figure 1C-1D and Figure 2C They were then co-transfected with reporter gene-1 into HEK293T cells. Compared to gACA19 embedded in host gene introns and HBB introns, respectively, the PTC readthrough efficiency of gACA19 driven by the hU6 promoter increased by 1.9 and 1.3 times (…). Figure 1C-1D and Figure 2A-2B The efficiency of gsnoRNA driven by the hU6 promoter is similar to that driven by the hU1 promoter. Figure 2C The effect of PTC readthrough was further characterized by the lengthening or shortening of gsnoRNA: no significant effect was observed with the lengthened gACA19 (gACA19-L, 9nt at 5' and 9nt at 3'), while the shortened gACA19 (gACA19-S, 3nt at 5') reduced the reporter gene-1 PTC readthrough efficiency to 35% compared to the full-length gACA19. Figure 1D (See Figures 2E-2F). Since gsnoRNA driven by small RNA promoters is more efficient in mediating PTC readthrough than gsnoRNA embedded in introns, the inventors chose gsnoRNA driven by the hU6 promoter for subsequent analysis.
[0287] To determine whether endogenous DKC1 protein was responsible for the above observations, the inventors performed RESTART v1.1 on DKC1 stably knocked-down (DKC1 KD) HEK293T cells. Figure 1ENo PTC readthrough was observed for gsnoRNAs from DKC1-KD cells, while these gsnoRNAs activated Venus expression in the control group. Figure 1F This supports the role of endogenous DKC1 in mediating PTC readthrough of reporter gene-1. Figure 1A Overall, these observations confirm that the aforementioned gsnoRNA can induce PTC readthrough of the targeted transcript.
[0288] Example 2: Optimization of gsnoRNA scaffold improves PTC readthrough efficiency
[0289] To identify the optimal gsnoRNA scaffold, the inventors selected scaffolds with RNAfold structure. 3 Five snoRNAs (gACA3, gACA17, gACA19, gACA2b, and gACA36) with predicted stable secondary structures were used as candidate scaffolds for further characterization (RESTART v1.2). Figure 4A and Figure 5A The inventors designed a snoRNA expression construct consisting of a gsnoRNA driven by the hU6 promoter and a BFP gene driven by the CMV promoter to standardize transfection efficiency. Figure 4B Among them, gACA36 and gACA2b were superior to gACA19 and showed the highest efficiency of PTC readthrough (relative to Venus-positive cells: 13.7% and 12.2%, respectively). Figure 4C-4D gACA19 has a minimum free energy of -37.10 kcal / mol, gACA2b has a minimum free energy of -54.90 kcal / mol, and gACA36 has a minimum free energy of -43.50 kcal / mol. There appears to be no direct relationship between the stability of the gsnoRNA scaffold and its editing efficiency.
[0290] To study the role of the two hairpins in gsnoRNA, the inventors introduced mutations in the 5' and 3' guide elements, respectively. Figure 4A , Figure 5A and Figure 6A The editing efficiency of gACA19 5' hairpin mutation (gACA19-5m) is comparable to that of gACA19, while gACA19-3m shows reduced efficiency. Figure 4E-4F Regarding gACA36, the editing efficiency of gACA36-3m is comparable to that of gACA36, while gACA36-5m displays negligible signal. Figure 6B These results show that only one hairpin of gACA19 / gACA36 plays a dominant role, and the two hairpins targeting the same gsnoRNA do not compete with each other.
[0291] The inventors then attempted to further improve PTC readthrough efficiency by engineering the gsnoRNA scaffold (RESTART v1.3). Figure 4E-4F (See Figures 3-4). Given that RNA polymerase III terminates transcription at the small U-extension, the inventors introduced a single-base mutation into the "UUUU" sequence in the terminal loop of gACA19. Figure 4A and Figure 5C It is worth noting that both gACA19-UUCU and gACA19-UGUU showed improvement (). Figure 4E-4F Unbound by theory, the inventors realized that altering the distance within the gsnoRNA hairpin to make the distance between the nucleotide in the guide region that hybridizes with the target uridine and the H / ACA box 14 nucleotides would increase the editing efficiency of the gsnoRNA. In one example, the inventors inserted a single base after U115 of gACA19, making the distance between the nucleotide in the guide region that hybridizes with the target uridine and the H / ACA box 14 nucleotides. Figure 4A and Figure 5D Compared to unmodified gACA19, gACA19-3addG increases efficiency by up to 1.4 times. Figure 4E-4F Furthermore, unconstrained by theory, the inventors discovered that making the guide element of the gsnoRNA hairpin more open (e.g., reducing the base pairing probability of secondary structures within the guide region) increases the editing efficiency of the gsnoRNA. To make the guide element more open, the inventors inserted a dinucleotide after U8 in the 5' hairpin of gACA19 (…). Figure 4A and Figure 5D It is worth noting that gACA19-5addCU increases the read efficiency of the aforementioned PTC by 60%. Figure 4E-4F However, the engineered gACA36 stent did not further improve the efficiency of PTC readout. Figures 6A-6B The inventors also combined optimized mutations of gACA19 and expressed two tandem gsnoRNAs, but these did not further improve the efficiency.
[0292] Example 3: Spatial proximity effect between gsnoRNA and target PTC site
[0293] The inventors then inquired whether the spatial proximity of the gsnoRNA and the target PTC site affected the efficiency of PTC readthrough. The inventors designed two new reporter genes: (1) Reporter gene-2 contains a PTC site located between the mCherry and EGFP coding regions and is activated by gsnoRNA from RESTART v1.3. The mCherry is used to normalize transfection efficiency. (2) In reporter gene-3 (RESTART v1.4), the gsnoRNA is tandemly arranged with the PTC reporter gene, which is the same PTC reporter gene as reporter gene-2. This gsnoRNA showed comparable efficiency in suppressing the PTC of both reporter gene-2 and reporter gene-1, indicating that the gsnoRNA targets different reporter genes. Unexpectedly, the gsnoRNA increased PTC readthrough efficiency in reporter gene-3 (relative to EGFP-positive cells: ~30%, compared to RESTART v1.3, ~2-fold).
[0294] Example 4: RESTART enables PTC readout in multiple cell lines.
[0295] The inventors tested RESTART v1.4 in four different cell lines derived from different tissues (including three human cell lines and one murine cell line). Highly efficient PTC readthrough events were observed in all tested cell lines, indicating that the gsnoRNA design disclosed herein is a universal strategy for suppressing PTC in different mammalian cell types.
[0296] Example 5: Increased DKC1-isotype 3 expression significantly improved PTC readthrough.
[0297] Notably, neither optimizing the combination of mutations nor increasing gsnoRNA expression levels through transfection of two tandem gsnoRNA constructs further increased PTC readthrough, indicating that RESTART v1.3 provides gsnoRNA with optimal structure and expression levels. Based on the inventors' recognition that the engineered gsnoRNA of this disclosure provides optimized gsnoRNA structure and expression levels, the inventors wanted to know whether enzyme levels and accessibility, rather than gsnoRNA stability and expression, could be limiting factors for the rate. DKC1 is responsible for snoRNA-guided pseudouridine deposition and the accompanying PTC readthrough in RESTART. Figure 1A , 1F Two DKC1 isotypes exist in human cells: DKC1 isotype 1 is the typical DKC1 form containing both bisecting N and C-terminal nuclear localization signals (NLS); DKC1 isotype 3 is an alternative splicing variant that arises through the retention of intron 12 and lacks C-terminal NLS. Figure 7AThe endogenous mRNA expression level of isotype 1 was approximately 20 times higher than that of isotype 3. 4 .
[0298] First, the inventors generated a cell line stably overexpressing DKC1, and then transfected the aforementioned DKC1-isotype 1 overexpressing cells with reporter gene-3. Figures 7B-7C DKC1-isotype 1 overexpression only slightly increased the relative fraction and relative EGFP intensity of EGFP-positive cells by 1.2 and 1.3 times, respectively, compared to control cells. Figure 7D-7F Surprisingly, in cells overexpressing isotype 3, the relative fraction of EGFP-positive cells and the relative EGFP intensity increased significantly by 2.5 and 5.2 times, respectively. Figure 7D-7F These observations were further confirmed by co-transfection of reporter gene-1 and gsnoRNA constructs into cells stably overexpressing DKC1. Figures 8A-8C To further investigate the transient effects of DKC1, reporter gene-3 was co-transfected with a DKC1-expressing construct. Similarly, transient overexpression of isotype 3 significantly increased PTC readthrough. Figure 8D We also removed the N-terminal NLS of DKC1-type 3, and these truncated NLS have similar PTC read efficiency to type 3. Figure 9 These unexpected results confirmed that exogenous DKC1-isotype 3 significantly improved PTC readthrough efficiency, achieving 61.4% EGFP-positive cells (relative to the control reporter gene) and 13.2% EGFP intensity (relative to the control reporter gene). The aforementioned gsnoRNA and DKC1-isotype 3 were used as the second-generation RESTART (RESTART v2).
[0299] To better characterize RESTART, another set of reporter genes-3 was constructed to include all three types of stop codons, and the resulting reporter gene constructs were transfected into HEK293T cells with and without exogenous DKC1-isotype 3. Figure 7B The efficiency of RESTART-mediated readthrough is positively correlated with basal or drug-induced translational readthrough. 5 The highest readout was at the opal codon (UGA), followed by the amber codon (UAG) and then the ochre codon (UAA). Figure 7G-7H ,and Figure 10A-10DFor this UGA (opal) codon, the relative percentage and relative EGFP intensity of EGFP-positive cells were 45.3% and 5.8% (RESTART v1.4), and 72.3% and 28.6% (RESTART v2), respectively; while in the absence of exogenous DKC1 (RESTART v1.4), the UAA codon showed a negligible signal, and in the case of DKC1-isotype 3 overexpression (RESTART v2), the relative percentage of EGFP-positive cells was 2.9% and the relative EGFP intensity was 0.2%. Figure 7G-7H ,and Figures 10A-10B Increasing the amount of the DKC1-isotype 3 construct expressing the UAA (ochre) codon improved PTC readthrough to 14.8% of EGFP-positive cells, compared to 25% and 19% for UAG (amber) and UGA (opal), respectively. Figure 10C-10E In summary, RESTART facilitates the reading of all three meaningless codons.
[0300] Next, reporter gene-3 constructs for each of the three stop codons were co-transfected separately into HEK293T cells along with 200 ng of DKC1-isotype 3 expression constructs. Target locus-specific pseudouridine modification was achieved using a label-free qPCR-based method. 6 Detected ( Figure 7I and Figure 11A-11C Changes in the unwinding curve were observed for all three stop codons. Figure 7I ), while negligible changes were observed for the gCtrl group lacking the Ψ modifier. Figure 11A Conversely, the unwinding curve changes at position Ψ1045 in 18S rRNA are comparable between the gACA19 and gCtrl groups. Figure 11A-11C Overall, these results confirm that DKC1-isotype 3 gsnoRNA-guided pseudouridinelation efficiently promotes readthrough of all three PTC codons.
[0301] Example 6: RESTART inhibits disease-related PTC
[0302] This example demonstrates the use of RESTART to correct disease-associated early stop codons (PTCs). Disease-associated PTCs lead to the expression of the full-length gene product through RNA-guided pseudouridineization via the RESTART system. Furthermore, RESTART was demonstrated to restore protein function in the CFTR gene containing disease-associated PTCs. In the following examples, "X" indicates a stop codon mutation. The sequences of the tested gsnoRNAs are provided in Table 4.
[0303] A PTC disease reporter gene was constructed, in which the disease gene containing the PTC site was followed by EGFP (as shown in Figure 12A). gsnoRNAs based on gACA19 and gACA36 were designed and tested to target seven disease-related nonsense mutations from six pathogenic genes: PEX7, SMN1, ALDOB, C8orf37, PCCB, and CBS (Figures 12B-13). By co-expressing the gsnoRNA / PTC disease gene pair (RESTARTv1) in HEK293T cells, PTC readthrough was achieved at all sites: compared to the positive control, 6.7% (cells expressing ALDOB-W148), 25.2% (SMN1-W190X), 33.8% (PEX7-R232X), 1.7% (C8orf37-W185X), 38.8% (PCCB-R111X), 22.1% (CBS-C275X), and 8.0% (CBS-W390X) of EGFP-positive cells were detected. Figure 14A Next, PTC readout of the disease gene was tested for RESTARTv2 (DKC1-isotype 3 overexpression). Figure 14B Compared to RESTARTv1, DKC1-isotype 3 overexpression (RESTARTv2) increased the relative fraction of EGFP-positive cells (indicating PTC readthrough) by an average of ~2.8-fold. Figures 14A-14B The results showed that, against cells expressing ALDOB-W148X and CBS-W390X, DKC1-isotype 3 overexpression significantly increased the relative fraction of EGFP-positive cells (4.8 and 6.3 times, respectively).
[0304] like Figure 14C Further validation showed that RESTART inhibited disease-associated PTC LMNA-R225X (associated with familial dilated cardiomyopathy (DCM) with conduction disorder (DCM-CD)), F9-Y22X and F9-G21X (associated with hemophilia B), ABCA4-R408X (associated with Starfardt's disease), RS1-Y65X (associated with X-linked retinoschisis) and Rpe65-R44X (associated with Leber congenital amaurosis).
[0305] Finally, RESTART demonstrated the ability to restore protein function in the CFTR (cystic fibrosis transmembrane conduction regulator) gene, which contains disease-associated PTCs. Mutations in CFTR cause the monogenic disease cystic fibrosis, affecting 1 in 2500 live-born Caucasian infants. The ability of RESTART to repair the CFTR R553X (CGA-TGA) and W1282X (TGG-TGA) PTC sites and restore protein function was tested using electrophysiological analysis, the "gold standard" for evaluating CFTR functional recovery. After RESTART delivery, the function of PTC-containing CFTR was restored to approximately 30% of the WT CFTR level, demonstrating the therapeutic potential of RESTART in targeting certain monogenic diseases.
[0306] Example 7: Delivery of RESTART via clinically relevant forms of gsnoRNA
[0307] This example demonstrates the design and synthesis of functional oligonucleotides for delivering gsnoRNA to cells.
[0308] Full-length gsnoRNA oligonucleotides were prepared via in vitro transcription (IVT). To increase the stability of gsnoRNA oligonucleotides in cells, a 5' cap was modified (m 7 A 5' cap analogue (G(5′)ppp(5′)G cap) is added to the above gsnoRNA oligonucleotide. This 5' cap modification is not present in endogenous intronic snoRNAs. As an example, the full-length gACA19 oligonucleotide (rACA19) with a 5' cap modification targeting reporter gene-2 was prepared by in vitro transcription. Figure 15A -C). It should be noted that compared to the gACA19 expression construct vector, rACA19 increases the efficiency of PTC readthrough for both RESTARTv1 and RESTARTv2. Figure 15D (Data are displayed as mean ± standard deviation).
[0309] A chemically synthesized semi-rACA19 oligonucleotide modified with 2'-O-methyl and phosphate thioester linkages was prepared, and its ability to achieve efficient PTC readthrough in cells was tested, such as... Figure 15E The symbols indicate ("P" represents a phosphate thioester linkage and "2'O-methyl" represents a 2'O-methyl modified nucleoside). gsnoRNA was delivered to cells via transfection.
[0310] Advantageously, compared to full-length gsnoRNAs (~130 nt) that are too long to be synthesized efficiently, semi-gsnoRNA oligonucleotides facilitate chemosynthesis. Furthermore, the rH5 and rH3 oligonucleotides, synthesized with only six phosphate-thioester linkages and four 2'O-methyl modifications per oligonucleotide, demonstrate that minimal modification is sufficient to promote the stability and function of chemosynthesized semi-gsnoRNAs. Compared to gACA19 oligonucleotides prepared via IVT, the 5' hairpin (gH5, with an H-box) and 3' hairpin (gH3, with an ACA-box) constructs reduce PTC readthrough efficiency. However, both rH5 and rH3 oligonucleotides, having the same sequences as gH5 and gH3, show comparable efficiency to the full-length gACA19 construct. Figure 15D ).
[0311] These results demonstrate that gsnoRNA can be efficiently delivered to cells as full-length RNA oligonucleotides prepared via in vitro transcription (e.g., with a 5' cap for increased stability) or as semi-oligonucleotides containing 5' or 3' hairpins prepared via chemical synthesis. Furthermore, the data confirm that chemically synthesized rH3 or rH5 oligonucleotides with six phosphate thioester linkages and only four 2'O-methyl modifications are stable and functional in cells. Advantageously, using chemically synthesized rH3 and rH5 oligonucleotides with minimal modifications reduces the cost of preparing chemically synthesized oligonucleotides. The delivered RNA oligonucleotides function better than the same constructs delivered to cells in the form of DNA vectors encoding the same gsnoRNA constructs.
[0312] References
[0313] 1.Dieci,G.,Preti,M.&Montanini,B.,Eukaryotic snoRNAs: a paradigm forgene expression flexibility.Genomics 94,83-8(2009).
[0314] 2. Jorjani, H. et al. An updated human snoRNAome. Nucleic Acids Res 44, 5068-82 (2016).
[0315] 3.Gruber,AR,Lorenz,R.,Bernhart,SH, R. & Hofacker, IL TheVienna RNAwebsuite. Nucleic Acids Res 36, W70-4 (2008).
[0316] 4. Angrisani, A., Turano, M., Paparo, L., Di Mauro, C. & Furia, M., A new humandyskerin isoform with cytoplasmic localization. Biochim Biophys Acta 1810, 1361-8 (2011).
[0317] 5. Dabrowski, M., Bukowy-Bieryllo, Z. & Zietkiewicz, E., Translational readthrough potential of natural termination codons in eukaryotes---The impact of RNA sequence. RNA Biol 12,950-8(2015).
[0318] 6. Lei, Z. & Yi, C., A Radiolabeling-Free, qPCR-Based Method for Locus-Specific Pseudouridine Detection. Angew Chem Int Ed Engl 56, 14878-14882 (2017). <110> Modit Therapeutics Beijing Limited <120> Compositions, systems, and methods for RNA editing using DKC1 <130> PG03455A-FF00634CN <140> Not yet Assigned <141> Concurrently Herewith <150> PCT / CN2021 / 096122 <151> 2021-05-26 <160> 197 <170> FastSEQ for Windows Version 4.0 <210> 1 <211> 514 <212> PRT <213> Homo sapiens <400> 1 Met Ala Asp Ala Glu Val Ile Ile Leu Pro Lys Lys His Lys Lys Lys 1 5 10 15 Lys Glu Arg Lys Ser Leu Pro Glu Glu Asp Val Ala Glu Ile Gln His 20 25 30 Ala Glu Glu Phe Leu Ile Lys Pro Glu Ser Lys Val Ala Lys Leu Asp 35 40 45 Thr Ser Gln Trp Pro Leu Leu Leu Lys Asn Phe Asp Lys Leu Asn Val 50 55 60 Arg Thr Thr His Tyr Thr Pro Leu Ala Cys Gly Ser Asn Pro Leu Lys 65 70 75 80 Arg Glu Ile Gly Asp Tyr Ile Arg Thr Gly Phe Ile Asn Leu Asp Lys 85 90 95 Pro Ser Asn Pro Ser Ser His Glu Val Val Ala Trp Ile Arg Arg Ile 100 105 110 Leu Arg Val Glu Lys Thr Gly His Ser Gly Thr Leu Asp Pro Lys Val 115 120 125 Thr Gly Cys Leu Ile Val Cys Ile Glu Arg Ala Thr Arg Leu Val Lys 130 135 140 Ser Gln Gln Ser Ala Gly Lys Glu Tyr Val Gly Ile Val Arg Leu His 145 150 155 160 Asn Ala Ile Glu Gly Gly Thr Gln Leu Ser Arg Ala Leu Glu Thr Leu 165 170 175 Thr Gly Ala Leu Phe Gln Arg Pro Pro Leu Ile Ala Ala Val Lys Arg 180 185 190 Gln Leu Arg Val Arg Thr Ile Tyr Glu Ser Lys Met Ile Glu Tyr Asp 195 200 205 Pro Glu Arg Arg Leu Gly Ile Phe Trp Val Ser Cys Glu Ala Gly Thr 210 215 220 Tyr Ile Arg Thr Leu Cys Val His Leu Gly Leu Leu Leu Gly Val Gly 225 230 235 240 Gly Gln Met Gln Glu Leu Arg Arg Val Arg Ser Gly Val Met Ser Glu 245 250 255 Lys Asp His Met Val Thr Met His Asp Val Leu Asp Ala Gln Trp Leu 260 265 270 Tyr Asp Asn His Lys Asp Glu Ser Tyr Leu Arg Arg Val Val Tyr Pro 275 280 285 Leu Glu Lys Leu Leu Thr Ser His Lys Arg Leu Val Met Lys Asp Ser 290 295 300 Ala Val Asn Ala Ile Cys Tyr Gly Ala Lys Ile Met Leu Pro Gly Val 305 310 315 320 Leu Arg Tyr Glu Asp Gly Ile Glu Val Asn Gln Glu Ile Val Val Ile 325 330 335 Thr Thr Lys Gly Glu Ala Ile Cys Met Ala Ile Ala Leu Met Thr Thr 340 345 350 Ala Val Ile Ser Thr Cys Asp His Gly Ile Val Ala Lys Ile Lys Arg 355 360 365 Val Ile Met Glu Arg Asp Thr Tyr Pro Arg Lys Trp Gly Leu Gly Pro 370 375 380 Lys Ala Ser Gln Lys Lys Leu Met Ile Lys Gln Gly Leu Leu Asp Lys 385 390 395 400 His Gly Lys Pro Thr Asp Ser Thr Pro Ala Thr Trp Lys Gln Glu Tyr 405 410 415 Val Asp Tyr Ser Glu Ser Ala Light Light Glu Val Val Ala Glu Val Val 420 425 430 Lys Ala Pro Gln Val Val Ala Glu Ala Ala Lys Thr Ala Lys Arg Lys 435 440 445 Arg Glu Ser Glu Ser Glu Ser Asp Glu Thr Pro Pro Ala Ala Pro Gln 450 455 460 Leu Ile Lys Lys Glu Lys Lys Ser Lys Lys Asp Lys Lys Ala Lys 465 470 475 480 Ala Gly Leu Glu Ser Gly Ala Glu Pro Gly Asp Gly Asp Ser Asp Thr 485 490 495 Thr Lys Lys Lys Lys Lys Lys Lys Lys Ala Lys Glu Val Glu Leu Val 500 505 510 Ser Glu <210> 2 <211> 420 <212> PRT <213> Homo sapiens <400> 2 Met Ala Asp Ala Glu Val Ile Ile Leu Pro Lys Lys His Lys Lys Lys 1 5 10 15 Lys Glu Arg Lys Ser Leu Pro Glu Glu Asp Val Ala Glu Ile Gln His 20 25 30 Ala Glu Glu Phe Leu Ile Lys Pro Glu Ser Lys Val Ala Lys Leu Asp 35 40 45 Thr Ser Gln Trp Pro Leu Leu Leu Lys Asn Phe Asp Lys Leu Asn Val 50 55 60 Arg Thr Thr His Tyr Thr Pro Leu Ala Cys Gly Ser Asn Pro Leu Lys 65 70 75 80 Arg Glu Ile Gly Asp Tyr Ile Arg Thr Gly Phe Ile Asn Leu Asp Lys 85 90 95 Pro Ser Asn Pro Ser Ser His Glu Val Val Ala Trp Ile Arg Arg Ile 100 105 110 Leu Arg Val Glu Lys Thr Gly His Ser Gly Thr Leu Asp Pro Lys Val 115 120 125 Thr Gly Cys Leu Ile Val Cys Ile Glu Arg Ala Thr Arg Leu Val Lys 130 135 140 Ser Gln Gln Ser Ala Gly Lys Glu Tyr Val Gly Ile Val Arg Leu His 145 150 155 160 Asn Ala Ile Glu Gly Gly Thr Gln Leu Ser Arg Ala Leu Glu Thr Leu 165 170 175 Thr Gly Ala Leu Phe Gln Arg Pro Pro Leu Ile Ala Ala Val Lys Arg 180 185 190 Gln Leu Arg Val Arg Thr Ile Tyr Glu Ser Lys Met Ile Glu Tyr Asp 195 200 205 Pro Glu Arg Arg Leu Gly Ile Phe Trp Val Ser Cys Glu Ala Gly Thr 210 215 220 Tyr Ile Arg Thr Leu Cys Val His Leu Gly Leu Leu Leu Gly Val Gly 225 230 235 240 Gly Gln Met Gln Glu Leu Arg Arg Val Arg Ser Gly Val Met Ser Glu 245 250 255 Lys Asp His Met Val Thr Met His Asp Val Leu Asp Ala Gln Trp Leu 260 265 270 Tyr Asp Asn His Lys Asp Glu Ser Tyr Leu Arg Arg Val Val Tyr Pro 275 280 285 Leu Glu Lys Leu Leu Thr Ser His Lys Arg Leu Val Met Lys Asp Ser 290 295 300 Ala Val Asn Ala Ile Cys Tyr Gly Ala Lys Ile Met Leu Pro Gly Val 305 310 315 320 Leu Arg Tyr Glu Asp Gly Ile Glu Val Asn Gln Glu Ile Val Val Ile 325 330 335 Thr Thr Lys Gly Glu Ala Ile Cys Met Ala Ile Ala Leu Met Thr Thr 340 345 350 Ala Val Ile Ser Thr Cys Asp His Gly Ile Val Ala Lys Ile Lys Arg 355 360 365 Val Ile Met Glu Arg Asp Thr Tyr Pro Arg Lys Trp Gly Leu Gly Pro 370 375 380 Lys Ala Ser Gln Lys Lys Leu Met Ile Lys Gln Gly Leu Leu Asp Lys 385 390 395 400 His Gly Lys Pro Thr Asp Ser Thr Pro Ala Thr Trp Lys Gln Glu Tyr 405 410 415 Val Asp Tyr Arg 420 <210> 3 <211> 104 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 38, 62, 91 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 3 gugcacanga ccugcuuucu uuuaugugag uaguguungu gcuauacaaa uaauugaagg 60 cngcaguaua acuauaaaua guaaugcugc nccuucagac aaaa 104 <210> 4 <211> 110 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 6, 36, 61, 95 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 4 cagcangggc uguggcuggu cauagccaug ggaucngcau gcaagagcaa ccuggaaaga 60 nacagcgcag gucaguacaa uaccugcaag cugcnagcuu uccuauaaug 110 <210> 5 <211> 98 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 7, 35, 59, 86 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 5 uaccccngcc aguuggacuu augucuuuau uggunagugg ggcaaaggaa auauccuunu 60 caggcaaacu ggguguuugu cuguangagg aaacaaau 98 <210> 6 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 52, 73, 112 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 6 ugugcacang cuuggaguug aggcuacuga cuggccgaug aacucgcaag ungugcuaca 60 ugaggggcaa gunacaccac aagggucucu ggcccaauga guggaguuug anauucuugc 120 uacaagua 128 <210> 7 <211> 102 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 5, 35, 60, 89 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 7 cacangaccu gcuuucuuuu augugaguag uguunugugc uauacaaaua auugaaggcn 60 gcaguauaac uauaaauagu aaugcugcnc cuucagacaa aa 102 <210> 8 <211> 122 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 17, 47, 71, 100 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 8 60 aauugaaggc ngcaguauaa cuauaaauag uaaugcugcn ccuucagaca aaaauucau 120 aa 122 <210> 9 <211> 106 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 6, 37, 59, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 9 aucganacgc uuggguaucg gcuauugccu gagugunccu cgaagaguaa cugcugacna 60 cuggcugugg gccuuauggc acagucagun cagguuagag acaugc 106 <210> 10 <211> 103 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 10, 39, 65, 92 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 10 acugccccun gcagcugugg cugccguguc acaucugung uggcagagau uagagaggcu 60 auguncaagc guucugcccc gugaacguuu gngucucaca cuc 103 <210> 11 <211> 116 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 38, 65, 102 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 11 uuggcucung gccagcaguu ugcugaagcu guuggccnca ggagccuaaa gaauugucuu 60 ucuanuuggc cauuucauaa cuuuggaaau guaaugguca anagaaagaa acauga 116 <210> 12 <211> 107 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 66, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 12 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcuaa 60 gggagngggu agaaaguauu aucuauucn ccucccagcc uacaaaa 107 <210> 13 <211> 400 <212> PRT <213> Homo sapiens <400> 13 Met Leu Pro Glu Glu Asp Val Ala Glu Ile Gln His Ala Glu Glu Phe 1 5 10 15 Leu Ile Lys Pro Glu Ser Lys Val Ala Lys Leu Asp Thr Ser Gln Trp 20 25 30 Pro Leu Leu Leu Lys Asn Phe Asp Lys Leu Asn Val Arg Thr Thr His 35 40 45 Tyr Thr Pro Leu Ala Cys Gly Ser Asn Pro Leu Lys Arg Glu Ile Gly 50 55 60 Asp Tyr Ile Arg Thr Gly Phe Ile Asn Leu Asp Lys Pro Ser Asn Pro 65 70 75 80 Ser Ser His Glu Val Val Ala Trp Ile Arg Arg Ile Leu Arg Val Glu 85 90 95 Lys Thr Gly His Ser Gly Thr Leu Asp Pro Lys Val Thr Gly Cys Leu 100 105 110 Ile Val Cys Ile Glu Arg Ala Thr Arg Leu Val Lys Ser Gln Gln Ser 115 120 125 Ala Gly Lys Glu Tyr Val Gly Ile Val Arg Leu His Asn Ala Ile Glu 130 135 140 Gly Gly Thr Gln Leu Ser Arg Ala Leu Glu Thr Leu Thr Gly Ala Leu 145 150 155 160 Phe Gln Arg Pro Pro Leu Ile Ala Ala Val Lys Arg Gln Leu Arg Val 165 170 175 Arg Thr Ile Tyr Glu Ser Lys Met Ile Glu Tyr Asp Pro Glu Arg Arg 180 185 190 Leu Gly Ile Phe Trp Val Ser Cys Glu Ala Gly Thr Tyr Ile Arg Thr 195 200 205 Leu Cys Val His Leu Gly Leu Leu Leu Gly Val Gly Gly Gln Met Gln 210 215 220 Glu Leu Arg Arg Val Arg Ser Gly Val Met Ser Glu Lys Asp His Met 225 230 235 240 Val Thr Met His Asp Val Leu Asp Ala Gln Trp Leu Tyr Asp Asn His 245 250 255 Lys Asp Glu Ser Tyr Leu Arg Arg Val Val Tyr Pro Leu Glu Lys Leu 260 265 270 Leu Thr Ser His Lys Arg Leu Val Met Lys Asp Ser Ala Val Asn Ala 275 280 285 Ile Cys Tyr Gly Ala Lys Ile Met Leu Pro Gly Val Leu Arg Tyr Glu 290 295 300 Asp Gly Ile Glu Val Asn Gln Glu Ile Val Val Ile Thr Thr Lys Gly 305 310 315 320 Glu Ala Ile Cys Met Ala Ile Ala Leu Met Thr Thr Ala Val Ile Ser 325 330 335 Thr Cys Asp His Gly Ile Val Ala Lys Ile Lys Arg Val Ile Met Glu 340 345 350 Arg Asp Thr Tyr Pro Arg Lys Trp Gly Leu Gly Pro Lys Ala Ser Gln 355 360 365 Lys Lys Leu Met Ile Lys Gln Gly Leu Leu Asp Lys His Gly Lys Pro 370 375 380 Thr Asp Ser Thr Pro Ala Thr Trp Lys Gln Glu Tyr Val Asp Tyr Arg 385 390 395 400 <210> 14 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 14 cagcaagcau cgaggggcug uggcugguca uagccauggg aucguacucc gcaugcaaga 60 gcaaccugga aagacaguga cagcgcaggu caguacaaua ccugcaagcu gcaugccagc 120 uuuccuauaa ug 132 <210> 15 <211> 104 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 38, 62, 91 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 15 gugcacanga ccugcuuucu ucuaugagag uaguguungu gcuauacaaa uaauugaagg 60 cngcaguaua acuauaaaua guaaugcugc nccuucagac aaaa 104 <210> 16 <211> 104 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 38, 62, 91 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 16 gugcacanga ccugcuuucu guuaugugag uaguguungu gcuauacaaa uaauugaagg 60 cngcaguaua acuauaaaua guaaugcugc nccuucagac aaaa 104 <210> 17 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 38, 62, 91 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 17 gugcacanga ccugcuuucu uuuaugugag uaguguungu gcuauacaaa uaauugaagg 60 cngcaguaua acuauaaaua guaaugcugc ngccuucaga caaaa 105 <210> 18 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 38, 63, 92 <223> n=A, U, G or C, and is present in repeats of 4, 5, 6, 7, 8, 9, 10, 11, or 12 <400> 18 gugcacanga ccugcuuucu uuuaugugag uaguguungu gcuauacaaa uaauugaagg 60 cungcaguau aacuauaaau aguaaugcug cnccuucaga caaaa 105 <210> 19 <211> 107 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 11, 41, 65, 94 <223> n = A, U, G, or C, and exists as a repeating sequence of 44, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 19 gugcacaucu ngaccugcuu ucuuuuaugu gaguaguguu ngugcuauac aaauaauuga 60 aggcngcagu auaacuauaa auaguaaugc ugcnccuuca gacaaaa 107 <210> 20 <211> 117 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 10, 39, 66, 103 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 20 uuggcucucn ggccagcagu uugcugaagc uguuggccnc aggagccuaa agaauugucu 60 uucuanuugg ccauuucaua acuuuggaaa uguaaugguc aanagaaaga aacauga 117 <210> twenty one <211> 114 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 7, 36, 63, 100 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> twenty one uuggcunggc cagcaguuug cugaagcugu uggccncagg agccuaaaga auugucuuuc 60 uanuuggcca uuucauaacu uuggaaaugu aauggucaan agaaagaaac auga 114 <210> twenty two <211> 107 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 66, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> twenty two uuccaaanuc aguccagggc agcuucccug uacuganuuu gggacauuaa aaugggcuaa 60 gggagngggu agaaaguauu aucuauucn ccucccagcc uacaaaa 107 <210> twenty three <211> 107 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 66, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> twenty three uuccaaanuc aguccagggc agcuucccug gacuganuuu gggacauuaa aaugggcuaa 60 gggagngggu agaaaguauu aucuauucn ccucccagcc uacaaaa 107 <210> twenty four <211> 107 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 66, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> twenty four uuccaagnuc aguccagggc agcuucccug uucugancuu gggacauuaa aaugggcuaa 60 gggagngggu agaaaguauu aucuauucn ccucccagcc uacaaaa 107 <210> 25 <211> 106 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 65, 89 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 25 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcuaa 60 gggangggua gaaaguauua uucuauucnc cuccagccu acaaaa 106 <210> 26 <211> 104 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 63, 87 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 26 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcugg 60 ganggguaga aaguauuauu cuauucnccu cccagccuac aaaa 104 <210> 27 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 64, 88 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 27 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcuag 60 gganggguag aaaguauuau ucuauucncc ucccagccua caaaa 105 <210> 28 <211> 107 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 66, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 28 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcuaa 60 gggagngggu agaaaguauu aucuauccn ccucccagcc uacaaaa 107 <210> 29 <211> 104 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 63, 87 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 29 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcugg 60 gangguaga aaguauuauu cuauccnccu cccagccuac aaaa 104 <210> 30 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 65, 89 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 30 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aauggcuaag 60 ggagngggua gaaaguauua uucuauucnc cuccagcua caaaa 105 <210> 31 <211> 106 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 66, 90 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 31 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcuaa 60 gggagngggu agaaaguauu aucuauucn cuccagccu acaaaa 106 <210> 32 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 37, 65, 89 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 32 uuccaaanuc aguccagggc agcuucccug uucuganuuu gggacauuaa aaugggcuaa 60 gggangggua gaaaguauua uucuauucnc ucccagccua caaaa 105 <210> 33 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 8, 38, 63, 92 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 33 gugcacanga ccugcuuucu ucuaugagag uaguguunug ugcuauacaa auaauugaag 60 gcngcaguau aacuauaaau aguaaugcug cnccuucaga caaaa 105 <210> 34 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 39, 63, 92 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 34 gugcacaung accugcuuuc uucuauguga guaguguung ugcuauacaa auaauugaag 60 gcngcaguau aacuauaaau aguaaugcug cnccuucaga caaaa 105 <210> 35 <211> 105 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 39, 63, 92 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 35 gugcacaung accugcuuuc uuuuauguga guaguguung ugcuauacaa auaauugaag 60 gcngcaguau aacuauaaau aguaaugcug cnccuucaga caaaa 105 <210> 36 <211> 106 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 10, 40, 64, 93 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 36 gugcacaucn gaccugcuuu cuucuaugug aguaguguun gugcuauaca aauaauugaa 60 ggcngcagua uaacuauaaa uaguaaugcu gcnccuucag acaaaa 106 <210> 37 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 37 60 caaauaauug aaggcgaucc cgcaguauaa cuauaaauag uaaugcugcu ccgguccuuc 120 agacaaaa 128 <210> 38 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 38 cagcagcuga ucccgggcug uggcugguca uagccauggg aucuccggug gcaugcaaga 60 gcaaccugga aagaauccca cagcgcaggu caguacaaua ccugcaagcu gcuccggagc 120 uuuccuauaa ug 132 <210> 39 <211> 126 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 39 uacccccggc ugaucccgcc aguuggacuu augucuuuau ugguuccggu aguggggcaa 60 aggaaauauc cuuugauccc ucaggcaaac uggguguuug ucuguauccg gugagaggaa 120 acaaau 126 <210> 40 <211> 154 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 40 uggcacacu gaucccgcuu ggaguugagg cuacugacug gccgaugaac ucgcaaguuc 60 cggugaug cuacaugagg ggcaagucug aucccacacc acaagggucu cuggcccaau 120 gaguggaguu ugauccggau ucuugcuaca agua 154 <210> 41 <211> 125 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 41 cacaugaucc cgaccugcuu ucuuuuaugu gaguaguguu uccggugaug ugcuauacaa 60 auaauugaag gcgaucccgc aguauaacua uaaauaguaa ugcugcuccg guccuucaga 120 caaaa 125 <210> 42 <211> 146 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 42 ucaguauuug ugcacaugau cccgaccugc uuucuuuuau gugaguagug uuuccgguga 60 ugugcuauac aaauaauuga aggcgauccc gcaguauaac uauaaauagu aaugcugcuc 120 cgguccuucagacaaaaauucuauaa 146 <210> 43 <211> 131 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 43 aucgaggcug aucccacgcu uggguaucgg cuauugccug aguguuccgg ugaccucgaa 60 gaguaacugc ugacugaucc cacuggcugu gggccuuaug gcacagucag uuccgcaggu 120 uagagacaug c 131 <210> 44 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 44 acugccccuc ugaucccgca gcuguggcug ccgugucaca ucuguuccgg ugaguggcag 60 agauuagaga ggcuauguug auccccaagc guucugcccc gugaacguuu guccggugau 120 agucucacac uc 132 <210> 45 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 45 uuggcucuug aucccggcca gcaguuugcu gaagcuguug gccuccggca ggagccuaaa 60 120 gguagaaaga aacauga 137 <210> 46 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 46 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauucu ccgccucccca 120 gccuacaaaa 130 <210> 47 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 47 gugcacaugu gcuagaccug cuuucuuuua ugugaguagu guugcuguua augugcuaua 60 caaauaauug aaggcgaucc cgcaguauaa cuauaaauag uaaugcugcu ccgguccuuc 120 agacaaaa 128 <210> 48 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 48 60 caaauaauug aaggcgaccg ggcaguauaa cuaaaauag uaaugcugca gcuauccuuc 120 agacaaaa 128 <210> 49 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 49 gugcacauga ucccgaccug cuuucuucua ugugaguagu guuuccggug augugcuaua 60 caaauaauug aaggcgaucc cgcaguauaa cuauaaauag uaaugcugcu ccgguccuuc 120 agacaaaa 128 <210> 50 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 50 gugcacauga ucccgaccug cuuucuguua ugugaguagu guuuccggug augugcuaua 60 caaauaauug aaggcgaucc cgcaguauaa cuauaaauag uaaugcugcu ccgguccuuc 120 agacaaaa 128 <210> 51 <211> 129 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 51 60 caaauaauug aaggcgaucc cgcaguauaa cuauaaauag uaaugcugcu ccggugccuu 120 cagacaaaa 129 <210> 52 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 52 60 caaauaauug aaggcugauc ccgcaguaua acuauaaaua guaaugcugc uccggugccu 120 ucagacaaaa 130 <210> 53 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 53 60 uacaaauaau ugaaggcgau cccgcaguau aacuauaaau aguaaugcug cuccgguccu 120 ucagacaaaa 130 <210> 54 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 54 uuccaaagcu cuaagaucag uccagggcag cuucccuguu cugaguaauu gauuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauucu ccgccucccca 120 gccuacaaaa 130 <210> 55 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 55 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuaagggag guaagagggu agaaaguauu aucuauucg uauccuccca 120 gccuacaaaa 130 <210> 56 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 56 uuccaaagcu gaucccucag uccagggcag cuucccugua cugauccggu gauuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauucu ccgccucccca 120 gccuacaaaa 130 <210> 57 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 57 uuccaaagcu gaucccucag uccagggcag cuucccugga cugauccggu gauuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauucu ccgccucccca 120 gccuacaaaa 130 <210> 58 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 58 gacuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauucu ccgccucccca 120 gccuacaaaa 130 <210> 59 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 59 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuaagggau gaucccgggu agaaaguauu aucuauucu ccgccucccca 120 gccuacaaaa 130 <210> 60 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 60 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcugggauga ucccggguag aaaguauuau ucuauucucc gccucccagc 120 cuacaaaa 128 <210> 61 <211> 129 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 61 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuagggaug aucccgggua gaaaguauua uucuauucuc cgccucccag 120 ccuacaaaa 129 <210> 62 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 62 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauccu ccgccucccca 120 gccuacaaaa 130 <210> 63 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 63 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcugggauga ucccggguag aaaguauuau ucuauccucc gccucccagc 120 cuacaaaa 128 <210> 64 <211> 128 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 64 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg cuaagggagg aucccgggua gaaaguauua uucuauucuc cgccucccag 120 cuacaaaa 128 <210> 65 <211> 129 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 65 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuaagggag gaucccgggu agaaaguauu auucuauucu ccgcucccag 120 ccuacaaaa 129 <210> 66 <211> 129 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 66 uuccaaagcu gaucccucag uccagggcag cuucccuguu cugauccggu gauuugggac 60 auuaaaaugg gcuaagggau gaucccgggu agaaaguauu auucuauucu ccgcucccag 120 ccuacaaaa 129 <210> 67 <211> 129 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 67 gugcacauga ucccgaccug cuuucuucua ugugaguagu guuuccggug augugcuaua 60 caaauaauug aaggcgaucc cgcaguauaa cuauaaauag uaaugcugcu ccggugccuu 120 cagacaaaa 129 <210> 68 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 68 gugcacaucu gaucccgacc ugcuuucuuc uaugugagua guguuuccgg ugaugugcua 60 uacaaauaau ugaaggcgau cccgcaguau aacuauaaau aguaaugcug cuccgguccu 120 ucagacaaaa 130 <210> 69 <211> 131 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 69 60 uacaaauaau ugaaggcgau cccgcaguau aacuauaaau aguaaugcug cuccggugcc 120 uucagacaaa a 131 <210> 70 <211> 131 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 70 gugcacaucu gaucccgacc ugcuuucuuc uaugugagua guguuuccgg ugaugugcua 60 uacaaauaau ugaaggcgau cccgcaguau aacuauaaau aguaaugcug cuccggugcc 120 uucagacaaa a 131 <210> 71 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 71 gugcacaucg gcacgugacc ugcuuucuuc uaugugagua guguucuucc caaugugcua 60 uacaaauaau ugaaggcgca cgugcaguau aacuauaaau aguaaugcug ccuuccuccu 120 ucagacaaaa 130 <210> 72 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 72 uuccaaagcg gcacguucag uccagggcag cuucccuguu cugauuuccu aauuugggac 60 auuaaaaucg gcuggugagc acguggacua agaaaguauu auucauaguc ccuucccgac 120 cagccuacaa aa 132 <210> 73 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 73 gugcacauaa gaguuugacc ugcuuucuuc uaugugagua guguuuggag uaaugugcua 60 uacaaauaau ugaaggcgag uuugcaguau aacuauaaau aguaaugcug cuggaguccu 120 ucagacaaaa 130 <210> 74 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 74 uuccaaagaa gaguuuacag uccagggcag cuucccuguu cuguuggagu aguuugggac 60 auuaaaaucg gcuggugaga guuuggacua agaaaguauu auucauaguc cuggagcgac 120 cagccuacaa aa 132 <210> 75 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 75 gugcacauug guucuugacc ugcuuucuuc uaugugagua guguugcuau acaugugcua 60 uacaaauaau ugaaggaguu cuugcaguau aacuauaaau aguaaugcug cgcuauaccu 120 ucagacaaaa 130 <210> 76 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 76 60 auuaaaaugg gcuaagggua guucuugggu agaaaguauu auucuauucg cuauauccca 120 gccuacaaaa 130 <210> 77 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 77 gugcacaucg auucuugacc ugcuuucuuc uaugugagua guguuuccgg ggaugugcua 60 uacaaauaau ugaagggauu cuugcaguau aacuauaaau aguaaugcug cuccgggccu 120 ucagacaaaa 130 <210> 78 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 78 60 auuaaaaucg gcugguggau ucuuggacua agaaaguauu auucauaguc cuccgggaac 120 cagccuacaa aa 132 <210> 79 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 79 gugcacaugu aguguugacc ugcuuucuuc uaugugagua guguuccugu ccaugugcua 60 uacaaauaau ugaagguagu guugcaguau aacuauaaau aguaaugcug cccugucccu 120 ucagacaaaa 130 <210> 80 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 80 60 auuaaaaucg gcugguguag uguuggacua agaaaguauu auucauaguc cccugucgac 120 cagccuacaa aa 132 <210> 81 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 81 gugcacauuu cggccugacc ugcuuucuuc uaugugagua guguuucuag ugaugugcua 60 uacaaauaau ugaaguucgg ccugcaguau aacuauaaau aguaaugcug cucuaguacu 120 ucagacaaaa 130 <210> 82 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 82 uuccaaauau cggccuucag uccagggcag cuucccuguu cugauuuagu ucuuuggaac 60 auuaaaaucg gcuggaaucg gccuggacua agaaaguauu auucauaguc cuucaguguc 120 cagccuacaa aa 132 <210> 83 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 83 gugcacaucu gguugugacc ugcuuucuuc uaugugagua gugucuauau ugaugugcua 60 uacaaauaau ugaagguggu ugugcaguau aacuauaaau aguaaugcug cuauauuucu 120 ucagacaaaa 130 <210> 84 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 84 uuccaaaucu gguuguucag uccagggcag cuucccuguu cugauauauu ucuuugggac 60 auuaaaaucg gcuggucugg uuguggacua agaaaguauu auucauaguc cuauguuuac 120 cagccuacaa aa 132 <210> 85 <211> 419 <212> PRT <213> Homo sapiens <400> 85 Met Ala Asp Ala Glu Val Ile Ile Leu Pro Lys Lys His Lys Lys Lys 1 5 10 15 Lys Glu Arg Lys Ser Leu Pro Glu Glu Asp Val Ala Glu Ile Gln His 20 25 30 Ala Glu Glu Phe Leu Ile Lys Pro Glu Ser Lys Val Ala Lys Leu Asp 35 40 45 Thr Ser Gln Trp Pro Leu Leu Leu Lys Asn Phe Asp Lys Leu Asn Val ]>50 55 60 Arg Thr Thr His Tyr Thr Pro Leu Ala Cys Gly Ser Asn Pro Leu Lys 65 70 75 80 Arg Glu Ile Gly Asp Tyr Ile Arg Thr Gly Phe Ile Asn Leu Asp Lys 85 90 95 Pro Ser Asn Pro Ser Ser His Glu Val Val Ala Trp Ile Arg Arg Ile 100 105 110 Leu Arg Val Glu Lys Thr Gly His Ser Gly Thr Leu Asp Pro Lys Val 115 120 125 Thr Gly Cys Leu Ile Val Cys Ile Glu Arg Ala Thr Arg Leu Val Lys 130 135 140 Ser Gln Gln Ser Ala Gly Lys Glu Tyr Val Gly Ile Val Arg Leu His 145 150 155 160 Asn Ala Ile Glu Gly Gly Thr Gln Leu Ser Arg Ala Leu Glu Thr Leu 165 170 175 Thr Gly Ala Leu Phe Gln Arg Pro Pro Leu Ile Ala Ala Val Lys Arg 180 185 190 Gln Leu Arg Val Arg Thr Ile Tyr Glu Ser Lys Met Ile Glu Tyr Asp 195 200 205 Pro Glu Arg Arg Leu Gly Ile Phe Trp Val Ser Cys Glu Ala Gly Thr 210 215 220 Tyr Ile Arg Thr Leu Cys Val His Leu Gly Leu Leu Leu Gly Val Gly 225 230 235 240 Gly Gln Met Gln Glu Leu Arg Arg Val Arg Ser Gly Val Met Ser Glu 245 250 255 Lys Asp His Met Val Thr Met His Asp Val Leu Asp Ala Gln Trp Leu 260 265 270 Tyr Asp Asn His Lys Asp Glu Ser Tyr Leu Arg Arg Val Val Tyr Pro 275 280 285 Leu Glu Lys Leu Leu Thr Ser His Lys Arg Leu Val Met Lys Asp Ser 290 295 300 Ala Val Asn Ala Ile Cys Tyr Gly Ala Lys Ile Met Leu Pro Gly Val 305 310 315 320 Leu Arg Tyr Glu Asp Gly Ile Glu Val Asn Gln Glu Ile Val Val Ile 325 330 335 Thr Thr Lys Gly Glu Ala Ile Cys Met Ala Ile Ala Leu Met Thr Thr 340 345 350 Ala Val Ile Ser Thr Cys Asp His Gly Ile Val Ala Lys Ile Lys Arg 355 360 365 Val Ile Met Glu Arg Asp Thr Tyr Pro Arg Lys Trp Gly Leu Gly Pro 370 375 380 Lys Ala Ser Gln Lys Lys Leu Met Ile Lys Gln Gly Leu Leu Asp Lys 385 390 395 400 His Gly Lys Pro Thr Asp Ser Thr Pro Ala Thr Trp Lys Gln Glu Tyr 405 410 415 Val Asp Tyr <210> 86 <211> 407 <212> PRT <213> Homo sapiens <400> 86 Met Ala Asp Ala Glu Val Ile Ile Leu Pro Glu Glu Asp Val Ala Glu 1 5 10 15 Ile Gln His Ala Glu Glu Phe Leu Ile Lys Pro Glu Ser Lys Val Ala 20 25 30 Lys Leu Asp Thr Ser Gln Trp Pro Leu Leu Leu Lys Asn Phe Asp Lys 35 40 45 Leu Asn Val Arg Thr Thr His Tyr Thr Pro Leu Ala Cys Gly Ser Asn 50 55 60 Pro Leu Lys Arg Glu Ile Gly Asp Tyr Ile Arg Thr Gly Phe Ile Asn 65 70 75 80 Leu Asp Lys Pro Ser Asn Pro Ser Ser His Glu Val Val Ala Trp Ile 85 90 95 Arg Arg Ile Leu Arg Val Glu Lys Thr Gly His Ser Gly Thr Leu Asp 100 105 110 Pro Lys Val Thr Gly Cys Leu Ile Val Cys Ile Glu Arg Ala Thr Arg 115 120 125 Leu Val Lys Ser Gln Gln Ser Ala Gly Lys Glu Tyr Val Gly Ile Val 130 135 140 Arg Leu His Asn Ala Ile Glu Gly Gly Thr Gln Leu Ser Arg Ala Leu 145 150 155 160 Glu Thr Leu Thr Gly Ala Leu Phe Gln Arg Pro Pro Leu Ile Ala Ala 165 170 175 Val Lys Arg Gln Leu Arg Val Arg Thr Ile Tyr Glu Ser Lys Met Ile 180 185 190 Glu Tyr Asp Pro Glu Arg Arg Leu Gly Ile Phe Trp Val Ser Cys Glu 195 200 205 Ala Gly Thr Tyr Ile Arg Thr Leu Cys Val His Leu Gly Leu Leu Leu 210 215 220 Gly Val Gly Gly Gln Met Gln Glu Leu Arg Arg Val Arg Ser Gly Val 225 230 235 240 Met Ser Glu Lys Asp His Met Val Thr Met His Asp Val Leu Asp Ala 245 250 255 Gln Trp Leu Tyr Asp Asn His Lys Asp Glu Ser Tyr Leu Arg Arg Val 260 265 270 Val Tyr Pro Leu Glu Lys Leu Leu Thr Ser His Lys Arg Leu Val Met 275 280 285 Lys Asp Ser Ala Val Asn Ala Ile Cys Tyr Gly Ala Lys Ile Met Leu 290 295 300 Pro Gly Val Leu Arg Tyr Glu Asp Gly Ile Glu Val Asn Gln Glu Ile 305 310 315 320 Val Val Ile Thr Thr Lys Gly Glu Ala Ile Cys Met Ala Ile Ala Leu 325 330 335 Met Thr Thr Ala Val Ile Ser Thr Cys Asp His Gly Ile Val Ala Lys 340 345 350 Ile Lys Arg Val Ile Met Glu Arg Asp Thr Tyr Pro Arg Lys Trp Gly 355 360 365 Leu Gly Pro Lys Ala Ser Gln Lys Lys Leu Met Ile Lys Gln Gly Leu 370 375 380 Leu Asp Lys His Gly Lys Pro Thr Asp Ser Thr Pro Ala Thr Trp Lys 385 390 395 400 Gln Glu Tyr Val Asp Tyr Arg 405 <210> 87 <211> 387 <212> PRT <213> Homo sapiens <400> 87 Met Glu Phe Leu Ile Lys Pro Glu Ser Lys Val Ala Lys Leu Asp Thr 1 5 10 15 Ser Gln Trp Pro Leu Leu Leu Lys Asn Phe Asp Lys Leu Asn Val Arg 20 25 30 Thr Thr His Tyr Thr Pro Leu Ala Cys Gly Ser Asn Pro Leu Lys Arg 35 40 45 Glu Ile Gly Asp Tyr Ile Arg Thr Gly Phe Ile Asn Leu Asp Lys Pro 50 55 60 Ser Asn Pro Ser Ser His Glu Val Val Ala Trp Ile Arg Arg Ile Leu 65 70 75 80 Arg Val Glu Lys Thr Gly His Ser Gly Thr Leu Asp Pro Lys Val Thr 85 90 95 Gly Cys Leu Ile Val Cys Ile Glu Arg Ala Thr Arg Leu Val Lys Ser 100 105 110 Gln Gln Ser Ala Gly Lys Glu Tyr Val Gly Ile Val Arg Leu His Asn 115 120 125 Ala Ile Glu Gly Gly Thr Gln Leu Ser Arg Ala Leu Glu Thr Leu Thr 130 135 140 Gly Ala Leu Phe Gln Arg Pro Pro Leu Ile Ala Ala Val Lys Arg Gln 145 150 155 160 Leu Arg Val Arg Thr Ile Tyr Glu Ser Lys Met Ile Glu Tyr Asp Pro 165 170 175 Glu Arg Arg Leu Gly Ile Phe Trp Val Ser Cys Glu Ala Gly Thr Tyr 180 185 190 Ile Arg Thr Leu Cys Val His Leu Gly Leu Leu Leu Gly Val Gly Gly 195 200 205 Gln Met Gln Glu Leu Arg Arg Val Arg Ser Gly Val Met Ser Glu Lys 210 215 220 Asp His Met Val Thr Met His Asp Val Leu Asp Ala Gln Trp Leu Tyr 225 230 235 240 Asp Asn His Lys Asp Glu Ser Tyr Leu Arg Arg Val Val Tyr Pro Leu 245 250 255 Glu Lys Leu Leu Thr Ser His Lys Arg Leu Val Met Lys Asp Ser Ala 260 265 270 Val Asn Ala Ile Cys Tyr Gly Ala Lys Ile Met Leu Pro Gly Val Leu 275 280 285 Arg Tyr Glu Asp Gly Ile Glu Val Asn Gln Glu Ile Val Val Ile Thr 290 295 300 Thr Lys Gly Glu Ala Ile Cys Met Ala Ile Ala Leu Met Thr Thr Ala 305 310 315 320 Val Ile Ser Thr Cys Asp His Gly Ile Val Ala Lys Ile Lys Arg Val 325 330 335 Ile Met Glu Arg Asp Thr Tyr Pro Arg Lys Trp Gly Leu Gly Pro Lys 340 345 350 Ala Ser Gln Lys Lys Leu Met Ile Lys Gln Gly Leu Leu Asp Lys His 355 360 365 Gly Lys Pro Thr Asp Ser Thr Pro Ala Thr Trp Lys Gln Glu Tyr Val 370 375 380 Asp Tyr Arg 385 <210> 88 <211> 381 <212> PRT <213> Homo sapiens <400> 88 Met Glu Ser Lys Val Ala Lys Leu Asp Thr Ser Gln Trp Pro Leu Leu 1 5 10 15 [[ID=2Glu Tyr Val Gly Ile Val Arg Leu His Asn Ala Ile Glu Gly Gly Thr 115 120 125 Gln Leu Ser Arg Ala Leu Glu Thr Leu Thr Gly Ala Leu Phe Gln Arg 130 135 140 Pro Pro Leu Ile Ala Ala Val Lys Arg Gln Leu Arg Val Arg Thr Ile 145 150 155 160 Tyr Glu Ser Lys Met Ile Glu Tyr Asp Pro Glu Arg Arg Leu Gly Ile 165 170 175 Phe Trp Val Ser Cys Glu Ala Gly Thr Tyr Ile Arg Thr Leu Cys Val 180 185 190 His Leu Gly Leu Leu Leu Gly Val Gly Gly Gln Met Gln Glu Leu Arg 195 200 205 Arg Val Arg Ser Gly Val Met Ser Glu Lys Asp His Met Val Thr Met 210 215 220 His Asp Val Leu Asp Ala Gln Trp Leu Tyr Asp Asn His Lys Asp Glu 225 230 235 240 Ser Tyr Leu Arg Arg Val Val Tyr Pro Leu Glu Lys Leu Leu Thr Ser 245 250 255 His Lys Arg Leu Val Met Lys Asp Ser Ala Val Asn Ala Ile Cys Tyr 260 265 270 Gly Ala Lys Ile Met Leu Pro Gly Val Leu Arg Tyr Glu Asp Gly Ile 275 280 285 Glu Val Asn Gln Glu Ile Val Val Ile Thr Thr Lys Gly Glu Ala Ile 290 295 300 Cys Met Ala Ile Ala Leu Met Thr Thr Ala Val Ile Ser Thr Cys Asp 305 310 315 320 His Gly Ile Val Ala Lys Ile Lys Arg Val Ile Met Glu Arg Asp Thr 325 330 335 Tyr Pro Arg Lys Trp Gly Leu Gly Pro Lys Ala Ser Gln Lys Lys Leu 340 345 350 Met Ile Lys Gln Gly Leu Leu Asp Lys His Gly Lys Pro Thr Asp Ser<00021<220> <223> Synthetic constructs <400> 90 ugguugca guauaacuau aaauaguaau gcugcuauau uuc 43 <210> 91 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 91 cugguuguuc aguccagggc agcuucccug uucugauaua uuuc 44 <210> 92 <211> 42 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 92 cugguugugg acuaagaaag uauuauucau aguccuaugu uu 42 <210> 93 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 93 cggcacguga ccugcuuucu ucuaugugag uaguguucuu ccca 44 <210> 94 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 94 cgcacgugca guauaacuau aaauaguaau gcugccuucc u 41 <210> 95 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 95 cggcacguuc aguccagggc agcuucccug uucugauuuc cuaa 44 <210> 96 <211> 40 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 96 agcacgugga cuaagaaagu auuauucuaagucccuuccc 40 <210> 97 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 97 aagaguuuga ccugcuuucu ucuaugugag uaguguuugg agua 44 <210> 98 <211> 40 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 98 gaguugcag uauaacuaua aauaguaaug cugcuggagu 40 <210> 99 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 99 aagaguuuac aguccagggc agcuucccug uucuguugga guag 44 <210> 100 <211> 40 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 100 agaguuugga cuaagaaagu auuauucua guccuggagc 40 <210> 101 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 101 aauguaugac aaccag 16 <210> 102 <211> 17 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 102 ggaauguaug acaacca 17 <210> 103 <211> 18 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 103 ggaauguaug acaaccag 18 <210> 104 <211> 17 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 104 gaauguauga caaccag 17 <210> 105 <211> 17 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 105 ugggaaguga cgugcug 17 <210> 106 <211> 14 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 106 gggaagugac gugc 14 <210> 107 <211> 18 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 107 uugggaagug acgugcug 18 <210> 108 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 108 gggaagugac gugcu 15 <210> 109 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 109 gcuccaugaa acucuu 16 <210> 110 <211> 14 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 110 gcuccaugaa acuc 14 <210> 111 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 111 cuccaugaaa cucuu 15 <210> 112 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 112 gcuccaugaa acucu 15 <210> 113 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 113 ugguucuuga ccugcuuucu ucuaugugag uaguguugcu auac 44 <210> 114 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 114 aguucuugca guauaacuau aaauaguaau gcugcgcuau a 41 <210> 115 <211> 45 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 115 uucggccuga ccugcuuucu ucuaugugag uaguguuucu aguga 45 <210> 116 <211> 42 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 116 uucggccugc aguauaacua uaaauaguaa ugcugcucua gu 42 <210> 117 <211> 43 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 117 uaguucuuuc aguccagggc agcuucccug uucugagcua cac 43 <210> 118 <211> 37 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 118 uaguucuugg guagaaagua uuauucuauu cgcuaua 37 <210> 119 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 119 ucggccuuca guccagggca gcuucccugu ucugauuuag u 41 <210> 120 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 120 ucggccugga cuaagaaagu auuauucaua guccuucagu g 41 <210> 121 <211> 42 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 121 gauucuugac cugcuuucuu cuaugugagu aguguuucg gg 42 <210> 122 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 122 gauucuugca guauaacuau aaauaguaau gcugcuccgg g 41 <210> 123 <211> 43 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 123 guaguguuga ccugcuuucu ucuaugugag uaaguguuccu guc 43 <210> 124 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 124 uaguguugca guauaacuau aaauaguaau gcugcccugu c 41 <210> 125 <211> 44 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 125 ugaucuuuuc aguccagggc agcuucccug uucugauccg ggac 44 <210> 126 <211> 41 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 126 gauucuugga cuaagaaagu auuauucua guccuccggg a 41 <210> 127 <211> 43 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 127 gugguguuuc aguccagggc agcuucccug uucugaccug ucg 43 <210> 128 <211> 42 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 128 guaguguugg acuaagaaag uauuauucau aguccccugu cg 42 <210> 129 <211> 17 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 129 guguagcuga agaacua 17 <210> 130 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 130 uguagcugaa gaacu 15 <210> 131 <211> 18 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 131 ucacuggaug aggccgaa 18 <210> 132 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 132 acuggaugag gccgaa 16 <210> 133 <400> 133 000 <210> 134 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 134 uguagcugaa gaacua 16 <210> 135 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 135 acuggaugag gccga 15 <210> 136 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 136 cacuggauga ggccga 16 <210> 137 <400> 137 000 <210> 138 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 138 ccuggaugaa ggauc 15 <210> 139 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 139 gacaggugaa ugcugc 16 <210> 140 <211> 15 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 140 gacaggugaa ugcug 15 <210> 141 <211> 17 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 141 guccuggaug aaggauc 17 <210> 142 <211> 16 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 142 uccuggauga aggauc 16 <210> 143 <211> 17 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 143 cgacagguga augcugc 17 <210> 144 <400> 144 000 <210> 145 <211> 114 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 38, 63, 100 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 145 uuggcucung gccagcaguu ugcugaagcu guuggccngg agccuaaaga auugucuuuc 60 uanuuggcca uuucauaacu uuggaaaugu aauggucaan agaaagaaac auga 114 <210> 146 <211> 115 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 10, 39, 64, 101 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 146 uuggcucucn ggccagcagu uugcugaagc uguuggccng gagccuaaag aauugucuuu 60 cuanuuggcc auuucauaac uuuggaaaug uaauggucaa nagaaagaaa cauga 115 <210> 147 <211> 112 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 7, 36, 61, 98 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 147 uuggcunggc cagcaguuug cugaagcugu uggccnggag ccuaaagaau ugucuuucua 60 nuuggccauu ucauaacuuu ggaaauguaa uggucaanag aaagaaacau ga 112 <210> 148 <211> 115 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 38, 64, 101 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 148 uuggcucung gccagcaguu ugcugaagcu guuggccnca ggagccuaaa gaauugucuu 60 ucunuuggcc auuucauaac uuuggaaaug uaauggucaa nagaaagaaa cauga 115 <210> 149 <211> 115 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 38, 64, 101 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 149 uuggcucung gccagcaguu ugcugaagcu guuggccnca ggagccuaaa gaauugucgu 60 ucunuuggcc auuucauaac uuuggaaaug uaauggucaa nagaacgaaa cauga 115 <210> 150 <211> 116 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 38, 65, 102 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 150 uuggcucung gccagcaguu ugcugaagcu guuggccnca ggagccuaaa gaauugucuu 60 ucuanguggc cauuucauaa cuuuggaaau guaaugguca cnagaaagaa acauga 116 <210> 151 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 151 uuggcucuug uaagaggcca gcaguuugcu gaagcuguug gccguacuca ggagccuaaa 60 120 gguagaaaga aacauga 137 <210> 152 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 152 uuggcucuug aucccggcca gcaguuugcu gaagcuguug gccuccggca ggagccuaaa 60 gaauugucuu ucuauguagg auuggccauu ucuaacuuu ggaaauguaa uggucaagug 120 aguagaaaga aacauga 137 <210> 153 <211> 138 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 153 uuggcucucu gaucccggcc agcaguuugc ugaagcuguu ggccuccggc aggagccuaa 60 agaauugucu uucuaugauc ccuuggccau uucauaacuu uggaaaugua auggucaauc 120 cgguagaaag aaacauga 138 <210> 154 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 154 uuggcugcug aucccggcca gcaguuugcu gaagcuguug gccuccggca ggagccuaaa 60 120 gguagaaaga aacauga 137 <210> 155 <400> 155 000 <210> 156 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 156 uuggcucuug aucccggcca gcaguuugcu gaagcuguug gccuccggug ggagccuaaa 60 120 gguagaaaga aacauga 137 <210> 157 <211> 138 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 157 60 agaauugucu uucuaugauc ccuuggccau uucauaacuu uggaaaugua auggucaauc 120 cgguagaaag aaacauga 138 <210> 158 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 158 uuggcugcug aucccggcca gcaguuugcu gaagcuguug gccuccggug ggagccuaaa 60 120 gguagaaaga aacauga 137 <210> 159 <211> 136 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 159 uuggcucuug aucccggcca gcaguuugcu gaagcuguug gccuccggca ggagccuaaa 60 gaauugucuu ucuugauccc uuggccauuu cauaacuuug gaaauguaau ggucaauccg 120 guagaaagaa acauga 136 <210> 160 <211> 136 <212> RNA <213> Artificial sequence <220> <223> Synthetic construct <400> 160 uuggcucuug aucccggcca gcaguuugcu gaagcuguug gccuccggca ggagccuaaa 60 gaauugucgu ucuugauccc uuggccauuu cauaacuuug gaaauguaau ggucaauccg 120 guagaacgaa acauga 136 <210> 161 <400> 161 000 <210> 162 <400> 162 000 <210> 163 <400> 163 000 <210> 164 <400> 164 000 <210> 165 <400> 165 000 <210> 166 <400> 166 000 <210> 167 <400> 167 000 <210> 168<00 000 <210> 169 <400> 169 000 <210> 170 <211> 137 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 170 uuggcucuug aucccggcca gcaguuugcu gaagcuguug gccuccggca ggagccuaaa 60 120 gguagaaaga aacauga 137 <210> 171 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 171 gugcacaucu gauccugacc ugcuuucuuc uaugugagua guguuuccgg ugaugugcua 60 uacaaauaau ugaaggcgau ccugcaguau aacuauaaau aguaaugcug cuccgguccu 120 ucagacaaaa 130 <210> 172 <211> 164 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 172 gugcacaucu gauccugacc ugcuuucuuc uaugugagua guguuuccgg ugaugugcua 60 uacaaauaau ugaaggcgau ccugcaguau aacuauaaau aguaaugcug cuccgguccu 120 ucagacaaaa ucuaucuauc uagagcggac uucgguccgc uuuu 164 <210> 173 <211> 183 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 173 gugcucgcuu cggcagcaca uauacuagug cacaucugau ccugaccugc uuucuucuau 60 gugaguagug uuuccgguga ugugcuauac aaauaauuga aggcgauccu gcaguauaac 120 uauaaauagu aaugcugcuc cgguccuuca gacaaaaucu agagcggacu ucgguccgcu 180 uuu 183 <210> 174 <211> 65 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 174 gugcacaucu gauccugacc ugcuuucuuc uaugugagua guguuuccgg ugaugugcua 60 uacaa 65 <210> 175 <211> 65 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 175 ggaauugaag gcgauccugc aguauaacua uaaauaguaa ugcugcuccg guccuucaga 60 caaaa 65 <210> 176 <400> 176 000 <210> 177 <211> 139 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 39, 63, 92 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 177 gugcacaung accugcuuuc uucuauguga guaguguung ugcuauacaa auaauugaag 60 gcngcaguau aacuauaaau aguaaugcug cnccuucaga caaaaucuau cuaucuagag 120 cggacuucgg uccgcuuuu 139 <210> 178 <211> 158 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 36, 66, 90, 119 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 178 gugcucgcuu cggcagcaca uauacuagug cacaungacc ugcuuucuuc uaugugagua 60 guguungugc uauacaaaua auugaaggcn gcaguauaac uauaaauagu aaugcugcnc 120 cuucagacaa aaucuagagc ggacuucggu ccgcuuuu 158 <210> 179 <211> 50 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 9, 39 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 179 gugcacaung accugcuuuc uucuauguga guaguguung ugcuauacaa 50 <210> 180 <211> 55 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <220> <221> misc_feature <222> 13, 42 <223> n = A, U, G, or C, and exists as a repeating sequence of 4, 5, 6, 7, 8, 9, 10, 11, or 12. <400> 180 ggaauugaag gcngcaguau aacuauaaau aguaaugcug cnccuucaga caaaa 55 <210> 181 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 181 gugcacaucc auuagugccc ugcuuucuuc uaugugagua gugguggucu cgaugugcua 60 uacaaauaau ugaaggcauu aguggaguau aacuauaaau aguaaugcuu uggucucccu 120 ucagacaaaa 130 <210> 182 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 182 60 auuaaaaucg gcugguccau ugguugacua agaaaguauu auucauaguu gggucucaac 120 cagccuacaa aa 132 <210> 183 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 183 gugcacaucg aguagcgacc ugcuuucuuc uaugugagua guguuuccug agaugugcua 60 uacaaauaau ugaagcgagu agcgcaguau aacuauaaau aguaaugcug cuccugagcu 120 ucagacaaaa 130 <210> 184 <211> 133 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 184 uuccaaagug aguagcucag uccagggcag cuucccuguu cugaucca aauuugggac 60 auuaaaaucg gcugguugag uagcggacua agaaaguauu auucauaguc cuccugaaaa 120 ccagccuaca aaa 133 <210> 185 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 185 gugcacaucu agauaugacc ugcuuucuuc uaugugagua guguuuaaaa ugaugugcua 60 uacaaauaau ugaagcuaga uaugcaguau aacuauaaau aguaaugcug cuaaaaugcu 120 ucagacaaaa 130 <210> 186 <211> 133 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 186 uuccaaaggu agauauucag uccagggcag cuucccuguu cugauaaaag gauuugggac 60 auuaaaaucg gcugguguag auauggacua agaaaguauu auucauaguc cuaaaaggaa 120 ccagccuaca aaa 133 <210> 187 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 187 gugcacaucu auccuugacc ugcuuucuuc uaugugagua guguuugcug cgaugugcua 60 uacaaauaau ugaagcuauc cuugcaguau aacuauaaau aguaaugcug cugcugcgcu 120 ucagacaaaa 130 <210> 188 <211> 133 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 188 uuccaaagau auccuuucag uccagggcag cuucccuguu cugaugcugc aguuugggac 60 auuaaaaucg gcugguguau ccuuggacua agaaaguauu auucauaguc cugcugcaga 120 ccagccuaca aaa 133 <210> 189 <211> 130 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 189 gugcacaucc uugugcgacc ugcuuucuuc uaugugagua gugucugggu ugaugugcua 60 uacaaauaau ugaaggcuug ugcgcaguau aacuauaaau aguaaugcug cuggguuccu 120 ucagacaaaa 130 <210> 190 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 190 uuccagaucc uuggcugag uccagggcag cuucccuguu cucaugggua acucugggac 60 auuaaaaucg gcugguguuu gugcggacua agaaaguauu auucauaguc cugggugaac 120 cagccuacaa aa 132 <210> 191 <211> 131 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 191 gugcacaugg uuacaugacc ugcuuucuuc uaugugagua guguugagga gauaugugcu 60 auacaaauaa uugaagguuu acaugcagua uaacuauaaa uaguaaugcu gugagggucc 120 uucagacaaa a 131 <210> 192 <211> 132 <212> RNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 192 uuccaaagac uuacaugcag uccagggcag cuucccuguu cugugagggc gauuugggac 60 auuaaaaucg gcugguguuu acauggacua agaaaguauu auucauaguc ugagggaaac 120 cagccuacaa aa 132 <210> 193 <211> 30 <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 193 gcctatatca ccggataagg atcagcccca 30 <210> 194 <211> 30 <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 194 gcctatatca ccggatagggg atcagcccca 30 <210> 195 <211> 30 <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 195 gcctatatca ccggatcagg atcagcccca 30 <210> 196 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 196 ggatagggat cagcc 14 <210> 197 <211> 14 <212> DNA <213> Artificial sequence <220> <223> Synthetic constructs <400> 197 ggaggtggat cagcc 14
Claims
1. The use of both 1) and 2) below in the manufacture of a medicament for editing a target RNA in a host cell, wherein: 1) an engineered guide small nucleolar RNA (gsnoRNA) or a nucleic acid molecule encoding the engineered gsnoRNA, and 2) a nucleic acid molecule encoding a DKC1 protein, wherein the gsnoRNA comprises a guide sequence that hybridizes to a sequence comprising a target uridine residue in the target RNA, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue, wherein the DKC1 protein is selected from one of: a) a human DKC1 isoform 3 protein, b) a human DKC1 protein lacking an N-terminal and / or C-terminal NLS, or c) a fragment of a human DKC1 isoform 3 protein that comprises at least amino acid residues 41 to 420 in SEQ ID NO: 2 and does not contain an N-terminal NLS sequence or a C-terminal NLS sequence.
2. The use of an engineered gsnoRNA in the manufacture of a medicament for editing a target RNA in a host cell, wherein the gsnoRNA comprises: a) a guide sequence that hybridizes to a sequence comprising a target uridine residue in the target RNA, wherein the guide sequence is noted as Xn, X is any of A, U, G, or C, and n is 4, 5, 6, 7, 8, 9, 10, 11, or 12; b) a scaffold sequence that is the scaffold sequence in any one of the reference sequences selected from the group consisting of: SEQ ID NO: 3, 6-12, 15-19, 22-36, 71-84, 171, or 177-192, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA to a pseudouridine residue.
3. The use of claim 2, wherein the medicament further comprises a nucleic acid encoding the DKC1 protein.
4. The use of claim 1, wherein the DKC1 protein is a human DKC1 isoform 3 protein.
5. The use of claim 1, wherein the DKC1 protein is a fragment of a human DKC1 isoform 3 protein that comprises at least amino acid residues 41 to 420 in SEQ ID NO: 2 and does not contain an N-terminal NLS sequence or a C-terminal NLS sequence.
6. The use of claim 1, wherein the DKC1 protein is a human DKC1 protein lacking an N-terminal and / or C-terminal NLS.
7. The use of claim 3, wherein the DKC1 protein is selected from one of: a) a human DKC1 isoform 3 protein, b) a human DKC1 protein lacking an N-terminal and / or C-terminal NLS, or c) a fragment of a human DKC1 isoform 3 protein that comprises at least amino acid residues 41 to 420 in SEQ ID NO: 2 and does not contain an N-terminal NLS sequence or a C-terminal NLS sequence.
8. The use according to claim 2, wherein, The medicament further comprises a splice-switching antisense oligonucleotide (ASO), wherein the ASO enhances expression of a DKC1 protein that is an endogenous DKC1 isoform having cytoplasmic localization in the host cell.
9. The use of claim 7, wherein the DKC1 isoform is isoform 3 of the human DKC1 protein.
10. The use of claim 7, wherein the DKC1 protein is a human DKC1 protein lacking an N-terminal and / or C-terminal NLS.
11. The use of claim 1, wherein the target RNA is not a ribosomal RNA (rRNA).
12. The use of claim 1, wherein the gsnoRNA comprises a scaffold sequence derived from a wild-type H / ACA-snoRNA selected from the group consisting of: ACA19, ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17.
13. The use of claim 12, wherein the gsnoRNA comprises a scaffold sequence derived from ACA2b.
14. The use of claim 12, wherein the gsnoRNA comprises a scaffold sequence derived from ACA36.
15. The use of claim 1, wherein the gsnoRNA comprises a scaffold sequence comprising in SEQ ID NOs: 71-84 and 181-192.
16. The use of claim 12, wherein the gsnoRNA comprises a scaffold sequence derived from ACA19.
17. The use of claim 1, wherein the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA.
18. The use of claim 17, wherein at least one of the one or more guide sequences is located in a hairpin structure of a 3’-end portion of the wild-type H / ACA-snoRNA.
19. The use of claim 17, wherein at least one of the one or more guide sequences is located in a hairpin structure of a 5’-end portion of the wild-type H / ACA-snoRNA.
20. The use of claim 2, wherein the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA.
21. The use of any one of claims 1-20, wherein the engineered gsnoRNA comprises one or more substitution mutations in nucleotides of a poly-U sequence of the wild-type H / ACA-snoRNA, wherein the poly-U sequence comprises at least 4 consecutive U residues.
22. The use of any one of claims 1-20, wherein the engineered gsnoRNA comprises one or more insertion or deletion mutations between the nucleotide residues in the guide region that hybridize to the target uridine and the H / ACA box of the wild-type H / ACA snoRNA, such that the engineered gsnoRNA comprises 14 or 15 nucleotides between the nucleotide residues in the guide region that hybridize to the target uridine and the H / ACA box.
23. The use of claim 20, wherein the one or more mutations are selected from the following: substitution of residues 26-29 with UUCU, substitution of residues 26-29 with UGUU, addition of a G to the 3’ hairpin structure after residue 115, and addition of CU to the 5’ hairpin after residue 8, and wherein the numbering is according to SEQ ID NO:
37.
24. The use of any one of claims 1-20, wherein the medicament comprises a nucleic acid molecule encoding the gsnoRNA.
25. The use of claim 24, wherein the nucleic acid molecule encoding the gsnoRNA is driven by a promoter selected from the group consisting of a U6 promoter and a U1 promoter.
26. The use of claim 24, wherein the nucleic acid molecule encoding the gsnoRNA is embedded in an intron sequence located between a first exon sequence and a second exon sequence, and wherein the first exon sequence, the intron sequence, and the second exon sequence are derived from a naturally-occurring gene.
27. The use of any one of claims 1, 3, and 9-19, wherein the nucleic acid molecule encoding the DKC1 protein and / or the nucleic acid molecule encoding the gsnoRNA is present in a viral vector.
28. The use of any one of claims 1, 3, and 9-19, wherein the medicament comprises a vector comprising a first nucleic acid sequence encoding the DKC1 protein and a second nucleic acid sequence encoding the gsnoRNA.
29. The use of claim 28, wherein the vector is a viral vector.
30. The use of claim 28, wherein the vector is an adeno-associated virus (AAV) vector.
31. The use of any one of claims 1-20, wherein the gsnoRNA comprises one or more chemically-modified nucleosides and / or internucleoside linkages.
32. The use of claim 31, wherein the gsnoRNA comprises one or more nucleosides having a 2’-OMe or 2’-MOE modification.
33. The use of claim 31, wherein the gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically-modified nucleosides.
34. The use of claim 31, wherein the gsnoRNA comprises one or more phosphorothioate internucleoside linkages.
35. The use of claim 31, wherein the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphorothioate internucleoside linkages.
36. The use of any one of claims 1-20, wherein the gsnoRNA comprises a 5’ cap modification.
37. The use of claim 36, wherein the 5’ cap modification is a 7-methylguanosine (m7G) cap.
38. The use of any one of claims 1-20, wherein the efficiency of editing the target RNA is at least 10%.
39. The use of any one of claims 1-20, wherein the target RNA is an mRNA.
40. The use of any one of claims 1-20, wherein the sequence comprising the target uridine in the target RNA is a stop codon, and wherein modifying the target uridine to a pseudouridine results in the stop codon being translated as a coding codon.
41. The use of claim 40, wherein the stop codon is a premature termination codon (PTC).
42. The use of claim 41, wherein the PTC is associated with a genetic disease or disorder.
43. The use of any one of claims 1-20, wherein the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA.
44. The use of any one of claims 1-20, wherein the host cell is an archaeal or eukaryotic cell.
45. The use of claim 44, wherein the host cell is a mammalian cell.
46. The use of claim 44, wherein the host cell is a human cell.
47. An engineered gsnoRNA, comprising: a) a guide sequence that hybridizes to a sequence comprising a target uridine residue in a target RNA in a host cell, b) a scaffold sequence that is the scaffold sequence in any one of the reference sequences selected from the group consisting of: SEQ ID NOs: 3, 6-12, 15-19, 22-36, 71-84, 171, or 177-192.
48. The engineered gsnoRNA of claim 47, wherein the gsnoRNA comprises a 5’ cap modification.
49. The engineered gsnoRNA of claim 48, wherein the 5’ cap modification is a 7-methylguanosine (m7G) cap.
50. The engineered gsnoRNA of claim 47, wherein the gsnoRNA comprises one or more chemically modified nucleosides and / or internucleoside linkages.
51. The engineered gsnoRNA of claim 50, wherein the gsnoRNA comprises one or more nucleosides with a 2’-OMe or 2’-MOE modification.
52. The engineered gsnoRNA of claim 50, wherein the gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides.
53. The engineered gsnoRNA of claim 50, wherein the gsnoRNA comprises one or more phosphorothioate internucleoside linkages.
54. The engineered gsnoRNA of claim 50, wherein the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphorothioate internucleoside linkage bonds.
55. An isolated nucleic acid molecule comprising a sequence encoding the gsnoRNA of any one of claims 47-54.
56. An engineered RNA editing system comprising: (a) a gsnoRNA comprising a guide sequence that hybridizes to a sequence comprising a target uridine residue in a target RNA in a host cell, or a nucleic acid molecule encoding the gsnoRNA; and (b) a DKC1 protein, or a nucleic acid molecule encoding the DKC1 protein, wherein the gsnoRNA can recruit the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue, wherein the DKC1 protein is selected from one of: i) a human DKC1 isoform 3 protein, ii) a human DKC1 protein lacking an N-terminal and / or C-terminal NLS, or iii) a fragment of a human DKC1 isoform 3 protein, the fragment comprising at least amino acid residues 41 to 420 in SEQ ID NO: 2 and does not contain an N-terminal NLS sequence or a C-terminal NLS sequence.
57. A pharmaceutical composition comprising the gsnoRNA of any one of claims 47-54, the nucleic acid molecule of claim 55, or the engineered RNA editing system of claim 56, and a pharmaceutically acceptable carrier.
58. A kit for editing a target RNA in a host cell comprising the gsnoRNA of any one of claims 47-54, the nucleic acid molecule of claim 55, or the engineered RNA editing system of claim 56.
59. Use of the gsnoRNA of any one of claims 47-54, the nucleic acid molecule of claim 55, or the engineered RNA editing system of claim 56, in the manufacture of a medicament for treating a disease or disorder in a subject, wherein translational read-through of a PTC in a target RNA is caused by modification of a uridine residue in the target RNA to a pseudouridine residue, thereby treating the disease or disorder in the subject, wherein the disease or disorder is selected from the group consisting of cystic fibrosis, Hurler syndrome, amyotrophic lateral sclerosis, Charcot-Marie-Tooth disease, chronic obstructive pulmonary disease (COPD), distal spinal muscular atrophy (DSMA), Duchenne / Becker muscular dystrophy, Fabry disease, familial adenomatous polyposis, galactosemia, hemophilia, hereditary hemochromatosis, Leber congenital amaurosis, Marfan syndrome, mucopolysaccharidosis, muscular dystrophy, neurofibromatosis, Pompe disease, primary ciliary disease, (autosomal dominant) retinitis pigmentosa, Sandhoff disease, spinal muscular atrophy, Stargardt disease, Tay-Sachs disease, Usher syndrome, X-linked immunodeficiency, Sturge-Weber syndrome.
60. A non-therapeutic method for editing a target RNA in a host cell, comprising introducing into the host cell both: 1) an engineered guide small nucleolar RNA (gsnoRNA) or a nucleic acid molecule encoding the engineered gsnoRNA, and 2) a nucleic acid molecule encoding a DKC1 protein, wherein the gsnoRNA comprises a guide sequence that hybridizes to a sequence comprising a target uridine residue in the target RNA, and wherein the gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA to a pseudouridine residue, wherein the DKC1 protein is selected from one of: a) a human DKC1 isoform 3 protein, b) a human DKC1 protein lacking an N-terminal and / or C-terminal NLS, or c) a fragment of a human DKC1 isoform 3 protein that comprises at least amino acid residues 41 to 420 in SEQ ID NO: 2 and does not contain an N-terminal NLS sequence or a C-terminal NLS sequence.
61. A non-therapeutic method for editing a target RNA in a host cell, comprising introducing into the host cell a gsnoRNA, wherein the gsnoRNA comprises: a) a guide sequence that hybridizes to a sequence comprising a target uridine residue in the target RNA, wherein the guide sequence is noted as Xn, X is any of A, U, G, or C, and n is 4, 5, 6, 7, 8, 9, 10, 11, or 12; b) a scaffold sequence that is the scaffold sequence in any one of the reference sequences selected from the group consisting of: SEQ ID NOs: 3, 6-12, 15-19, 22-36, 71-84, 171, or 177-192, and wherein the gsnoRNA recruits a DKC1 protein in the host cell to modify the target uridine residue in the target RNA to a pseudouridine residue.
62. The non-therapeutic method of claim 61, wherein the method further comprises a nucleic acid encoding the DKC1 protein.
63. The non-therapeutic method of claim 60, wherein the DKC1 protein is a human DKC1 isoform 3 protein.
64. The non-therapeutic method of claim 60, wherein the DKC1 protein is a fragment of a human DKC1 isoform 3 protein that comprises at least amino acid residues 41 to 420 in SEQ ID NO: 2 and does not contain an N-terminal NLS sequence or a C-terminal NLS sequence.
65. The non-therapeutic method of claim 60, wherein the DKC1 protein is a human DKC1 protein lacking an N-terminal and / or C-terminal NLS.
66. The non-therapeutic method of claim 62, wherein the DKC1 protein is selected from one of: a) a human DKC1 isoform 3 protein, b) a human DKC1 protein lacking an N-terminal and / or C-terminal NLS, or c) a fragment of human DKC1 isoform 3 protein, the fragment comprising at least amino acid residues 41 to 420 in SEQ ID NO: 2, and not containing the N-terminal NLS sequence or the C-terminal NLS sequence.
67. The non-therapeutic method of claim 61, wherein, The method further comprises a splice-switching antisense oligonucleotide (ASO), wherein the ASO enhances expression of a DKC1 protein that is an endogenous DKC1 isoform having cytoplasmic localization in the host cell.
68. The non-therapeutic method of claim 66, wherein the DKC1 isoform is isoform 3 of human DKC1 protein.
69. The non-therapeutic method of claim 66, wherein the DKC1 protein is human DKC1 protein lacking the N-terminal and / or C-terminal NLS.
70. The non-therapeutic method of claim 60, wherein the target RNA is not a ribosomal RNA (rRNA).
71. The non-therapeutic method of claim 60, wherein the gsnoRNA comprises a scaffold sequence derived from a wild-type H / ACA-snoRNA selected from the group consisting of ACA19, ACA2b, ACA36, ACA44, ACA27, E2, ACA3, and ACA17.
72. The non-therapeutic method of claim 71, wherein the gsnoRNA comprises a scaffold sequence derived from ACA2b.
73. The non-therapeutic method of claim 71, wherein the gsnoRNA comprises a scaffold sequence derived from ACA36.
74. The non-therapeutic method of claim 60, wherein the gsnoRNA comprises a scaffold sequence comprising in SEQ ID NOs: 71-84 and 181-192.
75. The non-therapeutic method of claim 71, wherein the gsnoRNA comprises a scaffold sequence derived from ACA19.
76. The non-therapeutic method of claim 60, wherein the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA.
77. The non-therapeutic method of claim 76, wherein at least one of the one or more guide sequences is located in a hairpin structure of the 3’-end portion of the wild-type H / ACA-snoRNA.
78. The non-therapeutic method of claim 76, wherein at least one of the one or more guide sequences is located in a hairpin structure of the 5’-end portion of the wild-type H / ACA-snoRNA.
79. The non-therapeutic method of claim 61, wherein the gsnoRNA comprises one or more guide sequences, each located in a region corresponding to a hairpin structure of the wild-type H / ACA-snoRNA.
80. The non-therapeutic method of any one of claims 59-79, wherein the engineered gsnoRNA comprises one or more substitution mutations in nucleotides of a poly-U sequence of the wild-type H / ACA-snoRNA, wherein the poly-U sequence comprises at least 4 consecutive U residues.
81. The non-therapeutic method of any one of claims 59-79, wherein the engineered gsnoRNA comprises one or more insertion or deletion mutations between a nucleotide residue in a guide region that hybridizes to a target uridine and an H / ACA box of the wild-type H / ACA snoRNA, such that the engineered gsnoRNA comprises 14 or 15 nucleotides between the nucleotide residue in the guide region that hybridizes to the target uridine and the H / ACA box.
82. The non-therapeutic method of claim 79, wherein the one or more mutations are selected from the group consisting of: substitution of residues 26-29 with UUCU, substitution of residues 26-29 with UGUU, addition of a G to the 3’ hairpin structure after residue 115, and addition of CU to the 5’ hairpin after residue 8, and wherein the numbering is according to SEQ ID NO:
37.
83. The non-therapeutic method of any one of claims 60-79, wherein the agent comprises a nucleic acid molecule encoding the gsnoRNA.
84. The non-therapeutic method of claim 83, wherein the nucleic acid molecule encoding the gsnoRNA is driven by a promoter selected from the group consisting of a U6 promoter and a U1 promoter.
85. The non-therapeutic method of claim 83, wherein the nucleic acid molecule encoding the gsnoRNA is embedded in an intron sequence located between a first exon sequence and a second exon sequence, and wherein the first exon sequence, the intron sequence, and the second exon sequence are derived from a naturally-occurring gene.
86. The non-therapeutic method of any one of claims 60, 62, and 68-78, wherein the nucleic acid molecule encoding the DKC1 protein and / or the nucleic acid molecule encoding the gsnoRNA is present in a viral vector.
87. The non-therapeutic method of any one of claims 60, 62, and 68-78, wherein the agent comprises a vector comprising a first nucleic acid sequence encoding the DKC1 protein and a second nucleic acid sequence encoding the gsnoRNA.
88. The non-therapeutic method of claim 87, wherein the vector is a viral vector.
89. The non-therapeutic method of claim 87, wherein the vector is an adeno-associated virus (AAV) vector.
90. The non-therapeutic method of any one of claims 60-79, wherein the gsnoRNA comprises one or more chemically-modified nucleosides and / or internucleoside linkages.
91. The non-therapeutic method of claim 90, wherein the gsnoRNA comprises one or more nucleosides having a 2’-OMe or 2’-MOE modification.
92. The non-therapeutic method of claim 90, wherein the gsnoRNA comprises no more than 10, no more than 8, no more than 6, or no more than 4 chemically modified nucleosides.
93. The non-therapeutic method of claim 90, wherein the gsnoRNA comprises one or more phosphorothioate internucleoside linkage.
94. The non-therapeutic method of claim 90, wherein the gsnoRNA comprises no more than 10, no more than 9, no more than 8, or no more than 6 phosphorothioate internucleoside linkages.
95. The non-therapeutic method of any one of claims 60-79, wherein the gsnoRNA comprises a 5’ cap modification.
96. The non-therapeutic method of claim 95, wherein the 5’ cap modification is a 7- methylguanosine (m7G) cap.
97. The non-therapeutic method of any one of claims 60-79, wherein the efficiency of editing the target RNA is at least 10%.
98. The non-therapeutic method of any one of claims 60-79, wherein the target RNA is an mRNA.
99. The non-therapeutic method of any one of claims 60-79, wherein the sequence comprising the target uridine in the target RNA is a stop codon, and wherein modifying the target uridine to a pseudouridine results in the stop codon being translated as a coding codon.
100. The non-therapeutic method of claim 99, wherein the stop codon is a premature termination codon (PTC).
101. The non-therapeutic method of claim 100, wherein the PTC is associated with an inherited disease or disorder.
102. The non-therapeutic method of any one of claims 60-79, wherein the DKC1 protein is part of a ribonucleoprotein (RNP) complex that binds to the gsnoRNA.
103. The non-therapeutic method of any one of claims 60-79, wherein the host cell is an archaeal or eukaryotic cell.
104. The non-therapeutic method of claim 103, wherein the host cell is a mammalian cell.
105. The non-therapeutic method of claim 103, wherein the host cell is a human cell.
Citation Information
Patent Citations
Double-stranded antisense nucleic acid with exon-skipping effect
US10190117B2
Compositions and methods for synthesizing 5′-Capped RNAS
US10494399B2
Improvement in corn plows and planters
US107110A
Improvement in brick-kilns
US187217A
Splice Switching Oligomers for TNF Superfamily Receptors and Their Use in Treatment of Disease
US20120040917A1