Pseudouridine-modified rna single base editing system and method

By using a modified gsnoRNA system to bind to the PTC site of the target RNA, pseudouridine modification is achieved, which solves the accuracy and safety issues of existing RNA editing strategies in nonsense mutation therapy and significantly improves PTC readthrough efficiency and protein expression.

CN121380197BActive Publication Date: 2026-04-24PEKING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PEKING UNIV
Filing Date
2025-12-22
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing RNA editing strategies struggle to precisely insert specific amino acids into nonsense mutations, leading to poor treatment outcomes for nonsense mutation-related diseases. Furthermore, the delivery of exogenous proteins presents challenges and immunogenicity issues.

Method used

A modified gsnoRNA was developed that binds to the PTC site in the target RNA and recruits DKC1 protein to form an RNP complex through a guide sequence, CAB box, and scaffold sequence, thereby achieving pseudouridine modification and improving the readthrough efficiency of the PTC site.

Benefits of technology

It significantly improves the readthrough efficiency of PTC, restores the coding information of mRNA, expresses the complete protein, avoids the difficulties of exogenous protein delivery and immune response, and is suitable for in vivo and in vitro cell and organoid therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121380197B_ABST
    Figure CN121380197B_ABST
Patent Text Reader

Abstract

The present application relates to the field of biotechnology and the field of RNA editing. The present application relates to a RNA single base editing system and method based on pseudouridine modification. Specifically, the present application relates to a method for inhibiting a premature termination codon (PTC) in a target RNA in a host cell. The present application also relates to an engineered guide small nucleolar RNA (gsnoRNA), an isolated nucleic acid molecule comprising the same, a composition, a host cell. The present application also relates to the use thereof in the preparation of a medicinal product. The method of the present application can significantly improve the efficiency of pseudouridine modification, thereby restoring the coding information of mRNA, making the PTC site be decoded, and significantly improving the efficiency of PTC readthrough.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of biotechnology and RNA editing. Specifically, this application relates to a method for inhibiting premature stop codons (PTCs) in target RNA in host cells. This application also relates to an engineered guide small nucleolar RNA (gsnoRNA), isolated nucleic acid molecules containing it, compositions, and host cells. This application further relates to their use in the preparation of pharmaceuticals. Background Technology

[0002] Nonsense mutations are genetic mutations caused by single-base substitutions in the coding region of mRNA. These substitutions convert a codon encoding a normal amino acid into a stop codon, resulting in a premature termination codon (PTC) and premature termination of protein translation. According to the Human Gene Mutation Database (HGMD), nonsense mutations account for more than 20% of disease-related single nucleotide mutations and as much as 11% of all mutations causing human genetic diseases.

[0003] Nonsense mutations are associated with a variety of genetic diseases and can cause severe disease phenotypes. For example, nonsense mutations in the α-L-iduronidase (IDUA) gene can lead to Hurler syndrome, a mucopolysaccharidosis caused by lysosomal abnormalities that prevents the metabolism of glycosaminoglycans, resulting in toxic effects and damage and dysfunction in multiple systems of the body. In patients with cystic fibrosis, nonsense mutations in the cystic fibrosis transmembrane conductance regulator (CFTR) gene account for approximately 10%. Defects in this CFTR gene lead to ion transport disorders, the accumulation of thick mucus in epithelial cells, and consequently, chronic inflammation and irreversible lung damage. Another example is spinal muscular atrophy (SMA). Nonsense mutations in the Survivalmotor neuron 1 (SMN1) gene impair the structure and function of motor neurons, manifesting as progressive muscle weakness and atrophy, and patients exhibit a high mortality rate in infancy. Besides genetic diseases, nonsense mutations can also occur in certain cancer-related genes, such as the tumor suppressor gene p53 (TP53), leading to its dysfunction. Therefore, exploring strategies to suppress nonsense mutations is extremely important for the treatment of various diseases.

[0004] Given its significant harm to patients' health, developing effective treatment strategies has become a research focus. Currently, scientists are exploring various methods to address the challenges posed by nonsense mutations, including using translation-read-through inducing drugs to interfere with ribosome recognition of PTCs, using repressive tRNAs to recognize PTCs and incorporate specific amino acids, applying DNA editing strategies to correct mutations in gene sequences, and employing RNA editing strategies to restore the coding information of mRNA.

[0005] As one of the most widely studied RNA editing strategies, ADAR-based RNA editing tools play an important role in nonsense mutation suppression. Researchers have designed gRNAs that enable ADAR to target and edit the A protein in the PTC (protocol site), converting the UGA or UAG codon to the UGG codon, and then incorporating tryptophan at the PTC site, thereby restoring the expression of the full-length protein. However, since ADAR can only achieve A-to-G editing, this strategy can only insert tryptophan at the PTC site. For some critical PTC sites that cannot tolerate missense mutations, further exploration of other RNA editing strategies capable of precisely inserting specific amino acids is needed to achieve more effective repair.

[0006] Besides ADAR-based RNA editing strategies, researchers are also exploring other RNA editing strategies to achieve broader nonsense mutation repair. Among them, the U-to-Ψ editing strategy, due to its unique base modification characteristics and endogenous targeted modification mechanism, is gradually becoming an important direction in nonsense mutation research. In 2011, Yu's laboratory first discovered that pseudouridine modification of PTC could inhibit nonsense mutations and restore the expression of full-length proteins in in vitro experiments and yeast cells. This discovery sparked researchers' enthusiasm for exploring the inhibition of nonsense mutations by pseudouridine modification.

[0007] To avoid delivery difficulties and immunogenicity caused by exogenous protein overexpression, researchers first considered using the endogenous pseudouridine mechanism for targeted modification. To this end, they systematically explored the catalytic mechanism of H / ACAbox snoRNPs in modified rRNA and snRNA in a yeast system. The study found that pseudouridine modification requires three core sequence and structural elements: the stability of the pseudouridine pocket and hairpin structure of the snoRNA, a fixed distance of 14–15 nt between the target uridine and the H / ACA box, and the base pairing strength between the pseudouridine pocket sequence and the target sequence. Furthermore, the researchers also explored the mechanism by which pseudouridine modification of PTCs inhibits nonsense mutations. They proposed that pseudouridine modification can affect the codon-anticodon interaction at the PTC site, making the binding affinity of certain tRNAs to the PTC site stronger than that of the releasing factor, thereby inhibiting translation termination and incorporating specific amino acids. By identifying the amino acids incorporated into pseudouridine-containing PTC sites in yeast cells, researchers found that the ΨAA and ΨAG codons primarily incorporate threonine or serine, while the ΨGA codon primarily incorporates phenylalanine or tyrosine. This result further suggests that pseudouridine-containing can influence the coding rules at PTC sites in a specific way, thereby enabling the precise incorporation of certain amino acids into the PTC site.

[0008] Overall, the H / ACA box snoRNP-based U-to-Ψ editing system is a highly flexible strategy for suppressing nonsense mutations. Its flexibility in targeting sequence pairing and diversity in amino acid incorporation enable it to play a significant role in the treatment of a wide range of nonsense mutation-related diseases. Furthermore, this strategy regulates only endogenous transcripts without affecting ribosome function, demonstrating higher safety and precision. Therefore, further development and optimization of the U-to-Ψ editing system, and exploration of its effects in nonsense mutation disease scenarios, have significant scientific and practical value.

[0009] Prior to this, the applicant developed a novel programmable RNA pseudouridine editing system—RESTART. By synergistically utilizing the endogenous targeting modification mechanism of pseudouridine and the regulatory properties of stop codons, it achieved precise pseudouridine modification at PTC sites, promoting PTC readthrough and restoring full-length functional protein expression, thereby efficiently and specifically repairing nonsense mutations. Based on this, multiple generations of systems (RESTART v1, v2, v3, and v3-mini) were developed. However, the editing efficiency of these systems still has significant room for improvement. Therefore, it is urgent to optimize the design of components within the system and regulate the expression levels of endogenous proteins to comprehensively improve modification and readthrough levels, develop novel RNA editing systems, and mediate nonsense mutation repair. Summary of the Invention

[0010] This application optimizes and modifies the previously studied RESTART pseudouridine modification system, providing a modified gsnoRNA that can efficiently bind to the target site and improve the recruitment and assembly rate of pseudouridine synthase DKC1 and the core proteins NHP2, GAR1 and NOP10, thereby significantly improving the efficiency of pseudouridine modification.

[0011] Therefore, in a first aspect, this application provides a method for suppressing premature stop codons (PTCs) in target RNA in host cells, characterized in that the method comprises:

[0012] An engineered guide small nucleolar RNA (gsnoRNA) or a nucleic acid for expressing the gsnoRNA is introduced into the host cell, wherein the gsnoRNA comprises: (i) at least one guide sequence, (ii) at least one CAB box and / or CTE (constitutive transport element), and (iii) a scaffold sequence; wherein,

[0013] The guide sequence hybridizes with the target RNA containing the target uridine residue (U) of PTC;

[0014] The CAB box is derived from the loop region of a stem-loop structure of natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA).

[0015] The scaffold sequence is derived from natural H / ACA type snoRNA and / or natural H / ACA type scaRNA.

[0016] There are two types of natural snoRNAs and scaRNAs: one is a C / D type structure, and the other is an H / ACA type structure. The H / ACA type structure can recruit the DKC1 protein. Therefore, in this paper, the gsnoRNA described can recruit the DKC1 protein in this host cell and modify the target uridine residues in the target RNA into pseudouridine residues.

[0017] In some embodiments, the scaffold sequence is not limited to a specific length and sequence, as long as it retains the ability of snoRNA or scaRNA to interact with the DKC1 protein. In some embodiments, the length of the scaffold sequence is 30-50 nt, 50-80 nt, 80-100 nt, 100-130 nt, 130-150 nt, 150-200 nt, or longer.

[0018] In some embodiments, the scaffold sequence is a complete snoRNA derived from a natural H / ACA type structure (e.g., containing two hairpin structures, an H box, and an ACA box) or a portion thereof (e.g., containing one hairpin structure, an H box, and an ACA box). In some embodiments, the gsnoRNA contains a single hairpin and an H box, but not an ACA box. In some embodiments, the gsnoRNA contains a single hairpin and an ACA box, but not an H box.

[0019] In some embodiments, the scaffold sequence is a complete scaRNA derived from a natural H / ACA type structure (e.g., containing two hairpin structures, an H box, and an ACA box) or a portion thereof (e.g., containing one hairpin structure, an H box, and an ACA box).

[0020] In some embodiments, to improve in vivo stability and delivery efficiency, the gsnoRNA may contain one or more modified (e.g., chemically modified) nucleotides. In some embodiments, one or more nucleotides of the gsnoRNA contain a 2'-O-methyl (2'-OMe) modification. In some embodiments, the gsnoRNA contains one or more phosphate thioester (PS) nucleotide bonds. In some embodiments, the gsnoRNA contains a 5' cap modification (e.g., an m7G cap modification).

[0021] In some embodiments, the methods, gsnoRNAs, and compositions provided herein include modifications to target RNAs (e.g., mRNAs) in eukaryotic organisms (e.g., mammalian cells, such as human cells). In some aspects, the host cell can be a cell from any organ, such as skin, lung, heart, kidney, liver, pancreas, intestine, muscle, glands, eye, brain, blood, etc. In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a human cell. The host cell can be located in vitro or in vivo. In some embodiments, the host cell is an ex vivo cell.

[0022] One advantage of the methods, gsnoRNAs, and compositions provided herein is that they can be used both in situ in living organisms and in cells cultured in vitro. In some embodiments, the host cells are treated in vitro and then introduced into a living organism (e.g., reintroduced into the organism of their original source).

[0023] The methods, gsnoRNAs, and compositions presented in this article can also be used to read PTCs or re-encode Ψ-modified codons within so-called organoid cells. Organoids can be considered three-dimensional, in vitro-derived tissues, but driven by specific conditions to produce individual tissues. In a therapeutic setting, organoids can be synthesized in vitro and reintroduced into patients as autologous material, which is less likely to be rejected than normal transplantation.

[0024] In some embodiments, the host cell has a gene mutation. The mutation may be heterozygous or homozygous. In some embodiments, the methods, gsnoRNA, and compositions provided herein can be used to modify point mutations (e.g., for reading through point-induced PTCs, or for recoding point mutations in sense codons). In some embodiments, when a human subject has a PTC-related disease, the methods, gsnoRNA, and compositions provided herein are suitable for modifying RNA sequences in cells, tissues, or organs associated with the subject's (e.g., a human subject's) disease state.

[0025] In the method of this application, any suitable method can be used to introduce gsnoRNA into the host cell. Furthermore, regardless of the gsnoRNA introduction method, the method can suppress premature stop codons (PTCs) in the target RNA of the host cell, significantly improving the readthrough efficiency of PTCs.

[0026] In some embodiments, the gsnoRNA is introduced into the host cell via direct delivery. In other embodiments, the gsnoRNA is introduced into the host cell by transfecting the host cell with a vector containing a nucleotide sequence expressing the gsnoRNA.

[0027] The method of this application can effectively restore the coding information of mRNA without changing the DNA sequence of the host cell, so that the PTC site can be decoded, significantly improving the reading efficiency of PTC, thereby expressing the complete protein.

[0028] In some embodiments, PTC is caused by mutations in sense codons. In some embodiments, the above method can significantly improve PTC readthrough efficiency. In some embodiments, the PTC readthrough efficiency of the above method is at least 2-fold, at least 3-fold, at least 5-fold, at least 10-fold, or at least 20-fold higher than that of other methods that do not use the above gsnoRNA. For example, the above method of this application is at least 2-fold higher than the PTC readthrough efficiency of RESTART v3mini.

[0029] When using the method of this application, gsnoRNA is introduced into the host cell. The guide sequence in gsnoRNA pairs complementaryly with the bases of the target mRNA, causing the uridine residue U in the PTC to enter the pseudouridine pocket formed by the stem-loop structure of gsnoRNA. gsnoRNA recruits endogenous and / or exogenously introduced DKC1, NOP10, NHP2, and GAR1 proteins to assemble into the RNP complex. The uridine residue located in the pseudouridine pocket can enter the catalytic center of DKC1, thereby catalyzing the modification of uridine residue U into the pseudouridine residue Ψ. Further, endogenous and / or exogenously introduced nc-tRNA binds to the Ψ-modified PTC and translates it into amino acids, achieving readthrough of the PTC site.

[0030] The gsnoRNA of this application can co-modify and decode PTC with endogenously expressed DKC1 protein and / or endogenously expressed nc-tRNA.

[0031] In some embodiments, the gsnoRNA of this application can be delivered alone to achieve PTC readthrough. In this case, other elements required for modification and decoding (e.g., DKC1 protein, nc-tRNA) are provided solely by the cell's endogenous sources. That is, by introducing the engineered guide small nucleolar RNA (gsnoRNA) or the nucleic acid for expressing the gsnoRNA into the host cell, it is possible to suppress premature stop codons (PTCs) in the target RNA of the host cell.

[0032] The gsnoRNA of this application can co-modify and decode PTC with exogenously expressed DKC1 protein and / or exogenously expressed nc-tRNA.

[0033] In some embodiments, a nucleotide sequence encoding the DKC1 protein can be delivered, utilizing exogenously expressed DKC1 protein to achieve pseudouridine modification. In some embodiments, the method further includes introducing the nucleotide sequence encoding the DKC1 protein into a host cell.

[0034] In some embodiments, the method may deliver a nucleotide sequence encoding nc-tRNA or directly deliver an nc-tRNA molecule to achieve pseudouridine modification using exogenous nc-tRNA. In some embodiments, the method further includes introducing the nucleotide sequence encoding nc-tRNA or the nc-tRNA into a host cell.

[0035] Therefore, the source of DKC1 protein and nc-tRNA does not affect the implementation of the method in this application. For example, DKC1 protein and nc-tRNA can both be from intracellular expression, both from exogenous expression, or one from intracellular expression and the other from exogenous expression.

[0036] CAB box and support sequence

[0037] The CAB box (Cajal-body localization box) element is a short RNA motif that can be localized to the Cajal body.

[0038] In some embodiments, the CAB box has the following sequence: X1X2AG; wherein X1 and X2 are each independently selected from any one of A, U, C, G.

[0039] In some embodiments, the CAB box has a nucleotide sequence selected from the following: UGAG, AAAG, GAAG, UAAG, UCAG, CGAG, AUAG, GCAG, CUAG, CAAG, or AGAG.

[0040] In some implementations, the sequence of the CAB box is AGAG.

[0041] In some embodiments, the CAB box can enhance the assembly efficiency of the gsnoRNA with core proteins (such as DKC1, NOP10, NHP2, GAR1) and the ability to form complete snoRNPs (gsnoRNA, DKC1, NOP10, NHP2, and GAR1).

[0042] In some embodiments, the support sequence includes a first hairpin structure near the 5' end and a second hairpin structure near the 3' end;

[0043] Furthermore, the CAB box is located in the loop area of ​​the first hairpin structure, or in the loop area of ​​the second hairpin structure, or simultaneously in the loop areas of both the first and second hairpin structures.

[0044] In some implementations, the CAB box is located in the loop region of the first hairpin structure in the gsnoRNA.

[0045] In some embodiments, the gsnoRNA comprises a CAB box derived from a natural scaRNA, AluRNA, or htrRNA, and a sequence in the natural scaRNA, AluRNA, or htrRNA containing the entire loop region of the CAB box.

[0046] The introduction of CAB boxes can be further divided into different types.

[0047] In one implementation, the CAB box can be directly introduced into the engineered guide small nucleolar RNA (gsnoRNA), specifically, the four nucleotide sequences of the X1X2AG of the CAB box can be directly introduced. In some implementations, the four nucleotide sequences of the X1X2AG of the CAB box are inserted into the loop region of the first hairpin structure, or the loop region of the second hairpin structure, or simultaneously into the loop regions of both the first and second hairpin structures in the gsnoRNA.

[0048] In such implementations, while maintaining the basic structure of the gsnoRNA, the four nucleotide sequences of the CAB box X1X2AG can be inserted between any two nucleotides that meet the above positions.

[0049] In a second embodiment, a sequence containing the entire loop region of the CAB box can also be introduced into the engineered guide small nucleolar RNA (gsnoRNA). That is, a sequence containing the entire loop region of the CAB box from a natural scaRNA, AluRNA, or htrRNA derived from the CAB box is directly introduced. In some embodiments, the sequence containing the entire loop region of the CAB box from the natural scaRNA, AluRNA, or htrRNA is inserted into the loop region of the first hairpin structure, or into the loop region of the second hairpin structure, or simultaneously into the loop regions of both the first and second hairpin structures in the gsnoRNA.

[0050] In such implementations, while maintaining the basic structure of the gsnoRNA, the sequence containing the entire loop region of the CAB box can be inserted between any two nucleotides that meet the above-described positions.

[0051] In a third embodiment, in the engineered guide small nucleolar RNA (gsnoRNA), the entire or part of the loop region sequence of the gsnoRNA scaffold sequence may be replaced with a sequence containing the entire loop region of the CAB box. In some embodiments, the loop region of the first hairpin structure in the gsnoRNA, or the loop region of the second hairpin structure, or both loop regions of the first and second hairpin structures, may be replaced with a sequence of the entire loop region containing the CAB box of the natural scaRNA, AluRNA, or htrRNA.

[0052] In such implementations, the entire loop region sequence of the first hairpin structure of the gsnoRNA scaffold sequence is replaced with the sequence of the entire loop region containing the CAB box.

[0053] In some embodiments, the natural scaRNA is selected from scaRNA1, scaRNA4, scaRNA5, scaRNA6, scaRNA8, scaRNA11, scaRNA12, scaRNA13, scaRNA14, scaRNA15, scaRNA16, scaRNA18, scaRNA20, scaRNA21, scaRNA22, scaRNA23, scaRNA26, scaRNA85 and / or scaRNA27.

[0054] In some embodiments, the natural AluRNA is selected from AluACA2, AluACA5, AluACA7, AluACA8, AluACA9, AluACA13, AluACA15, AluACA17, AluACA21, AluACA24, AluACA43, AluACA48, AluACA91, AluACA97, AluACA177, AluACA208, AluACA214 and / or AluACA303.

[0055] In some implementations, the natural htrRNA is human telomerase RNA.

[0056] Guide sequence

[0057] The guide sequence is the part of the gsnoRNA responsible for target recognition. It consists of a nucleotide sequence complementary to a specific region in the target RNA (such as a PTC sequence containing the target uridine residue). The guide sequence pairs complementary bases with the target mRNA, allowing the uridine residue U (e.g., the uridine residue U in the PTC) to enter the pseudouridine pocket formed by the stem-loop structure of the gsnoRNA. The uridine residue U in the pseudouridine pocket can then enter the catalytic center of the DKC1 enzyme, thereby achieving editing.

[0058] In some embodiments, the guiding sequence is located in the stem region of the first hairpin structure, or in the stem region of the second hairpin structure, or in the stem regions of both the first and second hairpin structures.

[0059] In some embodiments, the gsnoRNA comprises a first guide sequence and a second guide sequence, wherein the first guide sequence is located in the stem region of the first hairpin structure and the second guide sequence is located in the stem region of the second hairpin structure.

[0060] In some embodiments, the gsnoRNA comprises, from the 5' end to the 3' end, the following sequence linked by a scaffold sequence: a first portion of a first guide sequence, a loop region of a CAB box or stem-loop structure derived from natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA), a second portion of the first guide sequence, a first portion of the second guide sequence, and a second portion of the second guide sequence.

[0061] In some embodiments, the guide sequence in each hairpin structure is divided into a first part sequence and a second part sequence, with the first part sequence located on one side of the 5' end of its stem and the second part sequence located on one side of the 3' end of its stem. That is, in a stem-loop structure of the gsnoRNA, from the 5' end to the 3' end, it sequentially includes, connected by a scaffold sequence, the first part sequence of the first guide sequence, the loop region of the CAB box or stem-loop structure derived from natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA), and the second part sequence of the first guide sequence.

[0062] In some embodiments, the guide sequence is not limited to a specific length and sequence, as long as it can hybridize with a sequence containing a target uridine residue (U) of PTC in the target RNA. In some embodiments, the guide sequence is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nt or longer.

[0063] In some implementations, the guide sequence is fully or highly complementary to a specific region of the target RNA containing the target uridine residue (such as a PTC sequence). For example, the number of bases complementary to this specific region in the guide sequence is at least 70%, at least 80%, at least 90%, or 100% of the total number of bases in the guide sequence to ensure hybridization specificity and efficiency.

[0064] In some implementations, the method of this application is highly specific and does not lead to pseudouridineization and readthrough of normal stop codons.

[0065] In some embodiments, the naturally occurring H / ACA-type snoRNA is selected from the group consisting of: ACA19, ACA2b, ACA36, ACA24, ACA5, ACA14a, ACA13, ACA20, ACA44, ACA27, E2, ACA3, and ACA17.

[0066] In some embodiments, the naturally occurring H / ACA-type scaRNA is selected from the group consisting of scaRNA11 and scaRNA15.

[0067] In some embodiments, the gsnoRNA comprises a nucleotide sequence selected from any one of SEQ ID NO: 2-25, SEQ ID NO: 28-33, SEQ ID NO: 68-72, SEQ ID NO: 73-78 or SEQ ID NO: 50-61.

[0068] CTE components

[0069] A CTE element is an RNA element, also known as a constitutive transport element. Some studies suggest that CTE sequences can interact with helicases, allowing the target RNA to open its complex secondary structures, thereby enhancing accessibility. In some embodiments, the CTE element can enhance the binding of the gsnoRNA to the target RNA.

[0070] In some embodiments, the CTE element is attached to the 5' end of the scaffold sequence of the gsnoRNA, or to the 3' end of the scaffold sequence of the gsnoRNA, or to both the 5' and 3' ends of the scaffold sequence of the gsnoRNA.

[0071] In some embodiments, the CTE element is a CTE element derived from a retrovirus or a truncated version thereof, wherein the truncated version retains or partially retains the function or activity of its derived CTE element. In some embodiments, the CTE element is a CTE element derived from Mason Fischer monkey virus (MPMV) or simian D retrovirus (SRV) or a truncated version thereof, wherein the truncated version retains or partially retains the function or activity of its derived CTE element.

[0072] In some implementations, the truncated body retains or partially retains the secondary structure of the CTE element from which it originates.

[0073] In some implementations, to reduce the molecular weight of gsnoRNA and improve the efficiency of direct delivery of RNA molecules to the host cell, a truncated form of a CTE element is attached to the gsnoRNA.

[0074] In some embodiments, the sequence of the CTE element is as shown in SEQ ID NO:79~88.

[0075] In some embodiments, the gsnoRNA comprises, from the 5' end to the 3' end, the following sequence linked by a scaffold sequence: a first portion of a first guide sequence, a CAB box derived from natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA) or a loop region containing the CAB box, a second portion of the first guide sequence, a first portion of the second guide sequence, a second portion of the second guide sequence, and a CTE element.

[0076] In some implementations, the CTE element is connected to the stent sequence via a linker.

[0077] In some implementations, the linker is 4-12 bp in length (e.g., 4 bp, 5 bp, 6 bp, 7 bp, 8 bp, 9 bp, 10 bp).

[0078] In some implementations, the CTE element is linked to the 3' end of the scaffold sequence of the gsnoRNA and is linked via a 4-6 bp or 6-8 bp linker.

[0079] In some embodiments, the gsnoRNA includes those selected from SEQ ID NO:2~25, SEQ ID NO:28~33, SEQ ID NO:68~72, SEQ ID NO:73~78 or SEQ ID NO:50~61.

[0080] DKC1 protein

[0081] In some embodiments, the gsnoRNA recruits core proteins (DKC1, NHP2, GAR1, and NOP10) and forms snoRNA with the core proteins to modify the target uridine residues in the target RNA into pseudouridine residues (Ψ). In some embodiments, the gsnoRNA recruits the DKC1 protein to modify the target uridine residues in the target RNA into pseudouridine residues (Ψ).

[0082] It is understandable that the DKC1 protein can be the endogenously expressed DKC1 iso1 and / or DKC1 iso3, or the exogenously expressed DKC1 iso1 and / or DKC1 iso3.

[0083] In some embodiments, gsnoRNA is introduced into the host cell, which hybridizes with the target RNA and recruits the DKC1 protein, and modifies the target uridine residue (U) of the PTC contained in the target RNA to a pseudouridine residue (Ψ).

[0084] In some embodiments, the DKC1 protein recruited by the gsnoRNA comprises: endogenous DKC1 protein of the host cell, and / or exogenous DKC1 protein of the host cell.

[0085] In some embodiments, the method further includes introducing a nucleic acid molecule encoding the DKC1 protein into the host cell.

[0086] In some embodiments, the DKC1 protein is overexpressed in the host cells.

[0087] In some embodiments, the DKC1 protein is a naturally occurring DKC1 subtype that has cytoplasmic localization in the host cell.

[0088] In some embodiments, the DKC1 is selected from human DKC1 protein iso1, human DKC1 protein iso3, or any combination thereof.

[0089] In some embodiments, the amino acid sequence of the human DKC1 protein iso3 (DKC1 iso3) is shown in SEQ ID NO: 66.

[0090] In some embodiments, the amino acid sequence of the human DKC1 protein iso1 (DKC1 iso1) is shown in SEQ ID NO: 67.

[0091] In some embodiments, the DKC1 protein comprises an amino acid sequence having at least 85%, or at least 85%, or at least 88%, or at least 90%, or at least 91%, or at least 92%, or at least 93%, or at least 94%, or at least 95%, or at least 96%, or at least 97%, or at least 98%, or at least 99%, or at least 99.1%, or at least 99.2%, or at least 99.3%, or at least 99.4%, or at least 99.5% identity with SEQ ID NO: 66 or 67.

[0092] Based on codon degeneracy known in the art, in some embodiments, the nucleotide sequence encoding the DKC1 protein (e.g., DKC1 iso1 or DKC1 iso3) can be substituted according to codon degeneracy. In some embodiments, the nucleotide sequence encoding the DKC1 protein (e.g., DKC1 iso1 or DKC1 iso3) is codon-optimized.

[0093] nc-tRNA

[0094] nc-tRNA, or near-cognate tRNA, has an anticodon that has a single base mismatch or wobble pair with a Ψ-modified codon in mRNA (such as ΨAA, ΨAG, ΨGA) (the remaining positions are Watson-Crick pairs). It recognizes and decodes the codon into a specific amino acid through non-canonical base pairing.

[0095] Decoding the Ψ-modified codon can utilize either endogenous or exogenous nc-tRNA. For example, the nucleotide sequence encoding nc-tRNA can be delivered or the nc-tRNA molecule can be delivered directly, utilizing exogenous nc-tRNA to achieve pseudouridine modification. In some embodiments, the method further includes introducing the nucleotide sequence encoding nc-tRNA or the nc-tRNA into the host cell.

[0096] In some embodiments, the method further includes:

[0097] The near-homologous transfer RNA (nc-tRNA) of the PTC or a nucleic acid molecule for expressing the nc-tRNA is introduced into the host cell.

[0098] In some embodiments, the nucleic acid molecule expressing the nc-tRNA contains 1, 2, 3, 4 or more copies of the nucleotide sequence encoding the nc-tRNA.

[0099] In some implementations, nc-tRNA is a mature tRNA. In some implementations, nc-tRNA is a complete tRNA containing introns.

[0100] In some embodiments, the target uridine residue in the PTC of the target RNA is modified into a pseudouridine residue to provide a Ψ-modified PTC. Then, the nc-tRNA of the Ψ-modified PTC or a nucleic acid molecule for expressing the nc-tRNA is introduced into the host cell to decode the PTC into amino acids, thereby inhibiting the PTC.

[0101] In some implementations, the nc-tRNA is modified.

[0102] In some embodiments, the modification comprises one or more modifications of mature tRNA. In some embodiments, the modification comprises one or more modifications selected from the group consisting of psiU, m1A, m1G, and m5C. In some embodiments, the modified tRNA significantly increases the readthrough efficiency of gsnoRNA compared to unmodified tRNA of the same type. In some embodiments, the modified nctRNA contains 5' phosphorylation. In some embodiments, the modified tRNA is a mature tRNA without introns.

[0103] In some embodiments, the nc-tRNA sequences are shown in SEQ ID NO:62~65 and SEQ ID NO:89~172.

[0104] In some embodiments, PTC is a UGA codon, wherein the Ψ-modified PTC is decoded to arginine or tryptophan. In some embodiments, the Ψ-modified PTC is decoded to arginine at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, or at least 95% of the time.

[0105] In some embodiments, PTC is a UAG codon, wherein Ψ-modified PTC is decoded as glutamine, leucine, or tyrosine. In some embodiments, nc-tRNA is tRNA-Q-CUG, tRNA-L-CAA, tRNA-Y-AUA, and / or tRNA-Y-GUA.

[0106] In some embodiments, PTC is a UAA codon, where Ψ-modified PTC is decoded as glutamine or tyrosine. In some embodiments, nc-tRNA is tRNA-Q-UUG, tRNA-Y-AUA, and / or tRNA-Y-GUA.

[0107] In some embodiments, the PTC is caused by a nonsense mutation in the codon encoding arginine. In some embodiments, the PTC is the UGA codon, and the nc-tRNA is tRNA R UCU, which decodes the Ψ-modified PTC into arginine.

[0108] In some embodiments, the PTC is caused by a nonsense mutation in the codon encoding glutamine. In some embodiments, the PTC is caused by a nonsense mutation in the codon encoding glutamine, and the nc-tRNA is tRNA-Q-CUG or tRNA-Q-UUG, which decodes the Ψ-modified PTC into glutamine.

[0109] The twenty common amino acids referred to herein are written in accordance with conventional usage. See, for example, Immunology-ASynthesis (2nd Edition, ES Golub and DR Gren, Eds., Sinauer Associates, Sunderland, Mass. (1991)), which is incorporated herein by reference. In this invention, the terms “polypeptide” and “protein” have the same meaning and are used interchangeably. Furthermore, in this invention, amino acids are generally represented by single-letter and three-letter abbreviations known in the art. For example, alanine can be represented by A or Ala.

[0110] gsnoRNA

[0111] On the other hand, this application provides an engineered guide small nucleolar RNA (gsnoRNA), characterized in that the gsnoRNA comprises:

[0112] (i) at least one guide sequence, (ii) at least one CAB box and / or CTE (constitutive transport element) element, and (iii) a stent sequence; wherein,

[0113] The guide sequence hybridizes with the target RNA containing the target uridine residue (U) of PTC;

[0114] The CAB box is derived from the loop region of a stem-loop structure of natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA).

[0115] The scaffold sequence is derived from natural H / ACA type snoRNA and / or natural H / ACA type scaRNA.

[0116] In some embodiments, the scaffold sequence is not limited to a specific length and sequence, as long as it retains the ability of snoRNA or scaRNA to interact with the DKC1 protein. In some embodiments, the length of the scaffold sequence is 30-50 nt, 50-80 nt, 80-100 nt, 100-130 nt, 130-150 nt, 150-200 nt, or longer.

[0117] In some embodiments, the scaffold sequence is a complete snoRNA derived from a natural H / ACA type structure (e.g., containing two hairpin structures, an H box, and an ACA box) or a portion thereof (e.g., containing one hairpin structure, an H box, and an ACA box).

[0118] In some embodiments, the scaffold sequence is a complete scaRNA derived from a natural H / ACA type structure (e.g., containing two hairpin structures, an H box, and an ACA box) or a portion thereof (e.g., containing one hairpin structure, an H box, and an ACA box).

[0119] In some implementations, the gsnoRNA comprises two stem-loop structures, which, from 5' to 3', sequentially include: a first stem near the 5' end of the stem-loop structure, a first loop near the 5' end of the stem-loop structure, a hinge structure (containing an "H box"), a second stem near the 3' end of the stem-loop structure, a second loop near the 3' end of the stem-loop structure, and a tail structure (containing an "ACA box"). The guide sequence may be located in the first stem and / or the second stem, and the CAB box may be located in the first loop and / or the second loop.

[0120] In some embodiments, to improve in vivo stability and delivery efficiency, the gsnoRNA may contain one or more modified (e.g., chemically modified) nucleotides. In some embodiments, the ribose moiety of one or more nucleotides of the gsnoRNA contains a 2′-O-methyl (2′-OMe) modification. In some embodiments, the gsnoRNA contains one or more phosphate thioester (PS) nucleotide internucleotide bonds. In some embodiments, the gsnoRNA contains a 5′ cap modification (e.g., an m7G cap modification). In some embodiments, the modifications are concentrated at the 5′ and 3′ ends of the gsnoRNA to resist nuclease degradation without significantly affecting function.

[0121] In some embodiments, the CAB box has the following sequence: X1X2AG; wherein X1 and X2 are each independently selected from any one of A, U, C, G.

[0122] In some embodiments, the support sequence includes a first hairpin structure near the 5' end and a second hairpin structure near the 3' end.

[0123] Furthermore, the CAB box is located in the loop area of ​​the first hairpin structure, or in the loop area of ​​the second hairpin structure, or simultaneously in the loop areas of both the first and second hairpin structures.

[0124] In some embodiments, the guiding sequence includes a stem region located in the first hairpin structure, or in the stem region of the second hairpin structure, or in the stem regions of both the first and second hairpin structures.

[0125] In some embodiments, the CTE element is a CTE element or a truncated form of a masonfish virus (MPMV) or a simian D retrovirus (SRV).

[0126] In some embodiments, the gsnoRNA comprises, from the 5' end to the 3' end, the following sequence linked by a scaffold sequence: a first portion of a first guide sequence, a CAB box derived from natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA) or a loop region containing the CAB box, a second portion of the first guide sequence, a first portion of the second guide sequence, a second portion of the second guide sequence, and a CTE element.

[0127] In some embodiments, the CAB box has the features described in the first aspect.

[0128] In some implementations, the bootstrap sequence has the features described in the first aspect.

[0129] In some embodiments, the stent sequence has the features described in the first aspect.

[0130] In some embodiments, the CTE element has the features described in the first aspect.

[0131] isolated nucleic acid molecules

[0132] On the other hand, this application provides an isolated nucleic acid molecule, characterized in that the isolated nucleic acid molecule contains a nucleic acid sequence for expressing gsnoRNA as described above.

[0133] In some implementations, the isolated nucleic acid molecule is DNA.

[0134] Composition

[0135] In another aspect, this application provides a composition characterized in that the composition comprises: gsnoRNA as described above or isolated nucleic acid molecules as described above.

[0136] In some embodiments, the composition comprises: gsnoRNA as described above or isolated nucleic acid molecules as described above; and a near-homologous transfer RNA (nc-tRNA) of PTC contained in the target sequence or nucleic acid for expressing the nc-tRNA.

[0137] In some embodiments, the composition comprises: gsnoRNA as described above or isolated nucleic acid molecules as described above; and DKC1 protein or nucleic acid molecules encoding DKC1 protein.

[0138] In some embodiments, the composition comprises: gsnoRNA as described above or isolated nucleic acid molecules as described above, near-homologous transfer RNA (nc-tRNA) of PTC contained in the target sequence or nucleic acid for expressing the nc-tRNA, and DKC1 protein or nucleic acid molecules encoding DKC1 protein.

[0139] In some embodiments, the composition comprises: the isolated nucleic acid molecule as described above, the nucleic acid for expressing the nc-tRNA, and the nucleic acid molecule encoding the DKC1 protein, wherein the nucleic acid or nucleic acid molecule is present in the same or different vectors.

[0140] In some embodiments, the DKC1 protein has the characteristics described in the first aspect.

[0141] In some embodiments, the nc-tRNA has the characteristics described in the first aspect.

[0142] Delivery and delivery composition

[0143] The gsnoRNA described in this article can be directly delivered into host cells after in vitro transcription and synthesis, and can be delivered by any method known in the art.

[0144] Such methods include, but are not limited to, electroporation, lipid transfection, nuclear transfection, microinjection, acoustic hole effect, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, magnetic transfection, puncture transfection, optical transfection, reagent-enhanced nucleic acid uptake, and delivery via liposomes, immunoliposomes, viral particles, artificial viruses, etc.

[0145] Therefore, in another aspect, this application provides a delivery composition comprising a delivery vector and one or more selected from the following: gsnoRNA as described above, isolated nucleic acid molecules as described above, or compositions as described above.

[0146] In some implementations, the delivery carrier is a particle.

[0147] In some embodiments, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns, or viral vectors (e.g., replication-defective retroviruses, lentiviruses, adenoviruses, or adeno-associated viruses).

[0148] In some embodiments, the delivery composition further comprises a pharmaceutically acceptable carrier and / or excipient.

[0149] As used herein, the term “pharmaceutically acceptable carrier and / or excipient” means a carrier and / or excipient that is pharmacologically and / or physiologically compatible with the subject and the active ingredient, which is well known in the art and includes, but is not limited to: pH adjusters, surfactants, adjuvants, ionic strength enhancers, diluents, agents for maintaining osmotic pressure, agents for delaying absorption, and preservatives.

[0150] host cells

[0151] In another aspect, this application provides a host cell characterized in that the host cell contains the gsnoRNA as described above, or the isolated nucleic acid molecule as described above, or the composition as described above, or the delivery composition as described above.

[0152] Such host cells include, but are not limited to, prokaryotic cells such as bacterial cells (e.g., Escherichia coli cells), eukaryotic cells such as fungal cells (e.g., yeast cells), insect cells, plant cells, and animal cells (e.g., mammalian cells, such as mouse cells, human cells, etc.).

[0153] In some embodiments, the host cell is a prokaryotic cell or a eukaryotic cell.

[0154] In some implementations, the host cell is a mammalian cell (e.g., a human cell).

[0155] In some implementations, the host cell contains a PTC (early stop codon) mutant gene.

[0156] Preparation method

[0157] In another aspect, this application provides a method for preparing gsnoRNA or the composition or the delivery composition as described above, characterized in that the method comprises culturing host cells as described above under conditions that allow nucleic acid and protein expression, and recovering the gsnoRNA or the composition or the delivery composition from the cultured host cell culture.

[0158] use

[0159] On the other hand, this application provides the use of the gsnoRNA as described above, or the isolated nucleic acid molecule as described above, or the composition as described above, or the delivery composition as described above, or the host cell as described above, in the preparation of a formulation for editing target RNA or for inhibiting premature stop codons (PTCs) in target RNA in a host cell.

[0160] In some embodiments, the formulation is used to modify uridine residues (U) in the target RNA into pseudouridine residues (Ψ).

[0161] In some embodiments, the host cell is a mammalian cell. In some embodiments, the host cell is a human cell.

[0162] On the other hand, this application provides the use of the gsnoRNA as described above, or the isolated nucleic acid molecule as described above, or the composition as described above, or the host cell as described above, or the delivery composition as described above in the preparation of a pharmaceutical product, characterized in that the pharmaceutical product is used to treat a subject for a disease and / or symptoms caused or resulting from a PTC mutation.

[0163] In some embodiments, the disease is selected from cystic fibrosis, spinal muscular atrophy, fructose intolerance, dilated cardiomyopathy, Heller syndrome, alpha-1-antitrypsin (A1AT) deficiency, Parkinson's disease, Alzheimer's disease, albinism, amyotrophic lateral sclerosis, asthma, β-thalassemia, Cadasil syndrome, and Charcot-Marie-Tooth disease. Diseases including chronic obstructive pulmonary disease (COPD), distal spinal muscular atrophy (DSMA), Duchenne muscular dystrophy, dystrophic epidermolysis bullosa, epidermolysis bullosa, Fabry disease, Leiden factor V-related disorders, familial adenomatous polyposis, galactosemia, Gaucher disease, glucose-6-phosphate dehydrogenase, hemophilia, hereditary hemochromatosis, Huntington's disease, inflammatory bowel disease (IBD), hereditary polyagglutination syndrome, Lesch-Nair syndrome, Lynch syndrome, Marfan syndrome, mucopolysaccharidosis, muscular dystrophy, and type I and II myotonic dystrophy. Neurofibromatosis, Niemann-Pick disease types A, B, and C, NY-esol-related cancers, Boytz-Yage syndrome, phenylketonuria, Pomper's disease, primary ciliary body disease, prothrombin mutation-related diseases (such as prothrombin G20210A mutation), pulmonary hypertension, (autosomal dominant) retinitis pigmentosa, Sandhoff's disease, severe combined immunodeficiency syndrome (SCID), sickle cell anemia, Staggart disease, Tay-Sachs disease, Usher syndrome, X-linked immunodeficiency, Sturge-Weber syndrome, or any combination thereof.

[0164] In some embodiments, the subject is a mammal. In some embodiments, the subject is a human, a cynomolgus monkey, or a mouse.

[0165] In some embodiments, the disease is fructose intolerance. In some embodiments, the PTC mutation is a nonsense mutation in the nucleotide encoding amino acid 148 of ALDOB (fructose diphosphate aldolase B).

[0166] In some embodiments, the disease is cystic fibrosis. In some embodiments, the PTC mutation is a nonsense mutation in the nucleotide encoding amino acid 553 of the CFTR (cystic fibrosis transmembrane conduction regulator).

[0167] In some embodiments, the disease is dilated cardiomyopathy. In some embodiments, the PTC mutation is a nonsense mutation in the nucleotide encoding amino acid 225 of LMNA (lamin A / C).

[0168] method

[0169] On the other hand, this application provides a method for editing target RNA in vitro or in vivo, characterized in that the method comprises, under conditions suitable for editing target RNA, contacting the target RNA with one or more of the following: gsnoRNA as described above, isolated nucleic acid molecules as described above, or a combination as described above, or a delivery composition as described above, thereby editing the target RNA.

[0170] In some embodiments, the method modifies the uridine residue (U) in the target RNA with a pseudouridine residue (Ψ).

[0171] In some embodiments, the method modifies one or more uridine residues in the target RNA with pseudouridine residues.

[0172] It should be understood that this application provides a universally applicable method for editing target RNA, which can be used to edit any uridine residues of target RNA in vitro or in vivo. By designing a suitable guide sequence for gsnoRNA, the desired uridine residues to be edited are introduced into a pseudouridine-enhanced pocket formed by gsnoRNA. Under conditions suitable for pseudouridine synthase catalysis, the pseudouridine synthase is contacted with the target RNA and gsnoRNA. The uridine residues located in the pseudouridine-enhanced pocket can then enter the catalytic center of DKC1, achieving editing. In some embodiments, the method does not involve the use of tRNA (e.g., nc-tRNA).

[0173] In some implementations, the target RNA is edited to suppress premature stop codons (PTCs) in the target RNA within the host cell.

[0174] Terminology Definition

[0175] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.

[0176] As used herein, the term “guide small nucleolar RNA (gsnoRNA)” refers to an engineered non-coding RNA molecule. gsnoRNA can specifically hybridize with specific sequences in target RNA (such as target uridine residues in premature stop codons PTC) and recruit pseudouridine synthase complexes (e.g., complexes containing DKC1 protein) to the target site, thereby catalyzing the conversion of target uridine residues into pseudouridine residues (Ψ).

[0177] As used herein, the term "guide sequence" refers to the portion of gsnoRNA responsible for target recognition, consisting of a nucleotide sequence complementary to a specific region of the target RNA (such as a PTC sequence containing target uridine residues). This sequence hybridizes to the specific region of the target RNA via Watson-Crick base pairing.

[0178] As used herein, the term "CAB box (Cajal-body localization box) element" is a short RNA motif capable of localizing to the Cajal body. Its sequence is a conserved sequence "X1X2AG," where the third and fourth AG positions are highly conserved, while the first two can vary. For example, conserved CAB box elements can be AAAG, GAAG, UAAG, UGAG, UCAG, CGAG, AUAG, GCAG, CUAG, CAAG, AGAG, etc. CAB boxes are generally found in scaRNAs, but can also be found in AluRNAs and HTRs (Human Teleomerase RNA).

[0179] As used herein, the term "CTE element" refers to an RNA element, also known as a constitutive transport enhancer. Some studies suggest that CTE sequences can interact with helicases, allowing the target RNA to open its complex secondary structures, thereby enhancing accessibility. CTE elements are present in Mason Fischer monkey virus (MPMV) or simian D retrovirus (SRV), and functional homologous modules have also been found in Ross sarcoma virus (RSV) and avian leukosis virus (ALV). In some embodiments, the CTE element used herein is the aforementioned functional homologous module.

[0180] As used in this article, the term "snoRNA" refers to nucleolar small RNA, a type of non-coding RNA. Based on their structure, they can be divided into three types: C / D type, H / ACA type, and MRP RNA type. C / D type snoRNA can guide the 2'-O-ribose methylation modification of rRNA precursors, while H / ACA type snoRNA can guide the pseudouracilization modification of rRNA precursors.

[0181] As used herein, the term "scaRNA" refers to small Cajal body-specific RNA, a type of non-coding RNA. Structurally equivalent to snoRNA (C / D or H / ACA type), it additionally carries the Cajal body localization signal CAB box. C / D type scaRNAs can guide 2'-O-ribose methylation modification of rRNA precursors, while H / ACA type scaRNAs can guide pseudouracilization modification of rRNA precursors.

[0182] As used herein, the term "H / ACA-type structure" refers to a secondary structure of a small RNA. Specifically, the H / ACA-type structure contains two hairpin structures, each with a conventional stem structure and a top loop structure, also known as a "stem-loop structure." From the 5' end to the 3' end, it sequentially includes a first hairpin structure, a hinge structure (containing an "H-box"), a second hairpin structure, and a tail structure (containing an "ACA-box"). Each hairpin structure forms a pseudouridine pocket, allowing a uridine residue U to enter the catalytic center. The H / ACA structure itself determines the recruitment of the DKC1 protein. The DKC1, NOP10, NHP2, and GAR1 quaternary complex is recruited to the H / ACA structure, forming an RNP complex and catalyzing the modification of the uridine residue U into a pseudouridine residue.

[0183] As used herein, the term "scaffold sequence" refers to a sequence in gsnoRNA other than the guide sequence and the CAB box and / or CTE (constitutive transport element). In some embodiments, the scaffold sequence is derived from snoRNA with a natural H / ACA structure. In some embodiments, the scaffold sequence is derived from scaRNA with a natural H / ACA structure. Therefore, in gsnoRNA, the scaffold sequence is responsible for maintaining the secondary structure of the RNA (such as hairpin conformation) and mediating interactions with core proteins (such as DKC1, NOP10, NHP2, GAR1). In some embodiments, the scaffold sequence comprises, from the 5' end to the 3' end, a first hairpin structure, a hinge structure (containing an "H box"), a second hairpin structure, and a tail structure (containing an "ACA box"). In some embodiments, the scaffold sequence is used to connect the CAB box, the guide sequence, and / or the CTE element. For example, in some embodiments, the CAB box and the guide sequence are located in the first hairpin structure of the gsnoRNA, and the CTE element is located at the 3' end of the scaffold sequence. The CAB box, the guide sequence and the CTE element are connected together through the scaffold sequence and constitute the gsnoRNA of this application.

[0184] As used in this article, the term "premature termination codon (PTC)" refers to a premature termination signal formed in the coding region of mRNA due to a nonsense mutation, leading to premature termination of translation and the production of truncated and often nonfunctional proteins. PTC is caused by point mutations in sense codons and is associated with a variety of genetic diseases, such as cystic fibrosis, Heller syndrome, and Duchenne muscular dystrophy. Therefore, PTC is a core target for therapeutic intervention in many diseases.

[0185] In this paper, the specific uridine residue in the PTC targeted by the guide sequence is referred to as the "target uridine residue". When the target uridine residue is modified by pseudouridineization, it is transformed into a "pseudouridine residue" (Ψ), forming the ΨAA, ΨAG, or ΨGA codons.

[0186] As used herein, the term “nearly homologous tRNA (nc-tRNA)” refers to a tRNA molecule whose anticodon has a single base mismatch or wobble pair with a Ψ-modified codon in mRNA (such as ΨAA, ΨAG, ΨGA) (the remaining positions are Watson-Crick pairs), and that the anticodon is recognized and decoded as a specific amino acid through non-canonical base pairing.

[0187] By decoding nc-tRNA, the premature stop codon is recoded as Arg, Gln, or Trp / Tyr, depending on the stop codon sequence and the type of nc-tRNA. This maintains or partially maintains protein folding and function, achieving full-length protein restoration.

[0188] For example, for the ΨGA codon, the nc-tRNA anticodon UCU (e.g., tRNA-R-UCU) is decoded as arginine (R). For example, for the ΨAG codon, the nc-tRNA anticodon CUG (e.g., tRNA-Q-CUG) is decoded as glutamine (Q). For example, for the ΨAA codon, the nc-tRNA anticodon UUG (e.g., tRNA-Q-UUG) is decoded as glutamine (Q).

[0189] As used herein, the term "DKC1 protein" is a highly conserved pseudouridine synthase. It is responsible for catalyzing the conversion of uridine residues in RNA to pseudouridine residues. In this invention, the DKC1 protein is recruited by gsnoRNA to a target RNA site to perform specific pseudouridine modification.

[0190] As used herein, the term "DKC1 iso1" refers to an isotype of the DKC1 protein, which is primarily located in the cell nucleus. Its accession number is DKC1 iso1: NP_001354.1.

[0191] As used herein, the term "DKC1 iso3" is an isotype of the DKC1 protein. Specifically, it is a splicing variant resulting from the retention of intron 12. This process leads to the absence of the C-terminal nuclear localization signal (NLS), thus DKC1 iso3 is primarily localized in the cytoplasm rather than the nucleus. In native endogenous mRNA expression, the expression level of DKC1 iso1 is much higher than that of DKC1 iso3. The accession number for DKC1 iso3 is NP_001275676.1.

[0192] Beneficial effects of the invention

[0193] This application optimizes and modifies the previously studied RESTART pseudouridine modification system, providing a modified gsnoRNA that can efficiently bind to target sites and improve the recruitment and assembly rates of pseudouridine synthase DKC1 and the core proteins NHP2, GAR1, and NOP10, thereby significantly improving the efficiency of pseudouridine modification. Furthermore, this modified gsnoRNA is suitable for various iterations of the RESTART system, efficiently recruiting both endogenously synthesized and exogenously expressed DKC1, as well as different DKC1 iso3 and DKC1 iso1. Moreover, this modified gsnoRNA can also be used in combination with other elements in the RESTART system (e.g., nc-tRNA) to significantly enhance the efficiency of pseudouridine modification.

[0194] For example, when applied to modify uridine residues in PTC mutations, systems containing this gsnoRNA can significantly improve the efficiency of pseudouridine modification, thereby restoring the coding information of the mRNA, enabling the PTC site to be decoded, significantly improving the readthrough efficiency of PTC, and thus expressing the complete protein. Therefore, this modified gsnoRNA has important significance and broad application prospects in the treatment of PTC mutation-related diseases.

[0195] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples. However, those skilled in the art will understand that the following drawings and examples are for illustrative purposes only and are not intended to limit the scope of the invention. Various objects and advantages of the present invention will become apparent to those skilled in the art from the following detailed description of the drawings and preferred embodiments. Attached Figure Description

[0196] Figure 1 This is a schematic diagram of one modification of the gsnoRNA in this application. The scaffold sequence of the gsnoRNA (blue) contains two stem-loop structures: a first hairpin structure near the 5' end and a second hairpin structure near the 3' end. From 5' to 3', the scaffold sequence sequentially includes: a first hairpin structure near the 5' stem-loop structure (composed of a first stem and a first loop), a hinge structure (containing an "H box"), a second hairpin structure near the 3' stem-loop structure (composed of a second stem and a second loop), and a tail structure (containing an "ACA box").

[0197] The guide sequence (green) can be located in the first and / or second stem. The CAB box (red) can be located in the first and / or second loop. The CTE sequence (orange) can be attached to the 5' and / or 3' end of the gsnoRNA scaffold sequence.

[0198] In a specific implementation, as shown in the figure, the first guide sequence is located in the first stem, the second guide sequence is located in the second stem, the CAB box is located in the first loop, and the CTE sequence is attached to the 3' end of the gsnoRNA scaffold sequence.

[0199] Figure 2 The mechanism by which the method of this application suppresses PTC is illustrated schematically.

[0200] Figure 2 As shown in Figure a, when a PTC is present in mRNA, mRNA translation terminates prematurely, resulting in a truncated protein. When using the method of this application, the guide sequence in the gsnoRNA pairs complementaryly with the target mRNA bases, allowing the uridine residue U in the PTC to enter the catalytic center formed by the stem-loop structure of the gsnoRNA. The gsnoRNA recruits DKC1, NOP10, NHP2, and GAR1 proteins to assemble into an RNP complex, catalyzing the modification of the uridine residue U into a pseudouridine residue Ψ. Furthermore, the Ψ-modified PTC can be decoded and translated into amino acids by nc-tRNA, achieving readthrough of the PTC site and restoring the production of the full-length protein. In some embodiments, the gsnoRNA of this application is delivered alone to achieve PTC readthrough, while other elements required for modification and decoding (e.g., DKC1 protein, nc-tRNA) are provided solely by endogenous sources. In some embodiments, the gsnoRNA of this application, together with exogenously expressed DKC1 protein and / or exogenously expressed nc-tRNA, co-modifies and decodes the PTC.

[0201] Figure 2 b further demonstrates the binding of gsnoRNA to the PTC target and its recruitment of the RNP complex, as well as the decoding function of nc-tRNA. Specifically, gsnoRNA binds to the PTC target and recruits the RNP complex, modifying the catalytic uridine residue U of the PTC (UGA, UAG, or UAA) to a pseudouridine residue Ψ, resulting in Ψ-modified PTC (ΨGA, ΨAG, or ΨAA). nc-tRNA then binds to the Ψ-modified PTC site and decodes it, translating the Ψ-modified PTC site into amino acids, thereby achieving full-length PTC readout and restoring the expression of the full-length protein.

[0202] Figure 3The introduction of CAB boxes improves the readability of the RESTART system. 3a-3b: Directly using scaRNA as the gsnoRNA backbone in UGA reporter systems with (a) no expression and (b) overexpression of DKC1 iso3 yields significant readability. 3c: In a UGA reporter system with no expression of DKC1 iso3, the RESTART readability before and after directly replacing the loop containing the CAB box of a known scaRNA with gACA19 was tested. 3d: Highly efficient CAB boxes were selected and their impact on readability was tested in a UGA reporter system with overexpression of DKC1 iso3. 3e: The effect of adding CAB boxes to other snoRNA backbones was tested in a UGA reporter system with overexpression of DKC1 iso3. 3f: In a UGA reporter system with no expression of DKC1 iso3, the effects of adding CAB boxes to the 5' stem-loop, the 3' stem-loop, and both ends of the gsnoRNA on readability were compared. 3g. Schematic diagram of the CAB box of scaRNA (U85) and the addition of a functionally inactivating mutant to gsnoRNA. 3h. Readthrough performance test of the CAB box of scaRNA (U85) and the addition of a functionally inactivating mutant to gsnoRNA. 3i. Testing the effects of adding the CAB box on RESTART v1, RESTART v2, and RESTART v3 before and after in a disease scenario with the CFTR-R553X nonsense mutation. 3j. Testing the effects of adding the CAB box on RESTART v1, RESTART v2, and RESTART v3 before and after in a disease scenario with the LMNA-R225X nonsense mutation. 3k. Testing the CAB box effect on other H(ACA)-containing RNAs in a UGA reporter system that does not express DKC1 iso3. HTR is human telomerase RNA; BIO is a specific stem loop result on AluRNA with a CAB box.

[0203] Figure 4The introduction of CTE improves the readability of the RESTART system. 4a, Structures of SRV CTE and mutants (M36 CTE and ΔCTE). 4b-4c, Testing the effects of adding CTE to the 5' and 3' ends of gsnoRNA and linkers of different lengths on RESTART readability in UGA reporter systems with (a) and (b) DKC1 iso3 underexpression. 4d-4e, Comparing the effects of CTE and inactivating mutant CTE on RESTART readability in UGA reporter systems with (a) and (b) DKC1 iso3 underexpression. 4f-4g, Testing the interaction between CTE and inactivating mutant CTE with helicase. 4h, Testing the effects of different CTEs and CTE mutants on RESTART readability in a UGA reporter system with DKC1-iso3 overexpression. Detailed Implementation

[0204] Sequence information

[0205] Information on a portion of the sequence involved in this invention is provided below.

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214] Exemplary gsnoRNAs are shown in SEQ ID NO: 68-72, wherein the guide sequence is denoted by (Xn), where Xn is a sequence of length n nucleotides of nucleotide X, where X is any one of A, U, G, or C, and n is an integer of appropriate length for the guide sequence. In some embodiments, n is 4, 5, 6, 7, 8, 9, 10, 11, or 12. Those skilled in the art will understand that the guide sequence (Xn) can be replaced to make the gsnoRNA target the desired target site.

[0215] In SEQ ID NO: 4~33, SEQ ID NO: 68~72 and SEQ ID NO: 73~78, the underlined sequence is a sequence containing the entire loop region of a CAB box derived from natural scaRNA, AluRNA or htrRNA, and the underlined and italicized sequence is a CAB box sequence.

[0216] The invention will now be described with reference to the following embodiments, which are intended to illustrate the invention (and not limit it). Unless otherwise specified, the experiments and methods described in the embodiments are generally carried out in accordance with conventional methods well known in the art and described in various references.

[0217] Furthermore, unless specific conditions are specified in the examples, conventional conditions or conditions recommended by the manufacturer should be followed. Reagents or instruments whose manufacturers are not specified are all commercially available conventional products. Those skilled in the art will understand that the examples are described by way of illustration and are not intended to limit the scope of protection claimed by the invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.

[0218] Example 1. Preparation of components in the RESTART system

[0219] 1. Plasmid preparation

[0220] The molecular clones constructed in this study mainly include two types: PTCreporter plasmid and DKC1 isoform3 (DKC1 iso3) plasmid using pLenti-CMV-MCS-BSD as vector, and gsnoRNA and nc-tRNA plasmids using Pcg2.0-BFP as vector.

[0221] Plasmids using pLenti-CMV-MCS-BSD and pAAV as vectors were primarily recombined via Gibson sequencing. The DKC1 iso3 gene sequence was amplified from HEK293T cDNA, and the disease reporter gene sequence was obtained from a human cDNA library from Peking University. Nonsense mutations in the disease reporter gene were induced by PCR primers, and the ligation sequence between mCherry and EGFP was adjusted via primer amplification. Specifically, this study used the TransStart FastPfu DNA polymerase kit to target and amplify the target fragment and vector backbone. The target fragment was obtained by agarose gel electrophoresis and purified using a universal DNA purification and recovery kit. After purification, homologous recombination ligation of the fragment and vector backbone was performed using NEBuilder® HiFi DNA Assembly Master Mix, and the ligation product was transformed into Trans T1 competent cells. Plasmids were extracted using the EndoFree Mini Plasmid Kit II and identified by Sanger sequencing.

[0222] The construction of gsnoRNA and nc-tRNA plasmids using Pcg2.0-BFP as vectors mainly involves two steps: primer bridging to construct the target sequence and Golden gate restriction enzyme ligation. First, using the TransStart FastPfu DNApolymerase kit, the designed gsnoRNA or nc-tRNA sequence was constructed via overlapping PCR with four primers (one forward primer and three reverse primers), and Golden gate restriction enzyme ligation sites were constructed at both ends of the sequence. The presence of Golden gate restriction enzyme ligation sites on the Pcg2.0-BFP backbone, along with the Bsd toxin protein in the middle of these ligation sites, prevents empty vector plasmids that fail to be digested from surviving during transformation. The product obtained by primer bridging was recovered using a universal DNA purification and recovery kit. 40 ng of the purified product and 20 ng of Pcg2.0-BFP plasmid were placed together in an enzyme digestion and ligation system containing the cutting enzymes BSMBI and T4 ligase (which also contains DTT and ATP) for reaction. The reaction product was then transformed into Trans T1 competent cells, and the plasmid was extracted for Sanger sequencing identification.

[0223] 2. Cell transfection and data analysis

[0224] HEK293T cells were cultured in DMEM medium containing 10% FBS and 1% penicillin-streptomycin at 37°C and 5% CO2. At cell passage, cells were washed with PBS, digested with 0.25% trypsin, and incubated at 37°C for 2 minutes. The trypsin was then neutralized with FBS-containing medium. After centrifugation at 600 rpm for 3 minutes, cells were counted and plated. The cell line was negative for mycoplasma contamination.

[0225] 20–24 hours before staining, cells were seeded in 24-well plates at a density of 2 × 10⁵ cells / well. The plasmid for transfection was extracted using a mini-prep kit, and the plasmid concentration was quantified using Nanodrop. The target plasmid and target RNA were transfected using Lipofectamine LTXwith PLUS reagent, following the recommended transfection protocol. The medium was changed 24 hours after transfection, and cell function was assessed at 48 or 72 hours.

[0226] To evaluate the efficiency of PTC readthrough in the fluorescence reporter system, cells were imaged using an ImageXpress® Micro 4 high-content imaging system (Molecular Devices LLC, Sunnyvale, CA) 48–72 hours post-transfection. Sixteen images from different sites in the same well were captured under a 10x microscope and then automatically analyzed using MetaXpress software. The percentage of EGFP-positive cells was calculated by dividing the number of EGFP-positive cells in the fluorescence image by the number of BFP / mCherry-positive cells in the corresponding image, and then normalized using data from the positive control. EGFP fluorescence intensity was calculated by multiplying the EGFP intensity per cell by the number of EGFP-positive cells in the fluorescence image, and then normalized using data from the positive control. The mean of the 16 images was considered an independent replicate, with each set of fluorescence analysis data having 2–3 biological replicates. Data are presented as the mean of 2 replicates or the mean ± standard deviation of 3 replicates.

[0227] 3. RNA extraction and modification detection

[0228] Discard the culture medium from the target cells in a clean bench, wash once with PBS, and aspirate dry. Add TRIzolreagent to the cells and aspirate 10 times until homogeneous. Transfer to an EP tube and incubate for 5 minutes. Add Chloroform (one-fifth the volume of TRIzolreagent) and vortex vigorously for 15 seconds. Incubate at room temperature for 15 minutes, then centrifuge at 12,000 rpm for 15 minutes at 4°C. Transfer the supernatant to a clean EP tube, add an equal volume of Isopropanol, mix by inverting, and incubate at -20°C for at least 1 hour. Centrifuge at 12,000 rpm for 30 minutes at 4°C. Discard the supernatant, wash twice with 1 mL of 75% ethanol solution, discard the supernatant again, centrifuge once empty, and incubate at room temperature for 10 minutes with the cap open. Once the RNA precipitate changes from white to clear, dissolve it in a certain amount of RNase-free water and determine the concentration using Nanodrop.

[0229] DNase was added to the target RNA to remove remaining genomic and plasmid fragments. After the reaction, the RNA sample was repurified. A labeling reaction solution was prepared: 85% K₂SO₃ / 15% NaHSO₃ solution was mixed with 100 mM Hydroquinone at a ratio of 100:1. 1 μg of purified RNA was mixed with 50 μL of the reaction solution and reacted at 70°C for 5 hours. After the reaction, the RNA was desalted and purified using a Micro Bio-spin 6 column, and then an equal volume of 1 M Tris-HCl [pH 9.0] was added and reacted at 75°C for 30 minutes. The labeled RNA sample was then purified. Reverse transcription was performed using a Maxima HMinus RT enzyme reaction system to obtain cDNA samples. End-specific PCR primers containing a ~130 nt sequence of the target site were designed. Library adapter sequences were added to the 5′ end of the primers, and specific amplification was performed using NEBNext Q5 Hot Start HiFi PCRMaster Mix. After amplification, PCR amplification was performed using Illumina primers, and sequencing adapter sequences were added to both sides of the target fragment. After the amplification reaction, the target DNA fragment was purified using AMPure XP beads, the concentration was determined, and it was identified using a 4150 microarray. Finally, next-generation sequencing analysis was performed.

[0230] RESTART v1-v3, v3 mini system

[0231] Prior to this, the applicant had developed multiple generations of RESTART systems, which contain different components and can achieve precise pseudouracil modification of PTC sites, promote PTC readthrough and restore full-length functional protein expression, thereby efficiently and specifically repairing nonsense mutations.

[0232] RESTART v1: The RESTART v1 system modifies the pseudouridine pocket of human snoRNA by designing a guide sequence, enabling the modified snoRNA to precisely bind to the target PTC site, resulting in gsnoRNA. Furthermore, this gsnoRNA recruits endogenous DKC1 iso1 to assemble snoRNPs for pseudouridine modification, achieving efficient PTC readthrough. In other words, the core component of the RESTART v1 system is gsnoRNA with a modified guide sequence. By introducing gsnoRNA into cells, endogenous DKC1 iso1 is recruited to perform pseudouridine modification of the PTC site in the target RNA.

[0233] RESTART v2: In the RESTART v2 system, DKC1 iso3 (Iso3) was found to enhance the modification efficiency of the RESTART system. Since DKC1 iso3 is a product of DKC1 aberrant splicing, cells under natural conditions generally express DKC1 iso1, expressing only a small amount of DKC1 iso3. Therefore, the RESTART v2 system, by overexpressing the catalytic enzyme DKC1 iso3, increased the modification efficiency of the RESTART system by 2-fold and the readthrough efficiency by 2–5-fold. In other words, the core components of the RESTART v2 system are gsnoRNA and DKC1 iso3. By introducing gsnoRNA and DKC1 iso3 into cells, pseudouracil modification of the PTC sites in the target RNA is achieved.

[0234] RESTART v3: The RESTART v3 system reveals that nc-tRNAs at PTC sites play a crucial role in decoding pseudouracil-modified PTC sites. Although these nc-tRNAs are naturally present, their expression levels in cells are low or they do not target PTC sites, thus requiring artificial introduction and overexpression. The RESTART v3 system significantly improves the readthrough efficiency of the RESTART system (1.3–8 times higher than RESTART v2) by overexpressing nc-tRNA, and also enhances the accuracy of amino acid incorporation at PTC sites, enabling approximately 50% of disease-related nonsense mutations to be accurately repaired to their original amino acids. In other words, the core components of the RESTART v3 system are gsnoRNA, DKC1 iso3, and nc-tRNA. By introducing gsnoRNA, DKC1 iso3, and nc-tRNA into cells, pseudouracil modification and readthrough of PTC sites in target RNA are achieved.

[0235] RESTART v3-mini: To simplify the RESTART v3 system and improve delivery flexibility, the RESTART v3-mini system delivers only gsnoRNA and nc-tRNA, without additional expression of DKC1 iso3. That is, the core components of the RESTART v3-mini system are gsnoRNA and nc-tRNA. By introducing gsnoRNA and nc-tRNA into the cell, pseudouracil modification and readthrough of the PTC site in the target RNA are performed.

[0236] In summary, it can be seen that the individual components of the RESTART system can function effectively, either individually or in combination. For example, gsnoRNA can be delivered alone to improve PTC readthrough efficiency, while other components required for modification can be provided solely by the cell's endogenous resources (similar to the RESTART v1 system). Similarly, in the presence of gsnoRNA, DKC1 iso3 can be overexpressed alone, or nc-tRNA can be delivered alone. Even without delivering the complete RESTART system, the RESTART system can be assembled in the cell using endogenous components to achieve PTC readthrough.

[0237] Construction and application of the new RESTART system

[0238] Furthermore, this application modifies the gsnoRNA in the previous RESTART system by introducing a CAB / CTE sequence to increase its efficiency in recognizing target sequences and recruiting RNPs, thereby improving its efficiency in ultimately reading nonsense mutations, and using it in the future fourth-generation RESTART system (i.e., RESTART V4). Specifically, in the RESTART system of this application, the gsnoRNA and nc-tRNA used are expressed by the Pcg2.0-BFP (purchased from Addgene) backbone plasmid, while the DKC1 iso3 protein component is expressed by the pLenti-CMV-MCS-BSD backbone plasmid. These constructed vectors can be used to assemble the RESTART system in cells to further evaluate its reading function in disease-related mRNAs containing nonsense mutations. In some embodiments, the constructed vector containing the nucleotide sequences encoding gsnoRNA, nc-tRNA, and DKC1 iso3 protein can be directly transfected into cells to express the desired RNA or protein, thereby assembling a new RESTART system containing the gsnoRNA optimized in this application in cells.

[0239] During the construction process, we introduced disease-related reporter genes carrying early stop codons via co-transfection and evaluated the reading efficiency and functional recovery effect of the RESTART system on target mRNA by using changes in the fluorescence expression level of the reporter system. In other words, a PTC disease model can be constructed by introducing a reporter system into cells, and the aforementioned vector can be introduced to detect the pseudouridine editing efficiency and PTC readthrough recovery effect of the RESTART system.

[0240] Alternatively, after obtaining the components of the system, it can be further constructed into a delivery system suitable for in vivo administration. Specific methods include integrating the system into a lentivirus or adeno-associated virus vector for delivery via in vitro transduction or in vivo injection; or using lipid nanoparticles to encapsulate the system's RNA components for targeted delivery via intravenous or local injection, etc.

[0241] Example 2. Modification of the localization sequence of snoRNA

[0242] Figure 1 The basic structure of the modified gsnoRNA of this application is shown. In this paper, we exemplarily list several sequences of gsnoRNA before modification of this application (the guide sequence is represented by Xn, and the remaining sequences are scaffold sequences), namely ACA19, ACA36, ACA24, ACA5 and ACA14a, and their nucleotide sequences are shown as SEQ ID NO: 68~72, respectively.

[0243] The RNA pseudouridine modification system mainly consists of snoRNA and snoRNPs formed by four core proteins: DKC1, NHP2, GAR1, and NOP10, which perform pseudouridine modification. Figure 2We attempted to improve the efficiency of pseudouridine modification by modifying snoRNA to increase the assembly rate of snoRNP. The sequences of the aforementioned core proteins can be found in DKC1 (DKC1 iso1: NP_001354.1 or DKC1 iso3: NP_001275676.1), NHP2: NP_060308.1, GAR1: NP_061856.1, and NOP10: NP_061118.1. We introduced CAB box elements from the hairpin loop present on natural scaRNAs (e.g., scaRNA11, scaRNA14, scaRNA15, scaRNA85 (i.e., U85)), AluRNA, and htrRNA into snoRNAs to form gsnoRNAs. The gsnoRNA of this application comprises two stem-loop structures, also known as hairpin structures, which, from 5' to 3', sequentially include: a first stem near the 5' end of the stem-loop structure, a first loop near the 5' end of the stem-loop structure, a second loop near the 3' end of the stem-loop structure, and a second stem near the 3' end of the stem-loop structure. The guide sequence may be located in the first stem and / or the second stem, and the CAB box may be located in the first loop and / or the second loop.

[0244] CAB boxes possess a conserved sequence "X1X2AG", where the AGs at positions 3 and 4 are highly conserved, while the first two can vary. For example, conserved sequences for CAB box elements can include AAAAG, GAAG, UAAG, UGAG, UCAG, CGAG, AUAG, GCAG, CUAG, CAAG, and AGAG. htrRNA is a Human Telomerase RNA, whose 3' hairpin contains a CAB box. AluRNA is a non-coding RNA belonging to the Alu elements expressed by intron sequences, and its 3' hairpin contains a CAB box.

[0245] This embodiment adds the CAB box in the following two ways:

[0246] 1. Construct gsnoRNA directly using scaRNA (containing CAB box) as the backbone.

[0247] We directly used natural scaRNA as the backbone of gsnoRNA, replacing its targeting sequence with a sequence targeting ALDOB-W148X to construct two gsnoRNAs (using natural scaRNA14 and scaRNA15 as backbones, with the replaced gsnoRNA sequences shown in SEQ ID NO: 2 and SEQ ID NO: 3, respectively). ALDOB-W148X refers to a PTC disease model gene in which the 148th leucine (W, corresponding to the codon UGG) in the protein expressed by the ALDOB (fructose diphosphate aldolase B) gene is mutated to a stop codon (UAG). The ALDOB database ID is NP_000026.2.

[0248] The results showed that scaRNA, as a gsnoRNA backbone, had a high readability level in the RESTART v1 system (i.e., only gsnoRNA was introduced into the host cell), and this readability level was higher than that of gsnoRNA (gACA19, whose sequence is SEQ ID NO: 1) constructed using ACA19 as a snoRNA backbone. Figure 3 (a-3b).

[0249] 2. Introduce a CAB box derived from scaRNA into gsnoRNA.

[0250] To improve the versatility of the CAB box in different scenarios, we directly added the CAB box or the entire loop sequence containing the CAB box from the scaRNA to the previously optimized gACA19. Depending on the scaRNA used, we constructed different gsnoRNAs as shown in SEQ ID NO: 4~22.

[0251] from Figure 3 As can be seen from c, this study screened CAB boxes on almost all scaRNAs. In the UGA reporter gene system that does not express DKC1 iso3 (only recruits endogenous DKC1 protein), the CAB boxes on most scaRNAs can improve the reading efficiency by about 50%.

[0252] We then tested the CABboxes of some scaRNAs in a UGA reporter gene system overexpressing DKC1 iso3. The results showed that these CABboxes derived from different scaRNAs could improve readthrough efficiency, and the gsnoRNA (SEQ ID NO: 4) constructed using a CABbox derived from U85 had the highest relative efficiency. Figure 3 d).

[0253] Next, we investigated the universality of the CAB box on different gsnoRNAs. The U85 CAB box was linked to different gsnoRNAs, and the resulting gsnoRNA sequences are shown in SEQ ID NO: 73-78. Figure 3 As can be seen, the CAB box of U85 has a 30%-50% improvement in readability across different gsnoRNAs.

[0254] Next, since the CAB box is added to the loop of the stem-loop structure of gsnoRNA, and gsnoRNA has two stem-loop structures, we further explored the effect of adding CAB at different positions on gsnoRNA. We constructed gsnoRNAs with CAB boxes added to the loop near the 5' stem-loop structure, the loop near the 3' stem-loop structure, and both of these positions (sequences shown in SEQ ID NO: 23-25, respectively). The results showed that adding the CAB box to all three positions significantly improved pseudouridineization efficiency and enhanced readability. Furthermore, adding the CAB box to the loop of the 5' stem-loop structure of the CAB box yielded the best results. Figure 3 f).

[0255] 3. Introduce an inactivated CAB box into gsnoRNA

[0256] Furthermore, to verify that the improved readthrough is directly related to the function of the CAB box, we compared the effects of a normal CAB box and a CAB box with a functionally inactivating mutation (i.e., one that does not conform to the conserved sequence of XXAG) on readthrough, and constructed gsnoRNAs (SEQ ID NO: 26 and SEQ ID NO: 27) containing CAB boxes with functionally inactivating mutations. Figure 3 As can be seen from g and 3h, a normally functioning CAB box significantly improves readability, while a CAB box with a functional inactivation mutation has little or no effect on readability, or even slightly reduces it. These results indicate that the CAB box does indeed improve the readability of the RESTART system.

[0257] Then, to enhance the application prospects of the CAB box, we tested its effectiveness using different versions of the RESTART system in two nonsense mutation disease scenarios: CFTR-R553X and LMNA-R225X. CFTR stands for Cystic fibrosis transmembrane conductance regulator. R553X means that the arginine at position 553 (R, codon CGA) has been mutated to a stop codon (UGA), and its NCBI number is NP_000483.3. LMNA (Lamin A / C) is a protein encoded by the human gene LMNA, belonging to the lamin family. R225X means that the arginine at position 225 (R, codon CGA) has been mutated to a stop codon (UGA), and its NCBI number is NP_001393912.1.

[0258] The results showed that ( Figure 3 (i and 3j) In all RESTART systems, the use of the CAB box effectively improves readthrough efficiency, with particularly significant effects in the RESTART v3mini and RESTART v3 systems. That is, whether the gsnoRNA linking the CAB box is used alone in the RESTARTv1 system, or in combination with DKC1 iso3 and / or nc-tRNA in the RESTART v1, v3, and v3 mini systems, this gsnoRNA can further enhance the pseudouracil modification level and readthrough efficiency of PTC.

[0259] 4. Introduce HTR or Alu-derived CAB boxes into gsnoRNA.

[0260] Finally, to explore the universality of the CAB box, we tested its effectiveness on CAB boxes carried on other non-scaRNAs. The study found ( Figure 3Adding the CAB box from the HTR directly to the 5' stem loop (5HTR) and 3' stem loop (3HTR) of gsnoRNA, respectively, resulted in gsnoRNAs (SEQ ID NO: 28 and SEQ ID NO: 29) with significantly improved readability. Furthermore, replacing half of the gsnoRNA backbone with the corresponding sequence containing the CAB box in the HTR (SEQ ID NO: 30 and SEQ ID NO: 31), both 5' and 3' end substitutions, showed higher readability than gsnoRNAs without the CAB box. Replacing half of the 5' end of the gsnoRNA backbone with the corresponding sequence containing the CAB box in the HTR yielded significantly better results. Similarly, adding the CAB box from AluRNA to gnoRNA (SEQ ID NO: 32 and SEQ ID NO: 33) also showed some improvement, with addition at the 5' end being more effective than at the 3' end. Both 5' and 3' end additions significantly improved readability compared to gsnoRNAs without the CAB box. Since the CAB boxes on different AluRNAs are highly similar, it can be expected that the CAB boxes on other AluRNAs will also produce similar effects.

[0261] In summary, regardless of the source of the CAB box (e.g., scaRNA, Alu, or hTR) or the connection position of the CAB box (5' end, 3' end, or 5' and 3' ends), gsnoRNAs carrying CAB boxes can significantly improve the pseudouridine esterification efficiency and PTC readthrough level of the RESTART system (v1-v3, v3 mini). Furthermore, while retaining targeting capabilities, directly using scaRNA as the backbone, or replacing a portion of the gsnoRNA sequence entirely with scaRNA, Alu, or hTR containing a CAB box, can also achieve the function of gsnoRNA, significantly improving the pseudouridine esterification efficiency and PTC readthrough level of the RESTART system. This also provides new possibilities for the selection and design of "gsnoRNA" or "guide RNA." That is, the key to recruiting the DKC1 enzyme lies in retaining the H / ACA structure; we can choose different small RNAs as the basic backbone for modification and design, or even design from scratch, to provide guide RNA with both "DKC1 enzyme recruitment capability" and "target sequence guidance capability."

[0262] Example 3. Design of CTE sequences for snoRNA

[0263] This embodiment attempts to open the structure of the substrate mRNA by linking a CTE sequence to gsnoRNA, thereby enhancing the targeting efficiency of snoRNA. The structure of the CTE is shown below. Figure 4 a.

[0264] 1. Introducing SRV-derived CTE into gsnoRNA

[0265] We selected the most commonly used sno-ACA19 to construct gsnoRNA (SEQ ID NO: 34). Furthermore, we linked the CTE element (SEQ ID NO: 79) of type D retrovirus (SRV) to the 5' and 3' ends of gsnoRNA (SEQ ID NO: 35 and SEQ ID NO: 36), respectively. The results showed that regardless of whether DKC1 iso3 was overexpressed, the reading efficiency was improved by 30%-80%. Figure 4 (b-4c), where connecting at the 3' end is more effective than connecting at the 5' end.

[0266] Next, we added linkers of different lengths between the gsnoRNA and the CTE. We experimented with the effects of linker sequences of 4-8 bp length attached to the 5' or 3' end, respectively. The experimental 4-8 bp linker sequences were UCUA, UCUAU, UCUAUC, UCUAUCU, and UCUAUCU, and the resulting gsnoRNA sequences are shown in SEQ ID NO: 37-46. The results showed that adding a linker at either the 5' or 3' end improved reading efficiency to some extent. Furthermore, when the linker was 6 bp and the CTE was attached to the 3' end of the gsnoRNA, the reading efficiency was the highest, approximately twice that of the original. Figure 4 b-4c).

[0267] 2. Introducing inactivated CTE into gsnoRNA

[0268] Next, we verified that the improved readthrough efficiency was indeed caused by the function of CTE by adding the inactivated CTE mutant to the gsnoRNA. The sequence of the gsnoRNA containing the native CTE is shown in SEQ ID NO: 47, and the sequences of the gsnoRNA containing the inactivated CTE mutant are shown in SEQ ID NO: 48 and SEQ ID NO: 49. Figure 4As shown in d-4e, regardless of whether DKC1 iso3 is overexpressed, ligation with a normally functional CTE significantly improved readability, while ligation with a functionally mutated CTE had virtually no effect on readability. These results indicate that the improved readability after CTE ligation is indeed related to CTE function. Then, to verify that gsnoRNA ligation with a CTE improves readability, we performed RIP experiments to analyze the interaction between helicase and ligated normal and inactivated CTEs. Figure 4 As can be seen from f-4g, only gsnoRNAs linked to normal CTEs can interact with helicase and be significantly enriched. These results further demonstrate that gsnoRNAs linked to CTEs can improve readthrough efficiency through interaction with helicase.

[0269] 3. Introducing MPMV-derived CTE into gsnoRNA

[0270] Finally, we attempted to find smaller CTEs to link into gsnoRNA to achieve gsnoRNA delivery as a small RNA. Mason Fischer monkey virus (MPMV) also has CTE elements; the nucleotide sequence of the MPMV CTE element is shown in SEQ ID NO:80. Furthermore, it has been reported that a truncated half-length MPMV CTE is sufficient to achieve the same function as the full-length MPMV CTE. Therefore, we explored the role of SRV CTE and MPMV CTE, as well as their corresponding truncated mutants, in readthrough.

[0271] The sequences of gsnoRNAs constructed using natural CTEs and truncated CTEs derived from SRV are shown in SEQ ID NO: 50 and SEQ ID NO: 51-52, respectively; the sequences of gsnoRNAs constructed using natural CTEs and truncated CTEs derived from MPMV are shown in SEQ ID NO: 53 and SEQ ID NO: 54-61, respectively. Figure 4 As can be seen, in the UGA reporter gene system without DKC1 iso3 overexpression, MPMV CTE and its partially truncated mutants can achieve similar or even higher reading efficiency improvements than SRV CTE. Among them, the SRV CTE truncated mutants SRV CTE-m1, SRV CTE-m2 and the MPMVCTE truncated mutants MPMV CTE-m1, MPMV CTE-m2, MPMV CTE-m3, MPMV CTE-m4, MPMV CTE-m6 and MPMVCTE-m8 have particularly significant effects on improving PTC reading efficiency. The sequences of these CTE element truncated mutants are shown in SEQ ID NO:81~88, respectively.

[0272] In summary, this application improves the efficiency of the RESTART system by enhancing the recruitment and targeting capabilities of gsnoRNA, and applies this technology to the future fourth-generation RESTART system (i.e., RESTART V4). Specifically, we added a CTE element linked to gsnoRNA to enhance its ability to bind to target mRNA, and inserted a CAB box element into the loop region of gsnoRNA to improve its ability to recruit RNPs. Furthermore, these modifications can be combined with other RESTART components to significantly improve the system's PTC readthrough efficiency.

[0273] Although specific embodiments of the invention have been described in detail, those skilled in the art will understand that various modifications and variations can be made to the details based on all the published teachings, and all such changes are within the scope of protection of the invention. The entire scope of the invention is given by the appended claims and any equivalents thereof.

Claims

1. A method for inhibiting premature stop codons (PTCs) in target RNA in host cells for non-therapeutic purposes, characterized in that, The method includes: An engineered guide small nucleolar RNA (gsnoRNA) or a nucleic acid for expressing the gsnoRNA is introduced into the host cell, wherein the gsnoRNA comprises: (i) at least one guide sequence, (ii) at least one CAB box and / or CTE (constitutive transport element), and (iii) a scaffold sequence; wherein, The guide sequence hybridizes with the target RNA containing the target uridine residue (U) of PTC; The CAB box is derived from the loop region of a stem-loop structure of natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA). The scaffold sequence is derived from naturally occurring H / ACA type snoRNA and / or naturally occurring H / ACA type scaRNA; and The support sequence includes a first hairpin structure near the 5' end and a second hairpin structure near the 3' end; The CAB box is located in the loop area of ​​the first hairpin structure, or in the loop area of ​​the second hairpin structure, or simultaneously in the loop areas of the first hairpin structure and the second hairpin structure. The guiding sequence is located in the stem region of the first hairpin structure, or in the stem region of the second hairpin structure, or simultaneously in the stem regions of both the first and second hairpin structures. The CTE element is attached to the 5' end of the scaffold sequence of the gsnoRNA, or to the 3' end of the scaffold sequence of the gsnoRNA, or to both the 5' and 3' ends of the scaffold sequence of the gsnoRNA.

2. The method according to claim 1, characterized in that, The CAB box has the following sequence: X1X2AG; wherein X1 and X2 are each independently selected from any one of A, U, C, G.

3. The method according to claim 1, characterized in that, The CAB box is located in the loop region of the first hairpin structure in gsnoRNA.

4. The method according to claim 1, characterized in that, The method has one or more features selected from the following: (1) The gsnoRNA contains a CAB box derived from natural scaRNA, AluRNA or htrRNA, and a sequence in natural scaRNA, AluRNA or htrRNA containing the entire loop region of the CAB box. (2) The natural scaRNA is selected from scaRNA1, scaRNA4, scaRNA5, scaRNA6, scaRNA8, scaRNA11, scaRNA12, scaRNA13, scaRNA14, scaRNA15, scaRNA16, scaRNA18, scaRNA20, scaRNA21, scaRNA22, scaRNA23, scaRNA26, scaRNA85 and / or scaRNA27; (3) The natural AluRNA is selected from AluACA2, AluACA5, AluACA7, AluACA8, AluACA9, AluACA13, AluACA15, AluACA17, AluACA21, AluACA24, AluACA43, AluACA48, AluACA91, AluACA97, AluACA177, AluACA208, AluACA214 and / or AluACA303; (4) The natural htrRNA is human telomerase RNA. (5) The natural H / ACA type snoRNAs are selected from the following group: ACA19, ACA2b, ACA36, ACA24, ACA5, ACA14a, ACA13, ACA20, ACA44, ACA27, E2, ACA3 and ACA17; (6) The natural H / ACA type scaRNA is selected from the following group: scaRNA11 and scaRNA15; (7) The gsnoRNA contains a first guide sequence and a second guide sequence, wherein the first guide sequence is located in the stem region of the first hairpin structure and the second guide sequence is located in the stem region of the second hairpin structure.

5. The method according to claim 1, characterized in that, The method has one or more features selected from the following: (1) The CTE element is a CTE element derived from a retrovirus; (2) The CTE element is a CTE element or a truncated form of a masonfish virus (MPMV) or a simian D retrovirus (SRV), wherein the sequence of the truncated form is shown in SEQ ID NO:81~88; (3) The sequence of the CTE element is as shown in SEQ ID NO:79 or 80; (4) The gsnoRNA comprises, from the 5' end to the 3' end, the following sequence connected by a scaffold sequence: the first part of the first guide sequence, a CAB box derived from natural scaRNA, AluRNA or htrRNA (Human Telomerase RNA) or a loop region containing the CAB box, the second part of the first guide sequence, the first part of the second guide sequence, the second part of the second guide sequence, and a CTE element.

6. The method according to claim 1, characterized in that, The method has one or more features selected from the following: (1) The CTE element and the stent sequence are connected via a linker; (2) The CTE element is connected to the 3' end of the scaffold sequence of the gsnoRNA and is connected by a 4-6 bp or 6-8 bp linker; (3) The nucleotide sequence of the gsnoRNA is shown in any one of SEQ ID NO:2~25, SEQ ID NO:28~33, SEQ ID NO:68~72, SEQ ID NO:73~78 or SEQ ID NO:50~61.

7. The method according to claim 1, characterized in that, The gsnoRNA recruits the DKC1 protein to modify the target uridine residue in the target RNA into a pseudouridine residue (Ψ). The DKC1 protein recruited by the gsnoRNA includes: endogenous DKC1 protein of the host cell, and / or exogenous DKC1 protein of the host cell.

8. The method of claim 7, characterized in that, The method has one or more features selected from the following: (1) Introduce gsnoRNA into the host cell, wherein the gsnoRNA hybridizes with the target RNA and recruits the DKC1 protein, and modifies the target uridine residue (U) of PTC contained in the target RNA into pseudouridine residue (Ψ). (2) The method further includes: introducing a nucleic acid molecule encoding the DKC1 protein into the host cell; (3) The DKC1 protein is overexpressed in the host cells; (4) The DKC1 protein is a naturally occurring DKC1 subtype that has cytoplasmic localization in the host cell; (5) The DKC1 is selected from human DKC1 protein iso1, human DKC1 protein iso3, or any combination thereof; (6) The amino acid sequence of the DKC1 protein is shown in SEQ ID NO: 66 or SEQ ID NO:

67.

9. The method according to claim 1, characterized in that, The method further includes: The near-homologous transfer RNA (nc-tRNA) of the PTC or a nucleic acid molecule for expressing the nc-tRNA is introduced into the host cell.

10. The method of claim 9, characterized in that, The method has the following characteristics: modifying the target uridine residue in the PTC of the target RNA into a pseudouridine residue to provide Ψ-modified PTC; then, introducing the Ψ-modified PTC nc-tRNA or a nucleic acid molecule for expressing the nc-tRNA into the host cell to decode the PTC into amino acids, thereby inhibiting the PTC.

11. The method of claim 9, characterized in that, The nc-tRNA is modified.

12. The method of claim 9, characterized in that, The nc-tRNA sequences are shown in SEQ ID NO: 62~65 and SEQ ID NO: 89~172.

13. An engineered guide small nucleolar RNA (gsnoRNA), characterized in that, The gsnoRNA contains: (i) at least one guide sequence, (ii) at least one CAB box and / or CTE (constitutive transport element) element, and (iii) a stent sequence; wherein, The guide sequence hybridizes with the target RNA containing the target uridine residue (U) of PTC; The CAB box is derived from the loop region of a stem-loop structure of natural scaRNA, AluRNA, or htrRNA (Human Telomerase RNA). The scaffold sequence is derived from naturally occurring H / ACA type snoRNA and / or naturally occurring H / ACA type scaRNA; and, The support sequence includes a first hairpin structure near the 5' end and a second hairpin structure near the 3' end; The CAB box is located in the loop area of ​​the first hairpin structure, or in the loop area of ​​the second hairpin structure, or simultaneously in the loop areas of the first hairpin structure and the second hairpin structure. The guiding sequence includes the stem region located in the first hairpin structure, or the stem region located in the second hairpin structure, or both the first and second hairpin structures. The CTE element is attached to the 5' end of the scaffold sequence of the gsnoRNA, or to the 3' end of the scaffold sequence of the gsnoRNA, or to both the 5' and 3' ends of the scaffold sequence of the gsnoRNA.

14. The gsnoRNA of claim 13, characterized in that, The gsnoRNA has one or more of the following characteristics: (1) The CAB box has the following sequence: X1X2AG; wherein X1 and X2 are each independently selected from any one of A, U, C, G; (2) The CTE element is a CTE element or a truncated form of a masonfish virus (MPMV) or a simian D retrovirus (SRV), wherein the sequence of the truncated form is shown in SEQ ID NO:81~88; (3) The gsnoRNA comprises, from the 5' end to the 3' end, the following sequence connected by a scaffold sequence: the first part of the first guide sequence, a CAB box derived from natural scaRNA, AluRNA or htrRNA (Human Telomerase RNA) or a loop region containing the CAB box, the second part of the first guide sequence, the first part of the second guide sequence, the second part of the second guide sequence, and a CTE element.

15. An isolated nucleic acid molecule, characterized in that, The isolated nucleic acid molecule contains a nucleic acid sequence for expressing the gsnoRNA of claim 13 or 14.

16. A composition, characterized in that, The composition comprises: the gsnoRNA of claim 13 or 14 or the isolated nucleic acid molecule of claim 15.

17. The composition of claim 16, characterized in that, The composition comprises: (a) The gsnoRNA of claim 13 or 14 or the isolated nucleic acid molecule of claim 15; and the near-homologous transfer RNA (nc-tRNA) of the PTC contained in the target sequence or the nucleic acid for expressing the nc-tRNA; (b) the gsnoRNA of claim 13 or 14 or the isolated nucleic acid molecule of claim 15; and the DKC1 protein or a nucleic acid molecule encoding the DKC1 protein; or (c) The gsnoRNA of claim 13 or 14 or the isolated nucleic acid molecule of claim 15, the near homologous transfer RNA (nc-tRNA) of the PTC contained in the target sequence or the nucleic acid for expressing the nc-tRNA, and the DKC1 protein or the nucleic acid molecule encoding the DKC1 protein.

18. A delivery composition, characterized in that, The delivery composition comprises: a delivery vector, and one or more of the following: gsnoRNA as described in claim 13 or 14, or an isolated nucleic acid molecule as described in claim 15, or a composition as described in claim 16 or 17; wherein the delivery vector is a particle.

19. A host cell, characterized in that, The host cell contains the gsnoRNA of claim 13 or 14, the isolated nucleic acid molecule of claim 15, the composition of claim 16 or 17, or the delivery composition of claim 18.

20. A method for preparing the gsnoRNA of claim 13 or 14, the composition of claim 16 or 17, or the delivery composition of claim 18, characterized in that, The method includes culturing the host cell of claim 19 under conditions that allow for nucleic acid and protein expression, and recovering the gsnoRNA or the composition or the delivery composition from the cultured host cell culture.

21. Use in the preparation of a formulation of the gsnoRNA of claim 13 or 14, or the isolated nucleic acid molecule of claim 15, or the composition of claim 16 or 17, or the delivery composition of claim 18, or the host cell of claim 19, for editing a target RNA or for inhibiting a premature stop codon (PTC) in a target RNA in a host cell.

22. The use of the gsnoRNA of claim 13 or 14, or the isolated nucleic acid molecule of claim 15, or the composition of claim 16 or 17, or the delivery composition of claim 18, or the host cell of claim 19, in the preparation of a pharmaceutical product, characterized in that, The drug is used to treat diseases and / or symptoms caused by or resulting from PTC mutations in the subject.

23. A method for editing target RNA in vitro for non-therapeutic purposes, characterized in that, The method includes, under conditions suitable for editing the target RNA, contacting the target RNA with one or more of the following: the gsnoRNA of claim 13 or 14, the isolated nucleic acid molecule of claim 15, the composition of claim 16 or 17, or the delivery composition of claim 18, thereby editing the target RNA.

Citation Information

Patent Citations

  • RNA TARGETING OF MUTATIONS VIA SUPPRESSOR tRNAs AND DEAMINASES

    CN110612353A

  • Nucleic acid molecules for pseudouridylation

    CN112020557A