RNA editing target site and sequence that forms the editing substrate
By designing a double-stranded RNA editing substrate with specific structural features, the RNA editing efficiency is enhanced, addressing the inefficiencies in current RNA editing technologies and enabling precise RNA editing in host cells.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- RECORNA (GUANGZHOU) BIOTECHNOLOGY CO LTD
- Filing Date
- 2023-03-31
- Publication Date
- 2026-04-23
AI Technical Summary
Current RNA editing technologies, particularly those using endogenous ADAR recruitment, suffer from low editing efficiency due to suboptimal design of guide RNA structures, leading to inefficient RNA editing in vivo.
A sequence forming an RNA editing target site and editing substrate is designed to create a double-stranded secondary structure with specific features such as mismatches, fluctuating base pairs, internal loops, and deletions, optimized using the MASSA method to enhance editing efficiency.
The optimized RNA editing substrate significantly improves editing efficiency by mimicking natural human body structures, allowing for precise and effective RNA editing in host cells.
Smart Images

Figure 2026513318000001_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of biomedicine and relates to a sequence that forms an RNA editing target site and an editing substrate.
Background Art
[0002] Gene editing, as one of the most important technologies in recent life sciences, has infinite possibilities and advantages in the treatment of diseases, especially genetic diseases, various chronic diseases and tumor treatment. It is expected that the next 5-10 years will be a golden age for the sequential clinical promotion and rapid development of gene editing technologies. The treatment concepts of conventional genetic diseases and rare diseases, as well as some chronic diseases, geriatric diseases and tumors that threaten human health, will be completely overturned, and gene editing-based medicine is expected to help countless patients in the near future.
[0003] Gene editing can be divided into DNA editing and RNA editing according to the target. Compared with DNA editing, RNA editing can repair mutations at the RNA level, does not change the genomic sequence, can be reversibly controlled, has more safety characteristics, and has great advantages in certain disease scenarios.
[0004] RNA editing currently specifically refers to single-base editing technology. Its principle is to use deaminase to achieve the purpose of modifying specific bases by deaminated ribosomal nucleic acid. Target RNA editing currently has two main technical routes. The first technology is a two-molecule editing system (including one exogenously expressed deaminase and one guide RNA).
[0005] The second technology is a single-molecule editing system that can recruit endogenous deaminases and deliver only sequences that can target the editing site. This second technology has only been successful in recruiting ADARs used for A-to-G targeted RNA editing. Compared to exogenous ADAR expression systems, systems that recruit endogenous ADARs only require the delivery of antisense RNA strands to recruit endogenous ADARs and edit specific RNA sequence sites. Such antisense RNA-based technologies have a series of advantages, including faster drug discovery rates and relatively mature safety and delivery technologies, which are beneficial for clinical conversion and drug discovery. For example, this technology does not induce immune responses, can be applied to various developmental stages and tissue types, facilitates tissue-specific gene editing, and can achieve different levels of editing of diverse genes. Therefore, it has the advantage of not having the safety risks associated with DNA editing technologies.
[0006] In the second technology, there are currently two methods for designing guide RNA. The first type of guide RNA consists of two segments: one segment is an ADAR substrate sequence that can recruit ADARs and has secondary structural properties, and the other segment is a customizable target sequence used to target and edit a specific site. The target sequence and target site form a double-stranded RNA structure in which they are perfectly complementary, and the target site is set as an A:C mismatch. There is an optimized version of this method to promote editing of the target site (RESTORE method) and reduce off-target editing (CLUSTER method). The guide RNA designed by the second method does not contain a recruitment sequence and is a target sequence (LEAPER method) that is complementary to the target sequence and is approximately 100-150 nt in length. In both of these methods, the dsRNA double-stranded structure is in which the region complementary to the target site is almost perfectly paired. Studies have shown that a nearly perfectly paired dsRNA double-stranded structure is not the best ADAR editing substrate, so both of these design methods have low editing efficiency in vivo.
[0007] RNA molecular structure is described in three hierarchical levels, including primary sequence, secondary structure, and tertiary spatial structure. RNA tertiary spatial structure is a stable structure formed in space, arising from interactions, deformation, and folding of secondary structural units. Furthermore, without RNA secondary structure, RNA tertiary structure formation becomes difficult. Predicting RNA secondary structure requires considering not only sequence arrangement but also stable pairing methods, including temporary knots and hairpins. RNA secondary structure has two important functions. First, it explains RNA function. RNA function is usually related to RNA structure, and secondary structure is the most important of all RNA structures (primary, secondary, and tertiary structures). Once RNA is formed, it undergoes changes to form a specific tertiary structure. Tertiary structure formation depends on the matching of base pairs in the secondary structure. Second, understanding secondary structure is useful for exploring new RNA functions.
[0008] Therefore, finding a suitable secondary structure for the editing substrate in appropriate RNA editing is of great importance for improving RNA editing efficiency. [Overview of the Initiative]
[0009] In one embodiment, according to this application, a sequence that forms an RNA editing target site and an editing substrate, The sequence forms a double-stranded secondary structure with the RNA editing target site in the form of a complementary strand, wherein the structure of the complementary strand includes at least one set of features shown in ix below. i. The edited site has 0 to 4 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 12 deletions at positions -100 to -81. ii. The edited site has 0 to 5 mismatches, 0 to 7 fluctuating base pairs, 0 to 5 internal loops, 0 to 5 protrusions, and 0 to 15 deletions at positions -80 to -61. iii. The edited site has 0 to 5 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 18 deletions at positions -60 to -41. iv. At the -40 to -21 position of the editing site, there are 0 to 4 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 8 protrusions, and 0 to 11 deletions. v. The edited site has 0-5 mismatches, 0-6 fluctuating base pairs, 0-4 internal loops, 0-5 protrusions, and 0-9 deletions at positions -20 to -1. vi. The edited site has 0-5 mismatches, 0-7 fluctuating base pairs, 0-3 internal loops, 0-3 protrusions, and 0-5 deletions at positions 0-19. vii. The edited region has 0-4 mismatches, 0-7 fluctuating base pairs, 0-3 internal loops, 0-6 protrusions, and 0-9 deletions at positions 20-39. viii. The edited site has 0-5 mismatches, 0-6 fluctuating base pairs, 0-3 internal loops, 0-5 protrusions, and 0-17 deletions at positions 40-59. ix. The edited site has 0-5 mismatches, 0-8 fluctuating base pairs, 0-4 internal loops, 0-5 protrusions, and 0-19 deletions at positions 60-79. x. A sequence is provided having 0 to 4 mismatches, 0 to 6 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 19 deletions at positions 80 to 100 of the editing site. In some embodiments, the type of mismatch on the complementary strand is one of AA, AG, AC, UU, UC, GG, GA, CA, CC, or CU.
[0010] In some embodiments, the fluctuating base pair on the complementary strand is a GU base pair. In some embodiments, the projection on the complementary strand has one base upstream and one downstream of the projection region that are complementaryally paired with the target edited strand, and the corresponding two bases of the target edited strand are consecutive. In some embodiments, the deletion on the complementary strand has consecutive bases upstream and downstream of the deletion region, and there are a number of bases corresponding to the target edited strand of the deletion region. In some embodiments, the internal loop on the complementary strand has one base upstream and one downstream of the internal loop that are complementaryally paired with the target edited strand and the complementary strand, and there are a corresponding number of mismatched bases in the internal loop.
[0011] In some embodiments, the double-stranded secondary structure is one selected from the Struc1 to Struc1196 structures. In some embodiments, the structural features of the double-stranded secondary structure are shown in Table 1.
[0012] In some embodiments, the structural features of the double-stranded secondary structure are as follows: A mismatch is formed between the bases represented by M. A fluctuating base pair represented by W, Internal loop represented by Li, Deletion represented by D, The protrusion represented by li, and Watson-Crick base pairs represented by P This includes one or more of the following: In some embodiments, in the features of the double-stranded secondary structure, the numerical values X, X=[n1,n2,n3…], n1, n2, etc., indicate the position of the X structure, where 0 is the position of the editing site, -n indicates that this position is the nth base upstream of the editing site, and n indicates that this position is the nth base downstream of the editing site.
[0013] Preferably, if the double-stranded secondary structure X is M, then M=[n1,n2,n3…] indicates that mismatch structures exist at positions such as n1, n2, n3 from the editing site.
[0014] If the double-stranded secondary structure X is W, then W=[n1,n2,n3…] indicates that fluctuation base pair structures exist at positions such as n1, n2, and n3 from the editing site.
[0015] If the double-stranded secondary structure X is D, then D=[n1,n2,n3…] indicates that deletion structures exist at positions such as n1, n2, n3, etc., from the editing site.
[0016] If X is Ii, then Ii = [n1, n2], where i is any positive integer, indicating that i bases are inserted between n1 and n2 from the editing site.
[0017] If the double-stranded secondary structure X is Li, then Li = [n1~n2, n3~n4…], where i is any positive integer, and the editor shows that internal loops of size i nt are formed in positional intervals such as n1 and n2, n3 and n4.
[0018] In any of the double-stranded secondary structures, unless X is one or more of M, W, D, Ii, or Li, its position is always P.
[0019] In some embodiments, the base motifs of the target RNA editing site and one site upstream and downstream thereof are NAN, preferably the NAN is UAN, AAN, CAN or GAN, preferably the UAN is UAG, UAU, UAC or UAA, preferably the AAN is AAU, AAC, AAG or AAA, preferably the CAN is CAU, CAG, CAC or CAA, and preferably the GAN is GAU, GAC, GAA or GAG.
[0020] In some embodiments, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are NAN, the double-stranded secondary structure is Struc1 to Struc1196.
[0021] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is UAA, the double-stranded secondary structure is Struc1 to Struc143.
[0022] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is UAG, the double-stranded secondary structure is Struc144 to Struc260.
[0023] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is UAU, the double-stranded secondary structure is Struc261 to Struc320.
[0024] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is UAC, the double-stranded secondary structure is Struc321 to Struc395.
[0025] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is CAA, the double-stranded secondary structure is Struc396 to Struc515.
[0026] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is CAC, the double-stranded secondary structure is Struc516 to Struc541.
[0027] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is AAG, the double-stranded secondary structure is Struc542 to Struc684.
[0028] In some embodiments, when the motif of the bases at one site each upstream and downstream of the target RNA editing site is AAC, the double-stranded secondary structure is Struc685 to Struc767.
[0029] In some embodiments, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are GAA, the double-stranded secondary structure is Struc768~Struc809.
[0030] In some embodiments, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are GAGs, the double-stranded secondary structure is Struc810~Struc875.
[0031] In some embodiments, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are AAU, the double-stranded secondary structure is Struc876~Struc982.
[0032] In some embodiments, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are CAG, the double-stranded secondary structure is Struc983 to Struc1085.
[0033] In some embodiments, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are CAU, the double-stranded secondary structure is Struc1086~Struc1196.
[0034] In some embodiments, the complementary strand has bases that pair complementarily at non-mismatched, non-deleted, non-protrusion, non-internal loop, or non-fluctuating base pairing sites.
[0035] In some embodiments, the length of the sequence forming the RNA editing target site and editing substrate is 25 to 300 bp, preferably 25 to 151 bp.
[0036] In some embodiments, the RNA is an RNA that forms a specific secondary structure with the target RNA to be edited. Preferably, the target RNA to be edited is messenger RNA (mRNA), Preferably, the low-molecular-weight RNA is miRNA, pri-miRNA, pre-miRNA, piRNA, siRNA, snoRNA, snRNA, exRNA, or scaRNA.
[0037] In one embodiment, the present application provides an mcRNA, which includes a simulation domain. The sequence of the simulation domain is a sequence that forms an RNA editing target site and an editing substrate, and forms the double-stranded secondary structure with the target RNA to be edited.
[0038] Preferably, the mcRNA further includes a binding domain, the binding domain being located at both ends of the simulation domain and pairing complementaryly with regions other than the double-stranded secondary structure of the target RNA to be edited. Preferably, the length of one side of the binding domain is 0-55 bp. Preferably, the mcRNA sequence is one sequence selected from sequence numbers 1-16083.
[0039] In one embodiment, according to the present application, the method for designing the mcRNA is as follows: Obtaining the reference sequence: Step S1 involves obtaining the nucleotide sequences of two strands, with the strand containing the editing site in the secondary structure being designated as the reference secondary structure editing strand, and the other strand as the reference secondary structure complementary strand. Step S2 involves obtaining the mcRNA sequence for preparing the target editing site: starting from the site to be edited, the target editing strand, which is the template sequence of the simulation domain, is obtained by extending it upstream and downstream by a predetermined length, and a sequence that is complementary to the template sequence is designed to become the prepared mcRNA sequence. Design of mcRNA simulation domain: Step S3 involves comparing the target edited strand sequence and the reference secondary structure sequence by extending the edited strand from the editing site to both ends, comparing the prepared mcRNA sequence and the reference secondary structure complementary strand sequence, and if the comparison results of the target edited strand sequence and the reference secondary structure edited strand sequence match, the corresponding position of the mcRNA simulation domain is set to the corresponding reference secondary structure complementary strand sequence, and if the comparison results of the target edited strand sequence and the reference secondary structure edited strand sequence do not match, the corresponding position of the mcRNA simulation domain is designed as a sequence that can form the same structure as the corresponding reference secondary structure together with the target editing site. Design of the mcRNA binding domain: Step S4 involves extending complementary pairing sequences of a predetermined length from both ends of the simulation domain, based on the mcRNA simulation domain designed in step S3, to form a binding domain. A design method including this is provided.
[0040] In this application, a method for designing MASSA (Mimic ADAR Substrates Structure Approach) is provided by simulating the structural characteristics of high-editing level double-stranded RNA substrates naturally present in the human body. This enables efficient targeted RNA editing by forming mcRNA with specific structural properties in relation to the upstream and downstream sequences of the target site.
[0041] In one embodiment, the present application provides the use of a sequence or mcRNA that forms an RNA editing target site and an editing substrate in the preparation of an RNA editing reagent.
[0042] In one embodiment, the present application provides an RNA editing reagent or composition comprising an RNA editing target site and a sequence or mcRNA that forms an editing substrate.
[0043] In one embodiment, the present application provides a method for editing a target RNA in a host cell. The method includes the step of introducing an oligonucleotide or mcRNA having a sequence that forms an RNA editing target site and an editing substrate into the host cell.
[0044] In some embodiments, the target RNA is the ATP7B gene, FGFR gene, TP53 gene, APC gene, SERPINA1 gene, MECP2 gene, or SCN1A gene.
[0045] In one embodiment, the present application provides the use of an RNA editing target site and a sequence or mcRNA that forms an editing substrate in the preparation of a pharmaceutical product for treating a disease or condition. The disease or condition is a hereditary genetic disorder or a disease or condition associated with one or more acquired genetic mutations. Preferably, the disease or condition is a single-gene disorder or condition. Preferably, the disease or condition is a polygenic disease or condition. Preferably, the genes are ATP7B, FGFR, TP53, APC, SERPINA1, MECP2, and SCN1A. Preferably, the diseases or conditions include Wilson's disease, Rett syndrome, Dravet syndrome, cystic fibrosis, Harler syndrome, α-1-antitrypsin (A1AT) deficiency, Parkinson's disease, Alzheimer's disease, albinism, amyotrophic lateral sclerosis, asthma, β-thalassemia, Kadasil syndrome, Charcot-Marie-Tooth disease, chronic obstructive pulmonary disease (COPD), distal spinal muscular atrophy (DSMA), Duchenne / Becker muscular dystrophy, dystrophic epidermolysis bullosa, epidermolysis bullosa, Fabry disease, factor V Leiden-related disease, familial adenoma, polyposis, galactosemia, Gaucher disease, glucose-6-phosphate dehydrogenase, hemophilia, hereditary hemochromatosis, Hunter syndrome, Huntington's disease These include nephrosis, inflammatory bowel disease (IBD), hereditary polyaggregation syndrome, Leber congenital amaurosis, Lesch-Nyhan syndrome, Lynch syndrome, Marfan syndrome, mucopolysaccharidosis, muscular dystrophy, myotonic dystrophy type I and II, neurofibromatosis, Niemann-Pick disease type A, B, and C, NY-esol-associated cancer, Peutz-Jeghers syndrome, phenylketonuria, Pompe disease, primary eyelash disorders, prothrombin mutation-related diseases such as the prothrombin G20210A mutation, pulmonary hypertension, retinitis pigmentosa, Sandhoff disease, severe combined immunodeficiency syndrome (SCID), sickle cell anemia, spinal muscular atrophy, Stargardt disease, Tay-Sachs disease, Usher syndrome, X-linked immunodeficiency, and cancer.
[0046] In some embodiments, the mcRNA may include one or more modifications. In some embodiments, the mcRNA has one or more modified nucleotides, including nucleic acid base modifications and / or cytoskeletal modifications. Exemplary modifications to RNA include, but are not limited to, phosphorothioate cytoskeletal modifications and ribose modifications of LNA.
[0047] This application provides a basic platform for targeting and editing RNA, such as messenger RNA, ribosomal RNA, transfer RNA, long non-coding RNA, and small RNA (e.g., miRNA). The methods described herein can introduce the mcRNA or structures described herein into host cells by using many delivery systems, including but not limited to viruses, liposomes, electroporation, microinjection, and conjugation. Nucleic acids can be introduced into mammalian cells or target tissues by conventional viral and non-viral-based gene transfer methods. Such methods can apply nucleic acids encoding the sequences described herein to cultured cells or host organisms. Non-viral vector delivery systems include DNA plasmids, RNA (e.g., transcripts of the constructs described herein), naked nucleic acids, and nucleic acids compounded with delivery media such as liposomes. Viral vector delivery systems include DNA viruses and RNA viruses and have free or integrated genomes for delivery to host cells.
[0048] The terms “nucleotide,” “nucleotide sequence,” and “nucleic acid” are interchangeable and refer to polymeric forms of nucleotides, deoxyribonucleotides, ribonucleotides, or their analogues of any length.
[0049] In the context of this application, “target RNA” means an RNA sequence designed such that the sequence of this application is fully or fundamentally complementary to it, and which hybridizes between the target sequence and the sequence of this application to form a double-stranded secondary structure containing target adenosine, and which recruits an adenosine deaminase (ADAR) that acts on the RNA, the enzyme deamination of target adenosine. In some embodiments, the ADAR is naturally present in host cells such as eukaryotic cells (preferably mammalian cells, more preferably human cells). In some embodiments, the ADAR is introduced into host cells.
[0050] The term "contains" includes both "contains" and "consists of." For example, a composition "containing X" may consist exclusively of X or may contain several additional components (e.g., X + Y). The term "about" in relation to the numerical value x is optional and refers to, for example, x ± 10%.
[0051] The expression "basically" does not exclude "completely"; for example, a composition that "basically does not contain Y" does not necessarily have to contain Y completely. In such cases, the expression "basically" may be omitted in the definitions of this application.
[0052] As used herein, the term “complementary” means that an oligonucleotide hybridizes to a target sequence under physiological conditions. This does not mean that each nucleotide in the oligonucleotide has a perfect pair with its corresponding nucleotide in the target sequence. In other words, an oligonucleotide may be complementary to a target sequence, and may have mismatches, fluctuations, and / or protrusions between the oligonucleotide and the target sequence, and may hybridize to the target sequence under physiological conditions, allowing cellular RNA editing enzymes to edit the target adenosine. Therefore, the term “basically complementary” means that, although mismatches, fluctuations, and / or protrusions are present, there are sufficient matching nucleotides between the oligonucleotide and the target sequence under physiological conditions in which the oligonucleotide hybridizes with the target RNA. As shown herein, an oligonucleotide may be complementary insofar as it can hybridize to a target under physiological conditions, but may contain one or more mismatches, fluctuations, and / or protrusions relative to the target sequence.
[0053] In relation to nucleic acid sequences, the term "downstream" means continuing along the sequence in the 3' direction, while the term "upstream" has the opposite meaning. Therefore, in any polypeptide sequence, the start codon is upstream of the stop codon in the sense strand, but downstream of the stop codon in the antisense strand.
[0054] "Hybridization" usually refers to specific hybridization, excluding nonspecific hybridization. Specific hybridization can be performed under experimental conditions selected by art known techniques, such that most stable interactions between the probe and target ensure that the probe and target have at least 70%, preferably at least 80%, and more preferably at least 90% sequence identity.
[0055] As used herein, the term "mismatch" refers to a situation where relative nucleotides in a double-stranded RNA complex do not form a complete base pair according to the Watson-Crick base pairing rule. Mismatched nucleotides are AA, AG, AC, UU, UC, GG, GA, CA, CC, and CU base pairs. [Brief explanation of the drawing]
[0056] [Figure 1] This is an example of the secondary structural features of double-stranded RNA. Figure 1 shows an example of the secondary structural features of a single double-stranded RNA, where the edited strand is the strand where the editing site is located, and the complementary strand is the strand where the sequence that forms a specific structure with the edited strand is located. In Figure 1, position 0 is the position where the base of editing site A is located, positions -1 to -22 are the positions where the upstream sequence of the editing site is located, and positions 1 to 16 are the positions where the downstream sequence of the editing site is located. The structural features shown in Figure 1 include mismatches, protrusions, deletions, fluctuating bases, and internal loops, and the remaining positions are all Watson-Crick base pair structures.
[0057] [Figure 2]This is a schematic diagram of the MASSA method of this application. Figure 2 shows a schematic diagram of the MASSA method used in this application. The left side of Figure 2 is a schematic diagram of the secondary structure of the ADAR substrate, and the right side shows the design of mcRNA by the MASSA method based on the structure on the left side. The asterisks in the figure indicate the location of the editing site. The regions with predetermined structural features upstream and downstream are simulation domains, which simulate the secondary structural features of the substrate and the binding domains at both ends. The binding domains form a Watson-Crick complementary base pair with the sequence of the target site.
[0058] [Figure 3] Figure 3 shows the mcRNA editing levels for the GAPDH gene UAG motif design in one embodiment of this application. Figure 3 shows the results of the mcRNA editing levels for designing combinations of simulation domains and binding domains of different lengths on the GAPDH gene UAA motif in one embodiment of this application. As shown, the lengths of the double-ended binding domains corresponding to the simulation domain lengths of 91nt, 61nt, and 41nt are 30nt, 45nt, and 55nt, respectively. In Figure 3, the black dots indicate the editing levels of mismatched mcRNA designs that lack structural features and have only one A and one C base at the editing site, while the other dots indicate the editing levels of mcRNAs with structural features, and the horizontal and vertical coordinates indicate the editing levels of each mcRNA in two biological replication experiments.
[0059] [Figure 4]This describes the design modes of the simulation domain and binding domain of different motif mcRNAs of the GAPDH gene in one embodiment of this application. Figure 4 shows the design modes of mcRNAs on different motifs of the GAPDH gene in one embodiment of this application, mainly showing the position and length of the simulation domain and binding domain. As shown in Figure 4, the design mainly has three modes. In the first mode, the length of the simulation domain is extended by 6 bases upstream and 14 bases downstream from the editing site. In the second mode, the length of the simulation domain is extended by 10 bases upstream and 10 bases downstream from the editing site. In the third mode, the length of the simulation domain is extended by 14 bases upstream and 6 bases downstream from the editing site. In the first mode, there is a 20-base binding domain downstream of the editing site. In the second mode, there are 10-base binding domains at each end. In the third mode, there is a 20-base binding domain upstream of the editing site. The length of the reference design is 41 nt in all cases.
[0060] [Figure 5] Figure 5 shows the mcRNA editing levels of the GAPDH gene UAN motif design in one embodiment of this application. The included motifs are UAG, UAU, UAC, and UAA. The black dots indicate the editing levels of mismatched mcRNA designs that lack structural features and have only one A and one C base at the editing site, while the other dots indicate the editing levels of mcRNAs with structural features designed in this application. The horizontal and vertical coordinates indicate the editing levels of each mcRNA in two biological replication experiments.
[0061] [Figure 6]This figure shows the mcRNA target site and detargeting edit of the GAPDH gene UAN motif design in one embodiment of this application. Figure 6 shows a heatmap of the mcRNA target site and detargeting (off-target) editing of the upstream and downstream regions of the target site in the GAPDH gene UAN motif design in one embodiment of this application. The UAN motif includes UAG, UAU, UAC, and UAA. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, and the blue to red color indicates the trend of change from low to high editing levels at the A base site, starting from the highest edited sequence at the target site and moving downwards.
[0062] [Figure 7] Figure 7 shows the mcRNA editing levels of the GAPDH gene CAN motif design in one embodiment of this application. The included motifs are CAA and CAC. The black dots indicate the editing levels of mismatched mcRNA designs that lack structural features and have only one A and one C base at the editing site, while the other dots indicate the editing levels of mcRNAs with structural features designed in this application. The horizontal and vertical coordinates indicate the editing levels of each mcRNA in two biological replication experiments.
[0063] [Figure 8] This invention relates to the mcRNA targeting and detargeting of the CAN motif design of the GAPDH gene in one embodiment of this application. Figure 8 shows a heatmap of the target site and detargeting of the upstream and downstream regions of the mcRNA target site in the CAN motif design of the GAPDH gene in one embodiment of this application. The CAN motif includes CAA and CAC. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base sites.
[0064] [Figure 9] This shows the mcRNA editing levels for the GAPDH gene AAN motif design in one embodiment of this application. The distribution of mcRNA editing levels for the GAPDH gene AAN motif design in one embodiment of this application is shown, with the included motifs being AAG and AAC. Black dots indicate the editing levels of mismatched mcRNA designs that lack structural features and have only one A and one C base at the editing site, while other dots indicate the editing levels of mcRNAs with structural features designed in this application. The horizontal and vertical coordinates represent the editing levels of each mcRNA in two biological replication experiments.
[0065] [Figure 10] This figure shows the mcRNA target site and detargeting edit of the AAN motif design of the GAPDH gene in one embodiment of this application. Figure 10 shows a heatmap of the mcRNA target site and detargeting edit of the upstream and downstream regions of the target site in the AAN motif design of the GAPDH gene in one embodiment of this application. The AAN motif includes AAG and AAC. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base sites.
[0066] [Figure 11] Figure 11 shows the mcRNA editing levels of the GAPDH gene GAN motif design in one embodiment of this application. The included motifs are GAG and GAA. The black dots indicate the editing levels of mismatched mcRNA designs that lack structural features and have only one A and one C base at the editing site, while the other dots indicate the editing levels of mcRNAs with structural features designed in this application. The horizontal and vertical coordinates indicate the editing levels of each mcRNA in two biological replication experiments.
[0067] [Figure 12] This figure shows the mcRNA target site and detargeting edit of the GAN motif design of the GAPDH gene in one embodiment of this application. Figure 12 shows a heatmap of the mcRNA target site and detargeting edit of the upstream and downstream regions of the target site in the GAN motif design of the GAPDH gene in one embodiment of this application. The GAN motif includes GAG and GAA. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the blue to red color indicates the trend of change from low to high editing levels at the A base sites.
[0068] [Figure 13] Figure 13 shows the mcRNA editing levels of the ATP7B gene UAG motif design in one embodiment of this application. The figure shows mcRNA editing levels of different lengths in the ATP7B gene UAG motif in one embodiment of this application, where black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0069] [Figure 14] This figure shows the mcRNA target site and detargeting edit of the ATP7B gene UAG motif design in one embodiment of this application. Figure 14 shows a heatmap of target site editing and upstream / downstream detargeting edit of all mcRNAs in the UAG motif of the ATP7B gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base sites.
[0070] [Figure 15] Figure 15 shows the mcRNA editing levels of the ATP7B gene AAU motif design in one embodiment of this application. The figure shows mcRNA editing levels of different lengths in the AAU motif of the ATP7B gene in one embodiment of this application, where black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0071] [Figure 16] This invention relates to mcRNA targeting and detargeting editing of the AAU motif of the ATP7B gene in one embodiment of this application. Figure 16 shows a heatmap of targeted site editing and upstream / downstream detargeting editing of all mcRNAs in the AAU motif of the ATP7B gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base sites.
[0072] [Figure 17] Figure 17 shows the mcRNA editing levels of the FGFR gene CAG motif design in one embodiment of this application. The figure shows mcRNA editing levels of different lengths in the CAG motif of the FGFR gene in one embodiment of this application, where black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0073] [Figure 18]This describes the mcRNA target site and detargeting edit of the FGFR gene CAG motif design in one embodiment of this application. Figure 18 shows a heatmap of target site editing and upstream / downstream detargeting edit of all mcRNAs in the CAG motif of the FGFR gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base site.
[0074] [Figure 19] Figure 19 shows the mcRNA editing levels of the APC gene GAG motif design in one embodiment of this application. The black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0075] [Figure 20] This describes the mcRNA target site and detargeting edit of the APC gene GAG motif design in one embodiment of this application. Figure 20 shows a heatmap of target site editing and upstream / downstream targeting detargeting edit of all mcRNAs in the GAG motif of the APC gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base site.
[0076] [Figure 21]Figure 21 shows the mcRNA editing levels of the APC gene UAG motif design in one embodiment of this application. The black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0077] [Figure 22] This document describes the mcRNA target site and detargeting edit of the APC gene UAG motif design in one embodiment of this application. Figure 22 shows a heatmap of target site editing and upstream / downstream detargeting edit of all mcRNAs in the UAG motif of the APC gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base site.
[0078] [Figure 23] Figure 23 shows the mcRNA editing levels of a different GAG motif design for the APC gene in one embodiment of this application. The black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, and the horizontal and vertical axes show the editing levels in two biological replication experiments for each mcRNA, where len41 indicates an mcRNA length of 41 nt and len31 indicates an mcRNA length of 31 nt.
[0079] [Figure 24]This invention relates to mcRNA targeting and detargeting editing of an alternative GAG motif design for the APC gene in one embodiment of this application. Figure 24 shows a heatmap of targeted site editing and upstream / downstream detargeting editing of all mcRNAs in an alternative GAG motif for the APC gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base sites.
[0080] [Figure 25] Figure 25 shows the mcRNA editing levels of the TP53 gene CAC motif design in one embodiment of this application. The figure shows mcRNA editing levels of different lengths in the TP53 gene CAC motif in one embodiment of this application, where black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0081] [Figure 26] This figure shows the mcRNA target site and detargeting edit of the TP53 gene CAC motif design in one embodiment of this application. Figure 26 shows a heatmap of target site editing and upstream / downstream detargeting edit of all mcRNAs in the CAC motif of the TP53 gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base site.
[0082] [Figure 27]Figure 27 shows the mcRNA editing levels of the TP53 gene CAG motif design in one embodiment of this application. The black dots indicate that only the editing site of the control mcRNA without secondary structural features has a single A-C base mismatch, the horizontal and vertical coordinates indicate the editing levels in two biological replication experiments for each mcRNA, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0083] [Figure 28] This figure shows the mcRNA target site and detargeting edit of the TP53 gene CAG motif design in one embodiment of this application. Figure 28 shows a heatmap of target site editing and upstream / downstream detargeting edit of all mcRNAs in the CAG motif of the TP53 gene in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base sites.
[0084] [Figure 29] Figure 29 shows the mcRNA editing levels of the TP53 gene CAU motif design in one embodiment of this application. The figure shows mcRNA editing levels of different lengths on the TP53 gene CAU motif in one embodiment of this application, where black dots indicate that the control mcRNA without secondary structural features has a single A-C base mismatch only at the editing site, the horizontal and vertical coordinates indicate the editing level of each mcRNA in two biological replication experiments, len41 indicates that the mcRNA length is 41 nt, and len31 indicates that the mcRNA length is 31 nt.
[0085] [Figure 30]This figure shows the mcRNA target site and detargeting edit of the TP53 gene CAU motif design in one embodiment of this application. Figure 30 shows a heatmap of target site editing and upstream / downstream detargeting edit of all mcRNAs of the TP53 gene CAU motif in one embodiment of this application. Here, the horizontal coordinate 0 is the target site, -1 to -20 is the upstream section of the editing site, and 1 to 20 is the downstream section of the editing site. The vertical coordinate is the editing level at different A bases of each mcRNA, arranged downwards from the highest edited sequence at the target site, and the color from blue to red indicates the trend of change from low to high editing levels at the A base site.
[0086] [Figure 31] This is a verification analysis of the mcRNA editing level by high-throughput sequencing in one embodiment of this application. Figure 31 shows the verification of the mcRNA editing level obtained by high-throughput sequencing in one embodiment of this application using a single low-throughput method. The horizontal axis shows the editing level obtained by measuring each mcRNA using the high-throughput sequencing method, and the vertical axis shows the editing level obtained by synthesizing each mcRNA as a modified RNA individually, transfecting it into cells, and performing Sanger sequencing. A larger correlation coefficient R2 indicates a higher correlation between the horizontal and vertical axes.
[0087] [Figure 32] Figure 32 shows the original Sanger sequencing results for the editing level at GAPDH of a single mcRNA in one embodiment of this application. Figure 32 shows the original peak chart of the editing level obtained by transfection of a single modified mcRNA into cells and sequencing in one generation in one embodiment of this application. As shown in Figure 32, different bases correspond to different colored peaks, with A bases corresponding to green peaks and G bases corresponding to black peaks. When different colored peaks appear at a site, it indicates the presence of different base compositions at that site. In this result, it indicates the editing level of A bases. The editing level is calculated by dividing the peak area of the G base by the sum of the peak areas of the A base and G base at that site.
[0088] [Figure 33] This shows different mcRNA design modes for the mouse GAPDH gene UAG motif in one embodiment of this application. Figure 33 shows that in one embodiment of this application, mcRNAs with different combinations of simulation domains and binding domains at different lengths were designed for the mouse GAPDH gene UAG motif. As shown in Figure 33, the 31nt mcRNA has binding domains of 5nt at both ends and a 21nt simulation domain in the middle; the 41nt mcRNA has binding domains of 10nt at both ends and a 21nt simulation domain in the middle; the 51nt mcRNA has binding domains of 10nt at both ends and a 31nt simulation domain in the middle; the 61nt mcRNA has binding domains of 10nt at both ends and a 41nt simulation domain in the middle; and the 71nt mcRNA has binding domains of 10nt at both ends and a 51nt simulation domain in the middle.
[0089] [Figure 34] Figure 34 shows the editing levels of mcRNAs of different lengths for the mouse GAPDH gene UAG motif in one embodiment of this application. The editing levels of a single mcRNA were verified by randomly selecting different numbers of mcRNAs from groups of different lengths. The cells used were primary mouse hepatocytes. The horizontal axis represents the mcRNAs of different lengths, and the vertical axis represents the editing level of each mcRNA.
[0090] [Figure 35]Figure 35 shows the original peak chart (biological repeat 1) of the first-generation sequencing Sanger model of the mcRNA editing level of the mouse GAPDH gene UAG motif design in one embodiment of this application. The original peak chart (biological repeat 1) of the editing level obtained by first-generation sequencing after transfecting cells with one modified mcRNA of the mouse GAPDH gene UAG motif design in one embodiment of this application. As shown in Figure 35, different bases correspond to peaks of different colors, with A bases corresponding to green peaks and G bases corresponding to black peaks. When peaks of different colors appear at one site, it indicates the presence of different base compositions at that site. In this result, it indicates the editing level of A bases. The editing level is calculated by dividing the peak area of the G base by the sum of the peak areas of the A base and G base at that site.
[0091] [Figure 36] Figure 36 shows the Sanger original peak chart (biological repeat 2) of the one-generation sequencing of the mcRNA editing level of the mouse GAPDH gene UAG motif design in one embodiment of this application. Figure 36 shows the original peak chart (biological repeat 2) of the editing level obtained by one-generation sequencing after transfecting cells with one modified mcRNA of the mouse GAPDH gene UAG motif design in one embodiment of this application. As shown in Figure 36, different bases correspond to peaks of different colors, A bases correspond to green peaks, and G bases correspond to black peaks. When peaks of different colors appear at one site, it indicates that different base compositions exist at that site. In this result, it indicates the editing level of the A base, and the editing level is calculated by dividing the peak area of the G base by the sum of the peak areas of the A base and G base at that site. [Modes for carrying out the invention]
[0092] The technical concept of this application will be further explained below with specific examples, and these examples will not limit the scope of protection of this application. Non-essential modifications and adjustments made by others based on the concepts of this application will still fall within the scope of protection of this application.
[0093] Unless otherwise specified, the reagents, methods, and apparatus used in this application are ordinary reagents, methods, and apparatus in the art.
[0094] Unless otherwise specified, all reagents and materials used in the following examples are commercially available.
[0095] Materials and methods I. Construction of a human ADAR1 expression plasmid The full-length sequence of human ADAR1 is amplified by PCR, homologous arms are added, the pcDNA3.1 vector is linearized by double enzyme cleavage with NheI and EcoRI, and the inserted fragment is ligated to the vector by homologous recombination.
[0096] II. Construction of mcRNA vector libraries by high-throughput screening First, the full-length GFP fluorescent gene and the gene fragment to be edited are amplified by PCR, homologous recombination sequences are added to both ends, and simultaneous double enzyme digestion with NheI and HindIII is performed without pcDNA3.1 loading to recover the linear vector, and homologous recombination is performed on the vector and insertion fragment by multiple fragment recombination to complete the construction of the reporter gene vector. The mcRNA mixed sequence pool is amplified by PCR, homologous arms are added, and the reporter gene vector is linearized by double enzyme digestion with BamHI and EcoRI. The mcRNA sequence pool is inserted into the gene vector using the homologous recombination method, transformed and coated, and then all monoclonal antibodies are collected and plasmid extraction is performed to finally obtain a plasmid library of the mcRNA sequence pool.
[0097] III. Culture of the HEK293 cell line Cell subculturing experiment steps: (1) The super clean bench was disinfected by UV irradiation for 30 minutes. (2) Cells awaiting passage were removed from the incubator, washed three times with PBS, lightly added and gently aspirated to prevent cell detachment, and then 0.25% trypsin was added in an amount sufficient to submerge the surface of the cells. After the cells became round, a culture medium containing serum was added to complete digestion. (3) The digested cell suspension was transferred to a 15 ml centrifuge tube and centrifuged at room temperature at 1000 rpm for 4 minutes. (4) Carefully aspirate and discard the supernatant, add fresh culture medium, and pipette the cells into a single suspension. (5) Take 20 μl of cell suspension and add it to a 96-well cell culture plate. Add another 20 μl of trypan blue stain, mix uniformly by pipetting, then aspirate 20 μl and add it to a cell counting plate, where the cells are counted under a microscope. (6) Based on the size of the inoculated cell culture plate, the required volume of cell suspension for a specific number of cells was calculated, the corresponding volume of fresh cell medium was added, and the cells were cultured in a cell incubator.
[0098] IV. Steps for cell line transfection experiments: (1) The necessary cell culture plates were set up according to the amount of cells required for the experiment, and a certain amount of culture medium and cells were added to the corresponding culture plate according to the cell passage method. When the cells grew for 24 hours and reached 70% confluence, the cell transfection experiment was performed. (2) The super clean bench was disinfected by UV irradiation for 30 minutes. (3) The transfection reagent mixture was prepared according to the instructions for the transfection reagent Lipofectamine 3000 and allowed to stand at room temperature for 5 minutes. (4) The cells were removed, the transfection solution was carefully added to the cells, and after being gently and uniformly mixed, they were placed in an incubator.
[0099] Construction of a V.ADAR1 overexpressing HEK293 cell line Following the cell passage and transfection steps described above, HEK293 cells were transfected with the ADAR1 expression vector. After 48 hours, the selection drug G418 was added, and the cell solution was changed every two days. Screening of HEK293 cell lines stably transfected with ADAR1 was completed when no viable cells remained in the non-transfected group.
[0100] Total cellular RNA could be extracted using a standard RNA extraction kit, ensuring that the RNA was not significantly degraded. Reverse transcription was performed using a standard reverse transcription kit. During the first PCR amplification, sequencing adapter sequences were added to both ends of the amplification primers. The number of cycles for the first amplification was controlled to less than 20. The first PCR product was verified by agarose gel electrophoresis, and library fragments of the corresponding size were recovered. After purification, a second PCR reaction was performed. The primers for the second PCR contained sequencing adapter sequences and specific base barcode sequences, and the second PCR was generally controlled to within 10 cycles. Agarose gel electrophoresis was performed again on the PCR product to confirm the size of the library fragments, and the gel was cut to recover the purified library. Finally, two-generation sequencing was performed on the target sequencing library.
[0101] VII. Target Sequencing Data Analysis First, the raw data was processed using cutapt software to remove the adapter. The original readings after adapter removal were then compared with the reference sequence. Based on the linker sequence, the reporter gene edits were mapped one-to-one with each gRNA, and the editing level of the reporter gene was calculated to obtain the editing level of the target site of each gRNA.
[0102] VIII. Isolation and culture of primary mouse hepatocytes: (1) Preparation of 1×HBSS solution: Take 100 mL of 10×HBSS and put it in a beaker, add 850 mL of redistilled water, add sodium bicarbonate powder to adjust the pH of the solution to 7.5, and finally bring the volume up to 1000 mL. (2) Preparation of pre-perfusion solution: 500 μL of 0.5 M EGTA solution (final concentration: 0.5 mM) was added to 500 mL of 1 × HBSS solution, filtered for sterilization, and prepared immediately before use. (3) Preparation of perfusion solution: Calcium chloride (final concentration: 3 mM) was added to 500 mL of 1 × HBSS solution, collagenase IV (final concentration: 100 U / mL) was added, dissolved thoroughly, filtered to remove sterilization, and prepared immediately before use. (4) Isotonic Percoll separation solution: 1 mL of 10×HBSS solution was taken and added to 10 mL of Percoll solution to prepare an isotonic Percoll separation solution with a concentration of 90%. (5) The electric constant-temperature water bath was started and preheated to a preheated temperature of 42°C, and the reagents that needed to be preheated included the pre-perfusion solution and the perfusion solution. (6) Surface disinfection: Surgical instruments and operating table were disinfected with 75% alcohol. (7) DMEM was pre-cooled on ice. (8) The mice were anesthetized by intraperitoneal injection. (9) After the mice were completely anesthetized, they were fixed to a foam plate and disinfected with 75% alcohol. (10) The skin and peritoneum of the mouse's abdomen were cut along the midline of the abdomen and fixed with a needle. (11) Using a cotton swab, the mouse's large and small intestines were rotated to the right to expose the mouse's portal vein (a vein located between the liver and stomach, through which blood flows into the liver) and inferior vena cava (a vein through which blood flows out of the liver). (12) The mouse liver was gently shaken with a cotton swab to locate the superior vena cava at the boundary between the liver and the diaphragm, and the superior vena cava was clamped with a venous clamp to block the circulating blood flow. (13) The retaining needle was inserted into the hepatic portal vein, and when blood reflux was observed in the tube as it exited the needle core, it indicated that insertion into the blood vessel was successful. The tube of the peristaltic pump was then connected to the retaining needle, and the inferior vena cava was severed. (14) The peristaltic pump was activated and the pre-perfusion fluid was perfused into the mouse liver. The pre-perfusion fluid was kept in a 42°C water bath, the perfusion rate was 5 mL / min, and the perfusion time was 5 min. During the perfusion process, the inferior vena cava was compressed with a cotton swab every 30 seconds to fill the mouse liver with the pre-perfusion fluid (the liver swelled due to the filling of water). The liver should have turned white after perfusion was complete, and the fluid that drained from the inferior vena cava did not appear red. (15) The peristaltic pump was stopped, the pre-perfusion fluid was replaced with the perfusion fluid (the perfusion fluid was kept in a 42°C water bath), and the liver was continuously perfused. The flow rate was adjusted to 3 mL / min, and the perfusion time was 8 min. During the perfusion process, the inferior vena cava was compressed with a cotton swab every 30 seconds to fill the mouse liver with the perfusion fluid. The liver after injection was flexible and inelastic, and a crack pattern was observed on the surface of the liver. (16) Gently separate the liver with scissors and tweezers (at this time, the liver is sufficiently fragile to prevent the envelope from breaking and hepatocytes from leaking out), place it in a culture dish containing DMEM, gently wash the surface of the liver, and remove the gallbladder (all steps are performed on ice). (17) The liver was transferred to another culture dish containing clean DMEM, and the liver capsule was gently torn with elbow forceps to vibrate the liver, causing hepatocytes to leak out and the entire solution to become cloudy (all steps were performed on ice). (18) The liver was vibrated until all hepatocytes had been released, at which point only fibrous tissue remained in the liver. The washing solution was aspirated with a pipette and filtered through a 70 μm cell mesh into a 50 mL centrifuge tube (the entire procedure was performed on ice). (19) When 50g was centrifuged at 4°C for 2 minutes, a large amount of liver cells settled at the bottom. (20) Remove the supernatant, resuspend the hepatocytes in 7.5 mL of DMEM, then add 5 mL of isotonic Percoll separation solution, mix thoroughly and homogeneously, and centrifuge at 200 g at 4°C for 15 min. (21) After density separation, the hepatocytes that settled at the bottom were active cells, while the cells suspended in the solution were dead cells and inactive cells, and these were removed by vacuum aspiration of the supernatant. (22) The hepatocytes were rinsed twice with 15 mL of DMEM, the hepatocytes were placed on a plate, the precipitate was resuspended with culture medium, the cells were counted, the cell density was adjusted to 2 × 10^5 cells / mL, and the cells were inoculated into a 12-well cell culture plate. (23) After 4 hours, once the hepatocytes had fully adhered, the medium was replaced with hepatocyte maintenance medium and used for the next experiment.
[0103] IX. Preparation of first-generation sequencing Sanger samples First, total RNA from cells was extracted using a conventional commercially available RNA extraction kit and quality control was performed. Before the reverse transcription reaction, the purified total RNA was digested with DNA enzymes. After inversion, the downstream sequence of the editing site was amplified and purified using a standard PCR amplification enzyme, and finally, Sanger sequencing was performed.
[0104] X.Sanger Data RNA Editing Level Measurement The original sequence files obtained by Sanger sequencing required measurement of the editing level of the edited sites. This application uses EditR software (https: / / moriarity lab.shinyaps.io / editr_v10 / ) to perform the calculation and analysis. This software analyzes the peak status of each base (A, T, G, and C) for each base in the upstream and downstream regions of the edited site, and the editing level was calculated as 100 * (1 - A peak value / sum of A + T + C + G peak values).
[0105] XI. Synthesis of modified mcRNA In several examples, the chemically synthesized modified skeleton and base sequence were used. The main types included phosphorothioate skeleton modification (structure represented by formula I) and ribose LNA modification (structure represented by formula II). [ka]
[0106] Example 1: A 151nt mcRNA was designed using the MASSA method, and a specific region of the human GAPDH gene was edited. In this embodiment, by simulating the structural characteristics of naturally occurring human high-editing double-stranded RNA substrates, we developed a MASSA (Mimic ADAR Substrates Structure Approach) design method (Figure 2). By forming mcRNA with specific structural characteristics in relation to the upstream and downstream sequences of the target site, we achieved efficient targeted RNA editing.
[0107] In this embodiment, the design of the present application was verified by selecting a site where the motif at position 1323 of the human-derived GAPDH gene is UAA. An mcRNA was designed using a total length of 151 nt, which is an extension of 75 nt upstream and downstream from the editing site. In this embodiment, simulation domains and binding domains of different lengths were designed, with simulation domain lengths of 41 nt, 61 nt, and 91 nt, and the corresponding binding domain lengths at both ends of the domain were 55 nt, 45 nt, and 30 nt, respectively. In this embodiment, a control mcRNA without secondary structure characteristics was also designed. A C base was placed only at the position corresponding to the editing site, and all other positions were base complements. The specific steps of this embodiment are as follows.
[0108] 1) Preparation of mcRNA reference substrate structure Based on the UAA editing motif selected in this embodiment, the reference substrate secondary structure of the corresponding UAA motif was selected, and secondary structure features of the corresponding length were extracted (shown in Table 1). The chain in which the reference substrate editing site is located was the reference substrate editing chain, and the other was the reference substrate complementary chain.
[0109] 2) Construction of mcRNA sequences that pair complementaryly with the target editing site. We selected simulation domains of corresponding lengths by extending the A base of the UAA motif in the GAPDH that needed editing both upstream and downstream. The strand containing the site to be edited was designated as the target editing strand, and the strand containing the mcRNA was designated as the mcRNA strand.
[0110] 3) Design of mcRNA sequence simulation domain Each target edited strand and mcRNA sequence strand were compared one-to-one with the base pairs of the reference substrate edited strand and its complementary strand. If the bases of both edited strands were the same, the mcRNA strand sequence was directly set as the reference substrate complementary strand sequence. If the bases of both strands were different, they were designed to have the same structure. For example, if there was a mismatch of one base at a certain site, the base at the corresponding position in the mcRNA sequence strand was changed to create the mismatch. In this way, the sequence design of the entire simulation domain was completed.
[0111] 4) Design of mcRNA sequence binding domains Based on Step 3, the sequence design of the entire mcRNA was completed by directly extending the complementary pairing sequences at both ends of the mcRNA. In this embodiment, the lengths of the corresponding binding domains at both ends were 55 nt, 45 nt, and 30 nt, respectively.
[0112] 5) Derivation of mcRNA sequence library The minimum free energy of the pairing region between the mcRNA and the target editing sequence was calculated using RNAfold software, sequence information was derived, and the design of the mcRNA sequence library was completed (see Table 2 for the mcRNA sequences designed in this example, specifically mcRNA661 to mcRNA1205, and the mcRNA sequences are shown in Table 1 with reference to their secondary structures).
[0113] 6) Verification of the effectiveness of high-throughput mcRNA sequence editing The synthesized mcRNA sequence library was constructed in a plasmid vector, and the plasmid was then transfected into HEK293 cells overexpressing ADAR1. After 48 hours, RNA was extracted, and the target-constructed library was sequenced to analyze the editing level of the target site corresponding to each mcRNA (shown in Figure 3).
[0114] Table 1: Names and Characteristics of Reference Secondary Structures * [Table 1-1] [Table 1-2] Table 1-3 Table 1-4 Table 1-5 Table 1-6 Table 1-7 Table 1-8 Table 1-9 Table 1-10 Table 1-11 Table 1-12 Table 1-13 Table 1-14 Table 1-15 Table 1-16 Table 1-17 Table 1-18 Table 1-19 Table 1-20 Table 1-21 Table 1-22 Table 1-23 Table 1-24 Table 1-25 Table 1-26 Table 1-27 Table 1-28 Table 1-29 Table 1-30 Table 1-31 Table 1-32 Table 1-33 Table 1-34 Table 1-35 Table 1-36 [Table 1-37] [Table 1-38] [Table 1-39] [Table 1-40] *Note: The structural features of the double-stranded secondary structures described in Table 1 include mismatches (indicated by M), fluctuating base pairs (indicated by W), internal loops (indicated by Li), deletions (indicated by D), protrusions (indicated by Ii), and Watson-Crick base pairs (indicated by P). The aforementioned double-stranded secondary structure features are represented by X. X = [n1, n2, n3…]. The numerical values such as n1, n2 indicate the position of the X structure, where the editing site is 0, -n indicates the nth base upstream of the editing site, and n indicates the nth base downstream of the editing site.
[0115] If the double-stranded secondary structure X is M, then M=[n1,n2,n3…] indicates that mismatch structures exist at positions such as n1, n2, and n3 from the editing site. If the double-stranded secondary structure X is W, then W=[n1,n2,n3…] indicates that fluctuation base pair structures exist at positions such as n1, n2, and n3 from the editing site. If the double-stranded secondary structure X is D, then D=[n1,n2,n3…] indicates that deletion structures exist at positions such as n1, n2, and n3 from the editing site. If X is Ii, then Ii=[n1,n2] (where i is any positive integer) indicates that i bases are inserted between n1 and n2 from the editing site. If the double-stranded secondary structure X is Li, then Li=[n1~n2,n3~n4…] (where i is any positive integer) indicates that internal loops of size i nt are formed in the positional intervals such as n1 and n2 and n3 and n4 from the editing site. In any of the double-stranded secondary structures, unless X is one or more of M, W, D, Ii, or Li, the position is always P.
[0116] Table 2 is shown at the end of the specification and is part of this application, and is incorporated herein by reference in its entirety.
[0117] As the results show, the mcRNAs designed by simulating the secondary structure of this application performed better than mcRNAs without structure in all combinations of simulation domains and binding domains. Simulation domains of different lengths from 41nt to 91nt did not show a significant effect on the editing level, and the overall mcRNA editing level was high on average, ranging from 30% to 70%. This embodiment demonstrates that a 151nt mcRNA designed based on the secondary structure of this application can screen mcRNA sequences without structure.
[0118] Example 2: A 41nt mcRNA was designed using the MASSA method to edit different regions of the human GAPDH gene. The objective of this embodiment is to design a series of mcRNAs using the MASSA method based on the secondary structure characteristics of this application for different base motifs and to verify their editing effects. The names of the reference secondary structures and mcRNA sequences are shown in Table 2. The reference secondary structure characteristics are shown in Table 1. In this embodiment, 10 A bases at GAPDH sites were randomly selected, and the editing effects of the mcRNAs designed based on the secondary structure characteristics were verified. The aforementioned 10 sites are as follows: the motif located at position 787 of GAPDH is the UAG motif, and the mcRNAs designed using this motif are mcRNA1206 to mcRNA1871, with the referenced substrate secondary structure and characteristics shown in Table 1; the motif located at position 1257 is the UAU motif, and the mcRNAs designed using this motif are mcRNA3881 to mcRNA4540, with the referenced substrate secondary structure and characteristics shown in Table 1; the motif located at position 1275 is the UAC motif, and the mcRNAs designed using this motif are mcRNA4541 to mcRNA5200, with the referenced substrate secondary structure and characteristics shown in Table 1; and the motif located at position 1323 is the UAA motif, and the mcRNAs designed using this motif are mcRNA1 to mcRNA660. Table 1 shows the referenced substrate secondary structures and characteristics. The motif located at position 816 is the AAC motif, and the mcRNAs designed using this motif are mcRNA8157 to mcRNA8789. The referenced substrate secondary structures and characteristics are shown in Table 1. The base located at position 873 is the AAG motif, and the mcRNAs designed using this motif are mcRNA7516 to mcRNA856. The referenced substrate secondary structures and characteristics are shown in Table 1. The mcRNAs designed using the CAC motif at position 860 are mcRNA5861 to mcRNA6520. The referenced substrate secondary structures and characteristics are shown in Table 1. The mcRNAs designed using the CAA motif at position 971 are mcRNA5201 to mcRNA5860. The referenced substrate secondary structures and characteristics are shown in Table 1.
[0119] In this example, different editing modes were designed to further verify the editing effect of mcRNA designed based on secondary structure. The specific designs are shown in Figure 4. There are mainly three modes in this design. In the first mode, the length of the simulation domain was determined by extending 6 bases upstream and 14 bases downstream from the editing site. In the second mode, the length of the simulation domain was determined by extending 10 bases upstream and 10 bases downstream from the editing site. In the third mode, the length of the simulation domain was determined by extending 14 bases upstream and 6 bases downstream from the editing site. In the first mode, there was a 20-base binding domain downstream of the editing site, in the second mode, there were 10-base binding domains at each end, and in the third mode, there was a 20-base binding domain upstream of the editing site. In each mode, there was a structureless mcRNA of the same length with one base AC mismatch at the editing site. The length of the reference design was 41 nt in all cases. The sequence was designed and the mcRNA editing level verified by referring to the method of Example 1 (Figure 5-12).
[0120] As the results show, the combinations of simulation domain and binding domain lengths in the different modes of this embodiment allow for screening a larger number of sequences than non-structured mcRNA sequences at different sites and motifs. Specifically, as shown in Figure 5, in the UAN motif, the non-structured mcRNA of the UAU motif had editing levels ranging from 30% to 55% in the three modes described above, while the editing levels of non-structured mcRNAs of the other three motifs were all below 30%. In the UAG motif, the mcRNA simulating the designed secondary structure features had a 59% higher rate than one of the three non-structured mcRNAs, and in the UAA motif, the mcRNA simulating the designed secondary structure features had a 76% higher rate than one of the three non-structured mcRNAs. To further analyze the obtained mcRNA editing effects, parallel comparisons were performed on editing at the target site and upstream / downstream A base positions of the mcRNA at each site. The editing effects were compared using heatmaps for editing and detargeting of each sequence (Figure 6). The results showed that the UAN motif had very high editing efficiency at the target site, with very weak detargeting upstream and downstream. This indicated that mcRNAs designed by secondary structure have relatively high specificity and target editing ability, and the same was true for the UAU and UAC motifs. The difference was that in the UAA motif, highly edited mcRNAs showed some editing even at a single A base downstream of the editing site.
[0121] The CAN motif shown in Figure 7 includes the CAA motif and the CAC motif. The mcRNA editing level without the CAA motif structure was approximately 5%, and the mcRNA editing level without the CAC motif structure was less than 5%. McRNAs with structural characteristics designed based on the secondary structure of this application achieved editing of over 20%, and the proportion of mcRNAs with editing levels exceeding one of the three types of mcRNAs without structure accounted for 55% in the CAA motif and 94% in the CAC motif. Editing of the target site in the CAA motif was significantly higher than upstream and downstream detargeting, while there was some degree of detargeting in the CAC motif (shown in Figure 8).
[0122] The AAN motif shown in Figure 9 includes the AAG and AAC motifs. In both motifs, the editing level of mcRNA without structure was approximately 5%. McRNAs with structural properties designed based on the secondary structure of this application could achieve up to 40% editing, and the percentage of mcRNAs with an editing level exceeding one of the three types of mcRNAs without structure accounted for 95% in the AAG motif and 61% in the AAC motif. The mcRNAs in the AAG motif showed almost no detargeting edit, while the mcRNAs in the AAC motif showed some detargeting edit at two sites downstream of the editing site (shown in Figure 10).
[0123] The GAN motifs shown in Figure 11 include GAG and GAA motifs. In both motifs, the editing level of mcRNA without structure was less than 5%. The mcRNA without structure in the GAA motif had little editing effect. mcRNAs with structural characteristics designed based on the secondary structure of this application achieved a maximum editing level of 20% in the GAG motif and a maximum editing level of 8% in the GAA motif. The proportion of mcRNAs with an editing level exceeding one of the three types of mcRNA without structure accounted for 74% in the GAG motif and 94% in the GAA motif. Both the mcRNAs in the GAG and GAA motifs showed some degree of detargeting editing. In the GAA motif, the editing level of specific mcRNA target sites was not high, but the detargeting after the editing site was higher than the editing effect of the target site (shown in Figure 12).
[0124] The results of this embodiment demonstrate that the mcRNAs obtained by the secondary structure and mcRNA design method of this application can efficiently edit target sites of different motifs, and that the editing level of mcRNAs with a structure of more than 95% at specific sites is higher than that of mcRNAs without a structure.
[0125] Example 3: Different lengths of mcRNA were designed using the MASSA method to edit different regions of the human ATP7B gene. The objective of this embodiment is to verify that mcRNAs with high editing activity can be designed in the human-derived ATP7B gene based on the secondary structure features of this application. Mutations in the ATP7B gene are the main cause of Wilson's disease. In this embodiment, the A base at position 1532 on the ATP7B gene was selected, and a single base displacement, a C base to a T base, was mutated upstream of this site, resulting in a UAG motif where the edit site after the mutation is located. Additionally, a G to A mutation site at position 3316 was selected, resulting in an AAU motif. In this embodiment, a total of 31nt and 41nt upstream and downstream of the edit site were selected, and mcRNAs were designed using the MASSA method corresponding to the substrate secondary structure of a specific motif. The editing level at the target site of each mcRNA was measured using a high-throughput method. In this embodiment, the mcRNA sequences designed for the UAG motif on the ATP7B gene are mcRNA2870 to mcRNA5860 (Table 2), and the reference secondary structure is shown in Table 1.
[0126] As can be seen from the results, as shown in Figure 13, mcRNAs with a 41nt structural feature in the UAG motif achieved a maximum editing level of 50%, while mcRNAs without the 41nt structural feature had an editing level of less than 5%. The editing ability of mcRNAs with a 31nt structural feature in the UAG motif was significantly lower than that of mcRNAs shorter than 41nt, reaching a maximum editing level of around 20%, and the editing level of mcRNAs without the 31nt structure was also less than 5%. Compared to mcRNAs without structure, the proportion of mcRNAs with structure that had a higher editing level than those without structure was 51% for 41nt and 17% for 31nt, respectively. As shown in Figure 14, all mcRNAs in the ATP7B gene UAG motif had a much more efficient editing level at the target site than upstream and downstream detargeting, with little detargeting editing only upstream of the editing site.
[0127] As shown in Figure 15, in the AAU motif of the ATP7B gene, both 41nt and 31nt mcRNAs exhibited low editing ability, with editing levels of 10% or less in both cases. Compared to mcRNAs with structure, mcRNAs without structure had almost no editing ability. The proportion of mcRNAs exceeding the editing ability of mcRNAs without structural features was 53% for 41nt mcRNAs, but only 14% for 31nt mcRNAs. As shown in Figure 16, there was some detargeting edit at the A base upstream of the editing site in the AAU motif of the ATP7B gene. This example demonstrates that even when mcRNAs without structure cannot be edited, it is possible to perform up to approximately 15% editing on difficult-to-edit sites by simulating the secondary structure of this application.
[0128] These results further demonstrate that the secondary structure of this application can be applied to different genes to edit different motifs.
[0129] Example 4: mcRNA is designed using the MASSA method to edit a human-derived FGFR gene. The objective of this embodiment is to verify that mcRNAs can be designed to be effectively edited in the human-derived FGFR gene based on the secondary structure of this application. FGFR gene mutations are the cause of diseases such as achondroplasia, dysplasia of the ossicles, and dwarfism. In this embodiment, the site where a G-to-A mutation occurred at the G358R site on the FGFR gene was selected as the editing site, so the motif where the edited site is located after the mutation is the CAG motif. In this embodiment, a total of 31nt and 41nt upstream and downstream of the editing site were selected, and mcRNAs were designed using the MASSA method corresponding to the substrate secondary structure of a specific motif, and the editing level of the target site of each mcRNA was measured using a high-throughput method. In this embodiment, the mcRNA sequences designed for this site were mcRNA13097 to mcRNA14086 (shown in Table 2). The reference secondary structure is shown in Table 2.
[0130] As can be seen from the results, as shown in Figure 17, the editing level of mcRNAs without structure at 41nt and 31nt lengths in the CAG motif of the FGFR gene was only about 2%, while mcRNAs designed based on the secondary structure of this application achieved up to 35% editing for 41nt lengths and 10% editing for 31nt mcRNAs. In all gRNAs, there was significant detargeting at two sites upstream of the editing site, but no detargeting editing downstream of the editing site. Overall, gRNAs designed based on the 41nt target sequence had editing levels of less than 20% for most sequences, gRNAs designed based on the 31nt target sequence had editing levels of less than 10% for most sequences, and the proportion of 41nt mcRNAs with higher editing levels than mcRNAs without structure was about 50%. As shown in Figure 18, the mcRNAs of this CAG motif of FGFR had no detargeting downstream of the editing site, but there was low detargeting editing upstream of the editing site. This embodiment further demonstrates that it is possible to design mcRNAs that can be edited with high efficiency by the secondary structure of this application in different genes and different motifs.
[0131] Example 5: mcRNA is designed using the MASSA method and the human-derived APC gene is edited. The objective of this embodiment is to verify that it is possible to design mcRNAs that can be edited with high efficiency based on the secondary structure of this application in the human-derived APC gene. The APC gene is a tumor suppressor gene, and in some cases, mutations result in a loss of protein function. For example, if the G base at position 2627 of the gene is mutated to a T base, that position and the next two bases form a stop codon, affecting protein translation. In this embodiment, the A base at the position immediately following that site is selected for editing, and the motif located therein is GAG. Similarly, a mutation from a C base to a T base at position 4099 of the gene also forms a stop codon, and the A base immediately following that site is selected as the second editing site, with the motif located there being UAG. Another high-frequency mutation site is a mutation from a C base to a T base at position 4348, which similarly forms a stop codon, and the A base at the second position immediately following that site is selected as the editing site, with the motif located being GAG. In this example, a total of 31nt and 41nt were selected upstream and downstream of the editing site, and mcRNAs were designed using the MASSA method corresponding to the substrate secondary structure of a specific motif. The editing level at the target site of each mcRNA was measured using a high-throughput method. The mcRNA sequences designed by the first editing site were mcRNA10110 to mcRNA11094 (shown in Table 2), and their corresponding reference secondary structure characteristics are shown in Table 1. The mcRNA sequences designed by the second editing site were mcRNA1872 to mcRNA2869 (shown in Table 2), and their reference secondary structure characteristics are shown in Table 1. The mcRNA sequences designed by the third editing site were mcRNA11095 to mcRNA12090 (shown in Table 2), and their reference secondary structure characteristics are shown in Table 1.
[0132] As can be seen from the results, as shown in Figures 19 to 24, the 41nt and 31nt mcRNAs without structure at the UAG motif of the APC gene both showed only about 10% editing, while the mcRNAs designed based on the secondary structure of this application achieved up to about 80% editing for the 41nt length and up to 60% editing for the 31nt mcRNA (shown in Figure 21). The editing levels at two other GAG sites different from this site were generally lower than those of the UAG motif, with the 41nt mcRNA at the first site showing less than 30% overall editing and the 31nt mcRNA showing less than 10% overall editing, and the mcRNAs without structure at this site showing almost no editing for both the 41nt and 31nt lengths (shown in Figure 19). The third editing site showed an overall editing status similar to that of the first editing site (shown in Figure 23). At these three sites, mcRNA edited at the target site had a low detargeting effect, with the UAG motif having the lowest detargeting effect, and the two GAG motif sites having detargeting editing on a small amount of mcRNA upstream and downstream (shown in Figures 20, 22, and 24). At the three editing sites in the APC gene, the 41nt mcRNA designed by the first GAG motif accounted for 36% of the total mcRNA with an editing level higher than mcRNA with no structure, while the 31nt mcRNA accounted for a lower proportion, only 8%. The 41nt mcRNA designed by the second UAG motif accounted for a higher editing level than mcRNA with no structure, accounting for 81% of the total mcRNA and 63% of the 31nt mcRNA. The 41nt mcRNA designed by the third GAG motif accounted for a higher editing level than mcRNA with no structure, accounting for 61% of the total mcRNA and 29% of the 31nt mcRNA. This embodiment further demonstrates that it is possible to design mcRNAs that efficiently edit different genes and different motifs based on the secondary structure of this application.
[0133] Example 6: The human-derived TP53 gene was edited by designing mcRNA using the MASSA method. The objective of this embodiment is to verify that it is possible to design mcRNAs that efficiently edit the human-derived TP53 gene based on the secondary structure of this application. The TP53 gene is a tumor suppressor gene, and in some cases, mutations result in a loss of protein function. For example, if the G base at position 524 of the gene mutates to an A base, the protein's coding sequence changes. In this embodiment, this site is selected for editing, and the resulting motif is CAC. Additionally, a mutation from a G base to an A base always occurs at position 743 of the gene, and this site is selected as the second editing site, with the resulting motif being CAG. Another high-frequency mutation site is the mutation of the C base to a T base at position 818, and this site is selected as the editing site, with the resulting motif being CAU. In this example, a total of 31nt and 41nt were selected upstream and downstream of the editing site, and mcRNAs were designed using the MASSA method corresponding to the substrate secondary structure of a specific motif. The editing level at the target site of each mcRNA was measured using a high-throughput method. The mcRNA sequences designed by the first editing site were mcRNA6521 to mcRNA7515 (shown in Table 2), and their corresponding reference secondary structure characteristics are shown in Table 1. The mcRNA sequences designed by the second editing site were mcRNA14087 to mcRNA15086 (shown in Table 2), and their reference secondary structure characteristics are shown in Table 1. The mcRNA sequences designed by the third editing site were mcRNA15087 to mcRNA16083 (shown in Table 2), and their reference secondary structure characteristics are shown in Table 1.
[0134] As can be seen from the results, the editing motif in the TP53 gene is CAN, as shown in Figures 25 to 30. The CAG and CAU motifs showed a higher overall editing effect than the CAC motif. With the CAC motif, both 41nt and 31nt mcRNAs without structure showed only about 2% editing. mcRNAs designed based on the secondary structure of this application achieved a maximum editing of 35% at 41nt length and a maximum editing of 10% at 31nt length (shown in Figure 25). With the CAG motif, mcRNAs without structure showed less than 10% editing under different length conditions, while mcRNAs designed based on the secondary structure achieved a maximum editing of 70% at 41nt length and a maximum editing of 20% at 31nt length (shown in Figure 27). With the CAU motif, mcRNAs without structure showed almost no editing, while mcRNAs designed based on the secondary structure achieved a maximum editing of 60% at 41nt length and a maximum editing of 30% at 31nt length (shown in Figure 29). These three sites showed that mcRNA edited at the target site had a low detargeting effect, with the CAG motif showing the lowest detargeting effect. Both the CAG and CAU motifs resulted in detargeting editing of only a small amount of mcRNA upstream (shown in Figures 26, 28, and 30). At the three editing sites in the TP53 gene, the 41nt mcRNA designed by the first CAC motif had a higher editing level than mcRNA without structure, accounting for 40% of the total mcRNA, while the proportion of 31nt mcRNA was low, at only 10%. The 41nt mcRNA designed by the second CAG motif had a higher editing level than mcRNA without structure, accounting for 57% of the total mcRNA, with a proportion of 10% for 31nt mcRNA. The 41nt mcRNA designed by the third CAU motif had a higher editing level than mcRNA without structure, accounting for 57%, with a proportion of 13% for 31nt mcRNA. This embodiment further demonstrates that it is possible to design mcRNAs that efficiently edit different genes and different motifs based on the secondary structure of this application.
[0135] The multiple examples described above demonstrate that different secondary structures can design highly edited mcRNAs in different genes. As can be seen from the comparison of the data, the same secondary structure can design highly edited mcRNAs in different genes. For example, the UAG motif-based sites include the GAPDH, APC, and ATP7B genes. As shown in Table 3, each of the listed secondary structures can design higher sequences than mcRNAs without the structure in different sites of the three genes, further demonstrating that the secondary structures of this application can design highly edited mcRNAs at any target site.
[0136] Table 3: Levels of mcRNA editing in different genes with the same secondary structure [Table 2-1] [Table 2-2]
[0137] Example 5: The editing effect of mcRNA obtained by high-throughput in HEK293 cells is verified. The measurement of the editing level at the target site of mcRNAs designed based on the secondary structure of this application was completed by high-throughput sequencing. To further demonstrate the truthfulness of the results of mcRNAs with different editing levels obtained by high-throughput sequencing, the mcRNA sequences measured on the GAPDH gene UAG motif in Example 2 were each validated in this example. In this example, mcRNA sequences with editing levels ranging from 3.93% to 21.66% were selected by high-throughput screening. The sequence information is GAPDH-UAG-S1 to GAPDH-UAG-S13 in Table 4, and the chemically modified gRNA sequences correspond to GAPDH-UAG-M1 to GAPDH-UAG-M13. For the sequences screened by high-throughput, the individual chemically modified mcRNAs were transfected into cells and the editing level was validated (Figures 31, 32). The stability of the modified mcRNA is higher than that of mcRNA expressed on a plasmid, and the editing levels are all higher than those screened by high-throughput methods. The modification strategy for the chemically modified mcRNA involves modifying the bases with thiophosphate and adding three LNA modifications to each end. Analysis of the editing levels screened by high-throughput methods and those of single-modified mcRNA showed a strong correlation, with a correlation coefficient of 0.83 (Figure 31), further demonstrating the accuracy of the method described in this application.
[0138] The editing levels of GAPDH-UAG-S1 to GAPDH-UAG-S12 in Table 4 were measured using a high-throughput mixed system. The editing levels of GAPDH-UAG-M1 to GAPDH-UAG-M12 were measured using a single RNA system.
[0139] Table 4: Sequence information table verified by high-throughput screening [Table 3]
[0140] Example 6: Design mcRNA using the MASSA method and edit the mouse-derived GAPDH gene. To verify that this application is applicable to different species, this embodiment designs five different mcRNA lengths of 31nt, 41nt, 51nt, 61nt, and 71nt for the mouse-derived GAPDH1222 site using the MASSA method, performs single-mcRNA validation, and the specific combinations of binding domain length and simulation domain length are shown in Figure 33. For the 31nt mcRNA, two were simulated and selected based on the secondary structure, with the editing site located at its 3' end and binding domains at both ends consisting of 5 bases each. For the 41nt mcRNA, six were selected, with the editing site located at its center and binding domains at both ends consisting of 10 bases each. For the 51nt mcRNA, four were selected, with the editing site located at its center and binding domains at both ends consisting of 10 bases each. For the 61nt mcRNA, four were selected, with the editing site at its center and binding domains at both ends consisting of 10 bases each. For 71nt mcRNAs, three were selected, with the editing site located at the center and each end-binding domain consisting of 10 nucleotides. Table 5 shows the reference secondary structure, mcRNA sequence, and modified sequence of the above-mentioned mcRNAs. The purpose of this example is to verify the editing status of mcRNAs designed by the MASSA method in mouse cells and to compare the differences in editing levels of mcRNAs of different lengths. The above method includes the following steps.
[0141] 1. Preparation of the mcRNA reference substrate structure Depending on the upstream and downstream lengths of the editing site that needs to be simulated in the mcRNA, the reference substrate RNAhybrid software extracted a length segment corresponding to the folded structure. The sequence structure information within this segment is the reference secondary structure data, the strand where the reference secondary structure editing site is located is the reference substrate editing strand, and the other strand of the substrate is the reference substrate complementary strand.
[0142] 2. Construction of mcRNA sequences that pair complementaryly with the target editing site. Simulated domains of corresponding lengths upstream and downstream are selected, centered around the mouse-derived GAPDH1222 site. For example, if a 31nt mcRNA contains a 5nt binding domain, the editing site is a complementary pairing sequence extending 1nt downstream from position 25 and 19nt upstream. The strand where the site to be edited is located is the target editing strand, and the strand where the mcRNA is located is the mcRNA strand. If a 41nt mcRNA contains a 10nt binding domain, the editing site is at the center and extends 10nt upstream and downstream.
[0143] 3. Design of mcRNA sequence simulation domain structure Each target edited strand and mcRNA sequence strand are compared one-to-one with the base pairs of the reference substrate edited strand and its complementary strand. If the base pairs of the two edited strands are the same, the mcRNA strand sequence is directly set to the reference substrate complementary strand sequence. If the base pairs are different, the design is made to produce a similar structure. For example, if there is a mismatch of one base at a certain site, the base at the corresponding position in the mcRNA sequence strand is designed to be mismatched. Following these principles, the sequence design of the entire simulation domain is completed.
[0144] 4. Design of mcRNA sequence binding domain structure Based on Step 3, sequences that directly and complementaryly pair with both ends of the mcRNA are extended to complete the overall sequence design of the mcRNA. The length of the binding domain can be customized according to the needs. In this embodiment, 5nt binding domains are designed at each end of the 31nt sequence, and 10nt binding domains are designed at each end of the 41nt, 51nt, 61nt, and 71nt sequences (Figure 33).
[0145] 5. Selection and validation of mcRNA sequences RNAfold software is used to calculate the minimum free energy of the pairing region between mcRNA and the target editing sequence to derive sequence information, complete the design of the mcRNA sequence library, and then mcRNAs of different lengths are randomly selected to verify the editing effect of a single molecule.
[0146] 6. Selection and validation of mcRNA sequences Step 5 randomly selects and validates several mcRNAs of different lengths from the sequence library. In this example, the effect of the relative position of different lengths and editing sites on the editing level is mainly tested. The original mcRNA sequence and the modified mcRNA sequence are shown in Table 5. The synthesized mcRNA was transfected into primary mouse hepatocytes in a 12-well plate at a transfection rate of 40 pmol / well. After 48 hours, total RNA was extracted, and then reverse transcription was performed using random primers. Finally, the upstream and downstream sequences of the target site were amplified by PCR and Sanger sequencing was performed.
[0147] As can be seen from the results shown in Figures 34, 35, and 36, mcRNAs of different lengths derived from mouse can all perform high-level editing of the GAPDH1222 site. All five different lengths of mcRNAs can achieve editing levels exceeding 90%.
[0148] These results indicate that mcRNAs designed using the MASSA method can edit specific sites of specific genes in primary mouse hepatocytes, and that different combinations of simulation domain and binding domain lengths result in different editing effects.
[0149] Table 5: Five different mcRNA lengths designed for the mouse-derived GAPDH1222 site in this example. [Table 4]
[0150] Table 2: List of mcRNA sequences and reference secondary structure names [Table 5-1] [Table 5-2] [Table 5-3] Table 5-4 Table 5-5 Table 5-6 Table 5-7 Table 5-8 Table 5-9 Table 5-10 Table 5-11 Table 5-12 Table 5-13 Table 5-14 Table 5-15 Table 5-16 Table 5-17 Table 5-18 Table 5-19 Table 5-20 Table 5-21 Table 5-22 Table 5-23 Table 5-24 Table 5-25 Table 5-26 Table 5-27 Table 5-28 Table 5-29 Table 5-30 Table 5-31 Table 5-32 Table 5-33 Table 5-34 Table 5-35 Table 5-36 Table 5-37 Table 5-38 Table 5-39 Table 5-40 Table 5-41 Table 5-42 Table 5-43 Table 5-44 Table 5-45 Table 5-46 Table 5-47 Table 5-48 Table 5-49 Table 5-50 Table 5-51 Table 5-52 Table 5-53 Table 5-54 Table 5-55 Table 5-56 Table 5-57 Table 5-58 Table 5-59 Table 5-60 Table 5-61 Table 5-62 Table 5-63 Table 5-64 Table 5-65 Table 5-66 Table 5-67 Table 5-68 Table 5-69 Table 5-70 Table 5-71 Table 5-72 Table 5-73 Table 5-74 Table 5-75 Table 5-76 Table 5-77 Table 5-78 Table 5-79 Table 5-80 Table 5-81 Table 5-82 Table 5-83 Table 5-84 Table 5-85 Table 5-86 Table 5-87 Table 5-88 Table 5-89 Table 5-90 Table 5-91 Table 5-92 Table 5-93 Table 5-94 Table 5-95 Table 5-96 Table 5-97 Table 5-98 Table 5-99 Table 5-100 Table 5-101 Table 5-102 Table 5-103 Table 5-104 Table 5-105 Table 5-106 Table 5-107 Table 5-108 Table 5-109 Table 5-110 Table 5-111
Claims
1. A sequence that forms an RNA editing target site and an editing substrate, The sequence forms a double-stranded secondary structure with the RNA editing target site in the form of a complementary strand, wherein the structure of the complementary strand includes at least one set of features shown in i-x below. i. The edited region has 0 to 4 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 12 deletions at positions -100 to -81. ii. The edited region has 0 to 5 mismatches, 0 to 7 fluctuating base pairs, 0 to 5 internal loops, 0 to 5 protrusions, and 0 to 15 deletions at positions -80 to -61. iii. The edited region has 0 to 5 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 18 deletions at positions -60 to -41. iv. At the -40 to -21 position of the edited site, there are 0 to 4 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 8 protrusions, and 0 to 11 deletions. v. The edited site has 0 to 5 mismatches, 0 to 6 fluctuating base pairs, 0 to 4 internal loops, 0 to 5 protrusions, and 0 to 9 deletions at positions -20 to -1. vi. The edited region has 0 to 5 mismatches, 0 to 7 fluctuating base pairs, 0 to 3 internal loops, 0 to 3 protrusions, and 0 to 5 deletions at positions 0 to 19. vii. The edited region has 0-4 mismatches, 0-7 fluctuating base pairs, 0-3 internal loops, 0-6 protrusions, and 0-9 deletions at positions 20-39. viiii. The edited region has 0 to 5 mismatches, 0 to 6 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 17 deletions at positions 40 to 59. ix. The editing site has 0-5 mismatches, 0-8 fluctuating base pairs, 0-4 internal loops, 0-5 protrusions, and 0-19 deletions at positions 60-79. x. A sequence characterized by having 0 to 4 mismatches, 0 to 6 fluctuating base pairs, 0 to 3 internal loops, 0 to 5 protrusions, and 0 to 19 deletions at positions 80 to 100 of the editing site.
2. The type of mismatch on the complementary chain is one of A-A, A-G, A-C, U-U, U-C, G-G, G-A, C-A, C-C, and C-U. Preferably, the fluctuating base pair on the complementary strand is a G-U base pair. Preferably, the projection on the complementary chain is One base upstream and downstream of the protrusion region are paired complementaryly with the target edited strand, and the two corresponding bases of the target edited strand are consecutive. Preferably, the deletion on the complementary strand has consecutive bases upstream and downstream of the deletion region, and there are a number of bases corresponding to the target edited strand of the deletion region. Preferably, the sequence according to claim 1, characterized in that the internal loop on the complementary strand has one base upstream and one downstream of the internal loop that complementarily pairs with the target edited strand and the complementary strand, and a corresponding number of bases are mismatched in the internal loop.
3. The aforementioned double-stranded secondary structure is one selected from the Struc1 to Struc1196 structures. Preferably, the structural characteristics of the double-stranded secondary structure are shown in Table 1. Preferably, the structural features of the double-stranded secondary structure are: A mismatch is formed between bases represented by M. W represents a fluctuating base pair, An internal loop represented by Li, Deletion represented by D, The protrusion represented by li, and Watson-Crick base pairs represented by P This includes one or more of the following: Preferably, in the characteristics of the double-stranded secondary structure, the numerical values X, X = [n1, n2, n3…], n1, n2, etc. indicate the position of the X structure, where the position of the editing site is 0, -n indicates that this position is the nth base upstream of the editing site, and n indicates that this position is the nth base downstream of the editing site. Preferably, when the double-stranded secondary structure X is M, M = [n1, n2, n3...] indicates that mismatch structures exist at positions such as n1, n2, n3 from the editing site. If the double-stranded secondary structure X is W, then W = [n1, n2, n3...] indicates that fluctuation base pair structures exist at positions such as n1, n2, n3 from the editing site. If the double-stranded secondary structure X is D, then D = [n1, n2, n3…] indicates that deletion structures exist at positions such as n1, n2, n3 from the editing site. If X is Ii, then Ii = [n1, n2], where i is any positive integer, and we show that i bases are inserted between n1 and n2 from the editing site. If the double-stranded secondary structure X is Li, then Li = [n1 to n2, n3 to n4...], where i is any positive integer, and the editorial department shows that an internal loop of size int is formed in the positional intervals such as n1 and n2, n3 and n4. The sequence according to claim 1, characterized in that, in any of the double-stranded secondary structures, the position is always P, except when X is one or more of M, W, D, Ii, and Li.
4. The base motifs of the target RNA editing site and one site upstream and downstream thereof are NAN. Preferably, the NAN is a UAN, AAN, CAN, or GAN. Preferably, the UAN is UAG, UAU, UAC, or UAA. Preferably, the AAN is AAU, AAC, AAG, or AAA. Preferably, the CAN is CAU, CAG, CAC, or CAA. Preferably, the arrangement according to claim 1, characterized in that the GAN is GAU, GAC, GAA, or GAG.
5. When the base motifs of the target RNA editing site and one site upstream and downstream thereof are NAN, the double-stranded secondary structure is Struc1 to Struc1196. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are UAA, the double-stranded secondary structure is Struc1 to Struc143. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are UAG, the double-stranded secondary structure is Struc144 to Struc260, Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are UAU, the double-stranded secondary structure is Struc261 to Struc320. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are UAC, the double-stranded secondary structure is Struc321 to Struc395. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are CAA, the double-stranded secondary structure is Struc396 to Struc515. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are CAC, the double-stranded secondary structure is Struc516 to Struc541. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are AAG, the double-stranded secondary structure is Struc542 to Struc684. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are AAC, the double-stranded secondary structure is Struc685 to Struc767. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are GAA, the double-stranded secondary structure is Struc768 to Struc809. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are GAGs, the double-stranded secondary structure is Struc810 to Struc875. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are AAU, the double-stranded secondary structure is Struc876 to Struc982. Preferably, when the base motifs of the target RNA editing site and one site upstream and downstream thereof are CAG, the double-stranded secondary structure is Struc983 to Struc1085. Preferably, the sequence according to claim 1, wherein the base motifs of the target RNA editing site and one site upstream and downstream thereof are CAU, and the double-stranded secondary structure is Struc1086 to Struc1196.
6. The sequence according to claim 1, characterized in that the complementary strand has bases that pair complementarily at non-mismatch, non-deletion, non-protrusion, non-internal loop, or non-fluctuating base pair sites.
7. The length of the RNA editing target site and the sequence forming the editing substrate is 25 to 300 bp. Preferably, the sequence according to claim 1, characterized in that the length of the sequence is 25 to 151 bp.
8. The RNA in question is an RNA that forms a specific secondary structure with the target RNA to be edited. Preferably, the target RNA to be edited is messenger RNA, ribosomal RNA, transfer RNA, long non-coding RNA, or small RNA. Preferably, the sequence according to claim 1 is characterized in that the low molecular weight RNA is miRNA, prim-miRNA, pre-miRNA, piRNA, siRNA, snoRNA, snRNA, exRNA, or scaRNA.
9. mcRNA containing a simulation domain, The sequence of the simulation domain is the sequence described in any one of claims 1 to 8, and forms the double-stranded secondary structure with the target RNA to be edited. Preferably, the mcRNA further includes a binding domain, the binding domain being located at both ends of the simulation domain and pairing complementaryly with regions other than the double-stranded secondary structure of the target RNA to be edited. Preferably, the length of one side of the binding domain is 0-55 bp. Preferably, the mcRNA is characterized in that the sequence of the mcRNA is one sequence selected from sequence numbers 1-16083.
10. A method for designing mcRNA according to claim 9, Obtaining the reference sequence: Step S1 involves obtaining the nucleotide sequences of the two strands, with the strand containing the editing site in the secondary structure being designated as the reference secondary structure editing strand, and the other strand being designated as the reference secondary structure complementary strand. Step S2 involves obtaining the mcRNA sequence for preparing the target editing site: starting from the site to be edited, the target editing strand, which is the template sequence of the simulation domain, is obtained by extending it upstream and downstream by a predetermined length, and a sequence that pairs complementaryly with the template sequence is designed to become the prepared mcRNA sequence. Design of the mcRNA simulation domain: Step S3 involves comparing the target edited strand sequence and the reference secondary structure edited strand sequence by extending the edited strand from the editing site to both ends, comparing the prepared mcRNA sequence and the reference secondary structure complementary strand sequence, and if the comparison results of the target edited strand sequence and the reference secondary structure edited strand sequence match, the corresponding position of the mcRNA simulation domain is set to the corresponding reference secondary structure complementary strand sequence, and if the comparison results of the target edited strand sequence and the reference secondary structure edited strand sequence do not match, the corresponding position of the mcRNA simulation domain is designed as a sequence that can form the same structure as the corresponding reference secondary structure together with the target editing site, Design of the mcRNA binding domain: Step S4 involves extending complementary pairing sequences of a predetermined length from both ends of the simulation domain to form a binding domain based on the mcRNA simulation domain designed in step S3, A design method characterized by including the following.
11. Use of a sequence or mcRNA according to any one of claims 1 to 9 in the preparation of an RNA editing reagent or composition.
12. An RNA editing reagent or composition characterized by comprising the sequence or mcRNA described in any one of claims 1 to 9.
13. A method for editing target RNA in host cells, The step includes introducing an oligonucleotide having the sequence described in any one of claims 1 to 8 or the mcRNA described in claim 9 into the host cell, Preferably, the target RNA is the ATP7B gene, FGFR gene, TP53 gene, APC gene, SERPINA1 gene, MECP2 gene, or SCN1A gene, and the method is characterized in that.
14. The use of a sequence or mcRNA according to any one of claims 1 to 9 in the preparation of a pharmaceutical product for treating a disease or symptom, The aforementioned disease or condition is a hereditary genetic disorder or a disease or condition associated with one or more acquired gene mutations. Preferably, the disease or condition is a single-gene disorder or condition. Preferably, the disease or condition is a polygenic disease or condition. Preferably, the genes are ATP7B, FGFR, TP53, APC, SERPINA1, MECP2, and SCN1A. Preferably, the diseases or conditions include Wilson's disease, Rett syndrome, Dravet syndrome, cystic fibrosis, Harler syndrome, α-1-antitrypsin (A1AT) deficiency, Parkinson's disease, Alzheimer's disease, albinism, amyotrophic lateral sclerosis, asthma, β-thalassemia, Kadasil syndrome, Charcot-Marie-Tooth disease, chronic obstructive pulmonary disease (COPD), distal spinal muscular atrophy (DSMA), Duchenne / Becker muscular dystrophy, dystrophic epidermolysis bullosa, epidermolysis bullosa, Fabry disease, factor V Leiden-related disease, familial adenoma, polyposis, galactosemia, Gaucher disease, glucose-6-phosphate dehydrogenase, hemophilia, hereditary hemochromatosis, Hunter syndrome, Huntington's disease, inflammatory Use is characterized by the following conditions: intestinal disease (IBD), hereditary polyaggregation syndrome, Leber congenital amaurosis, Lesch-Nyhan syndrome, Lynch syndrome, Marfan syndrome, mucopolysaccharidosis, muscular dystrophy, myotonic dystrophy type I and II, neurofibromatosis, Niemann-Pick disease type A, B, and C, NY-esol-associated cancer, Peutz-Jeghers syndrome, phenylketonuria, Pompe disease, primary eyelash disorder, diseases associated with prothrombin mutations such as the prothrombin G20210A mutation, pulmonary hypertension, retinitis pigmentosa, Sandhoff disease, severe combined immunodeficiency syndrome (SCID), sickle cell anemia, spinal muscular atrophy, Stargardt disease, Tay-Sachs disease, Usher syndrome, X-linked immunodeficiency, and cancer.