Method for screening guide RNA

CN121586774APending Publication Date: 2026-02-27RECORNA (GUANGZHOU) BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380099852.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

In the existing technology, the CRISPR/Cas system has off-target effects and genome instability during the gene editing process, which limits its broad clinical application. It also lacks efficient guide RNA (gRNA) screening methods, making it difficult to design efficient Target sequences for A-to-I editing.

Method used

By designing a guide RNA containing a linker sequence and connecting it to the target RNA to form a transcript, and editing it in cells expressing ADAR proteins, second-generation sequencing technology is used to efficiently and high-throughput detect the editing level of the target RNA, thereby screening out efficient Guide RNA sequence.

Benefits of technology

It achieves efficient A-to-I editing of specific sites, fills the gap in the existing technology that lacks high-throughput screening methods, improves gRNA screening efficiency and editing accuracy, and enhances the clinical application potential of gene therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121586774A_ABST
    Figure CN121586774A_ABST
Patent Text Reader

Abstract

The invention relates to a screening method of guide RNA (Ribonucleic Acid) for editing RNA. The screening method comprises the following steps: (1) connecting the guide RNA and target RNA to the same transcript through a linker; (2) transferring the transcript into a cell through a vector; and (3) extracting RNA of the cells, analyzing the editing level of the guide RNA to a target site, and screening out the target guide RNA. According to the screening method, a substrate sequence and a guide RNA sequence are designed into the same transcript, target RNA editing is directly completed in a cell, and the sequence of each guide RNA and the editing efficiency of a corresponding target site are sequenced through high-throughput sequencing, so that the guide RNA with high editing efficiency is rapidly screened out.
Need to check novelty before this filing date? Find Prior Art

Description

A method for screening guide RNA Technical Field

[0001] The present invention belongs to the field of biomedicine and relates to a method for screening guide RNA. Background Art

[0002] With the rapid development of third-generation sequencing technology in recent years, the final piece of the human genome sequencing puzzle, hailed as the "moon landing project" of life sciences, has been completed. A complete, gapless genome map of nearly 200 million base pairs has been mapped, which means that more and more disease-related gene mutation sites can be more thoroughly uncovered. Previously, over 3,000 gene mutation sites have been found to be closely associated with the occurrence and progression of diseases, but many mutation sites are difficult to identify as suitable drug targets. Furthermore, the same disease may correspond to multiple mutation sites, and these sites are unevenly distributed across the population, making the development of traditional large-molecule and small-molecule drugs even more challenging. At this time, gene therapy is gradually demonstrating its enormous clinical value, directly repairing mutation sites to achieve the purpose of disease treatment.

[0003] CRISPR (Clustered regularly interspaced short palindromic repeats) was first discovered in the bacterial immune system, where it binds to Cas proteins to combat invading viruses. Since the 2012 report that this system can also perform genome editing and repair in mammalian cells, research on CRISP / Cas has mushroomed, with the specific mechanisms and applicable species and tissues being increasingly understood. Currently, CRISPR / Cas technology has been widely applied in biomedical and agricultural research. The CRISPR / Cas system can be designed to insert entire genes, knock down gene expression, or even delete genes, or repair single mutations. However, this system also has its fatal flaws. It requires the simultaneous expression of Cas protein and guide RNA (gRNA) and the provision of homologous repair DNA sequences. In addition, multiple research reports have detected significant off-target effects at the DNA or RNA levels. CRISPR / Cas protein cuts genomic DNA, causing DSB (double-strand break), and off-target effects can cause genomic instability. In addition, the CRISPR / Cas system may also trigger cellular innate immune responses, all of which greatly limit its clinical application.

[0004] A-to-I editing at the RNA level occurs under the mediation of ADAR family proteins. The I base will be recognized as the G base during translation, so the G-to-A mutation in the genome can be repaired at the RNA level through ADAR editing. There are three main subtypes of ADAR family proteins, ADAR1, ADAR2 and ADAR3, of which ADAR3 does not have deaminase activity. The editing events are mainly completed by ADAR1 and ADAR2, and the broad expression of ADAR1 protein in human tissues can ensure its application in multiple tissues or organs. Currently, a variety of methods have been developed that use overexpressed ADAR family proteins or endogenously expressed ADAR family proteins in combination with guide RNA to perform A-to-I editing at specific sites at the RNA level. From the perspective of clinical application, the most valuable method is to use endogenously expressed ADAR family proteins to repair mutation sites, because only the gRNA needs to be designed, and the two more mature technologies are as follows: the first gRNA consists of two segments, one is a 54nt ADAR substrate with structural characteristics plus another customizable targeting sequence (RESTORE method), and the optimized version of this method can reduce off-target editing (CLUSTER method); the second gRNA only has a targeting sequence longer than 100nt that is complementary to the target sequence (LEAPER method). ADAR family proteins prefer to edit double-stranded RNA, but there will also be multiple mutations within the editing region to form a certain secondary structure, which leads to the need to screen the gRNA when designing it, and no more efficient screening method has been reported.

[0005] Summary of the Invention

[0006] In some embodiments, the present invention provides a method for screening guide RNAs that edit RNA, comprising the steps of: (1) each guide RNA to be screened is connected to a linker and is connected to a target RNA on a nucleic acid chain that can form the same transcript, forming a plurality of said nucleic acid chains; (2) the said nucleic acid chains are transferred into cells expressing ADAR proteins; (3) the RNA in the cells is extracted and measured to determine the editing level of the target RNA, and the editing level of the corresponding guide RNA is determined by the barcode sequence.

[0007] In some embodiments, the linker is located between the target RNA and the guide RNA.

[0008] In some embodiments, the barcode sequence is contained in the linker.

[0009] In some embodiments, the nucleic acid chain is sequentially connected to a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, and a guide RNA nucleic acid fragment from upstream to downstream.

[0010] The linker sequence plays the role of a label in the present invention. The linker can preferably be located between the target RNA and the guide RNA. At this time, the linker also plays the role of separating the target RNA and the guide RNA to a certain extent, which is conducive to the folding of the two into a secondary structure. In addition, this connection method is also conducive to application in second-generation sequencing, thereby achieving efficient and high-throughput detection. In order to adapt to the sequence length limit of second-generation sequencing, the sequencing can be cut into two parts: target RNA + linker and linker + guide RNA. Through the labeling effect of the linker, the editing level of the target RNA under the action of the corresponding guide RNA can be learned.

[0011] In some embodiments, the transcript is linked to a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, a guide RNA nucleic acid fragment, and a polyA tail nucleic acid fragment in sequence from upstream to downstream.

[0012] In some embodiments, the reporter gene is selected from the group consisting of yfp gene, gus gene, rfp gene, gfp gene, cfp gene, bfp gene or kanamycin resistance gene, but is not limited thereto.

[0013] In some embodiments, the linker has a length of 26 nt to 2121 nt.

[0014] In some embodiments, the linker comprises SEQ ID NO:549 or SEQ ID NO:557.

[0015] In some embodiments, long linker sequences are added to the screening method to simulate the formation of double strands between intermolecular RNAs for editing.

[0016] In some embodiments, a mixed library of gRNAs of different lengths in the case of short linkers can also be screened for gRNAs with different editing efficiencies using the methods of the present invention.

[0017] In some embodiments, step (3) includes establishing a targeted sequencing library for next-generation sequencing.

[0018] In some embodiments, contacting the nucleic acid strand with the ADAR protein is achieved by introducing the nucleic acid strand into a cell that expresses the ADAR protein.

[0019] In some embodiments, the cell expresses an ADAR protein.

[0020] In some embodiments, the ADAR protein comprises ADAR1 or ADAR2.

[0021] In some embodiments, the guide RNA is capable of recruiting ADAR protein to edit the target RNA in the cell.

[0022] In some embodiments, the cell overexpresses an ADAR protein.

[0023] In some embodiments, the editing is A-to-I editing.

[0024] In some embodiments, the guide RNA is a guide RNA extended upstream and downstream with base A as the center.

[0025] In some embodiments, the guide RNA is a guide RNA with a length of 30 nt to 151 nt obtained by extending 15 nt to 75 nt upstream and downstream of the A base.

[0026] In some embodiments, the guide RNA is a 151nt guide RNA that is extended 75nt upstream and downstream from the A base.

[0027] In some embodiments, the guide RNA is a 28-89 nt guide RNA extending upstream and downstream from base A.

[0028] In some embodiments, the screening method comprises adding a 20-200 nt sequence upstream and downstream of the editing site to the downstream of the reporter gene and inserting the sequence into a vector to obtain a reporter gene vector, and then transferring the reporter gene vector into cells.

[0029] In some embodiments, the screening method is a high-throughput screening method.

[0030] In some embodiments, the present invention provides a method for high-throughput screening of guide RNA (gRNA) sequences that can recruit ADAR proteins and perform A-to-I editing on specific sites, filling the gaps in the prior art.

[0031] In some embodiments, the present invention provides a gRNA screening method established by the above method for any RNA editing system based on ADAR.

[0032] In some embodiments, the target RNA sequence is longer than the guide RNA sequence.

[0033] In some embodiments, the vector comprises a plasmid or a viral vector.

[0034] In some embodiments, the vector is a plasmid or viral vector for expression in higher eukaryotic cells or prokaryotic cells. In some embodiments, the vector backbone is a pcDNA3.1 vector.

[0035] In some embodiments, the screening method comprises: connecting a plurality of different guide RNAs to the target RNA to form a transcription system, and transferring the system into cells via the vector.

[0036] In some embodiments, the target guide RNA is a guide RNA that targets a site for efficient editing.

[0037] In some embodiments, the method further comprises: after obtaining the editing level of the guide RNA from step (3), selecting guide RNAs with high editing levels, synthesizing single-stranded DNA with the same sequence, and then designing error-prone PCR primers to randomly introduce mutations into the targeted region of the guide RNA, and then repeating steps (1) to (3).

[0038] In some embodiments, the present invention provides a method for screening guide RNA for editing RNA, comprising: (a) designing a guide RNA library; (b) inserting a gene fragment to be edited and a reporter gene into a vector to obtain a reporter gene vector; (c) inserting the guide RNA library into the reporter gene vector to obtain a plasmid library to be screened; (d) transferring the plasmid library to be screened into cells, and after transfection, extracting RNA from the cells and sequencing them; (e) using bioinformatics to analyze the editing level of the guide RNA; wherein the order of (a) and (b) can be swapped.

[0039] In some embodiments, step (b) comprises: adding sequences upstream and downstream of the editing site to the downstream of the reporter gene and inserting them into the vector to obtain a reporter gene vector. In some embodiments, step (b) comprises: adding sequences 20-200 nt upstream and downstream of the editing site to the downstream of the reporter gene and inserting them into the vector to obtain a reporter gene vector.

[0040] In some embodiments, step (b) includes: amplifying the full-length fluorescent gene and the gene fragment to be edited by PCR, adding homologous sequences for homologous recombination at both ends or introducing restriction sites, and performing double digestion of the empty vector with NheI and HindIII and recovering the linear vector, and inserting the multiple fragments amplified by PCR into the linear vector to obtain a reporter gene vector.

[0041] In some embodiments, step (c) comprises: amplifying the guide RNA in step (a) and adding homology arms or restriction sites to obtain a guide RNA sequence pool, and linearizing the reporter gene vector in step (b) with BamHI and EcoRI; then inserting the guide RNA sequence pool into the linearized reporter gene vector, and obtaining a plasmid library to be screened after transformation.

[0042] In some embodiments, the screening method further comprises: (f) selecting a guide RNA with high editing efficiency from the guide RNA, directly synthesizing a single-stranded DNA with the same sequence, designing error-prone PCR primers to randomly introduce mutations into the guide RNA targeting region, and repeating steps (b) to (e).

[0043] In some embodiments, the present invention provides a method for high-throughput screening of guide RNA sequences that can recruit ADAR proteins and perform A-to-I editing on specific sites, comprising the following steps: S1. Designing and synthesizing a linker sequence for the gRNA library to be screened. Specifically, one end of the gRNA library is a fixed sequence required for chip synthesis, and the other end is a 20nt linker sequence that can distinguish different libraries; S2. Designing and constructing a screening vector. Specifically, the present invention designs a strategy of adding target RNA and gRNA sequences downstream of the green fluorescent protein GFP coding region. The three are implemented on the same transcript. The three can be directly connected or can be directly connected. Intervals are made by linker sequences of different lengths; S3. Construction of the gRNA plasmid library to be screened, specifically, PCR amplification of the mixed library through the linker sequences corresponding to different libraries and further insertion into the vector system of S2; S4. Transfection of the plasmid library into specific cells by transfection; S5. RNA extraction and construction of targeted sequencing library and second-generation sequencing; S6. Bioinformatics analysis of the editing efficiency of each gRNA corresponding to the target RNA; S7. Selection of gRNA sequences with high editing levels, synthesis of single-stranded DNA with the same sequence and error-prone PCR to construct the plasmid library and a second round of screening according to the methods of S3 to S6.

[0044] In some embodiments, the present invention provides a method for high-throughput screening of guide RNA sequences that can recruit ADAR proteins and perform A-to-I editing at specific sites. This invention addresses the gap in high-throughput screening of guide RNA sequences for ADAR-based target RNA editing by developing experimental and bioinformatic methods, and accomplishes this screening through a series of optimized experimental design methods.

[0045] In some embodiments, the method described herein includes 1) designing the target RNA and gRNA onto the same transcript by simulating different regions of the gene. The introduction of a fluorescent gene facilitates observation of transfection efficiency and can be replaced with other fluorescent genes; 2) constructing a targeted sequencing library and analysis method that can accurately match each gRNA with its editing efficiency; 3) randomly introducing mutations based on the selected gRNA sequences with high editing levels to further screen for gRNAs with higher editing levels.

[0046] In some embodiments, the present invention provides a guide RNA obtained by any of the screening methods.

[0047] In some embodiments, the present invention provides a guide RNA, the sequence of which is shown in any one of SEQ ID NO:557, SEQ ID NO:559, SEQ ID NO:561, SEQ ID NO:563, SEQ ID NO:565, SEQ ID NO:567, SEQ ID NO:569, SEQ ID NO:571, SEQ ID NO:573, SEQ ID NO:575, SEQ ID NO:577, SEQ ID NO:579, SEQ ID NO:583 to SEQ ID NO:602, SEQ ID NO:603 to SEQ ID NO:622, or SEQ ID NO:623 to SEQ ID NO:642.

[0048] In some embodiments, the present invention provides a method for obtaining the guide RNA or the use of the obtained guide RNA in preparing a drug for treating a disease.

[0049] In some embodiments, the drug is used to site-specifically edit nucleotides in target genes in eukaryotic cells to achieve the purpose of treating diseases.

[0050] In some embodiments, the disease is a disease caused by a mutation in the GAPDH gene, ATP7B gene, or FGFR gene.

[0051] In some embodiments, the GAPDH gene specific site motif is AAC, AAU, AAG, UAA or UAG, but is not limited thereto.

[0052] In some embodiments, the ATP7B gene specific site motif is UAG, but is not limited thereto.

[0053] In some embodiments, the FGFR gene specific site motif is CAG, but is not limited thereto.

[0054] In some embodiments, the disease includes, but is not limited to, Wilson's disease, achondroplasia, thanatoskeletal dysplasia, or short stature.

[0055] In some embodiments, the use is through the action of an RNA editing entity that is naturally present in the cell and is capable of editing the nucleotide sequence.

[0056] In some embodiments, the present invention provides an engineered RNA, the structure of which includes: a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, and a guide RNA nucleic acid fragment connected in sequence from upstream to downstream.

[0057] In some embodiments, the present invention provides an engineered RNA, the structure of which includes: a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, a guide RNA nucleic acid fragment and a polyA tail nucleic acid fragment connected in sequence from upstream to downstream.

[0058] In some embodiments, the reporter gene is selected from the group consisting of yfp gene, gus gene, rfp gene, gfp gene, cfp gene, bfp gene, or kanamycin resistance gene.

[0059] In some embodiments, the linker is 26 nt to 2121 nt in length.

[0060] In some embodiments, the linker comprises SEQ ID NO:549 or SEQ ID NO:557.

[0061] In some embodiments, the guide RNA is a guide RNA extended upstream and downstream with base A as the center.

[0062] In some embodiments, the guide RNA is a 151nt guide RNA that is extended 75nt upstream and downstream from the A base.

[0063] In some embodiments, the present invention provides an engineered RNA structure library comprising a plurality of different engineered RNAs.

[0064] In some embodiments, the different engineered RNAs are obtained by linking target RNAs to different guide RNAs via different linkers.

[0065] In some embodiments, in the engineered RNA structure library, the target RNA nucleic acid segments contained in the engineered RNAs in the library are identical, and the guide RNA nucleic acid segment sequences are different.

[0066] In some embodiments, the present invention provides a vector comprising the engineered RNA or the engineered RNA structure library.

[0067] In some embodiments, the vector is obtained by adding the target RNA sequence downstream of the reporter gene sequence and inserting it into the vector backbone to obtain a reporter gene vector, and then inserting the amplified guide RNA library into the reporter gene vector.

[0068] In some embodiments, the present invention provides a cell comprising the engineered RNA or the engineered RNA structure library or the vector. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a schematic diagram of the high-throughput screening method.

[0070] Figure 1 shows the overall process of high-throughput screening of guide RNAs, including the design of sequences to be screened; construction of the synthesized sequences to be screened into a screening vector; introduction into cells for editing; RNA extraction, construction of a second-generation sequencing library, and high-throughput sequencing; analysis of sequencing data using bioinformatics methods to obtain the editing efficiency of each gRNA; selection of gRNAs with high editing efficiency for verification; and selection of multiple gRNAs from the high-editing gRNAs for random introduction of error-prone mutations and repetition of the screening process to obtain gRNAs with higher editing efficiency.

[0071] FIG2 is a schematic diagram of the transcript information of the screening vector.

[0072] Figure 2 shows the specific transcript composition information of the screening vector. Taking GFP as an example, a transcript that can simultaneously detect editing levels and gRNA is constructed. This transcript simulates normal gene expression, with GFP as the coding region, the target sequence to be edited, the linker, and the gRNA simulate the 3'UTR region. There is a linker sequence with a preset length in the target RNA (Reporter) and gRNA. On the one hand, it can separate the two to simulate the formation of a secondary structure between molecules. On the other hand, adding a barcode sequence (for distinguishing different gRNAs) in the linker region can distinguish different gRNAs and corresponding editing levels during subsequent high-throughput sequencing data analysis.

[0073] Figure 3. Schematic diagram of screening vector transcripts for RNA editing.

[0074] Figure 3 shows that after the screening vector expresses RNA, it will self-fold to form a double-stranded RNA region of Reporter and gRNA, which will be recognized and bound by the ADAR protein and edit the target site.

[0075] Figure 4 shows the results of gRNA screening at different sites in the GAPDH gene. (Figure 4a) Results for editing site 1 with an AAC motif. (Figure 4b) Results for editing site 2 with an AAU motif. (Figure 4c) Results for editing site 3 with an AAG motif. (Figure 4d) Results for editing site 4 with a UAA motif.

[0076] Figure 4 shows the editing performance of each gRNA, obtained through high-throughput screening of four editing sites with different motifs on the GAPDH gene. The horizontal and vertical axes represent two biological replicates. Black dots in the figure represent gRNAs with a C base aligned at the editing site and complementary pairs at other positions. Non-black dots represent sequences generated by introducing different mutations with other gRNAs based on this sequence. The editing level at each site was determined by high-throughput sequencing.

[0077] Figure 5 is a comparison of gRNA editing and off-target editing at different sites of the GAPDH gene.

[0078] Figure 5 shows a heat map comparison of on-target editing and off-target editing upstream and downstream of gRNAs at different sites in the GAPDH gene. The horizontal axis (0) represents the position of the edited A base at the target site, 0 to -25 represents the interval upstream of the editing site, and 0 to 25 represents the interval downstream of the editing site. The vertical axis shows the editing level of each gRNA at different A bases, arranged downwards from the highest editing sequence at the target site. The trend from blue to red indicates the change in editing level of the A base site from low to high.

[0079] Figure 6 shows the editing and off-target editing efficiencies of gRNAs of different lengths targeting site 787 of the human GAPDH gene.

[0080] Figure 6 shows a heat map of on-target editing and off-target editing upstream and downstream of all gRNAs at position 787 of the GAPDH gene. The horizontal axis represents the target site, with the range between -1 and 1 representing the upstream range of the editing site, and the range between 1 and 25 representing the downstream range of the editing site. The vertical axis shows the editing level of each gRNA at different A bases, arranged downwards from the highest editing sequence at the target site. The trend from blue to red indicates the change in editing level of the A base site from low to high.

[0081] Figure 7 High-throughput screening gRNA validation analysis.

[0082] Figure 7 shows that the gRNAs selected by high-throughput screening were validated using a low-throughput method. The horizontal axis represents the editing level of each gRNA measured in the high-throughput screening method, and the vertical axis represents the editing level measured by synthesizing the modified RNA of each gRNA and transfecting it into cells. Correlation coefficient R 2 Higher values ​​indicate higher correlation between the horizontal and vertical axes.

[0083] Figure 8 shows the original results of first-generation sequencing of the editing level of a single gRNA on the GAPDH gene.

[0084] Figure 8 shows the raw peak plot of editing levels obtained by next-generation sequencing after transfecting a single modified gRNA into cells. As shown, different bases correspond to different colored peaks: A base corresponds to a green peak, and G base corresponds to a black peak. The presence of different colored peaks at a site indicates the presence of different base compositions at that site. In this case, the editing level for A base is indicated. The editing level is calculated by dividing the peak area for G base by the sum of the peak areas for A and G bases at that site.

[0085] Figure 9 shows the gRNA screening results for the human ATP7B gene.

[0086] Figure 9 shows the editing effect of each gRNA obtained by high-throughput screening of selected editing sites on the ATP7B gene, where the horizontal and vertical axes represent two biological replicates. The black dots in the figure are gRNAs with C bases at the editing site and complementary pairs at other positions. The non-black dots are sequences obtained by introducing different mutations into other gRNAs based on this sequence. The editing level corresponding to each dot is obtained by high-throughput sequencing. len41 indicates the gRNAs with different mutations designed for the target editing site and 41 nt upstream and downstream, and len31 indicates the gRNAs with different mutations designed for the 31 nt upstream and downstream of the target editing site.

[0087] Figure 10 shows the gRNA on-target editing and off-target editing efficiency of the human ATP7B gene.

[0088] Figure 10 shows a heat map of on-target editing and off-target editing upstream and downstream of the ATP7B gene for all gRNAs. The horizontal axis represents the target site, 0 to -20 represents the interval upstream of the editing site, and 0 to 20 represents the interval downstream of the editing site. The vertical axis represents the editing level of each gRNA at different A bases, arranged downwards from the highest editing sequence at the target site. The trend from blue to red indicates the change in editing level of the A base site from low to high.

[0089] Figure 11 shows the second round screening results of gRNA with the UAG motif of ATP7B gene.

[0090] Figure 11 shows the top 20 sequences of varying lengths with high editing efficiency in the ATP7B gene, obtained through the first round of screening, were subjected to error-prone random mutagenesis and a second round of screening. After the screening vectors were transfected into cells, the newly generated gRNAs and their corresponding editing efficiencies were measured using high-throughput sequencing. The horizontal and vertical axes in the figure represent two biological replicates. len41 represents the results of the second round of screening using the 41nt gRNA group, and len31 represents the results of the second round of screening using the 31nt gRNA group.

[0091] Figure 12 Results of gRNA screening for human FGFR gene.

[0092] Figure 12 shows the editing effect of each gRNA obtained by high-throughput screening of selected editing sites on the FGFR gene, where the horizontal and vertical axes represent two biological replicates. The black dots in the figure are gRNAs with C bases at the editing site and complementary pairs at other positions. The non-black dots are sequences obtained by introducing different mutations into other gRNAs based on this sequence. The editing level corresponding to each dot is obtained by high-throughput sequencing. len41 indicates gRNAs with different mutations designed for the target editing site and 41 nt upstream and downstream, and len31 indicates gRNAs with different mutations designed for the target editing site and 31 nt upstream and downstream.

[0093] Figure 13 shows the gRNA on-target editing and off-target editing efficiency of the human FGFR gene.

[0094] Figure 13 shows a heat map of on-target editing and off-target editing upstream and downstream of the FGFR gene for all gRNAs. The horizontal axis (0) represents the target site, 0 to -20 represents the interval upstream of the editing site, and 0 to 20 represents the interval downstream of the editing site. The vertical axis represents the editing level of each gRNA at different A bases, arranged downwards from the highest editing sequence at the target site. Blue to red indicates the trend of editing levels at A base sites from low to high.

[0095] FIG14 shows the second round screening results of gRNAs with the CAG motif of human FGFR gene.

[0096] Figure 14 shows the top 20 41nt gRNA sequences with high editing efficiency obtained in the first round of screening for the FGFR gene. Error-prone random mutations were introduced and then screened in a second round. After transfection of the screening vectors into cells, high-throughput sequencing was used to measure the newly generated gRNAs and their corresponding editing efficiencies. The horizontal and vertical axes in the figure represent two biological replicates. len41 represents the results of the second round of screening for gRNAs selected from the 41nt gRNA group. DETAILED DESCRIPTION

[0097] The following is a detailed description of the technical solution of the present invention, which does not limit the scope of protection of the present invention. Non-essential modifications and adjustments made by others based on the concept of the present invention still fall within the scope of protection of the present invention.

[0098] Unless otherwise specified, the reagents and materials used in the following examples are commercially available.

[0099] The process of the guide RNA high-throughput screening method in this article is shown in Figure 1.

[0100] As used herein and in the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes a plurality of such methods and reference to "the fragment" includes reference to one or more fragments and equivalents thereof known to those skilled in the art, and so forth.

[0101] Furthermore, the use of "or" means "and / or" unless stated otherwise. Similarly, "including," "comprising," and "comprising" are interchangeable and are not intended to be limiting.

[0102] It should be further understood that where the term "comprising" is used to describe various embodiments, those skilled in the art will understand that in some specific cases, the language "consisting essentially of" or "consisting of" may alternatively be used to describe the embodiment.

[0103] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although many methods and reagents are similar or equivalent to those described herein, exemplary methods and materials are disclosed herein.

[0104] It is to be understood that the present disclosure is not limited to the particular methodology, protocols, reagents, etc. described herein and that these may vary. The terminology used herein is for the purpose of describing particular embodiments or aspects only and is not intended to limit the scope of the present disclosure.

[0105] "Engineered polynucleotides" or "engineered guide RNAs" are used interchangeably with circular guide RNAs. Engineered polynucleotides can include recombinant polynucleotides of DNA or RNA or hybrid DNA / RNA constructs. Engineered polynucleotides can generate guide RNAs, more specifically circular guide RNAs.

[0106] As used herein, the term "mutation" can refer to a change in a nucleic acid sequence encoding a protein relative to the consensus sequence of the protein. A "missense" mutation causes a codon to be replaced by another codon; a "nonsense" mutation changes a codon from a codon encoding a specific amino acid to a stop codon. Nonsense mutations typically result in truncated translation of a protein. A "silent" mutation is a mutation that has no effect on the resulting protein. As used herein, the term "point mutation" can refer to a mutation that only affects one nucleotide in a gene sequence. A "splice site mutation" is a mutation present in pre-mRNA (before processing to remove introns) that causes mistranslation of a protein due to incorrect depiction of the splice site and typically results in truncation. Mutations can include single nucleotide variations (SNVs). Mutations can include sequence variants, sequence variations, sequence changes, or allelic variants. Reference DNA sequences can be obtained from reference databases. Mutations may affect function. Mutations may not affect function. Mutations can occur in one or more nucleotides at the DNA level, in one or more nucleotides at the ribonucleic acid (RNA) level, in one or more amino acids at the protein level, or any combination thereof. Reference sequences can be obtained from databases such as the NCBI Reference Sequence Database (RefSEQ ID NO:) database. Specific changes that may constitute mutations may include substitutions, deletions, insertions, inversions, or transitions in one or more nucleotides or one or more amino acids. Mutations may be point mutations.

[0107] The term "RNA editing" refers to a co-transcriptional or post-transcriptional modification process that introduces changes in the sequence of genomically encoded RNA, resulting in RNA mutations. Editing of adenosine to inosine (A-to-I) in double-stranded RNA (dsRNA) is catalyzed by adenosine deaminases acting on RNA (ADAR) enzymes and is a common type of RNA editing in mammals. In vertebrates, a family of three ADAR proteins, ADAR1, ADAR2, and ADAR3, has been previously characterized. ADAR1 and ADAR2 (ADAR) catalyze all currently known A-to-I editing sites. ADAR3 has no known deaminase activity. Inosine (I) mimics guanosine (G), so ADAR proteins introduce virtual A to G substitutions in transcripts. This change can lead to specific amino acid substitutions, alternative splicing, microRNA-mediated gene silencing, or changes in transcript localization and stability.

[0108] After deamination, different methods can be used to determine the modification of the target RNA and / or the protein encoded by the target RNA, depending on the position of the target adenosine in the target RNA. For example, in order to determine whether "A" has been edited to "I" in the target RNA, RNA sequencing methods known in the art can be used to detect the modification of the RNA sequence. When the target adenosine is located in the coding region of the mRNA, RNA editing can cause changes in the amino acid sequence encoded by the mRNA. For example, point mutations can be introduced into the mRNA, or innate or acquired point mutations in the mRNA are restored to produce a wild-type gene product because "A" is converted to "I". Amino acid sequencing by methods known in the art can be used to find any changes in amino acid residues in the encoded protein. Modification of the stop codon can be determined by assessing the presence of functional, elongated, truncated, full-length and / or wild-type proteins. For example, when the target adenosine is located in a UGA, UAG or UAA stop codon, modification of the target A (UGA or UAG) or multiple A's (UAA) can create a read-through mutation and / or an extended protein, or a truncated protein encoded by the target RNA can be restored to produce a functional, full-length and / or wild-type protein. Editing of the target RNA can also generate abnormal splicing sites and / or variable splicing sites in the target RNA, resulting in an extended, truncated or misfolded protein, or the encoded abnormal splicing or variable splicing sites in the target RNA are restored to produce a functional, correctly folded, full-length and / or wild-type protein. In some embodiments, the present application contemplates editing innate and acquired genetic changes, for example, missense mutations, premature stop codons, abnormal splicing or variable splicing sites encoded by the target RNA. Using known methods to evaluate the function of the protein encoded by the target RNA, it can be found whether the RNA editing has achieved the desired effect. Because the deamination of inosine (I) by adenosine (A) can correct the mutant A at the target position in a mutant RNA encoding a protein, the identification of deamination to inosine can provide an assessment of the presence of a functional protein or an assessment of whether RNA associated with a disease or drug resistance caused by mutant adenosine has been restored or partially restored. Similarly, because the deamination of inosine (I) by adenosine (A) can introduce point mutations in the resulting protein, the identification of deamination to inosine can provide functional indications of the cause of a disease or a factor associated with the disease.

[0109] In some embodiments, the target RNA is a regulatory RNA. In some embodiments, the target RNA to be edited is a ribosomal RNA, a transfer RNA, a long non-coding RNA or a small RNA (e.g., miRNA, pri-miRNA, pre-miRNA, piRNA, siRNA, snoRNA, snRNA, exRNA or scaRNA). The deamination of the target adenosine includes, for example, ribosomal RNA, transfer RNA, a long non-coding RNA or a small RNA (e.g., miRNA), including changes in three-dimensional structure and / or loss of function or gain of function. In some embodiments, the deamination of target A in the target RNA changes the expression level of one or more downstream molecules (e.g., proteins, RNA and / or metabolites) of the target RNA. The change in the expression level of the downstream molecule can be an increase or decrease in expression level.

[0110] The term "RNA editing entity" refers to a biological molecule that can cause chemical modification of nucleotides to change nucleotides into different nucleotides. In some embodiments, RNA editing entities can be recruited to specific sites in a polynucleotide to cause changes in the nucleic acid sequence at the desired site. Examples of RNA editing entities include APOBEC proteins (e.g., APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4 proteins) or ADAR proteins (e.g., ADAR1, ADAR2, or ADAR3 proteins).

[0111] In the context of the present application, "target RNA" refers to an RNA sequence to which a deaminase is recruited that is designed to have complete or substantial complementarity therewith, and a double-stranded RNA (dsRNA) region containing a target adenosine is formed by hybridization between the target sequence and the dRNA, which recruits an adenosine deaminase (ADAR) acting on RNA, which deaminates the target adenosine. In some embodiments, ADAR is naturally present in a host cell, such as a eukaryotic cell (preferably a mammalian cell, more preferably a human cell). In some embodiments, the ADAR is introduced into a host cell.

[0112] As used herein, the term "barcode sequence" refers to a unique nucleotide sequence used to identify and / or track the source of a polynucleotide in a reaction. The size and composition of the barcode sequence can vary greatly. In certain embodiments, the length of the barcode sequence can range from 4 to 36 nucleotides, or 6 to 30 nucleotides, or 8 to 20 nucleotides.

[0113] The term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid connected thereto. The vector includes, but is not limited to, single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, (e.g., circular) nucleic acid molecules without free ends; nucleic acid molecules comprising DNA, RNA, or both; and other various polynucleotides known in the art. One type of vector is a "plasmid," which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, for example, by standard molecular cloning techniques. Certain vectors are capable of autonomous replication in the host cell into which they are introduced (e.g., bacterial vectors and additional mammalian vectors having a bacterial origin of replication). Other vectors (e.g., non-additional mammalian vectors) are integrated into the genome of the host cell after being introduced into the host cell, thereby replicating together with the host genome. In addition, certain vectors are capable of directing the transcription or expression of the encoding nucleotide sequence to which they are operably connected. Such vectors are referred to herein as "expression vectors."

[0114] Editing entities are typically proteinaceous in nature, such as ADAR enzymes found in metazoans, including mammals. Editing entities can also include complexes of nucleic acids and proteins or peptides, such as ribonucleoproteins. Editing enzymes can include only nucleic acids or consist of nucleic acids, such as ribozymes. All of these editing entities are included in the present invention, as long as they are recruited by the oligonucleotide constructs according to the present invention. Preferably, the editing entity is an enzyme, more preferably an adenosine deaminase or a cytidine deaminase, and even more preferably an adenosine deaminase. When the editing entity is an adenosine deaminase, Y is preferably cytidine or uridine, most preferably cytidine. One of the most interesting is human ADAR, hADAR1 and hADAR2, including any subtypes thereof, such as hADAR1p110 and p150.

[0115] The target RNA can be any cellular or viral RNA sequence, but is more typically a protein-coding pre-mRNA or mRNA.

[0116] Materials and methods

[0117] 1. Construction of human ADAR1 expression plasmid

[0118] The CDS of human ADAR1 (the sequence of human ADAR1 CDS is shown in SEQ: ID NO: 2) was amplified by PCR and homology arms were added. The pcDNA3.1 vector (sequence is shown in SEQ ID NO: 1) was linearized using NheI and EcoRI double enzyme digestion and the insert fragment was ligated to the vector using homologous recombination.

[0119] 2. Construction of reporter gene plasmid

[0120] First, the full-length fluorescent gene GFP and the gene fragment to be edited were amplified by PCR, and homologous sequences for homologous recombination were added at both ends. At the same time, the pcDNA3.1 empty vector was double-digested with NheI and HindIII and the linear vector was recovered. The vector and the inserted fragment were homologously recombined through multi-fragment recombination to complete the construction of the reporter gene vector.

[0121] 3. Construction of gRNA sequence pool plasmid library

[0122] The gRNA mixed sequence pool was amplified by PCR and homology arms were added at the same time. The reporter gene vector was linearized by double enzyme digestion with BamHI and EcoRI and the gRNA sequence pool was inserted into the reporter gene vector by homologous recombination. After transformation and plasmid plasmid extraction, all single clones were collected together to finally obtain the gRNA sequence pool plasmid library.

[0123] 4. Cell culture

[0124] 4.1 Cell passaging experimental steps:

[0125] (1) Disinfect the clean bench with ultraviolet light for 30 minutes.

[0126] (2) Remove the cells to be passaged from the incubator and wash them three times with PBS. Gently add and aspirate to prevent the cells from falling off. Then add 0.25% trypsin in an amount that can submerge the cell surface. After the cells become round, add culture medium containing serum to terminate the digestion.

[0127] (3) Transfer the digested cell suspension to a 15 ml centrifuge tube and centrifuge at 1000 rpm for 4 minutes at room temperature.

[0128] (4) Carefully discard the supernatant, add fresh culture medium and pipette the cells into a single suspension.

[0129] (5) Take 20ul of cell suspension and add it to a 96-well cell culture plate. Then add 20ul of trypan blue staining solution. After mixing by pipetting, take 20ul and add it to a cell counting plate and count the cells under a microscope.

[0130] (6) Calculate the volume of cell suspension required for a specific number of cells based on the size of the inoculated cell culture plate, add the corresponding volume of fresh cell culture medium, and place the cells in a cell culture incubator for culture.

[0131] 4.2 Cell transfection experimental steps:

[0132] (1) Set the required cell culture plate according to the cell amount required for the experiment. Add a certain amount of culture medium and cells to the corresponding culture plate according to the cell passaging method. When the cells reach 70% confluence after 24 hours of growth, the cell transfection experiment can be carried out.

[0133] (2) Disinfect the clean bench under ultraviolet irradiation for 30 minutes.

[0134] (3) Prepare the transfection reagent mixture according to the instructions provided by the transfection reagent lipofectamine 3000 and let it stand at room temperature for 5 minutes.

[0135] (4) Remove the cells and carefully add the transfection solution to the cells, gently mix evenly, and place them in the incubator.

[0136] 4.3 Construction of ADAR1-overexpressing HEK293 cell line

[0137] The ADAR1 expression vector was transfected into HEK293 cells according to the above-mentioned cell passage and transfection steps. After 48 hours, the screening drug G418 was added and the medium was changed every two days. The screening of HEK293 cell lines stably transfected with ADAR1 was completed when there were no living cells in the non-transfected group.

[0138] 4.4 Targeted Sequencing Library Construction and Quality Control

[0139] Total cellular RNA can be extracted using a conventional RNA extraction kit to ensure that the RNA is not significantly degraded. Reverse transcription is then performed using a conventional reverse transcription kit. During the first round of PCR amplification, sequencing adapter sequences must be added to both ends of the amplification primers, and the number of cycles in the first round of amplification must be controlled to be less than 20. Agarose gel electrophoresis is used to verify the first-round PCR products and to recover library fragments of the corresponding size. After purification, a second round of PCR is performed. The second-round primers contain sequencing adapter sequences and specific base barcode sequences. The second round of PCR is generally controlled within 10 cycles. The PCR products are again subjected to agarose gel electrophoresis to confirm the library fragment size, and the gel is excised to recover and purify the library. Finally, the targeted sequencing library is subjected to next-generation sequencing.

[0140] 4.4 Error-prone PCR amplification

[0141] Agilent's second-generation error-prone PCR kit was used to amplify specific regions and randomly introduce mutations. Samples were prepared according to the instructions, with a template addition amount of 0.1 ng and 30 amplification cycles.

[0142] 4.5 Targeted Sequencing Data Analysis

[0143] First, the raw data was processed by removing the linkers using cutadapt software, and then the raw read length after removing the linkers was aligned to the reference sequence. The editing of the reporter gene was matched one-to-one with each gRNA according to the linker sequence, and then the editing level of the reporter gene was calculated to obtain the editing level of each gRNA on the target site.

[0144] 4.6 Synthesis of Chemically Modified RNA

[0145] The modified RNA is synthesized primarily by conventional solid-phase synthesis in the art, with phosphorothioate backbone modifications between bases and modifications at the 2' position of the bases. In some embodiments, phosphorothioate modification (hereinafter represented by *) is selected, while in some embodiments, LNA modification at the 2' position (hereinafter represented by L) is selected. The structural formulas for phosphorothioate and LNA modifications are shown below:

[0146] phosphorothioate modification

[0147] LNA modification

[0148] Example 1 Screening of gRNA sequences at different A base sites on the GAPDH gene in HEK293 cells

[0149] This example takes the gRNA screening method for editing multiple A base sites of the human endogenous gene GAPDH (NM_002046.7) as an example to further illustrate the present invention. Specifically, the method includes the following steps:

[0150] 1.1 Design of gRNA library and adapter sequence for the specific A base of human GAPDH gene

[0151] A 151nt gRNA was designed with a specific A base as the center and extended 75nt upstream and downstream. Mutations of different numbers and positions were introduced into this gRNA to obtain a library to be screened. In this example, the A bases at four sites were included, and the triplet motifs at these sites were AAC, AAU, AAG, and UAA, respectively. Multiple gRNA libraries for these four sites were screened. The library sequence information of the AAC motifs in these four sites is listed as an example herein, as shown in SEQ ID NO: 1 to SEQ ID NO: 540.

[0152] Sequence synthesis can be distinguished by designing adapter sequences that will not cause crossover between different groups in the downstream amplification step. The editing site information and adapter sequences corresponding to the four libraries are shown in Table 1 below.

[0153] Table 1 Editing site information and linker sequence

[0154] 1.2 Design and construction of the required screening vector

[0155] As described in the above-mentioned material method content, in the present embodiment, the 100bp sequence upstream and downstream of the editing site (a total of 201 bases) is added to the downstream of the GFP gene and recombined into the pcDNA3.1 vector, and then the linker 1 sequence (sequence such as SEQ ID NO: 549) is inserted downstream. The linker 1 used in the present embodiment is 2121bp in length to obtain a reporter gene plasmid vector, wherein linker 1 comprises a 10bp NNNYRNNNYR (Barcode sequence, wherein N represents one of four bases A, C, G, T, Y represents C or T base, and R represents A or G base) degenerate bases are used to distinguish the editing efficiency of each gRNA. After obtaining the vector into which linker 1 is inserted, the amplified gRNA library is recombined into the reporter gene plasmid vector to obtain a screening vector as shown in Figure 2. The transcript expressed by the vector performs RNA editing mode as shown in Figure 3.

[0156] 1.3 The obtained screening vector was then transformed into HEK293 cells overexpressing ADAR1. After 48 hours, RNA was extracted and targeted library sequencing was performed (library construction primers are shown in Table 2). The target site editing and off-target editing levels of each gRNA were analyzed, as shown in Figures 4 and 5. In this example, a long linker 1 sequence was added to simulate the formation of double-stranded editing between intermolecular RNAs. Two libraries were constructed during library construction: one was a regional library of reporters and linkers to establish the correspondence between the barcode sequence and the editing level on each linker. The first-round primers used were the cDNA-NGS-round1 primer pair, and the second-round primers were the NGS-round2 primer pair. The other was a regional library of linkers and gRNAs. Because the gap between linkers and gRNAs exceeded 2000 bp, direct library construction and sequencing were not possible. Therefore, the full-length sequences of linkers and gRNAs were first circularized. After circularization, the plasmid-NGS-round1 primer pair was used for the first round of amplification, and the NGS-round2 primer pair was used to complete the entire library construction. This library can establish the correspondence between the barcode on the linker and the gRNA sequence. By comparing the barcode sequences of the two libraries, the editing efficiency corresponding to each gRNA can be obtained.

[0157] The results show that the editing efficiency corresponding to each gRNA can be screened out by the experimental and data analysis methods of the present invention, and the biological replicates are relatively consistent. Editing site 1 is an AAC motif (Figure 4a), and the black dots represent the editing levels of gRNAs extending from the editing site to both ends for a total of 151nt (the bases in the editing site are C bases) in two repeated experiments. The remaining yellow dots are the editing levels of different gRNAs after the introduction of mutations in two repeated experiments. The gRNAs for the AAC motif have a uniform gRNA distribution within 50% of the editing efficiency, and only a few gRNAs can achieve an editing efficiency range of 50% to 75%. Analysis of the heat map (Figure 5a) shows that the gRNAs for the AAC motif designed in this embodiment have obvious off-target editing at multiple sites, but the editing level of the target site is higher than that of the off-target site as a whole.

[0158] The AAU motif gRNA at editing site 2 (Figure 4b) showed an overall editing effect of less than 40%, and the editing efficiency of the gRNAs extending upstream and downstream from the editing site, except for the AC mismatch at the editing site, was also less than 30%.

[0159] The gRNA editing efficiency of the AAG motif (Figure 4c) is similar to that of the AAC motif. Unlike the other three editing sites, the overall editing level of the gRNA of the UAA motif (Figure 4d) is between 25% and 75%, but the A base adjacent to the gRNA of the UAA motif has relatively serious off-target editing. Only the gRNA screened with the AAG motif has low overall off-target editing (Figure 5). The gRNA designed in this embodiment is about 151nt in length. In addition, the gRNAs corresponding to different sites A have different editing abilities, indicating that gRNA sequences with different editing levels can be screened out by the present invention.

[0160] Table 2 Primer sequences used for library construction and sequencing

[0161] Example 2 Screening of gRNA sequences with UAG motif at specific sites of GAPDH gene in HEK293 cells

[0162] This example aims to verify that short gRNA and short linkers can be screened for editing specific sites on the human GAPDH gene. The A base at site 787 on the GAPDH gene was selected, and the motif where it is located is the TAG motif. 1073 gRNAs (length range 28nt-89nt) with different lengths and mutations at different sites were designed for this site. In this example, a short linker was used to verify the screening system, where the sequence of linker 2 is NNNNNNNNTGGGTTGAGGGTAGTGAG (sequence number is SEQ ID NO: 556; N is any one of A, T, C, G bases, and 8 Ns are the barcode sequence of the linker).

[0163] The specific implementation plan is as follows: first, the gRNA library is directly synthesized, and the linker 2 is directly connected to the gRNA sequence to synthesize the gRNA library with the barcode sequence; secondly, a reporter vector is constructed, and a sequence extending 201bp upstream and downstream with the editing site as the center is recombined with the GFP sequence and inserted into the pcDNA3.1 vector; the third step is to insert the gRNA library in step 1 into the reporter vector to complete the construction of the screening vector; the fourth step is to transfect the screening vector into HEK293 cells overexpressing ADAR1 protein for 48 hours, extract RNA and target library sequencing, and analyze the target site editing and off-target editing levels of each gRNA. The results are shown in Figure 6. The results show that in this example, in the case of a short linker, a mixed library of gRNAs of different lengths can also be used to screen gRNAs with different editing efficiencies by this method.

[0164] The high-throughput screening established based on the method of the present invention is to transfect a mixed gRNA library into cells. There may be a certain degree of interference between gRNAs. Therefore, this example verifies a single gRNA based on high-throughput screening. This example selects gRNA sequences with editing levels ranging from 3.93% to 21.66% from high-throughput screening. The sequence information is shown in Table 3, such as GAPDH-TAG-S1 to GAPDH-TAG-S12. The chemically modified gRNA sequences correspond to GAPDH-TAG-M1 to GAPDH-TAG-M12. For the sequences screened out by high-throughput screening, chemically modified single RNAs were used to transfect into cells to verify the editing level (as shown in Figure 8 and Table 3). The stability of the modified gRNA is higher than that of the gRNA expressed by the plasmid, so the editing level is higher than that of the high-throughput screening. The chemically modified gRNA modification strategy is to modify the bases with phosphorothioate and add 3 LNA modifications at each end. Analysis of the editing levels screened by high-throughput screening and the editing levels of single modified gRNAs showed a strong correlation, with a correlation coefficient of 0.83 (as shown in Figure 7), confirming the accuracy of the method of the present invention. The editing levels of GAPDH-TAG-S1 to GAPDH-TAG-S13 in Table 3 are the gRNA editing levels measured by the high-throughput gRNA mixture system; the editing levels of GAPDH-TAG-M1 to GAPDH-TAG-M13 are the editing levels measured by the single gRNA system.

[0165] Table 3 High-throughput screening verification sequence information

[0166] Example 3 Screening of gRNA sequences with UAG motif at specific sites of ATP7B gene in HEK293 cells

[0167] This example aims to verify that short gRNAs for editing specific sites can be screened on the human ATP7B (NM_000053.4) gene. ATP7B gene mutations are the main cause of Wilson's disease. In this example, the A base at site 1532 of the ATP7B gene is selected. There is a mutation in the C base upstream of the site that mutates to a T base. Therefore, the motif of the editing site after mutation is the UAG motif. 1011 gRNAs of different lengths and mutations at different sites are designed for this site. This example selects gRNAs with lengths of 31nt and 41nt, including the editing site and its upstream and downstream. Introducing mutations or deletions or adding protrusions for these two lengths will cause the actual length to fluctuate a little, with a floating range of 27nt-59nt. In this example, connector 2 is used to verify the screening system.

[0168] The specific implementation plan is as follows: first, directly synthesize gRNA libraries with lengths of 31nt and 41nt, and directly connect linker 2 to the gRNA sequence to synthesize a gRNA library with a barcode sequence; secondly, construct a reporter vector, and recombinantly insert a 201bp sequence extending upstream and downstream of the editing site into the pcDNA3.1 vector with the GFP sequence; thirdly, insert the gRNA library in step one into the reporter vector to complete the construction of the screening vector; fourthly, transfect the screening vector into HEK293 cells overexpressing ADAR1 protein, extract RNA 48 hours later, conduct targeted library sequencing, and analyze the target site editing and off-target editing levels of each gRNA. The results are shown in Figures 9 and 10.

[0169] The results show that this example can screen out gRNA sequences with high editing efficiency from a library of 1011 gRNAs of different lengths. Overall, the editing ability of 41nt gRNA is higher than that of 31nt gRNA (Figure 9), where the black dots indicate that except for one base mismatch at the editing site, the other bases are complementary paired gRNAs.

[0170] Comparing the editing and off-target efficiencies of all gRNAs of different lengths (Figure 10) reveals that the gRNA designed with base A at ATP7B 1532 exhibits relatively weak off-target editing. Overall, gRNAs designed based on 41nt target sequences exhibited editing levels below 20% for most sequences, while gRNAs designed based on 31nt target sequences exhibited editing levels below 20%.

[0171] In this example, 20 sequences with high editing efficiency and long sequencing reads were selected from each of the two groups of high-throughput screening results for the next round of error-prone verification of the editing effect after the introduction of mutations. The selected sequences from the 41nt group are shown in SEQ ID NO: 583 to SEQ ID NO: 602, and the selected sequences from the 31nt group are shown in SEQ ID NO: 603 to SEQ ID NO: 622. Error-prone PCR primer sequences were added to both ends of all selected sequences, with the 5' end being TCGAAGGTCGCTTAGACGGC (SEQ ID NO: 581) and the 3' end being CGGTCGTGAGTGCAGACGTG (SEQ ID NO: 582) to directly synthesize single-stranded sequences. During error-prone PCR amplification, two groups were amplified. In the 41nt group, all sequences of equal mass were mixed and amplified as templates. The 31nt group was treated in the same way. The amplified products were reinserted into the reporter vector and screened for the second round according to the first round. The results are shown in Figure 11.

[0172] The results showed that both the 41nt group and the 31nt group achieved significant improvements after the second round of screening. The 31nt group improved its editing efficiency from less than 20% in the first round to 20% to 80%, mainly distributed within 40% editing efficiency, with a small proportion of 40% to 80% editing. The 41nt group also had a small proportion of 20% to 50% editing efficiency. After the second round of screening, it was obvious that the editing proportion of 40% to 90% was relatively high. The above data also confirmed that the two-round screening of the present invention can screen out gRNAs with relatively high editing effects, especially the method of introducing error-prone mutations in the second round.

[0173] Example 4 Screening of gRNA sequences with CAG motif at specific sites of FGFR genes in HEK293 cells

[0174] This example aims to verify that short gRNAs can be screened for editing specific sites on the human FGFR (NM_015850.4) gene. FGFR gene mutations are the cause of diseases such as achondroplasia, lethal bone dysplasia, and dwarfism. In this example, the G358R site on the FGFR gene produces a G to A mutation. Therefore, the motif at the editing site after the mutation is the CAG motif. 989 gRNAs of different lengths and mutations at different sites were designed for this site. In this example, gRNAs of 31nt and 41nt in length, including the editing site and its upstream and downstream, were selected. Introducing mutations, deletions, or adding protrusions at these two lengths will cause the actual length to fluctuate slightly, with a floating range of 23nt-62nt.

[0175] In this example, linker 2 was selected to verify the screening system. The specific implementation plan is as follows: first, 31nt and 41nt gRNA libraries were directly synthesized, and linker 2 was directly connected to the gRNA sequence to synthesize a gRNA library with a barcode sequence; second, a reporter vector was constructed, and a sequence extending 201bp upstream and downstream of the editing site was recombined with the GFP sequence and inserted into the pcDNA3.1 vector; third, the gRNA library in step 1 was inserted into the reporter vector to complete the construction of the screening vector; fourth, the screening vector was transfected into HEK293 cells overexpressing ADAR1 protein. After 48 hours, RNA was extracted and targeted library sequencing was performed to analyze the target site editing and off-target editing levels of each gRNA. The results are shown in Figures 12 and 13.

[0176] The results show that this embodiment can screen out efficient editing gRNA sequences from 989 gRNA libraries of different lengths. Overall, the editing ability of 41nt gRNA is higher than that of 31nt gRNA (Figure 12), where the black dots indicate that except for one base mismatch at the editing site, the other bases are complementary paired gRNAs. All gRNAs of different lengths are arranged together to compare the editing efficiency and off-target efficiency of the target site (Figure 13). It can be seen that among all gRNAs, there are two sites upstream of the editing site with obvious off-target, but there is no off-target editing at all downstream of the editing site. Overall, the editing level of most sequences of gRNA designed based on the 41nt target sequence is less than 20%, and the editing of gRNAs designed based on the 31nt target sequence is almost all less than 10%.

[0177] In view of the low overall editing level of the 31nt group in this embodiment, only 20 sequences with high editing levels and long sequencing reads were selected from the high-throughput screening results of the 41nt gRNA group for the next round of error prone verification of the editing effect after the introduction of mutations. The selected sequences of the 41nt group are shown in SEQ ID NO: 623 to SEQ ID NO: 642. Error prone PCR primer sequences were added to both ends of all selected sequences, with the 5-end being TCGAAGGTCGCTTAGACGGC (SEQ ID NO: 581) and the 3-end being CGGTCGTGAGTGCAGACGTG (SEQ ID NO: 582) to directly synthesize single-stranded sequences. During error prone PCR amplification, all sequences were mixed with equal mass and amplified as templates. The amplified products were reinserted into the reporter vector and screened for the second round according to the first round. The results are shown in Figure 14.

[0178] The results showed that after the second round of screening by introducing error-prone random mutations, a significant improvement was achieved, with the editing effect increased from less than 20% in the first round to 80%. The above data once again confirmed that the two-round screening of the present invention can screen out gRNAs with relatively high editing effects, especially the method of introducing error-prone mutations in the second round.

Claims

1. A method for screening guide RNA for editing RNA, characterized in that: The method comprises the following steps: (1) each guide RNA to be screened is connected with a linker and is respectively connected to a target RNA on a nucleic acid chain capable of forming a transcript, thereby forming a plurality of nucleic acid chains; the nucleic acid chain contains a barcode sequence; (2) the plurality of nucleic acid chains are transferred into cells expressing ADAR protein; (3) the RNA in the cells is extracted and measured to determine the editing level of the target RNA, and the editing level of the corresponding guide RNA is determined by the barcode sequence.

2. The screening method according to claim 1, characterized in that The linker is located between the target RNA and the guide RNA; Preferably, the barcode sequence is contained in the linker; The nucleic acid chain is connected with a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, and a guide RNA nucleic acid fragment in sequence from upstream to downstream; Preferably, a polyA tail nucleic acid fragment is further connected downstream of the guide RNA nucleic acid fragment; Preferably, the reporter gene is selected from the group consisting of yfp gene, gus gene, rfp gene, gfp gene, cfp gene, bfp gene, or kanamycin resistance gene; Preferably, the length of the linker is 26 nt to 2121 nt. Preferably, the linker comprises SEQ ID NO: 549 or SEQ ID NO: 557; Preferably, step (3) comprises establishing a targeted sequencing library for next-generation sequencing; Preferably, the nucleic acid chain contacts the ADAR protein by transferring the nucleic acid chain into a cell expressing the ADAR protein; Preferably, the ADAR protein comprises ADAR1 or ADAR2; Preferably, the guide RNA is capable of recruiting ADAR protein in the cell to edit the target RNA; Preferably, the cell overexpresses ADAR protein; Preferably, the editing is A-to-I editing; Preferably, the guide RNA is a guide RNA extended upstream and downstream with base A as the center; Preferably, the guide RNA is a guide RNA of 30 nt to 151 nt in length obtained by extending 15 nt to 75 nt upstream and downstream from the A base; Preferably, the guide RNA is a sequence of complete complementary pairing or incomplete complementary pairing; Preferably, the length of the guide RNA is 20 nt to 200 nt; Preferably, the screening method comprises adding upstream and downstream sequences of the editing site to the downstream of the reporter gene and inserting them into a vector to obtain a reporter gene vector, and then transferring the reporter gene vector into cells; Preferably, the length of the sequence added to the reporter gene upstream and downstream of the editing site is 20 to 200 nt; Preferably, the screening method is a high-throughput screening method; Preferably, the target RNA sequence is longer than the guide RNA sequence; Preferably, the vector comprises a plasmid or a viral vector; Preferably, the vector is a plasmid or viral vector for expression in higher eukaryotic cells or prokaryotic cells; Preferably, the vector backbone is a pcDNA3.1 vector; Preferably, the target guide RNA is a guide RNA that targets a site for efficient editing; Preferably, the method further comprises: after obtaining the editing level of the guide RNA from step (3), selecting guide RNAs with high editing levels to synthesize single-stranded DNA with the same sequence, and then designing error-prone PCR primers to randomly introduce mutations into the target region of the guide RNA, and then repeating steps (1) to (3); Preferably, the sequences of the error-prone PCR primers are shown in SEQ ID NO:581 and SEQ ID NO:

582.

3. A method for screening guide RNA for editing RNA, characterized in that: include: (a) Design of guide RNA library; (b) inserting the gene fragment to be edited and the reporter gene into a vector to obtain a reporter gene vector; (c) inserting the guide RNA library into the reporter gene vector to obtain a plasmid library to be screened; (d) transferring the plasmid library to be screened into cells, and after transfection, extracting RNA from the cells and sequencing; (e) analyzing the editing level of the guide RNA using bioinformatics; The order of (a) and (b) can be reversed; Preferably, step (b) comprises: adding upstream and downstream sequences of the editing site to the downstream of the reporter gene and inserting them into the vector to obtain a reporter gene vector; Preferably, the length of the sequence added to the reporter gene upstream and downstream of the editing site is 20 to 200 nt; Preferably, step (b) comprises: amplifying the full length of the fluorescent gene and the gene fragment to be edited by PCR, adding homologous sequences for homologous recombination or introducing restriction sites at both ends, performing NheI and HindIII double restriction digestion on the empty vector and recovering the linear vector, and inserting the multiple fragments into the linear vector to obtain a reporter gene vector; Preferably, step (c) comprises: amplifying the guide RNA in step (a) and adding homology arms or restriction sites to obtain a guide RNA sequence library, and linearizing the reporter gene vector in step (b) by double restriction digestion with BamHI and EcoRI; then inserting the guide RNA sequence library into the linearized reporter gene vector, and obtaining a plasmid library to be screened after transformation; Preferably, the screening method further comprises: (f) After selecting a guide RNA with high editing efficiency from the guide RNA, directly synthesizing a single-stranded DNA with the same sequence, designing error-prone PCR primers to randomly introduce mutations into the guide RNA targeting region, and then repeating steps (b) to (e).

4. The guide RNA obtained by the screening method according to any one of claims 1 to 3.

5. A guide RNA, characterized in that The sequence of the guide RNA is shown in any one of SEQ ID NO:557, SEQ ID NO:559, SEQ ID NO:561, SEQ ID NO:563, SEQ ID NO:565, SEQ ID NO:567, SEQ ID NO:569, SEQ ID NO:571, SEQ ID NO:573, SEQ ID NO:575, SEQ ID NO:577, SEQ ID NO:579, SEQ ID NO:583 to SEQ ID NO:602, SEQ ID NO:603 to SEQ ID NO:622, or SEQ ID NO:623 to SEQ ID NO:

642.

6. Use of the screening method according to any one of claims 1 to 3 or the guide RNA according to any one of claims 4 to 5 in the preparation of a drug for treating a disease; Preferably, the drug is used to edit nucleotides in target genes in eukaryotic cells to achieve the purpose of treating diseases; Preferably, the disease is a disease caused by mutations in the GAPDH gene, ATP7B gene or FGFR gene; Preferably, the GAPDH gene specific site motif is AAC, AAU, AAG, UAA or UAG; Preferably, the ATP7B gene specific site motif is UAG; Preferably, the FGFR gene specific site motif is CAG; Preferably, the disease comprises Wilson's disease, achondroplasia, thanatoskeletal dysplasia or dwarfism; Preferably, said use is by the action of an RNA editing entity naturally present in said cell and capable of editing said nucleotides.

7. An engineered RNA, characterized in that It includes: a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, and a guide RNA nucleic acid fragment are sequentially connected from upstream to downstream; Preferably, a polyA tail nucleic acid fragment is further connected downstream of the guide RNA nucleic acid fragment; Preferably, the reporter gene is selected from the group consisting of yfp gene, gus gene, rfp gene, gfp gene, cfp gene, bfp gene, or kanamycin resistance gene; Preferably, the length of the linker is 26 nt to 2121 nt; Preferably, the linker comprises SEQ ID NO: 549 or SEQ ID NO: 557; Preferably, the guide RNA is a guide RNA extended upstream and downstream with base A as the center; Preferably, the guide RNA is a guide RNA with a length of 30 nt to 151 nt obtained by extending 15 nt to 75 nt upstream and downstream with base A as the center.

8. An engineered RNA structure library, characterized in that: Comprising a plurality of different engineered RNAs as described in claim 7; Preferably, the different engineered RNAs are obtained by connecting the target RNA to different guide RNAs via different linkers; Preferably, in the engineered RNA structure library, the target RNA nucleic acid fragments contained in the engineered RNA in the library are the same, and the guide RNA nucleic acid fragment sequences are different.

9. A carrier, characterized in that Comprising the engineered RNA of claim 7 or the engineered RNA structure library of claim 8; Preferably, the vector is obtained by adding the target RNA sequence downstream of the reporter gene sequence and inserting it into the vector backbone to obtain a reporter gene vector, and then inserting the amplified guide RNA library into the reporter gene vector.

10. A cell, characterized in that Comprising the engineered RNA of claim 7 or the engineered RNA structure library of claim 8 or the vector of claim 9.