RNA-guided endonuclease systems and gene editing applications thereof
By developing the RNA-guided RAD (IS607 TnpB) system, the problems of vector delivery difficulties and off-target effects in existing gene editing technologies have been solved, more efficient and safer gene editing has been achieved, the recognition range of gene editing tools has been expanded, and the ability to clear viruses has been acquired.
Patent Information
- Application Number
- CN202480001311.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2024-01-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-01-02
AI Technical Summary
Existing gene editing technologies face the challenges of low gene therapy efficiency, difficult vector delivery, unresolved off-target issues, and safety risks. In particular, the size and specificity of the CRISPR/Cas system limit its application in gene therapy.
An RNA-guided endonuclease system, called the RAD (IS607 TnpB) system, was developed. It consists of ORFA and ORFB coding sequences and transposition-related motifs. It has a smaller molecular weight and higher specificity, and can recognize novel motifs and perform efficient gene editing.
The RAD system achieves smaller molecular weight and higher specificity gene editing, expands the recognition range of gene editing tools, has single-stranded DNA cutting activity and virus clearance capabilities, and provides a more efficient and safer gene editing tool.
Smart Images

Figure BSB0000208903750000161 
Figure BSB0000208903750000171 
Figure BSB0000208903750000181
Abstract
Description
[0001] This application claims priority to PCT application PCT / CN2023 / 070108 with the filing date of January 3, 2023, and the title of “A Class of RNA-guided Endonuclease System and Its Gene Editing Applications”. TECHNICAL FIELD
[0002] The present application relates to the field of gene editing technology, in particular to endonuclease complexes and their applications. BACKGROUND
[0003] Genes are the most important genetic material in living individuals, which not only determine the phenotypes of various life forms, but also are closely related to various diseases. With the progress of science and technology in recent years, the comprehensive development of sequencing technology and the completion of various genome projects, people can analyze and study the vast amount of genetic data, not only to understand genes, but also to hope that through technical means, genes can be modified to solve major genetic disease problems, optimize and improve plant and animal genes, etc. Gene editing technology has developed rapidly under this background and has been innovated and replaced.
[0004] Gene editing technology refers to a technology that changes the sequence of a biological genome through human means to achieve targeted genetic modification. This technology uses genetically engineered nucleases to produce double-stranded DNA breaks at specific sites in the genome, and then uses the cell's own DNA repair system, including non-homologous end joining and homologous recombination, to repair double-stranded DNA breaks. The repair process will cause random base insertion and deletion, resulting in mutations at specific sites in the genome. Finally, the repaired genome changes due to sequence changes, and the final function of the gene changes.
[0005] The development of gene editing technology can be traced back to the 20th century. Before the 1990s, researchers mainly obtained gene mutations in model organisms through random mutations, which not only took a long time, but also randomly mutated, making it difficult to obtain the desired genetic changes. After the 1990s, Chandrasegaran et al. developed the first generation of gene editing tools, zinc finger nucleases (ZFN). It is the first new type of endonuclease in gene editing technology by artificially fusing zinc finger protein with Fok I endonuclease. In recent years, ZFN has been used for gene editing in crops, fruit flies, zebrafish, mice, human cells, etc. However, the zinc finger protein module that recognizes DNA sequences is a zinc finger that recognizes three consecutive bases, so it is time-consuming and laborious to design targeting. In addition, due to the problem of Fok I non-specific cutting, the large-scale application of zinc finger nuclease technology is limited. In 2009, Ulla Bonas' team found that transcription activator-like effector (TALE) has a very close relationship with the base pair of DNA. After researchers fused TALE and Fok I, they developed transcription activator-like effector nucleases (TALEN). TALEN is similar to ZFN in principle, but the preparation steps are simpler and the specificity is higher than ZFN. In 2010, TALEN protein was first applied in yeast, and TALEN technology has been widely used in crops, mice, zebrafish, human cells, etc. In 2012, TALEN technology was ranked as one of the top ten scientific breakthroughs by the top journal Science. In the same year, clustered regularly interspaced short palindromic repeats and their associated proteins (CRISPR / Cas) system began to attract widespread attention from researchers at home and abroad. Emmanuelle Charpentier and Jennifer A. Doudna first established a CRISPR / Cas9 system in vitro to cut double-stranded DNA. Zhang Feng reported in the journal Science that he used CRISPR / Cas9 to edit genes in human cells, applying CRISPR / Cas9 gene editing technology to eukaryotic research, and thus unveiled the veil of the third generation of gene editing technology. CRISPR / Cas, like ZFN and TALEN, can produce DNA double-strand breaks at specific gene locations, and thus can be applied to gene editing. Compared with the other two technologies, CRISPR / Cas technology has obvious advantages, such as simple operation, low cost, high efficiency and strong specificity, and is considered a gene editing tool with broad application prospects and a replacement for the previous two gene editing technologies.
[0006] CRISPR / Cas gene editing technology includes two important components: guide RNA and Cas effector protein that cuts DNA. The effector protein recognizes the protospacer adjacent motif (PAM) and targets specific locations in the genome to exert endonuclease function, causing DNA double-strand breaks and gene editing. The most widely used CRISPR / Cas system is the widely studied CRISPR / Cas9 system, and systems such as Cas12a in the CRISPR / Cas system are also continuously concerned.
[0007] Gene editing technology has a wide range of applications, including gene function research, plant breeding, and animal genetic improvement. However, the most attention is undoubtedly focused on the treatment of genetic diseases. Since the advent of gene editing technology, it has played an important role in the treatment of genetic diseases in animal disease models and even clinical diseases, becoming a key part of gene therapy. Currently, there are two main aspects of gene therapy: one is to deliver therapeutic genes to target organs or the whole body in vivo through viruses or materials; the other is to use patient or healthy cells as carriers to deliver successfully repaired cells or cells that can express therapeutic proteins to patients in cell therapy (ex vivo).
[0008] In vivo gene therapy methods have been developed for many years. Before the advent of gene editing technology, researchers began to use gene overexpression for gene therapy. After the discovery of adeno-associated virus (AAV), AAV-mediated gene therapy has been successful in various animal models of disease, especially in the clinical treatment of genetic diseases such as hemophilia. However, since AAV cannot completely change the patient's pathogenic gene at the genetic level, patients need to be injected with viruses multiple times to alleviate symptoms, which can cause immune reactions in the body and pose a serious threat to the patient's life. After the advent of gene editing technology, especially the advent of CRISPR / Cas9 technology, it has driven the development of gene therapy. Gene editing technology can destroy pathogenic genes or repair mutant genes in situ to achieve the goal of long-term treatment of diseases, and has been successfully applied to muscle diseases (Duchenne muscular dystrophy) and liver metabolic diseases (type I tyrosinemia, hemophilia, and phenylketonuria).
[0009] Cell therapy has developed most rapidly in recent years, mainly with hematopoietic stem cells and immune cells. The cells edited in vitro to restore function can be successfully treated by reinfusing the cells into patients with anemia and tumor cells, and currently several studies have been successfully applied to the clinic. Hematopoietic stem cell transplantation is the most promising method for treating thalassemia and sickle cell anemia. Immune cells, especially T cells, combined with gene editing technology are another hot spot in the field of tumor treatment. Chimeric antigen receptor T cell (CAR-T) therapy has saved the lives of many cancer patients in the clinic.
[0010] Gene editing technology has greatly promoted the development of gene therapy, but it also faces various problems, mainly focusing on the efficiency of gene therapy and the safety of cells in the body. The efficiency of gene therapy depends on the efficiency of gene editing, which is limited by two aspects: one is the delivery vector; the second is the gene integration efficiency. How to efficiently deliver gene editing tools to cells or the body has always been a problem for researchers to think about. For example, the current AAV delivery method suitable for clinical treatment, due to the maximum packaging capacity of AAV vector is less than 5 kb, so it is difficult to include CRISPR / Cas system and repair template in an AAV vector at the same time, mainly using segmented delivery to deliver CRISPR / Cas system and repair template separately, which greatly affects the effectiveness of gene therapy. The smaller the gene editing tool system, the more efficient it is, and the easier it is to construct the delivery vector. The size of the current mainstream CRISPR / Cas9, Cas12a, etc. system exceeds 5 kb, which brings great trouble to vector packaging and delivery. Therefore, developing smaller gene editing tool systems has always been the focus of researchers. Gene integration efficiency refers to the integration efficiency of repair templates in cells. High-efficiency site-specific integration is beneficial to the treatment of various genetic diseases, but integration efficiency is related to the DNA repair efficiency of cells themselves. Researchers have optimized to improve DNA damage repair efficiency, but there is no good technology to solve this problem. The safety problem of gene editing technology is the most important problem. Since the advent of gene editing technology, off-target problems have constantly appeared in research. Whether it is ZFN or CRISPR / Cas9 system, the off-target problem has not been solved. In addition to predictable non-specific cutting, genomic fragment deletion also occurs in vivo, which has an unpredictable impact on cell function and brings safety hazards. Moreover, the introduction of exogenous gene editing tool proteins also brings impact on cells due to the nature of the protein itself, which also prompts researchers to develop more accurate and safer gene editing tools.
[0011] In the past decade, the CRISPR / Cas system has developed into a variety of different gene editing systems as the mainstream gene editing technology, including single base editing technology based on the CRISPR / Cas system, which can also change the specific base conversion of the gene without introducing double-stranded DNA breaks. Although the gene editing technology is becoming more mature, the problems in its application have not been solved, and researchers have been exploring new gene editing tools and developing new gene editing technologies in the hope of solving the problem through new gene editing tools. In 2021, Zhang Feng et al. speculated that TnpB of the IS605 / IS200 family had RNA-guided nuclease activity, and then Virginijus et al. confirmed through a series of biochemical experiments that TnpB is a reprogrammable RNA-guided functional nuclease. TnpB does not have clustered regularly interspaced short palindromic repeats (CRISPR), which is guided by a new reRNA to cut human genomic DNA.
[0012] Meanwhile, in the development of gene editing technology, it is worth noting that whether it is zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN) or CRISPR / Cas system, it is basically based on foreign development, and domestic researchers optimize it.
[0013] The prior art needs new gene editing technology. SUMMARY
[0014] In view of the lack of gene editing tools with improved performance in the art, the inventors provide an endonuclease complex (also referred to herein as RAD or IS607 TnpB system) and the use of an endonuclease (also referred to herein as RAD or IS607 TnpB protein), which has a smaller molecular weight compared to existing products and can be more widely used. The RAD (IS607 TnpB) system of the present application is an RNA-guided endonuclease system.
[0015] The new gene editing system provided by the present application is derived from the IS607 transposase system, which consists of two protein-coding sequences ORFA and ORFB, and upstream and downstream transposition-related motifs. Current studies have shown that in the IS605 system, ORFB can be used for gene editing. Although the names IS605 and IS607 are similar, the mechanisms of transposition of the two are not the same. Similarly, in the IS607 system, ORFA plays a transposase activity, but ORFB has not been studied to show what the specific function is, so there is no gene editing field technical personnel attention and through biochemical experiments to prove whether the ORFB in the IS607 system has the characteristics of gene editing function.
[0016] In one aspect, the present application provides an endonuclease complex comprising an RNA-guided endonuclease and an RNA. The endonuclease complex can be used as an RNA-guided endonuclease system.
[0017] In one embodiment, the RNA-guided endonuclease comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191.
[0018] In one embodiment, the guide RNA comprises an RNA-guided endonuclease-associated RNA and a targeting RNA. The RNA-guided endonuclease-associated RNA is located at the 5’ end of the guide RNA, and the targeting RNA is located at the 3’ end, which can be directly connected or contiguous.
[0019] In one embodiment, the RNA-guided endonuclease-associated RNA comprises a sequence of any one of SEQ ID NOs: 58-114 and SEQ ID NOs: 192-233, or a sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of any one of SEQ ID NOs: 58-114 and SEQ ID NOs: 192-233. The RNA-guided endonuclease can be derived from IS607 TnpB.
[0020] In one embodiment, the targeting RNA is at the 3’ end of the RNA-guided endonuclease-associated RNA. In one embodiment, the targeting RNA is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical or complementary to a target sequence in a gene of interest (e.g., in a genome).
[0021] In one embodiment, the targeting RNA has a length of 12-40 bp. In one embodiment, the targeting RNA has a suitable length, e.g., 18-40 bp, and the first 14, 15, 16, 17, or 18 nucleotides of the targeting RNA are fully complementary to a target sequence in a gene of interest (e.g., in a genome).
[0022] In one embodiment, the RNA-guided endonuclease protein contains a RuvC domain. In one embodiment, the RNA-guided endonuclease protein comprises an amino acid at an amino acid position corresponding to the amino acid sequence of SEQ ID NO: 6 selected from the group consisting of Q18, G27, R30, N34, F50, W73, A92, F97, P104, K107, Y117, E145, P150, S173, G192, D194, G196, K198, A201, S204, N211, I212, N213, R227, Q229, S233, R234, E237, G242, N249, K252, N266, V280, K283, P284, E290, L292, N293, G296, M297, M298, K299, L303, S304, K305, Q310, K322, P338, S339, S340, C343, C346, L353, L355, R358, C362, C364, D369, R370, D371, A374, N377, and L378.
[0023] In one embodiment, the gene of interest (e.g., genome of interest) is single-stranded or double-stranded DNA. In one embodiment, the gene of interest (e.g., genome of interest) comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, for the system numbered 1-57, the gene of interest (e.g., genome of interest) comprises a motif selected from the group consisting of AGGAG or GAGGG.
[0024] In another aspect, the application provides a polynucleotide comprising a polynucleotide of an RNA-guided endonuclease herein and a polynucleotide for an RNA molecule. The polynucleotide for an RNA molecule can be a polynucleotide that transcribes an RNA molecule.
[0025] In another aspect, the application provides a vector comprising a polynucleotide molecule described herein. The vector can comprise an expression control sequence, e.g., a promoter or terminator. In one embodiment, the vector is an expression vector.
[0026] In another aspect, the application provides a viral particle (e.g., adeno-associated viral particle) comprising an endonuclease complex, polynucleotide, or vector described herein.
[0027] In another aspect, the application provides a method of specifically cleaving genomic DNA comprising the step of contacting an endonuclease complex described herein with the genomic DNA. In another aspect, the application provides a method of specifically cleaving genomic DNA comprising the step of contacting an endonuclease complex described herein with the genomic DNA.
[0028] In one embodiment, the DNA is genomic DNA of a eukaryotic cell, prokaryotic cell, or virus (e.g., a genome of interest).
[0029] In one embodiment, the DNA is single-stranded or double-stranded DNA.
[0030] In one embodiment, the method comprises the step of introducing a polynucleotide or vector herein into a cell.
[0031] In one embodiment, the targeted sequence of DNA (e.g., genomic DNA) comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, for systems numbered 1-57, the gene of interest (e.g., a genome of interest) comprises a motif selected from the group consisting of AGGAG or GAGGG. In one embodiment, the 3’ end of the motif is DNA that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical or complementary to the targeting RNA.
[0032] In one embodiment, the method is performed in vivo or in vitro.
[0033] In one embodiment, the method is performed under conditions selected from the group consisting of: 0.5 to 10 mM Mg 2+ , 25 to 250 mM Na + , a reaction temperature of 27-57 °C, and a reaction time of 3-60 minutes.
[0034] In another aspect, a method of performing gene editing is provided, comprising the step of contacting an endonuclease complex herein with genomic DNA of a cell to perform gene editing.
[0035] In one embodiment, the gene editing results in an insertion and / or deletion in the genomic DNA.
[0036] In one embodiment, the gene editing is performed in a mammalian cell to treat or prevent a disease, or to introduce a desired trait. In one embodiment, the gene of interest (e.g., a genome of interest) of the mammalian cell comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, for systems numbered 1-57, the gene of interest (e.g., a genome of interest) comprises a motif selected from the group consisting of AGGAG or GAGGG.
[0037] In another aspect, the present application provides a kit comprising one or more of the following: an endonuclease complex described herein, a viral particle, a polynucleotide described herein, and a vector described herein.
[0038] In another aspect, the present application provides a method of inactivating or identifying a virus comprising contacting the virus with an endonuclease complex or endonuclease described herein. In one embodiment, the endonuclease comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191. In one embodiment, a gene of interest (e.g., a genome of interest) of the virus comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, a gene of interest (e.g., a genome of interest) of the virus comprises a motif selected from the group consisting of AGGAG or GAGGG, e.g., for the systems numbered 1-57.
[0039] In another aspect, the present application provides an endonuclease complex or endonuclease described herein for use in inactivating or identifying a virus comprising contacting the virus with an endonuclease complex or endonuclease described herein. In one embodiment, the endonuclease comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191. In one embodiment, a gene of interest (e.g., a genome of interest) of the virus comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, a gene of interest (e.g., a genome of interest) of the virus comprises a motif selected from the group consisting of AGGAG or GAGGG, e.g., for the systems numbered 1-57.
[0040] In another aspect, the application provides use of an endonuclease complex or endonuclease described herein in the manufacture of a medicament for inactivating or identifying a virus. In one embodiment, the endonuclease comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191. In one embodiment, the gene of interest (e.g., the genome of interest) of the virus comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, the gene of interest (e.g., the genome of interest) of the virus comprises a motif selected from the group consisting of AGGAG or GAGGG, e.g., for systems numbered 1-57.
[0041] In another aspect, the application provides a method of treating a viral infection comprising administering to a patient infected with a virus an endonuclease complex described herein. In one embodiment, the gene of interest (e.g., the genome of interest) of the virus comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, the gene of interest (e.g., the genome of interest) of the virus comprises a motif selected from the group consisting of AGGAG or GAGGG, e.g., for systems numbered 1-57.
[0042] In another aspect, the application provides an endonuclease complex described herein for use in treating a viral infection, the use comprising administering to a patient infected with a virus an endonuclease complex described herein. In one embodiment, the gene of interest (e.g., the genome of interest) of the virus comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. In one embodiment, the gene of interest (e.g., the genome of interest) of the virus comprises a motif selected from the group consisting of AGGAG or GAGGG, e.g., for systems numbered 1-57.
[0043] In another aspect, the present application provides use of the endonuclease complex described herein in the manufacture of a medicament for treating a viral infection. In one embodiment, the gene of interest (e.g. genome of interest) of the virus comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN or GGGGN. In one embodiment, the gene of interest (e.g. genome of interest) of the virus comprises a motif selected from the group consisting of AGGAG or GAGGG, for example for systems numbered 1-57.
[0044] In another aspect, the present application provides a viral vector comprising the polynucleotide or vector herein. In one embodiment, the viral vector is an adeno-associated virus.
[0045] In another aspect, the present application provides a gene therapy for treating a disease, which comprises use of the viral vector or viral particle or adeno-associated virus particle described herein.
[0046] In another aspect, the present application provides the viral vector or viral particle or adeno-associated virus particle described herein for use in a gene therapy for treating a disease.
[0047] In another aspect, the present application provides use of the viral vector or viral particle or adeno-associated virus particle described herein in the manufacture of a medicament for a gene therapy for treating a disease.
[0048] In another aspect, the present application provides a cell comprising one or more of the following: the endonuclease complex, polynucleotide, vector or viral particle described herein. In one embodiment, the cell is a eukaryotic cell or a prokaryotic cell. In one embodiment, the eukaryotic cell is a fungal cell, a plant cell or an animal cell.
[0049] The benefits of the present application are that:
[0050] 1. The RNA sequence responsible for guiding cleavage in the endonuclease complex (i.e. RAD or IS607 TnpB system) described herein is independently found and confirmed by the applicant, which is the first discovery in the world. The function of the RAD (IS607 TnpB) protein in the endonuclease complex is also the first discovery in the world. (The sequence information of each component of the system is shown in Table 2)
[0051] 2. The RAD protein in the RAD (IS607 TnpB) system is small in size, for example, only consisting of 387 amino acids, which is even less than one third of the Cas9 protein in the most widely used CRISPR / Cas9 system at present. It has greater advantages for the delivery of gene therapy vectors.
[0052] 3. The RAD (IS607 TnpB) system possesses specific double-stranded DNA cleavage activity, a prerequisite for its development as a gene editing tool. It recognizes the 5' end motifs NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN (AGGAG or GAGGG for systems numbered 1-57). Currently, no gene editing tools recognize these motifs, expanding their repertoire. Furthermore, the RAD system exhibits higher specificity than the most widely used CRISPR / Cas9 system.
[0053] 4. The RAD (IS607 TnpB) system has both non-specific single-stranded DNA cleavage activity and single-stranded DNA trans-cleavage activity. This feature enables the RAD system to cleave and eliminate single-stranded DNA viruses existing in nature, and can also serve as a tool for single-stranded DNA virus identification.
[0054] 5. The RAD (IS607 TnpB) system successfully produced DNA interference in prokaryotes (Escherichia coli) and eukaryotes (human cells). This feature indicates that the RAD system can be used as a new gene editing tool, providing a basis for its development into an efficient and precise gene editing tool.
[0055] 6. RAD (IS607 TnpB) systems from 65 strains of 55 different bacterial genera exhibited in vitro DNA cleavage activity. RAD systems from 16 strains of 11 different bacterial genera also exhibited intracellular DNA cleavage activity. These diverse RAD systems are representative, demonstrating the widespread existence of DNA cleavage activity in RAD systems and suggesting that more RAD systems can be developed as gene editing tools. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 The acquisition process of the RAD (IS607 TnpB) system is shown.
[0057] Figure 2 The results of RNA sequencing matching to the loci are shown.
[0058] Figure 3 Results after treatment of the RAD complex with DNase and RNase, respectively, are shown.
[0059] Figure 4 The results of the analysis of transposon-associated motif (TAM sequence) sequences required by RAD systems from different sources are shown.
[0060] Figure 5 The results of the FbRADl protein TAM sequence and double-stranded DNA cleavage are shown.
[0061] Figure 6 The results of the FbRADl protein cleaving DNA after mutation of the active center of the RAD protein are shown.
[0062] Figure 7 The results of the secondary structure prediction of the RAD-RNA are shown.
[0063] Figure 8 The results of the alignment analysis of the RNA sequence in the RAD system with DNAse activity are shown.
[0064] Figure 9 The results of the alignment analysis of the protein sequence in the RAD system with DNAse activity are shown.
[0065] Figure 10 The results of the cleavage position analysis are shown.
[0066] Figure 11 The results of the cleavage of single-stranded DNA are shown.
[0067] Figure 12A The cleavage results of different motif sequences and mismatched bases are shown.
[0068] Figure 12B The results of the analysis of the in vitro cleavage of double-stranded DNA by the RAD protein are shown.
[0069] Figure 13 The results of the DNA interference experiment in E. coli are shown.
[0070] Figure 14 The results of the DNA interference in E. coli by different RAD systems (TAM sequence: AGGAG) are shown.
[0071] Figure 15 The results of the DNA interference in E. coli by different RAD systems (TAM sequence: GAGGG) are shown.
[0072] Figure 16 The results of the cleavage of DNA by the RAD system in human 293F cells are shown.
[0073] Figure 17 The results of the cleavage of DNA by different RAD systems in 293F cells are shown.
[0074] Figure 18 The results of the expression and purification of the RAD protein nucleic acid complex and the RNA sequencing process are shown.
[0075] Figure 19RAD plasmid library digestion and high throughput sequencing workflow is shown.
[0076] Figure 20 RAD protein in E. coli in vivo DNA cleavage and screening workflow is shown.
[0077] Figure 21 RAD protein in 293F cells DNA editing and detection workflow is shown. DETAILED DESCRIPTION
[0078] Unless specifically indicated otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Furthermore, any method or material similar or equivalent to those described herein can be used in the practice of the present disclosure. All publications cited herein are incorporated by reference in their entireties.
[0079] Unless defined otherwise herein, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. Definitions of common terms and techniques in the field of molecular biology and immunology can be found in Sambrook J. et al. (eds.), Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Press, Plainsview, New York (2001); Ausubel, F. M., et al. (eds.), Current Protocols in Molecular Biology, John Wiley & Sons, New York (2010); and Coligan, J. E. et al. (eds.), Current Protocols in Immunology, John Wiley & Sons, New York (2010). In addition, definitions for common terms in molecular biology can be found in Benjamin Lewin, Genes IX, published in Jones and Bartlett, 2008 (ISBN 0763752223); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, published by Blackwell Science Ltd., 1994 (ISBN 0632021829); and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, published by VCH Publishers, Inc., 1995 (ISBN 9780471185710). In the case of conflict between the definitions in the specification and those in the art, the definitions in the specification prevail.
[0080] The term "genome editing" refers to a type of genetic engineering in which DNA is inserted, replaced, or removed from a target DNA (e.g., the genome of a cell) using one or more nucleases. The nucleases create a specific double-stranded break (DSB) at a desired location in the genome and utilize endogenous mechanisms of the cell to repair the induced break by homology-directed repair (HDR) (e.g., homologous recombination) or non-homologous end joining (NHEJ). Any suitable nuclease can be introduced into a cell to induce genome editing of a target DNA sequence, including but not limited to CRISPR-associated protein (Cas) nucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), meganucleases, other endo- or exo-nucleases, variants thereof, fragments thereof, and combinations thereof. Nuclease-mediated genome editing can be performed using the stem-loop structure RNAs described herein in combination with RAD proteins.
[0081] The term "nucleic acid," "nucleotide," or "polynucleotide" refers to deoxyribonucleic acid (DNA), ribonucleic acid (RNA), and polymers thereof in single-, double-, or multi-stranded form. The term includes, but is not limited to, single-, double-, or multi-stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and / or pyrimidine bases, or other natural, chemically modified, biochemically modified, non-natural, synthetic, or derivatized nucleotide bases. The term nucleic acid is used interchangeably with gene, cDNA, and mRNA encoded by a gene.
[0082] The term "gene" or "polypeptide-encoding nucleotide sequence" refers to a segment of DNA involved in producing a polypeptide chain. The DNA segment can include regions preceding and following the coding region that are involved in the transcription / translation and regulation of transcription / translation of the gene product, as well as intervening sequences (introns) between individual coding segments (exons).
[0083] The term "complementary" refers to the capacity of one nucleic acid to hydrogen bond with another nucleic acid sequence by either traditional Watson-Crick or other non-traditional types of hydrogen bonding. Percent complementarity indicates the percentage of residues in a nucleic acid molecule that will hydrogen bond with a second nucleic acid sequence (e.g., 50%, 60%, 70%, 80%, 90%, and 100% complementary). "Perfectly complementary" refers to all consecutive residues of a nucleic acid sequence that will hydrogen bond with the same number of consecutive residues in a second nucleic acid sequence.
[0084] The term "expression vector" is a recombinantly or synthetically produced nucleic acid construct having a series of specific nucleic acid elements that allows for the transcription of a particular polynucleotide sequence in a host cell. An expression vector can be a plasmid, a viral genome, or a portion of a nucleic acid fragment. Typically, an expression vector includes a polynucleotide to be transcribed operably linked to a promoter.
[0085] In the present context, RAD (IS607 TnpB) system refers to the RAGATH-18 containing RNA Associated DNase system (RAD) is an abbreviation for the RAGATH18 related endonuclease, which is typically named according to the characteristics of the newly discovered system by the inventors. In the present context, RAD (IS607 TnpB) system is used interchangeably with endonuclease complex or RNA-guided endonuclease system. In the present context, RAD (IS607 TnpB) protein is the DNAse in the RAD (IS607 TnpB) system, which is also used interchangeably with endonuclease in the present context.
[0086] “Guide RNA” refers to a polynucleotide comprising: 1) a guide sequence (also referred to as “targeting RNA”) capable of hybridizing or complementary to a target sequence (also referred to herein as “targeting sequence”); and 2) a scaffold sequence capable of interacting with an endonuclease (referred to herein as “RNA-guided endonuclease-associated RNA”). The guide nucleic acid can be RNA. The guide RNA can be encoded by a DNA sequence on a polynucleotide molecule.
[0087] As used herein, the term “RuvC domain” refers to a conserved domain or motif of amino acids having nuclease (e.g., endonuclease) activity. As used herein, a protein having split RuvC domains refers to a protein having two or more RuvC motifs at different sequential positions within the sequence that interact in tertiary structure to form a RuvC domain.
[0088] The term “subject” includes a human or an animal. For example, an animal subject can be a mammal, a primate (e.g., a monkey), a livestock animal (e.g., a horse, a cow, a sheep, a pig, or a goat), a companion animal (e.g., a dog, a cat), a laboratory test animal (e.g., a mouse, a rat, a guinea pig, a bird), an animal of veterinary significance, or an animal of economic significance.
[0089] The term "sequence identity" or "identity" is used in the context of relatedness between two amino acid sequences or between two nucleotide sequences. The sequence identity between two amino acid sequences can be determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, J. Mol. Biol. 48: 443-453) as implemented in the Needle program of the EMBOSS software package (EMBOSS: The European Molecular Biology open software suite, Rice et al., 2000, Trends Genet. 16: 276-277, preferably version 6.6.0 or later), using the default parameters. The parameters used are a gap open penalty of 10 and a gap extension penalty of 0.5, and the EBLOSUM62 (EMBOSS version of BLOSUM62) substitution matrix. To produce results that report the "longest identity", the "nobrief" option must be specified in the command line. The output of the Needle program is computed as follows:
[0090] (identical residues x 100) / (length of alignment - total number of gaps in the alignment)
[0091] The sequence identity between two polynucleotide sequences can also be determined using the Needleman-Wunsch algorithm (Needleman and Wunsch, 1970, supra) as implemented in the Needle program of the EMBOSS package (EMBOSS: The European Molecular Biology open software suite, Rice et al., 2000, supra) (preferably version 6.6.0 or later), using the default parameters. The parameters used are a gap open penalty of 10 and a gap extension penalty of 0.5, and the EDNAFULL (EMBOSS version of NCBI NUC4.4) substitution matrix. To produce results that report the "longest identity", the "nobrief" option must be specified in the command line. The output of the Needle program is computed as follows:
[0092] (identical deoxyribonucleotides x 100) / (length of alignment - total number of gaps in the alignment,
[0093] Provided herein are endonuclease complexes comprising an RNA-guided endonuclease and an RNA. The RNA-guided endonuclease comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191. The guide RNA can comprise an RNA-guided endonuclease-associated RNA and a targeting RNA. The RNA-guided endonuclease-associated RNA can comprise a sequence of any one of SEQ ID NOs: 58-114 and SEQ ID NOs: 192-233 or a sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the sequence of any one of SEQ ID NOs: 58-114 and SEQ ID NOs: 192-233. The endonuclease systems of the present invention can comprise: an endonuclease of SEQ ID NO: 1 and a RNA-guiding sequence comprising SEQ ID NO: 58; an endonuclease of SEQ ID NO: 2 and a RNA-guiding sequence comprising SEQ ID NO: 59; an endonuclease of SEQ ID NO: 3 and a RNA-guiding sequence comprising SEQ ID NO: 60; an endonuclease of SEQ ID NO: 4 and a RNA-guiding sequence comprising SEQ ID NO: 61; an endonuclease of SEQ ID NO: 5 and a RNA-guiding sequence comprising SEQ ID NO: 62; an endonuclease of SEQ ID NO: 6 and a RNA-guiding sequence comprising SEQ ID NO: 63; an endonuclease of SEQ ID NO: 7 and a RNA-guiding sequence comprising SEQ ID NO: 64; and so on. That is, the endonuclease systems of the present invention can comprise a pair of an endonuclease of SEQ ID NO: N and a RNA-guiding sequence comprising SEQ ID NO: N+57, N being an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57.The endonuclease system of the present invention can comprise an endonuclease of SEQ ID NO: M and a pair comprising an RNA guide sequence of SEQ ID NO: M+42, M being an integer of 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, or 191. The endonuclease system of the present invention can comprise 99 pairs of RNA-guided endonucleases and RNAs listed herein.
[0094] The RNA guide sequence can comprise a RAGATH18-related RNA sequence or an endonuclease-related RNA sequence and a targeting RNA sequence. The endonuclease and the RNA guide sequence can have variations in sequence.
[0095] In one embodiment, the endonuclease complex comprises an RNA-guided endonuclease of SEQ ID NO: 6 and a guide RNA of SEQ ID NO: 63; an RNA-guided endonuclease of SEQ ID NO: 11 and a guide RNA of SEQ ID NO: 68; an RNA-guided endonuclease of SEQ ID NO: 13 and a guide RNA of SEQ ID NO: 70; an RNA-guided endonuclease of SEQ ID NO: 21 and a guide RNA of SEQ ID NO: 78; an RNA-guided endonuclease of SEQ ID NO: 43 and a guide RNA of SEQ ID NO: 100; an RNA-guided endonuclease of SEQ ID NO: 51 and a guide RNA of SEQ ID NO: 108; an RNA-guided endonuclease of SEQ ID NO: 52 and a guide RNA of SEQ ID NO: 109; an RNA-guided endonuclease of SEQ ID NO: 54 and a guide RNA of SEQ ID NO: 111; or an RNA-guided endonuclease of SEQ ID NO: 55 and a guide RNA of SEQ ID NO: 112.
[0096] The guide RNA can exhibit a stem-loop structure. The targeting RNA can be at the 3' end of the RNA-guided endonuclease-related RNA. The targeting RNA can be at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical or complementary to a target sequence of a genome of interest. The targeting RNA has a length of 12-40 bp. For example, the targeting RNA has a length of 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 bp.
[0097] The RNA-guided endonuclease protein can contain a functional RuvC domain. The RNA-guided endonuclease protein can comprise an amino acid selected from the group consisting of Q18, G27, R30, N34, F50, W73, A92, F97, P104, K107, Y117, E145, P150, S173, G192, D194, G196, K198, A201, S204, N211, I212, N213, R227, Q229, S233, R234, E237, G242, N249, K252, N266, V280, K283, P284, E290, L292, N293, G296, M297, M298, K299, L303, S304, K305, Q310, K322, P338, S339, S340, C343, C346, L353, L355, R358, C362, C364, D369, R370, D371, A374, N377, and L378 at an amino acid position corresponding to the amino acid sequence of SEQ ID NO: 6.
[0098] The genome of interest can be single-stranded or double-stranded DNA. The genome of interest can comprise a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN.
[0099] The present disclosure provides a method of specifically cleaving genomic DNA, comprising the step of contacting an endonuclease complex described herein with the genomic DNA. The genomic DNA can be genomic DNA of a eukaryotic cell, a prokaryotic cell, or a virus. The eukaryotic cell is not particularly limited, including but not limited to a mammalian (e.g., human) cell, a yeast cell. The prokaryotic cell is not particularly limited, including but not limited to an E. coli cell. The genomic DNA can be single-stranded or double-stranded DNA. The method can comprise the step of introducing a polynucleotide or a vector herein into the cell. The targeted sequence of the genomic DNA can comprise a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN. The 3’ end of the motif can be DNA that is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical or complementary to the targeting RNA. The method can be performed in vivo or in vitro. The method is performed under conditions selected from the group consisting of 0.5 to 10 mM divalent metal ions selected from the group consisting of Mg 2+ , Mn 2+ , and Ca 2+ , 25 to 250 mM Na + , the reaction temperature is 27-57 °C, and the reaction time is 3-60 minutes.
[0100] The concentration of divalent metal ions can be 0.5 to 10 mM, for example, 1, 2, 2.5, 3, 4, 5, 6, 7, 8, 9, 10 nM. The kind of salt that provides the divalent metal ions is not particularly limited, including but not limited to a chloride salt.
[0101] The concentration of Na + may be 25 to 250 mM, for example, 50, 100, 150, or 200 mM. The kind of sodium salt is not particularly limited, including but not limited to a chloride salt.
[0102] Provided is a method of performing gene editing, comprising the step of contacting an endonuclease complex herein with the genomic DNA of a cell to perform gene editing. The gene editing can generate an insertion and / or a deletion in the genomic DNA. The gene editing can be performed in a mammalian cell to treat or prevent a disease, or to introduce a desired trait.
[0103] Provided is a kit comprising one or more of the following: an endonuclease complex described herein, a polynucleotide described herein, and a vector described herein. Alternatively, the kit can comprise any component used herein.
[0104] Methods of inactivating or identifying a virus are provided, comprising contacting the virus with an endonuclease complex or endonuclease described herein. In one embodiment, the endonuclease comprises an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence that is at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191.
[0105] The RAD(IS607 TnpB) system of the present application can comprise the effector proteins of Table 1. However, effector proteins with similar functions can also be used as endonucleases for the system of the present application.
[0106] Table 1: Source strains and effector proteins of 57 RAD(IS607 TnpB) systems
[0107]
[0108]
[0109]
[0110]
[0111]
[0112] Exemplary endonucleases and their associated RNA sequences of the present application are listed in Table 2. The endonuclease system of the present application can comprise: an endonuclease of SEQ ID NO: 1 and an RNA guide sequence comprising SEQ ID NO: 58; an endonuclease of SEQ ID NO: 2 and an RNA guide sequence comprising SEQ ID NO: 59; an endonuclease of SEQ ID NO: 3 and an RNA guide sequence comprising SEQ ID NO: 60; an endonuclease of SEQ ID NO: 4 and an RNA guide sequence comprising SEQ ID NO: 61; an endonuclease of SEQ ID NO: 5 and an RNA guide sequence comprising SEQ ID NO: 62; an endonuclease of SEQ ID NO: 6 and an RNA guide sequence comprising SEQ ID NO: 63; an endonuclease of SEQ ID NO: 7 and an RNA guide sequence comprising SEQ ID NO: 64; and so on. That is, the endonuclease system of the present application can comprise an endonuclease of SEQ ID NO: N and an RNA guide sequence comprising SEQ ID NO: 57+N, N being an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57. The RNA guide sequence can comprise a RAGATH18-related RNA sequence or an endonuclease-related RNA sequence and a targeting RNA sequence. The endonuclease and the RNA guide sequence can have variations in sequence. The endonuclease can comprise an amino acid sequence selected from the group consisting of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, or an amino acid sequence having at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to any one of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191. The guide RNA can comprise an RNA-guided endonuclease-related RNA and a targeting RNA, the RNA-guided endonuclease-related RNA can comprise a sequence of any one of SEQ ID NOs: 58-114 and SEQ ID NOs: 192-233 or a sequence having at least 30%, 40%, 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identity to the sequence of any one of SEQ ID NOs: 58-114 and SEQ ID NOs: 192-233.The targeting RNA is at the 3' end of the RNA-guided endonuclease-related RNA and is at least 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% identical or complementary to the targeting sequence of the gene of interest.
[0113] Twelve RAD(IS607 TnpB) systems were found to successfully cleave in prokaryotes under the TAM sequences AGGAG and GAGGG, which are from 7 different genera, respectively, (6) FbRAD1 / ISFba1 (RHS75625.1), (10) CbRAD1 / ISCba1 (RHR46218.1), (11) RiRAD / ISRin (RGR69673.1), (13) ErRAD / ISEre1 (RHB06265.1), (15) WcRAD (RGR87203.1), (18) BsRAD / ISBsp4 (RHV03091.1), (20) CbRAD2 / ISClsp2 (RHR42937.1), (21) CbRAD3 / ISClsp3 (RGD88909.1), (24) CcRAD / ISCco (RGT87174.1), (30) FbRAD2 / ISFba2 (RGH01875.1), (42) FbRAD3 / ISFba3 (RHU26542.1), (43) FbRAD4 / ISFba4 (RGH37793.1). Different RAD systems were expressed in 293F cells to verify whether they all have in vivo editing function in mammalian cells, and it was found that 9 RAD(IS607 TnpB) systems can be edited in vivo, which are from 8 different genera, respectively, (6) FbRAD1 / ISFba1 (RHS75625.1), (11) RiRAD / ISRin (RGR69673.1), (13) ErRAD / ISEre1 (RHB06265.1), (21) CbRAD3 / ISClsp3 (RGD88909.1), (43) FbRAD4 (RGH37793.1), (51) DfRAD / ISDfo4 (GUT_GENOME239804_2_85784_86938_+), (52) ArRAD / ISAre (GUT_GENOME047820_36_478_1632_-), (54) HuRAD / ISHun (GUT_GENOME022845_45_3955_5109_+), (55) AfRAD / ISAfa (GUT_GENOME112567_131_1855_3009_-) Figure 17 )。
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128] Additional 42 proteins of RAD (IS607 TnpB) systems and corresponding RNA sequences. For RNA sequences, "U" in the following sequences can be shown as "T" in the sequence listing.
[0129] No. 58 (continued from Table 2 No.)
[0130] >ERR3307045_63958_30277_4897_6000_-(SEQ ID NO: 150)
[0131]
[0132] Related RNA sequence (SEQ ID NO: 192)
[0133]
[0134] No. 59:
[0135] >GUT_GENOME151180_45_9178_10326_+(SEQ ID NO: 151)
[0136]
[0137] Related RNA sequence (SEQ ID NO: 193)
[0138]
[0139] No. 60:
[0140] >ERR2017626_1022590_12863_8034_9173_+ (SEQ ID NO: 152)
[0141]
[0142] Related RNA sequence (SEQ ID NO: 194)
[0143]
[0144] No. 61:
[0145] > GUT_GENOME142162_18_64466_65659_ - (SEQ ID NO: 153)
[0146]
[0147]
[0148] Related RNA sequence (SEQ ID NO: 195)
[0149]
[0150] No. 62:
[0151] > NC_013216_1_421908_422903_+ (SEQ ID NO: 154)
[0152]
[0153] Related RNA sequence (SEQ ID NO: 196)
[0154]
[0155] No. 63:
[0156] > NC_004557_1_WP_035125047.1 (SEQ ID NO: 155)
[0157]
[0158] Related RNA sequence (SEQ ID NO: 197)
[0159]
[0160] No. 64:
[0161] > GUT_GENOME096560_11_244135_245259_-(SEQ ID NO: 156)
[0162]
[0163] Related RNA sequence (SEQ ID NO: 197)
[0164]
[0165]
[0166] No. 65:
[0167] > GUT_GENOME083533_37_11359_12597_-(SEQ ID NO: 157)
[0168]
[0169] Related RNA sequence (SEQ ID NO: 199)
[0170]
[0171] No. 66:
[0172] > NC_022878_1_WP_023518930.1 (SEQ ID NO: 158)
[0173]
[0174] Related RNA sequence (SEQ ID NO: 200)
[0175]
[0176] No. 67:
[0177] > NZ_CP020815_1_WP_012574395.1 (SEQ ID NO: 159)
[0178]
[0179] Related RNA sequence (SEQ ID NO: 201)
[0180]
[0181] No. 68:
[0182] > NZ CP024701_1 WP_099988772.1 (SEQ ID NO: 160)
[0183]
[0184] Related RNA sequence (SEQ ID NO: 202)
[0185]
[0186] No. 69:
[0187] > GUT_GENOME228548_94_18418_19641_+ (SEQ ID NO: 161)
[0188]
[0189] Related RNA sequence (SEQ ID NO: 203)
[0190]
[0191] No. 70:
[0192] > NC_010001_1 WP_012201916.1 (SEQ ID NO: 162)
[0193]
[0194] Related RNA sequence (SEQ ID NO: 204)
[0195]
[0196] No. 71:
[0197] > NC_010003_1 WP_081429131.1 (SEQ ID NO: 163)
[0198]
[0199]
[0200] Related RNA sequence (SEQ ID NO: 205)
[0201]
[0202] No. 72:
[0203] > GUT_GENOME249453_33_14687_15901_-(SEQ ID NO: 164)
[0204]
[0205] Related RNA sequence (SEQ ID NO: 206)
[0206]
[0207] No. 73:
[0208] > GUT_GENOME046003_7_13445_14386_-(SEQ ID NO: 165)
[0209]
[0210] Related RNA sequence (SEQ ID NO: 207)
[0211]
[0212] No. 74:
[0213] > GUT_GENOME043660_1_131655_132716_-(SEQ ID NO: 166)
[0214]
[0215] Related RNA sequence (SEQ ID NO: 208)
[0216]
[0217] No. 75:
[0218] > NZ_CP014176_1_WP_039636020.1 (SEQ ID NO: 167)
[0219]
[0220] Related RNA sequence (SEQ ID NO: 209)
[0221]
[0222] No. 76:
[0223] > NC_014614_1_WP_013360293.1 (SEQ ID NO: 168)
[0224]
[0225] Related RNA sequence (SEQ ID NO: 210)
[0226]
[0227] Number 77:
[0228] > GUT_GENOME243333_13_9207_10436_+ (SEQ ID NO: 169)
[0229]
[0230] Related RNA sequence (SEQ ID NO: 211)
[0231]
[0232] Number 78:
[0233] > GUT_GENOME228909_131_318_1547_- (SEQ ID NO: 170)
[0234]
[0235]
[0236] Related RNA sequence (SEQ ID NO: 212)
[0237]
[0238] Number 79:
[0239] > ref_RGS74140_1__44001_45231_+ QRWV01000004_1 (SEQ ID NO: 171)
[0240]
[0241] Related RNA sequence (SEQ ID NO: 213)
[0242]
[0243] Number 80:
[0244] > GUT_GENOME029159_118_4969_5874_- (SEQ ID NO: 172)
[0245]
[0246] Related RNA sequence (SEQ ID NO: 214)
[0247] 81:
[0249] > NC_010371_1_WP_012289953.1 (SEQ ID NO: 173)
[0250]
[0251] Related RNA sequence (SEQ ID NO: 215)
[0252]
[0253]
[0254] Number 82:
[0255] > NZ_CP009687_1_WP_044824965.1 (SEQ ID NO: 174)
[0256]
[0257] Related RNA sequence (SEQ ID NO: 216)
[0258]
[0259] Number 83:
[0260] > GUT_GENOME100288_478_819-1994_-(SEQ ID NO: 175)
[0261]
[0262] Related RNA sequence (SEQ ID NO: 217)
[0263]
[0264] Number 84:
[0265] > GUT_GENOME089678_61_13091_14269_-(SEQ ID NO: 176)
[0266]
[0267] Related RNA sequence (SEQ ID NO: 218)
[0268]
[0269] No. 85:
[0270] > GUT_GENOME 190643_17 8068 9183 - (SEQ ID NO: 177)
[0271]
[0272]
[0273] Related RNA sequence (SEQ ID NO: 219)
[0274]
[0275] No. 86:
[0276] > ERR1539620_45264_57524_39821_41020_ - (SEQ ID NO: 178)
[0277]
[0278] Related RNA sequence (SEQ ID NO: 220)
[0279]
[0280] No. 87:
[0281] > NZ CP013243 1 4129900 4131015 + WP 072587064.1 (SEQ ID NO: 179)
[0282]
[0283] Related RNA sequence (SEQ ID NO: 221)
[0284]
[0285] No. 88:
[0286] > CP029458 1 2205965 2207104 - AYF55140.1 (SEQ ID NO: 180)
[0287]
[0288] Related RNA sequence (SEQ ID NO: 222)
[0289]
[0290] No. 89:
[0291] >CP013700_1_58380_59501_-AUM93701.1 (SEQ ID NO: 181)
[0292]
[0293] Related RNA sequence (SEQ ID NO: 223)
[0294]
[0295] No. 90:
[0296] >SRR5892202_233507_16022_11195_12361_+ (SEQ ID NO: 182)
[0297]
[0298] Related RNA sequence (SEQ ID NO: 224)
[0299]
[0300] No. 91:
[0301] >NC_014171_1_3506713_3507831_-WP_000978855.1 (SEQ ID NO: 183)
[0302]
[0303] Related RNA sequence (SEQ ID NO: 225)
[0304]
[0305] No. 92:
[0306] >NC_017209_1_13469_14587_-WP_000604588.1 (SEQ ID NO: 184)
[0307]
[0308]
[0309] Related RNA sequence (SEQ ID NO: 226)
[0310]
[0311] Number 93:
[0312] > NZ CP012098_1_2114035_2115228_+ WP 077326813.1 (SEQ ID NO: 185)
[0313]
[0314] Related RNA sequence (SEQ ID NO: 227)
[0315]
[0316] Number 94:
[0317] > IS607_orfB_GUT_GENOME170581_2_14431_15588_-(SEQ ID NO: 186)
[0318]
[0319] Related RNA sequence (SEQ ID NO: 228)
[0320]
[0321] Number 95:
[0322] > NZ CP011114_1_2448296_2449408_- WP 025697990.1 (SEQ ID NO: 187)
[0323]
[0324] Related RNA sequence (SEQ ID NO: 229)
[0325]
[0326] Number 96:
[0327] > CP035280_1_2081523_2082695_- QAT40560.1 (SEQ ID NO: 188)
[0328]
[0329] Related RNA sequence (SEQ ID NO: 230)
[0330]
[0331] Number 97:
[0332] >NC_017078_1_59860_61161_-WP_014431089.1 (SEQ ID NO: 189)
[0333]
[0334] Related RNA sequence (SEQ ID NO: 231)
[0335]
[0336] Number 98:
[0337] >IS607_orfB_GUT_GENOME088555_122_512_1693_+ (SEQ ID NO: 190)
[0338]
[0339] Related RNA sequence (SEQ ID NO: 232)
[0340]
[0341] Number 99:
[0342] >CP030117_1_439976_441088_+AWX54000.1 (SEQ ID NO: 191)
[0343]
[0344] Related RNA sequence (SEQ ID NO: 233)
[0345]
[0346] Example
[0347] The present application is further described by the following examples which should not be construed as limiting the scope of the present application.
[0348] Example 1: Discovery and acquisition of gene sequences for the RAD (IS607 TnpB) system
[0349] Non-redundant, high-quality gut strain genome sequences of 1520 healthy human fecal samples were downloaded from NCBI (Project ID PRJNA482748), 19165 complete bacterial genome sequences were downloaded from NCBI FTP site (ftp: / / ftp.ncbi.nih.gov / genomes / Bacteria / ), 26097 bacterial genome sequences were downloaded from IMG database (https: / / img.jgi.doe.gov / ), 12715 assembled microbial genome sequences from marine metagenome were collected from the literature (Maria G Pachiadaki et al., Cell. 2019 Dec 12; 179(7): 1623-1635. el l. doi: 10.1016 / j.cell.2019.11.017; Charting the Complexity of the Marine Microbiome through Single-Cell Genomics), 204938 human gastrointestinal isolated microbial genome sequences were downloaded from MGnify database, and 141115 un-assembled metagenomic samples covering multiple species (plants, insects, birds, etc.) were downloaded from MGnify data, 96709 un-assembled metagenomic samples of multiple human tissue sites were downloaded from HumanMetagenomeDB, and further assembled to obtain genome sequences.
[0350] The inventors analyzed the intergenic region (IGR) near the defense-related genes one by one by enumeration method, and predicted the conserved secondary structure of the RNA near the IGR by Rfam Figure 1 ), and the prediction method is referred to Rfam 14: expanded coverage of metagenomic, viral and microRNA families. Hundreds of secondary structure conserved RNAs were found, and 10 upstream and downstream proteins were analyzed for gene family to explore other functional elements cooperating with these defense-related RNA candidate proteins. The analysis identified the RAGATH-18 related DNA nuclease system (RAD), which contains RAGATH-18 related RNA and an IS607 element, and the effector protein of RAD (IS607 TnpB) is the second coding sequence ORFB of IS607, which encodes about 387 amino acids.
[0351] The inventors selected a total of 99 (Table 1) source strains and effector proteins of the RAD (IS607 TnpB) system as representatives for subsequent analysis. Based on the effector proteins of the 99 RAD systems, the inventors obtained the DNA sequence information of the candidate proteins and the corresponding amino acid sequences and sequence information of RAGATH-18-related RNA from the above-mentioned multiple genomic data. The sequences of the candidate protein sequences and the corresponding RAGATH-18-related RNA are shown in Table 2 and the sequences below Table 2. Based on the strain ID, protein ID and amino acid information listed in this article, the corresponding DNA information of the protein can be obtained in the above-mentioned database.
[0352] Example 2: Expression and purification of RAD (IS607 TnpB) protein-nucleic acid complex
[0353] The obtained DNA coding sequences of the candidate proteins were sent to a gene synthesis company (Genwizhi Biotechnology Co., Ltd.) for gene synthesis. The Pgex-6p1 universal vector was selected as the expression vector. The gene synthesis company synthesized 99 expression plasmids PGEX-RAD 1 to 99 (containing the candidate protein coding sequences of SEQ ID NOs: 1-57 listed in Table 2 and the subsequent SEQ ID NOs: 150-191 and the corresponding RNA coding sequences, respectively). The expression cassettes were inserted into the BamH1 and EcoR1 sites, and the expression cassette structure was Rad protein-6HIS-RNA-N20-HDV (Rad protein was each of SEQ ID NOs: 1-57 and SEQ ID NOs: 150-191, such as FbRAD1 protein SEQ ID NO: 6).
[0354] The sequence of a typical plasmid PGEX-6Pl-FbRAD1-N20-HDV is shown below, wherein the RAD protein portion FbRAD1 protein is the bold underlined portion; 6HIS is the underlined portion, the RNA coding sequence is the italic underlined portion; N20 is the bold portion; and HDV is the italic portion.
[0355]
[0356]
[0357] The expression plasmids of other Rad proteins were synthesized by the gene synthesis company in a similar manner as PGEX-6P1-FbRAD1-N20-HDV, except that the FbRAD1 coding sequence was replaced with the coding sequence of other RAD proteins, and the RNA sequence of FbRAD1 was replaced with the RNA sequence of other RAD system, thereby obtaining the expression plasmids PGEX-RAD 1-99 of all 99 candidate proteins (including PGEX-6P1-FbRAD1-N20-HDV). All plasmids were verified by the gene synthesis company to be correctly cloned.
[0358] Plasmid transformation experiment : The E. coli C43 (DE3) competent cells stored in a refrigerator at -80°C were taken out and placed on ice. After thawing, 100 ng of each of the expression plasmids PGEX-RAD 1-99 was added to 50 μL of E. coli cells, mixed, incubated on ice for 30 minutes, heat shocked in a water bath at 42°C for 90 seconds, then immediately placed on ice for 2 minutes, and then operated aseptically in a clean bench. 500 μL of LB liquid medium without antibiotics was added to an EP tube, and cultured in a constant temperature shaker at 37°C for 45-60 minutes. After that, it was uniformly coated on an LB plate containing 100 μg / mL ampicillin resistance, and inverted in a 37°C incubator for overnight culture. The next day, single colonies were observed to grow.
[0359] RAD (IS607 TnpB) System expression and purification: Pick positive clones into 5 mL LB liquid medium with ampicillin, incubate at 37°C for 12 hours, transfer to LB liquid medium with ampicillin, incubate at 37°C until OD600 is about 0.6, then cool to 18°C, add IPTG to a final concentration of 0.6 mmol / L, induce for 16 hours, collect bacteria the next day. We use a gravity column with GST-Sepharose as the column material, and use the affinity of the GST tag and GST-Sepharose to extract the protein from the cell lysate. Centrifuge the overnight culture at 4°C, 4000 rpm for 10 minutes, add resuspension liquid according to the standard of 20 mL GST resuspension liquid per 1 L of bacterial liquid, shake well with a vortex shaker, and resuspend the bacteria. When resuspending, add PMSF to a final concentration of 2 mmol / L, and the resuspension liquid contains 25 mmol / L Tris (pH 8.0), 1 mol / L NaCl, and 3 mmol / L DTT. Use ultrasonic disruption to fully lyse the resuspended bacteria, and the entire ultrasonic disruption process should be performed on ice. To prevent the probe from overheating, pause for 3 seconds after every 3 seconds of ultrasonic disruption, and the ultrasonic disruption time is 1 minute, usually 5-8 times. Until the bacterial liquid is no longer viscous and becomes a homogenate. After lysing, the bacterial liquid is aliquoted into pre-cooled high-speed centrifuge tubes at 4°C, centrifuged at 15000 rpm and 4°C for 40-50 minutes, and the supernatant is collected. Pour the supernatant into a gravity column containing 2 mL of GST Sepharose column material, and make sure it flows through the column material. Use a salt (NaCl) gradient to wash it to remove non-specifically bound proteins and nucleic acids. First, use 40 mL of GST wash solution I (25 mmol / L Tris (pH 8.0), 1.5 mol / L NaCl, 3 mmol / L DTT) to wash, then use 15 mL of GST wash solution II (25 mmol / L Tris (pH 8.0), 0.5 mol / L NaCl, 3 mmol / L DTT) to wash, and finally use 10 mL of GST wash solution III (25 mmol / L Tris (pH 8.0), 0.3 mol / L NaCl, 3 mmol / L DTT) to wash, and all the wash solutions are pre-cooled at 4°C. Add 5 mL of GST wash solution III to seal the gravity column, add 80 μL of Prescission protease for overnight enzyme digestion at 4°C, resuspend the column 3-4 times with an interval of 30 minutes. The next day, add another 5 mL of GST wash solution III and collect the flow-through. Concentrate to 2 mL, centrifuge at 15000 rpm and 4°C for 5 minutes, and prepare for the next step of purification. According to the instrument specifications, connect the gel filtration chromatography column (Superdex200 100 / 300) to the FPLC fast protein liquid chromatography workstation.The gel filtration chromatography column (Superdex200 100 / 300) was equilibrated with pre-cooled gel filtration chromatography buffer (10 mmol / L Tris-HCl pH 8.0, 150 mmol / L NaCl, 3 mmol / L DTT, 2 mmol / L MgCl2). At least one column volume was equilibrated until the UV absorption at 280 nm was stable and no longer changed. 2 mL sample was injected into the workstation sample ring, and the workstation preset program was run. After the program was completed, the purity of the protein was observed by 10% SDS-PAGE, and the binding of nucleic acid was observed by 10% denaturing Urea-PAGE EB staining. For FbRAD1 complex, Figure 3 The results observed by 10% denaturing Urea-PAGE EB staining are shown in the leftmost lane of Figure 3 PAGE. After PAGE verification, the actual molecular weight of the RAD protein and RNA of all 99 RAD nucleic acid complexes was consistent with the expected molecular weight, and the RAD nucleic acid complex was correctly expressed. The following table shows experimental data for several typical complexes.
[0360]
[0361]
[0362] Nucleic acid identification : The RAD nucleic acid complex with a final concentration of 5 mg / ml was treated with DNase and RNase with a final concentration of 0.1 mg / ml in a 10 μL reaction system, and the results are shown in Figure 3 : Only RNase can completely degrade the nucleic acid on the complex, indicating that the nucleic acid bound by the RAD protein is RNA. All 99 RAD nucleic acid complexes were verified, and the results proved that the nucleic acid bound by the RAD protein of the 99 RAD nucleic acid complexes was RNA.
[0363] RNA sequencing : The purified complex was sent to a gene sequencing company (Pisenol Biotechnology Co., Ltd.) for processing, and sequence data information Figure 18 was obtained.
[0364] According to the sequencing results, the inventors found that the RNA bound by the RAD protein was the non-coding RNA containing RAGATH-18 downstream of the second coding frame of IS607 (as shown in Figure 2 ), indicating that the RAD system can express a protein RNA complex, and this RNA is unique to the system, and there is currently no related research involving the function of this non-coding RNA and the upstream protein.
[0365] Example 3: Determination of PAM sequence and enzyme activity
[0366] Restriction enzyme substrate construction PAM substrate plasmid library construction: a plasmid library covering 7 base arbitrary combinations was used, primers: PAM-I-F and PAM-I-R, PCR amplification (amplification time 3 minutes, primer annealing temperature 58°C, amplification cycle number 30) recombinant fragments; primers: PAM-V-F and PAM-V-R, PCR amplification (amplification time 3 minutes, primer annealing temperature 58°C, amplification cycle number 30) recombinant vector PUC19, recombine the fragments and the vector PUC19 by the recombination enzyme system of Vazyme (product number Vazyme.C112-01), transform into E. coli DH5α competent cells, directly amplify and culture, and extract the plasmid library using a gel extraction kit (product number CW2302; Kangwei Century Biotechnology Co., Ltd.) according to the manufacturer's instructions.
[0367] The PAM plasmid library is shown below, which is based on the modification of the PUC19 plasmid, N represents the four bases of ATCG, and the underlined sequence is the targeted sequence.
[0368]
[0369] 2700 bp double-stranded DNA substrate: PCR amplification of a 2700 bp double-stranded DNA fragment using primers DS2700-F and DS2700-R with the PAM substrate plasmid library as the template, amplification time 3 minutes, primer annealing temperature 58°C, amplification cycle number 30, use 1% agarose gel electrophoresis to separate the PCR product, and recover the product. The 2700 bp mismatched double-stranded DNA substrate (complementary to CTGATGGTCCATGTCTGTTA bases, one by one change to form mismatches), the construction method is the same. 5'FAM-labeled single-stranded DNA substrates 5FAM-55TAM-F, 5FAM-55TAM-R, 5FAM-55NOTAM-F, 5FAM-55NOTAM-R, and 5FAM-NTSSDNA were synthesized by Genesyn Biotech Co., Ltd., and the primer sequences and single-stranded DNA substrate sequences are shown below.
[0370] 2700 bp double-stranded DNA substrate (AGGAG)
[0371] The 2700 bp double-stranded DNA substrate sequence used in the in vitro enzyme cleavage experiment was based on the modification of the PUC19 plasmid and was obtained by amplification and recovery using primers DS2700-F and DS2700-R.
[0372] The underlined sequence is replaced with GAGGG, which is the GAGGG substrate.
[0373]
[0374] Table 3 Sequences of substrates and primers
[0375]
[0376]
[0377] In vitro enzyme cleavage experiment In a 20 μL reaction system (25 mM Tris 8.0, 50 mM NaCl, 2 mM DTT, 5 mM MgCl2), 5 nM of each RAD protein-nucleic acid complex was added, and different substrates were added for different reactions (200 ng PAM substrate plasmid library or 200 ng 2700 bp double-stranded DNA substrate or 20 ng 5' FAM-labeled single-stranded DNA substrate). The reaction time varied from 30 seconds to 2 hours, and the reaction temperature was 37°C. The reaction was terminated by adding an equal volume of 2x loading buffer (8 M urea, 0.05% bromophenol blue, 0.05% xylene cyanol, 1x TBE).
[0378] TAM sequence high-throughput sequencing After the RAD system protein-nucleic acid complex and the PAM substrate plasmid library were subjected to the in vitro cleavage experiment as described above, the reaction product was recovered, and the product sequence was amplified using TA cloning and adapters (see Table 3, Adapter-F / R) and primers TAMSEQ-F / R. The amplification time was 30 seconds, the primer annealing temperature was 58°C, and 20 amplification cycles were performed, and the product was sent to a sequencing company for high-throughput sequencing analysis. Figure 19 The data obtained were counted statistically to obtain the motif sequence (TAM) of the 5' end of the RAD system Figure 4 and Table 4). The results showed that the TAM of different RAD systems numbered 1-57 showed a preference for AGGAG, and different RAD proteins also had different preferences such as GAGGG. The RNA proteins numbered 58-99 had the TAM preferences shown in Table 4.
[0379] Table 4 Motif sequence (TAM) of the 5' end of the RAD system
[0380]
[0381]
[0382] In vitro cleavage activity of the RAD system
[0383] Double-stranded DNA substrate construction: After obtaining the sequence of AGGAG, the inventors constructed a double-stranded DNA substrate that can be specifically recognized by the RAD protein (the sequence is shown in the 2700 bp double-stranded DNA substrate above). Specifically, the 2700 bp double-stranded DNA substrate sequence used in the in vitro cleavage experiment was obtained by amplification and recovery using primers DS2700-F and DS2700-R based on the modification of the PUC19 plasmid.
[0384] Determination of the position where cleavage occurs : Using the FbRAD1 (RHS75625.1, SEQ ID No: 6) system derived from Firmicutes bacterium (AM43-11BH) as the experimental object, it was found that only when the TAM sequence and the targeted sequence coexist, the RAD system (containing the RAD protein SEQ ID No: 6, its related RNA sequence SEQ ID No: 63, and the targeted sequence) can specifically cleave the double-stranded DNA substrate (containing AGGAG and the targeted sequence) into two predetermined substrates (as shown in Figure 5 , the left panel is the analysis of the TAM sequence, and the right panel is the picture of the cleavage result). It is illustrated that the RAD system has the same basic mechanism as the conventional gene editing tool and can exert specific cleavage activity of double-stranded DNA in vitro.
[0385] After the in vitro cleavage experiment of the RAD system protein nucleic acid complex and the 2700 bp double-stranded DNA substrate according to the above Example 4, the reaction product was obtained. The PAM substrate plasmid library and the 2700 bp double-stranded DNA substrate were separated by 1% agarose gel electrophoresis, and the 5’FAM-labeled single-stranded DNA substrate was separated by 15% denaturing Urea-PAGE gel. After recovering the product, it was sent to a sequencing company for Sanger sequencing and high-throughput sequencing, and the sequencing map and the number of different cleavage positions were obtained, respectively.
[0386] Through high-throughput and Sanger sequencing, it was found that the RAD protein cleaved the double-stranded substrate DNA to produce a sticky end cut, and the cleavage mainly occurred at the 22nd base of the targeted strand and the 16th base of the non-targeted strand, Figure 10 (the experimental process is shown in the in vitro cleavage experiment). A small amount of cleavage also occurred at other positions. This experiment again proved that the RAD protein has specific double-stranded DNA endonuclease activity.
[0387] In vitro single-stranded enzyme cleavage experiment : According to the above in vitro cleavage experiment, the 5’ end of the FAM-labeled single-stranded DNA was used to perform the in vitro cleavage experiment. The results are shown in Figure 11As shown, the RAD system can not only specifically cleave double-stranded DNA, but also single-stranded DNA, and the cleavage of single-stranded DNA does not require TAM sequence. The activity of cleaving single-stranded DNA matching the RNA targeting sequence is significantly higher than that of non-matching. Single-stranded DNA can also be cleaved without matching the targeting sequence. After adding a 20-base single-stranded DNA matching the RNA end (20 activator in Table 3) to activate, the cleavage activity of the RAD protein on non-specific single-stranded DNA is significantly increased, indicating that it has trans-cleavage activity on single-stranded DNA. This single-stranded DNA cleavage activity can be developed as a single-stranded DNA virus identification tool. By constructing a sequence of a target virus on the RNA of the RAD system, a single-stranded DNA substrate is used, which generates a chemical or fluorescent signal after cleavage. When the RAD system recognizes the target virus, the trans-cleavage activity is activated, the substrate is cleaved to generate a signal, thereby identifying the target virus.
[0388] Determination of cleavage conditions
[0389] By changing different conditions of in vitro enzyme cleavage experiment, in vitro enzyme cleavage experiment was carried out, and the experimental process was the same as above. The suitable conditions of temperature, type of divalent cation, concentration of divalent cation, length of targeting sequence and salt concentration were obtained.
[0390] The results are shown in Table 2. Figure 12B The results show that for the RAD system containing FbRADl, the suitable conditions for in vitro enzyme cleavage are 5 to 10 mM MgCl2, 25 to 100 mM NaCl, 42°C, and the highest cleavage activity is obtained when the length of the targeting sequence is 20 bases.
[0391] Effect of motif and mismatch sequence
[0392] According to the cleavage substrate constructed above and the in vitro enzyme cleavage experiment, 5 nM FbRADl protein nucleic acid complex, 200 ng 2700 bp double-stranded DNA substrate were added in 20 μL reaction system (25 mM Tris 8.0, 50 mM NaCl, 2 mM DTT, 5 mM MgCl2), and the reaction was carried out at 37°C for 30 minutes. The DNA cleavage under different TAM sequences was verified, and it was found that for FbRADl protein of AGGAG sequence type, AGGAG is the most suitable TAM, and changing any base will greatly reduce the cleavage activity, and even lose the cleavage activity. Figure 12A
[0393] According to the above constructed enzyme cutting substrate and in vitro enzyme cutting experiment, 2700 bp mismatch double-stranded DNA substrate is used. The mismatch cutting experiment proves that the base mismatch at positions 14 to 20 at the 3' end of the target sequence can reduce the cutting activity, but the influence is not big. But the mismatch at positions 1 to 13 greatly reduces the cutting activity, and the experimental results show that the RAD system has low tolerance to the base mismatch at positions 1 to 13, and the RAD system has stronger specificity, because the conventional CRISPR / Cas9 and Cas12a can still produce cutting at positions 6 to 13, and the specificity is poorer than that of the RAD system. Figure 12B
[0394] Example 4: Domain analysis of RAD complex
[0395] Through domain analysis (first collect Cas proteins with RuvC that have solved structure, respectively take out three RuvC domain sequences, perform three times of iterative remote homology search in NCBI database, evalue: 1e-4, then perform multiple sequence alignment on the sequence of each domain and its homologous sequence, construct hmm based on the multiple sequence alignment result, use common tool hmmscan, set threshold value evalue = 1e-10, HMMER web server: interactive sequence similarity searching), it is found that the RAD protein has a RuvC domain, and the domain has three active centers, which are aspartic acid at amino acid No. 194, glutamic acid at No. 290 and aspartic acid at No. 371 in the FbRAD1 system. These three amino acids are highly conserved in all RAD proteins.
[0396] The wild type sequence is FbRAD1. The mutants D194A, E290A and D371A in the three active centers of the amino acid inactivation mutation to alanine in this embodiment are prepared according to the instructions of the Novozyme amino acid point mutation kit. The expression and purification of the RAD (IS607 TnpB) protein nucleic acid complex are shown in Example 1. The wild type, mutant D194A, E290A and D371A of the RAD (IS607 TnpB) protein nucleic acid complex are subjected to double-stranded DNA cutting determination, as described in Example 3.
[0397] Verification of the function of the active center : As shown in Figure 6 , after the amino acid inactivation mutation to alanine of the three active centers, the RAD protein loses the double-stranded DNA cutting activity, which shows that the RuvC domain of the RAD protein itself plays the DNA endonuclease activity. Figure 6
[0398] Structural analysis The RNA of RAGATH-18 forms a longest stem-loop structure in the whole complex, as found by the RNA secondary structure prediction website RNA folder web server. Figure 7
[0399] As shown in Figure 8 , by performing RNA multiple sequence alignment analysis on the MAFFT version 7 website with FbRAD1 (RHS75625.1) as the alignment reference for a plurality of RAD systems with DNase activity, the inventors found that the RAGATH18 RNA has an average length of 75 bp, contains a predicted long hairpin with an internal loop and a pseudoknot, and the RAGATH18 flanking sequences are also conserved, with the left flanking sequence having a length of about 50 bp and containing a predicted stem loop, and the right flanking sequence having a length of about 60 bp and having a predicted stem loop Figure 8 The high conservation of the RNA sequence of RAGATH-18 in the RAD system indicates that the RNA of RAGATH-18 is essential in the RAD system, and also indicates that the RAD system with RAGATH-18 related RNA has DNase activity.
[0400] As shown in Figure 9 , by performing protein amino acid multiple sequence alignment analysis on the MAFFT version 7 website with FbRAD1 (RHS75625.1) as the alignment reference for 20 RAD systems with DNase activity, we found that the following amino acids in the RAD system protein are highly conserved in different strains of different genera: Q18, G27, R30, N34, F50, W73, A92, F97, P104, K107, Y117, E145, P150, S173, G192, D194, G196, K198, A201, S204, N211, I212, N213, R227, Q229, S233, R234, E237, G242, N249, K252, N266, V280, K283, P284, E290, L292, N293, G296, M297, M298, K299, L303, S304, K305, Q310, K322, P338, S339, S340, C343, C346, L353, L355, R358, C362, C364, D369, R370, D371, A374, N377, and L378.
[0401] In combination with the above sequence alignment analysis results, since the RAD system is widely distributed in nature, the RAD protein with these conserved amino acids also has DNase activity.
[0402] Example 5: Prokaryotic in vivo cleavage activity
[0403] Addition of a targeting sequence The 3' end of the RNA sequence (SEQ ID NO: 63) bound by the RAD protein FbRAD1 was directly added with a 20-base sequence (CTGATGGTCCATGTCTGTTA, SEQ ID NO: 116) as a targeting sequence, which is complementary to the sequence of the PAM plasmid library and can target the PAM plasmid library. The method of addition is that, using primers N20-TAR-F and N20-TAR-R (see Table 4 for primer sequences), the PGEX-RAD 1-57 plasmid (e.g., PGEX-6P1-FbRAD1-N20-HDV) synthesized in Example 2 as a template, the expression plasmid was PCR amplified, the PCR amplification time was 6 minutes, the primer annealing temperature was 58°C, the amplification cycle number was 30, and the PCR product was digested with Dpn1 enzyme.
[0404] Obtaining of an E. coli in vivo expression vector The digested product was homologously recombined with the universal vector PET-28A linear vector, the recombination product was transformed into E. coli DH5a competent cells using the recombination enzyme system of Novagen, and the next day single colonies were picked for sequencing identification to obtain the RAD system expression plasmid pET-28a RAD 1-57 with a targeting sequence. The typical plasmid sequence pET-28a-FbRAD1-N20-HDV is as follows. This typical plasmid is a plasmid of the FbRAD1 system, which is an E. coli in vivo editing experimental vector. The RAD system is inserted into the Nco1 and Xho1 sites. The FbRAD1 protein-6HIS-RNA-N20-HDV structure is as follows, where bold underlined is FbRAD1 protein, underlined is 6HIS, bold is RNA, italic is N20, and italic bold is HDV:
[0405]
[0406]
[0407] E. coli in vivo DNA interference verification: 57 kinds of RAD systems in Table 2 were cloned into the PET28a(+) expression vector Nco1 and Xho1 enzyme digestion sites by homologous recombination, and a spectinomycin resistance plasmid PCS101ORI (pTargetF-aggag-n20-psc101ori) containing a TAM sequence and a targeting sequence (CTGATGGTCCATGTCTGTTA) was constructed. The specific sequence of the plasmid is as follows.
[0408] pTargetF-aggag-n20-psclOl ori (Target plasmid for spectinomycin resistance in vivo editing experiment in E. coli. Underlined is the motif sequence, bold is the targeted sequence)
[0409]
[0410]
[0411] Each of the two plasmids, spectinomycin resistance plasmid pTargetF-aggag-n20-psclOl ori and RAD system expression plasmid pET-28a RAD 1-57, was co-transformed into E. coli C43(DE3) with the same method as in Example 2, 100 ng of each of the two plasmids was added into the competent cells at the same time, and the cells were cultured on a double-antibiotic LB plate containing 100 μg / mL spectinomycin and kanamycin at 30°C overnight. A single colony on the double-antibiotic plate was picked and cultured in 2 mL LB medium containing 0.6 mM IPTG, spectinomycin and kanamycin at 30°C, 220 rpm, and 12 hours overnight. The next day, the bacteria were diluted ten-fold by ten-fold dilution, 5 μL of the bacteria was dropped on a double-antibiotic plate, and the plate was cultured at 30°C for 24 hours. The growth was observed and photographed for counting. Once the DNA-specific cleavage occurred, obvious growth difference could be observed on the plate with kanamycin and spectinomycin selection, and the cells with cleavage would die. -6 , 14, 15 show that the RAD system successfully performed specific DNA cleavage in E. coli. Figure 20 The RAD protein cleaves DNA in vivo in E. coli and the screening process is shown.
[0412] Results Figure 13 , 14, 15 show that the RAD system successfully performed specific DNA cleavage in E. coli.
[0413] The interfering DNA sequence used is the spectinomycin resistance sequence, which is an exogenous sequence (PUC19-TARGET1-AGGAG, PUC19-TARGET1-GAGGG, see Table 4). Under the conditions of TAM sequences AGGAG and GAGGG, 12 kinds of RAD systems were found to successfully cleave, which are from 7 different genera, respectively: (6) FbRAD1 / ISFba1 (RHS75625.1), (10) CbRAD1 / ISCba1 (RHR46218.1), (11) RiRAD / ISRin (RGR69673.1), (13) ErRAD / ISEre1 (RHB06265.1), (15) WcRAD (RGR87203.1), (18) BsRAD / ISBsp4 (RHV03091.1), (20) CbRAD2 / ISCisp2 (RHR42937.1), (21) CbRAD3 / ISCisp3 (RGD88909.1), (24) CcRAD / ISCco (RGT87174.1), (30) FbRAD2 / ISFba2 (RGH01875.1), (42) FbRAD3 / ISFba3 (RHU26542.1), (43) FbRAD4 / ISFba4 (RGH37793.1).
[0414] For DNA cleavage of the endogenous genome of E. coli, the targeting sequence for the E. coli genome was replaced, and the interfering DNA sequence used is the E. coli genome sequence, which is an endogenous sequence (ECOLI-TARGET-AGGAG and ECOLI-TARGET-GAGGG, see Table 4; replace the motif and targeting sequence in the above pTargetF-aggag-n20-psc101ori with the sequences shown in Table 4), and the ten-fold dilution titration screening was observed using kanamycin resistance plates Figure 14 , 15). Figure 14 and 15 are the DNA interference results of different RAD systems in E. coli under TAM sequences AGGAG and GAGGG, respectively, and the titration to plaque shows that the growth of the targeting group (T group) (i.e. FbRAD1, CbRAD1, RiRAD, ErRAD and BsRAD) is obviously inhibited compared with the non-targeting control group (NT group), and the difference in performance exceeds 10 4 .
[0415] Verification of enzyme cleavage activity after mutation of the active center : After the amino acid inactivation mutation of aspartic acid at position 194, glutamic acid at position 290, and aspartic acid at position 371 in the three active centers of FbRAD1 to alanine, DNA interference no longer occursFigure 13 ).
[0416] Table 5: Prokaryotic experimental target sequence
[0417]
[0418] Example 6: In vivo gene editing in human
[0419] Each protein in each RAD system in Table 2 is fused with a nuclear localization signal and a green fluorescent protein, and the expression is regulated using a CMV promoter. At the same time, the expression of the RNA in the RAD system is regulated using a U6 promoter, and the end of the RNA contains a 20-base sequence targeting the genome.
[0420] A typical plasmid sequence pcDNA3.1-u6-FbRAD1-cmv-3flag-sv40-FbRAD1-NLS-2a-EGFP for in vivo editing plasmid in human 293F cells is shown below, where bold is the U6 promoter sequence, italic bold is the RNA sequence related to the RAD protein FbRAD1, bold underline is the CMV sequence, italic is the 3flag sequence, bold italic underline is the SV40 sequence, bold italic double underline is the RAD protein FbRAD1, double underline is the NLS sequence, italic underline is the EGFP sequence, and bold double underline is the 2a sequence.
[0421]
[0422]
[0423] The coding nucleic acid sequence of the FbRAD1 protein and the related RNA in the typical sequence pcDNA3.1-u6-FbRAD1-cmv-3flag-sv40-FbRAD1-NLS-2a-EGFP is replaced by the coding nucleic acid of other effector proteins and corresponding RNAs, and other corresponding plasmid constructs are obtained.
[0424] Cell transfection Human 293F cells are selected as the transfection object, and in 10 mL of cell transfection system, the culture medium is SMM 293TII produced by Beijing Yiqiao Shenzhou Company, and the cell density is 7 x 10 5 5 minutes, 0.5 mL of OPTI-MEM is added with 50 μL of PEI for 5 minutes, and after incubation at room temperature for 25 minutes, 10 mL of cells are added, 5% CO2, 120 rpm in a cell culture shaker for 72 hours.
[0425] Flow cytometry sorting and high-throughput sequencingCells containing green fluorescent signal were sorted by BD FACSAria Fusion flow cytometer to remove dead and failed transfection cells. The whole genome of the cells was extracted using Kangweishidongyongzongyongzhunji Genomic Extraction Kit, and the product was directly sent to the company for high-throughput sequencing. The results were used to calculate the gene editing efficiency.
[0426] Through the construction of a eukaryotic expression vector (pcdna3.1-u6-FbRAD1-cmv-3flag-sv40-FbRAD1-NLS-2a-EGFP) and cell transfection experiments and high-throughput sequencing analysis, the inventors successfully verified in human cells that the FbRAD1 system can achieve specific DNA editing. Editing occurred at the EMX1, VEGFA, DNMT1, PITX1 gene sites (see Table 5 for sequence information), and the editing efficiency at the EMX1 site was as high as 32.8%. This efficiency is close to the widely concerned Cas9 and Cas12a, which were initially revealed to have in vivo editing activity, and provides a basis for further optimizing the RAD system into a high-efficiency gene editing tool. Figure 16 ). Figure 21 RAD protein editing DNA in 293F cells and detection process are shown.
[0427] Table 6: EMX1, VEGFA, DNMT1, PITX1 gene site information.
[0428]
[0429]
[0430] By the above experimental method, whether different RAD systems can be expressed in 293F cells to verify their in vivo editing function of mammalian cells, it is found that 9 RAD systems can be edited in vivo, which are from 8 different genera, which are: (6) FbRAD1 / ISFba1 (RHS75625.1), (11) RiRAD / ISRin (RGR69673.1), (13) ErRAD / ISEre1 (RHB06265.1), (21) CbRAD3 / ISClsp3 (RGD88909.1), (43) FbRAD4 (RGH37793.1), (51) DfRAD / ISDfo4 (GUT_GENOME239804_2_85784_86938_+), (52) ArRAD / ISAre (GUT_GENOME047820_36_478_1632_-), (54) HuRAD / ISHun (GUT_GENOME022845_45_3955_5109_+), (55) AfRAD / ISAfa (GUT_GENOME112567_131_1855_3009_-) Figure 17
[0431] So far, there are 16 RAD systems from 11 different genera of 16 strains that have intracellular DNA cleavage activity. As representatives of RAD systems, they confirm that the DNA cleavage activity in RAD systems is widespread, providing a theoretical basis for further research on RAD systems, and more RAD systems can be developed and optimized as gene editing tools.
[0432] RAD protein has a smaller protein size (about 387 amino acids), compared with the currently widely studied gene editing tool effector protein, such as Cas9 and Cas12a effector (generally more than 1000 amino acids), RAD protein size is the smallest (384-387 amino acids), has the value of further development of new gene editing tools. Smaller size is more advantageous than traditional large proteins for delivery vector construction. We identify the basic characteristics of the RAD system, establish a RAD system gene editing platform, and provide a basis for further optimizing it into a more precise and efficient gene editing tool.
[0433] The application described and claimed herein is not to be limited in scope by the specific aspects disclosed herein, since these aspects are intended as illustrations of several aspects of the application. Any equivalent aspects are intended to be within the scope of the application. Indeed, various modifications of the application in addition to those shown and described herein will become apparent to those skilled in the art from the foregoing description. Such modifications are also intended to fall within the scope of the appended claims. In the case of conflict between the specification and the appended claims, the latter control.
Claims
1. An endonuclease complex comprising an RNA-guided endonuclease and a guide RNA, wherein the RNA-guided endonuclease is an IS607 TnpB protein, and the guide RNA comprises a RAGATH-18-related RNA as an RNA-guided endonuclease-associated RNA; The amino acid sequence of the RNA-guided endonuclease is SEQ ID NO: N, and the RNA-guided endonuclease-associated RNA is the sequence of SEQ ID NO: N + 57, where N is an integer of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, or 57; or The amino acid sequence of the RNA-guided endonuclease is SEQ ID NO: M, and the RNA-guided endonuclease-associated RNA is the sequence of SEQ ID NO: M+42, where M is an integer of 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, or 191.
2. The endonuclease complex according to claim 1, wherein the guide RNA further comprises a targeting RNA.
3. The endonuclease complex of claim 2, wherein the targeting RNA is at the 3' end of the RNA-guided endonuclease-associated RNA and is identical or complementary to the targeting sequence in the target gene.
4. The endonuclease complex according to claim 3, wherein the targeting sequence has a length of 12-40 bp.
5. The endonuclease complex according to any one of claims 1 to 4, wherein the endonuclease complex comprises an RNA-guided endonuclease of SEQ ID NO: 6 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 63; an RNA-guided endonuclease of SEQ ID NO: 11 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 68; an RNA-guided endonuclease of SEQ ID NO: 13 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 70; an RNA-guided endonuclease of SEQ ID NO: 21 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 78; an RNA-guided endonuclease of SEQ ID NO: 43 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 100; an RNA-guided endonuclease of SEQ ID NO: 51 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 108; an RNA-guided endonuclease of SEQ ID NO: 52 and an RNA-guided endonuclease-associated RNA of SEQ ID NO: 109; The RNA-guided endonuclease of SEQ ID NO: 54 and the RNA-guided endonuclease-related RNA of SEQ ID NO: 111; or the RNA-guided endonuclease of SEQ ID NO: 55 and the RNA-guided endonuclease-related RNA of SEQ ID NO:
112.
6. The endonuclease complex of claim 1 , wherein the RNA-guided endonuclease contains a RuvC domain, and the RNA-guided endonuclease comprises an amino acid selected from the group consisting of Q18, G27, R30, N34, F50, W73, A92, F97, P104, K107, Y117, E145, P150, S173, G192, D194, G196, K198, A201, S204, N211, I212, N213, R227, Q229, S233, R234, E237, G242, N24 9. K252, N266, V280, K283, P284, E290, L292, N293, G296, M297, M298, K299, L303, S304, K305, Q310, K322, P338, S339, S340, C343, C346, L353, L355, R358, C362, C364, D369, R370, D371, A374, N377, and L378. The endonuclease complex according to claim 6 , wherein the target gene is a single-stranded or double-stranded DNA.
8. The endonuclease complex according to claim 7, wherein the target gene comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN or GGGGN.
9. A polynucleotide comprising a polynucleotide encoding the RNA-guided endonuclease according to any one of claims 1 to 8 and a polynucleotide encoding an RNA molecule, wherein the RNA molecule comprises an RNA related to the RNA-guided endonuclease according to any one of claims 1 to 8 that matches the endonuclease.
10. A vector comprising the polynucleotide molecule according to claim 9. The vector according to claim 10 , further comprising an expression control sequence.
12. The vector according to claim 10 or 11, wherein the RNA-guided endonuclease and the guide RNA are expressed under the control of the same or different regulatory elements.
13. The vector according to claim 12, wherein the regulatory element is a promoter.
14. A viral particle comprising the endonuclease complex according to any one of claims 1 to 8, the polynucleotide according to claim 9 or the vector according to any one of claims 10 to 13. The viral particle according to claim 14 , which is an adeno-associated viral particle.
16. A method for specifically cleaving DNA, comprising the step of contacting the endonuclease complex according to any one of claims 1 to 8 with DNA, wherein the method is a method for non-therapeutic or non-diagnostic purposes.
17. The method of claim 16, wherein the DNA is eukaryotic, prokaryotic, or viral DNA.
18. The method of claim 16, wherein the DNA is single-stranded or double-stranded DNA.
19. The method according to any one of claims 16 to 18, comprising the step of introducing the polynucleotide according to claim 9, the vector according to any one of claims 10 to 13, or the viral particle according to claim 14 or 15 into a cell.
20. The method of claim 19, wherein the DNA targeting sequence comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN, or GGGGN, and the 3' end of the motif is a DNA identical or complementary to the targeting RNA.
21. The method according to any one of claims 16 to 20, wherein the method is performed under conditions selected from the group consisting of: 0.5 to 10 mM selected from Mg 2+ 、Mn 2+ and Ca 2+ divalent metal ions, 25 to 250 mM Na + , the reaction temperature is 27-57°C, and the reaction time is 3-60 minutes.
22. A method for gene editing, comprising the step of contacting the endonuclease complex according to any one of claims 1 to 8 with the genomic DNA of a cell to perform gene editing, wherein the method is a method for non-therapeutic or non-diagnostic purposes.
23. The method of claim 22, wherein gene editing produces insertions and / or deletions in genomic DNA.
24. The method of claim 22, wherein the gene editing induces a desired trait in a mammalian cell.
25. A kit comprising one or more of the following: an endonuclease complex according to any one of claims 1 to 8, a polynucleotide according to claim 9, a vector according to any one of claims 10 to 13, or a viral particle according to claim 14 or 15.
26. A method for inactivating or identifying viruses in vitro, comprising contacting the viruses with the endonuclease complex according to any one of claims 1 to 8.
27. Use of the endonuclease complex according to any one of claims 1 to 8 in the preparation of a medicament for treating viral infection.
28. A viral vector comprising a polynucleotide according to claim 9 or a vector according to any one of claims 10 to 13.
29. The viral vector according to claim 28, which is an adeno-associated viral vector.
30. Use of the viral particle according to claim 14 or 15 or the viral vector according to claim 28 or 29 in the preparation of a medicament for gene therapy for treating a disease.
31. A cell comprising one or more of the following: an endonuclease complex according to any one of claims 1 to 8, a polynucleotide according to claim 9, a vector according to any one of claims 10 to 13, or a viral particle according to claim 14 or 15.
32. The cell of claim 31 , which is a eukaryotic cell or a prokaryotic cell.
33. The cell of claim 32, wherein the eukaryotic cell is a fungal cell, a plant cell, or an animal cell.
34. Use of the endonuclease complex according to any one of claims 1 to 8 in preparing a kit for a method for specifically cleaving DNA.
35. The use according to claim 34, wherein the DNA is eukaryotic, prokaryotic or viral DNA.
36. The use according to claim 34, wherein the DNA is single-stranded or double-stranded DNA.
37. Use according to any one of claims 34 to 36, wherein the method comprises the step of introducing a polynucleotide according to claim 9, a vector according to any one of claims 10 to 13 or a viral particle according to claim 14 or 15 into a cell.
38. The use according to claim 37, wherein the DNA targeting sequence comprises a motif selected from the group consisting of NGGAG, NGGNN, AGGAG, GAGGG, GGGGG, NNGGG, NNAGG, AGGNN, NTAAA, NGAGG, NNGGN or GGGGN, and the 3' end of the motif is a DNA identical or complementary to the targeting RNA.
39. The use according to any one of claims 34-38, wherein the method is performed under conditions selected from the group consisting of: 0.5 to 10 mM selected from Mg 2+ 、Mn 2+ and Ca 2+ divalent metal ions, 25 to 250 mM Na + , the reaction temperature is 27-57°C, and the reaction time is 3-60 minutes.
40. Use of the endonuclease complex according to any one of claims 1-8 in preparing a kit for a method for gene editing, which comprises the step of contacting the endonuclease complex according to any one of claims 1-8 with the genomic DNA of a cell to perform gene editing.
41. The use according to claim 40, wherein gene editing produces insertions and / or deletions in genomic DNA.
42. The method of claim 40, wherein gene editing is performed in mammalian cells to treat or prevent disease, or to induce a desired trait.
43. Use of the endonuclease complex or endonuclease according to any one of claims 1 to 8 in the preparation of a kit for inactivating or identifying viruses.