Cas protein, fusion protein, nucleic acid molecule, ribonucleoprotein complex, recombinant vector, transgenic cell, and use thereof
By developing the EDG-01Cas protein and its crRNA components, the existing CRISPR/Cas system has been solved, and the existing CRISPR/Cas system has limited target recognition and great off-target potential in eukaryotes is achieved, and efficient and specific gene editing is achieved, especially in plant cells.
Patent Information
- Application Number
- PCT/CN2024/138684
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-12
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-03
AI Technical Summary
The existing CRISPR/Cas gene editing system has problems in eukaryotes with limited target recognition, limited PAM and high off-target potential, resulting in insufficient editing efficiency and specificity.
A new Cas protein EDG-01 and its crRNA components are developed to identify PAMs as NTTV, TTTT and TTCN, with a broader target recognition capability and reduce off-target potential. SgRNA does not rely on tracrRNA and binds nuclear localization signal peptides for eukaryotic gene editing.
It achieves efficient gene editing in eukaryotes, especially in plant cells, and the editing efficiency reaches more than 70%, reducing the risk of off-targeting and improving the diversity and editing specificity of target recognition.
Smart Images

Figure CN2024138684_03072025_PF_FP_ABST
Abstract
Description
Cas protein, combined protein, nucleic acid molecule, ribonucleoprotein complex, recombinant vector, transgenic cell and its application
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of Chinese patent applications 202311861294.2 and 202411425800.8 filed on December 29, 2023 and October 12, 2024, the contents of which are incorporated herein by reference. Technical Field
[0003] The present invention relates to the field of nucleic acid editing, and specifically to a Cas protein, a combination protein, a gene encoding the Cas protein or the combination protein, a nucleic acid molecule, a ribonucleoprotein complex, a recombinant vector, a transgenic cell and its application, and a gene editing method. Background Art
[0004] CRISPR stands for Clustered Regularly Interspaced Short Palindromic Repeats, which resembles many identical short segments evenly inserted into a randomly arranged stretch of DNA. These sequences are called repeats, and the randomly arranged DNA that separates them is called spacer. Cas stands for CRISPR associated, abbreviated as the CRISPR / Cas system.
[0005] The CRISPR / Cas system, a sequence of CRISPR repeats and Cas proteins found in bacterial and archaeal genomes, primarily participates in bacterial immune responses, protecting them from foreign nucleic acid invasion. Cas proteins, guided by a guide RNA (sgRNA), can precisely target specific target regions within the genome. Therefore, they have been engineered into eukaryotic gene-editing tools for applications in gene function research, genetic improvement, and medical research in plants and animals. When the CRISPR / Cas system is functioning, the Cas9 protein and sgRNA form a protein-RNA complex. The Cas protein recognizes the protospacer adjacent motif (PAM) sequence, such as 5'-NGG-3' for Cas9. Assisted by the sgRNA, the Cas protein melts the DNA double strands complementary to the sgRNA near the PAM, cleaving the DNA double-strand break (DSB). Subsequently, endogenous NHEJ or HDR repair pathways introduce base insertions or deletions (indels) at the target site, leading to frameshift mutations or functional changes, ultimately resulting in knockout.
[0006] Based on the composition of the Cas protein, the CRISPR system can be divided into Class I and Class II, and according to its mode of action, it is mainly divided into six types (Type I to VI). Among these two types of CRISPR systems, Class I (Type I, III, IV) requires multiple protein subunits to form a complex to perform the nuclease function, while Class II (Type II, V, VI) only needs a single Cas protein and the corresponding sgRNA to perform the targeted nuclease function. Therefore, Class II Cas proteins are developed for gene editing. At present, the systems that can effectively act on chromosome engineering in eukaryotic cells are still mainly SpCas9 and AsCas12a, which belong to Type II and Type V respectively. When in use, SpCas9 is guided to the target site by an sgRNA composed of a segment of crRNA and tracrRNA for specific cutting, producing a blunt-end incision; while AsCas12a only requires an sgRNA composed of a segment of crRNA to cut the target site, and the resulting incision is usually a sticky end. Since DSBs with sticky ends usually need to be modified to flat ends with the help of exonucleases before ligation during the repair process, the resulting mutations will produce more nucleotide deletions compared to SpCas9.
[0007] Furthermore, the short sgRNA length of the Cas12 family makes it easier for the Cas12 family to enter cells with the aid of related delivery systems. Currently, the CRISPR / Cas editing system has been widely used in fields such as basic biological research, genetic modification, and medical research. Currently known CRISPR / Cas-based gene editing systems each have their own advantages and disadvantages due to factors such as editing efficiency, target PAM restriction, and specificity. Therefore, to meet the needs of various biological research and applications, the development of new and diverse CRISPR / Cas gene editing systems is still needed.
[0008] Gene editing technology can manipulate or modify genomic DNA in a targeted manner, thereby changing the characteristics of an organism. Currently, this technology mainly includes: (1) zinc finger nuclease (ZFN) technology, which uses a zinc finger domain to specifically recognize bases and then uses Fok I nuclease to cut the DNA double strand; (2) transcription activator-like effector nuclease TALEN technology, which uses transcription activator-like effectors to recognize bases and then cut double-stranded DNA through Fok I; (3) gene technology based on the CRISPR / Cas system, which uses sgRNA to guide Cas proteins to cut at designated target sites and introduce mutations with the help of the non-homologous end joining repair mechanism in the cell. Since CRISPR / Cas gene editing technology relies on sgRNA to recognize the target sequence through base complementary pairing, it has high specificity, and the system is simple to construct and has high editing efficiency. Therefore, this technology is currently the most widely used gene editing technology.
[0009] The CRISPR / Cas9 system, derived from bacteria, is the first technology developed for gene editing in eukaryotes. The invention stems from the bacterial immune system's ability to resist invasion by exogenous nucleic acids. When nucleic acids from other microorganisms, such as bacteriophages, are first injected into bacteria, the Class I Cas enzymes expressed by the CRISPR locus, such as Cas1, Cas2, and Cas4, cut the exogenous nucleic acid and then insert a DNA sequence into the CRISPR repeat cluster. When the bacteria are invaded again by the same exogenous nucleic acid, the CRISPR cluster sequence is transcribed and processed to form a short crRNA, which can complement and pair with a partial sequence of tracrRNA transcribed near the CRISPR cluster, forming a crRNA-tracrRNA complex called sgRNA. This guides the Cas9 protein to the 5' end of a specific PAM sequence (e.g., NGG for SpCas9, where N is one of A, T, C, or G), recognizing the target site through complementary pairing of a 20nt base sequence. The Cas9 protein, which contains HNH and RuvC nuclease domains, can cleave double-stranded DNA, creating double-strand breaks and subsequently introducing mutations through intracellular repair mechanisms. Using genetic engineering methods, by expressing bacterial Cas9 protein and artificially synthesized or transcribed sgRNA sequences, the Cas9 and sgRNA protein complexes are simultaneously introduced into cells, enabling targeted editing of specific genomic targets.
[0010] Since the CRISPR / Cas9 system can be used for eukaryotic gene editing, various types of CRISPR / Cas systems have been identified. The most widely used and most diverse Cas protein belongs to the Cas12 family. The Cas12 family protein components that were first discovered and developed as gene editing technology are Cas12a (also called Cpf1), including FnCpf1, LbCpf1, and AsCpf1. The recognition PAM of FnCpf1 is TTN, while that of LbCpf1 and AsCpf1 is TTTN. Unlike the aforementioned CRISPR / Cas9 system, the sgRNA of the Cpf1 protein only requires one short crRNA and is independent of tracrRNA. Cpf1 itself has the ability to process crRNA. In addition, the recognition target site of Cpf1 is 24nt and is located at the 3' end of the PAM. When using the CRISPR / Cas system for gene editing in eukaryotic organisms, one or more nuclear localization signals are typically added to the Cas protein to facilitate entry of the Cas-sgRNA protein complex into the cell nucleus. Furthermore, because the sgRNAs used in the Cpf1 system are shorter, they are easily delivered into cells. Summary of the Invention
[0011] The purpose of the present invention is to provide a new gene editing component with good performance and robustness.
[0012] In order to achieve the above object, the first aspect of the present invention provides a Cas protein, wherein the amino acid sequence of the Cas protein is an amino acid sequence selected from at least one of the following:
[0013] (1) the amino acid sequence shown in SEQ ID NO. 1 + SEQ ID NO. 24;
[0014] (2) An amino acid sequence derived from SEQ ID NO. 1 + SEQ ID NO. 24 that is at least 90% identical, preferably at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO. 1 + SEQ ID NO. 24, and whose enzyme activity remains unchanged.
[0015] The second aspect of the present invention provides a combination protein, which comprises the Cas protein described in the first aspect and a fragment a connected to the Cas protein; the fragment a is provided by at least one selected from a His tag and a nuclear localization signal peptide.
[0016] The third aspect of the present invention provides a gene I encoding the Cas protein described in the first aspect.
[0017] The fourth aspect of the present invention provides a gene II encoding the combination protein described in the second aspect.
[0018] The fifth aspect of the present invention provides a nucleic acid molecule comprising the nucleotide sequence shown in SEQ ID NO. 8 and / or SEQ ID NO. 9.
[0019] The sixth aspect of the present invention provides a ribonucleoprotein complex, which comprises a protein component and a nucleic acid component; the protein component is selected from at least one of the Cas protein described in the first aspect and the combination protein described in the second aspect; the nucleic acid component comprises a guide sequence consistent with the target sequence and a nucleic acid molecule; the nucleic acid molecule is the nucleic acid molecule described in the fifth aspect.
[0020] The seventh aspect of the present invention provides a recombinant vector comprising the nucleic acid molecule described in the fifth aspect.
[0021] The eighth aspect of the present invention provides a transgenic cell, which comprises the recombinant vector described in the seventh aspect.
[0022] The ninth aspect of the present invention provides the use of the Cas protein described in the first aspect, the combined protein described in the second aspect, the gene I described in the third aspect, the gene II described in the fourth aspect, the nucleic acid molecule described in the fifth aspect, the ribonucleoprotein complex described in the sixth aspect, or the recombinant vector described in the seventh aspect in gene editing.
[0023] The tenth aspect of the present invention provides a method for gene editing, which delivers the ribonucleoprotein complex described in the sixth aspect and / or the recombinant vector described in the seventh aspect into a cell containing a target gene; the target gene contains a target sequence.
[0024] The technical solution provided by the present invention has at least the following advantages:
[0025] The Cas protein (EDG-01) and its crRNA component provided by the present invention can specifically target and edit genomic DNA in eukaryotic organisms. The PAMs recognized by the Cas protein are mainly NTTV, TTTT, and TTCN, as well as weak TCTV and TGTA cleavage activities, which have a wider range of target recognition, and its four-base PAM feature can reduce its off-target potential. In addition, the gene editing system provided by the present invention does not rely on tracrRNA for sgRNA.
[0026] In addition, the EDG-01 identified in the present invention has low homology with the Cas12a family proteins currently reported for eukaryotic gene editing and is a new gene editing component system. In eukaryotic organisms such as plants, the editing efficiency of the ribonucleoprotein complex provided by the present invention can reach more than 70%. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of the secondary structure formed by crRNA;
[0028] FIG2 is an amino acid sequence alignment of EDG-01 with LbCpf1, Ascpf1, and Fncpf1 proteins;
[0029] FIG3 is an amino acid sequence identity analysis diagram of EDG-01 and LbCpf1, Ascpf1 and Fncpf1 proteins;
[0030] Figure 4 is a predicted three-dimensional structure of the EDG-01 and sgRNA complex with the substrate;
[0031] Figure 5 is a characteristic diagram of the PAM sequence recognized by the EDG-01 and sgRNA complex;
[0032] FIG6 is a schematic diagram of the depletion ability of EDG-01 at 64 PAM sites;
[0033] FIG7 is an electrophoretic pattern of EDG-01 cleaving 32 PAM substrates in vitro;
[0034] FIG8 is a schematic diagram comparing the depletion abilities of EDG-01 and Fncpf1 at 64 PAM sites;
[0035] FIG9 is an electrophoretic pattern of EDG-01 and Fncpf1 cleaving 8 PAM substrates in vitro;
[0036] FIG10 is an electrophoretic pattern of in vitro cleavage of double-stranded DNA by three mutants of EDG-01;
[0037] Figure 11 is a map of the pEGEDG-01Pubi-H and pEGLbCpf1Pubi-H expression vectors;
[0038] Figure 12 is a map of the pEGEDG-01Pubi-H-BEL-T and pEGLbCpf1Pubi-H-BEL-T editing vectors;
[0039] FIG13 is a statistical diagram of the editing efficiency of EDG-01 in plant cells;
[0040] Figure 14 is a schematic diagram of the specific base types deleted by EDG-01 gene editing mutation;
[0041] Figure 15 is a schematic diagram of the ratio of missing sequence lengths after EDG-01 editing.
[0042] Partial sequence information of the present invention is shown in Table 1.
[0043] Table 1 DETAILED DESCRIPTION
[0044] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0045] The terms "guide RNA" or "sgRNA" or "nucleic acid component" are used interchangeably and have the meanings generally understood by those skilled in the art. Generally speaking, any polynucleotide sequence that has sufficient complementarity to a target nucleic acid sequence to hybridize with the target nucleic acid sequence and guide sequence-specific binding of a nucleic acid-targeting complex to the target nucleic acid sequence is included. It can comprise, consist essentially of, or consist of a direct repeat sequence and a guide sequence (also referred to as a spacer in the context of an endogenous CRISPR system).
[0046] crRNA (CRISPR RNA): CRISPR RNA, or crRNA, is an RNA transcript derived from the CRISPR locus. In the Cpf1 (Cas12a) system, crRNA is a short, single-stranded RNA composed primarily of a direct repeat sequence. A direct repeat sequence consists of a stem-loop and a small loop. The stem-loop consists of a left arm (stem left) and a right arm (stem right), which pair to form the main structure of the direct repeat sequence. Direct repeat sequences are highly conserved, with most having the same stem-loop. The repeat sequence is a conserved region that binds to the Cpf1 protein and is responsible for forming a complex with Cpf1.
[0047] The endpoints of the ranges and any values disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoints of each range, the endpoints of each range and individual point values, and the individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered to be specifically disclosed herein.
[0048] In the present invention, EDG-01 or CasEDG-01 both represent the Cas protein described in the present invention.
[0049] In the present invention, "V" in the nucleotide sequence indicates that it is selected from A, C, and G; "N" in the nucleotide sequence indicates that it is selected from A, C, G, and T.
[0050] As mentioned above, the first aspect of the present invention provides a Cas protein, the amino acid sequence of which is an amino acid sequence selected from at least one of the following:
[0051] (1) the amino acid sequence shown in SEQ ID NO. 1 + SEQ ID NO. 24;
[0052] (2) An amino acid sequence derived from SEQ ID NO. 1 + SEQ ID NO. 24 that is at least 90% identical, preferably at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO. 1 + SEQ ID NO. 24, and whose enzyme activity remains unchanged.
[0053] Preferably, the Cas protein comprises a RuvC nuclease domain.
[0054] Preferably, the Cas protein is derived from a Flavobacterium branchiophilum strain.
[0055] As mentioned above, the second aspect of the present invention provides a combination protein, which comprises the Cas protein described in the first aspect and a fragment a connected to the Cas protein; the fragment a is provided by at least one selected from a His tag and a nuclear localization signal peptide.
[0056] Preferably, the His tag and / or the nuclear localization signal peptide are connected to the N-terminus and / or C-terminus of the Cas protein.
[0057] Preferably, the combination protein has an amino acid sequence selected from at least one of the following:
[0058] (1) the amino acid sequence shown in SEQ ID NO. 6 + SEQ ID NO. 28;
[0059] (2) An amino acid sequence derived from SEQ ID NO. 6 + SEQ ID NO. 28 that is at least 90% identical, preferably at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO. 6 + SEQ ID NO. 28, and whose enzyme activity remains unchanged.
[0060] As mentioned above, the third aspect of the present invention provides a gene I encoding the Cas protein described in the first aspect.
[0061] Preferably, the nucleotide sequence of the gene I is at least one nucleotide sequence selected from the group consisting of:
[0062] (1) the nucleotide sequence shown in SEQ ID NO. 7;
[0063] (2) A nucleotide sequence that is at least 90% identical, preferably at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO. 7.
[0064] As mentioned above, the fourth aspect of the present invention provides a gene II encoding the combination protein described in the second aspect.
[0065] Preferably, the nucleotide sequence of the gene II is at least one nucleotide sequence selected from the group consisting of:
[0066] (1) the nucleotide sequence shown in SEQ ID NO. 2;
[0067] (2) A nucleotide sequence that is at least 90% identical, preferably at least 91% identical, at least 92% identical, at least 93% identical, at least 94% identical, at least 95% identical, at least 96% identical, at least 97% identical, at least 98% identical, or at least 99% identical to SEQ ID NO. 2.
[0068] The fifth aspect of the present invention provides a nucleic acid molecule comprising the nucleotide sequence shown in SEQ ID NO.8 and / or SEQ ID NO.9.
[0069] Preferably, the nucleic acid molecule is RNA.
[0070] The sixth aspect of the present invention provides a ribonucleoprotein complex, which comprises a protein component and a nucleic acid component; the protein component is selected from at least one of the Cas protein described in the first aspect and the combination protein described in the second aspect; the nucleic acid component comprises a guide sequence consistent with the target sequence and a nucleic acid molecule; the nucleic acid molecule is the nucleic acid molecule described in the fifth aspect.
[0071] Preferably, the targeting sequence is linked to the 3' end of the nucleic acid molecule.
[0072] Preferably, the guide sequence is 20-28 nt in length.
[0073] Preferably, the target site PAM sequence recognized by the ribonucleoprotein complex is selected from at least one of NTTV, TTCN, TTTT, TCTV, and TGTA; wherein the V is selected from A, C, or G; and the N is selected from A, C, G, or T.
[0074] The PAM sequences described in the present invention are all in the order from 5' to 3'.
[0075] Preferably, the target site PAM sequence recognized by the ribonucleoprotein complex is selected from at least one of NTTV, TTCN, TTTT, and TCTV; wherein the V is selected from A, C, or G; and the N is selected from A, C, G, or T.
[0076] Preferably, the PAM sequence of the target site recognized by the ribonucleoprotein complex is TTTN, wherein the N is selected from A, C, G or T.
[0077] More preferably, the target site PAM sequence recognized by the ribonucleoprotein complex is TTTV, wherein the V is selected from A, C or G.
[0078] The seventh aspect of the present invention provides a recombinant vector comprising the nucleic acid molecule described in the fifth aspect.
[0079] According to a preferred embodiment, the recombinant vector of the present invention comprises the nucleic acid molecule described in the fifth aspect, and comprises the gene I described in the third aspect and / or the gene II described in the fourth aspect.
[0080] The eighth aspect of the present invention provides a transgenic cell, which comprises the recombinant vector described in the seventh aspect.
[0081] Preferably, the transgenic cell is a eukaryotic cell.
[0082] More preferably, the transgenic cell is a eukaryotic plant cell.
[0083] As mentioned above, the ninth aspect of the present invention provides the use of the Cas protein described in the first aspect, the combined protein described in the second aspect, the gene I described in the third aspect, the gene II described in the fourth aspect, the nucleic acid molecule described in the fifth aspect, the ribonucleoprotein complex described in the sixth aspect, or the recombinant vector described in the seventh aspect in gene editing.
[0084] Preferably, the gene editing is gene editing in eukaryotes.
[0085] Preferably, the gene editing method is gene deletion.
[0086] The tenth aspect of the present invention provides a method for gene editing, which delivers the ribonucleoprotein complex described in the sixth aspect and / or the recombinant vector described in the seventh aspect into a cell containing a target gene; the target gene contains a target sequence.
[0087] Preferably, the cell is a eukaryotic cell, more preferably a eukaryotic plant cell.
[0088] In the gene editing method described in the present invention, the gene I encoding the Cas protein or the gene II encoding the combined protein can be present in the same vector as the nucleic acid molecule and delivered to the cell containing the target gene for gene editing, or can be present in different vectors and delivered to the cell containing the target gene for gene editing. As long as the ribonucleoprotein complex described in the present invention can exist or can be synthesized and assembled in the system for performing the gene editing, the effect of gene editing can be exerted.
[0089] According to a particularly preferred embodiment of the present invention:
[0090] Using bioinformatics methods, the present invention screened bacterial genomes for sites with CRISPR cluster repeat sequence characteristics and adjacent sequences encoding a nuclease domain. This method identified a site with interspaced CRISPR repeat sequence characteristics from the genome of Flavobacterium branchiophilum. The repeat sequence fragment of this site is SEQ ID NO. 8: 5'-GTTTAAAACCACTTTAAAATTTCTACTATTGTAGAT-3'.
[0091] The repeat sequence fragment is repeated 38 times in the genome, and the specific sequence between each repeat sequence fragment (the specific sequence refers to the specific non-repeating sequence between two repeat sequences) is 28bp in length. An open reading frame is present next to the repeat sequence fragment, and its nucleotide sequence is shown in SEQ ID NO.7.
[0092] The open reading frame encodes a class II type V CRISPR protein containing a RuvC nuclease domain, also known as a Cas protein, named EDG-01 or CasEDG-01 in the present invention, and its amino acid sequence is shown in SEQ ID NO.1+SEQ ID NO.24.
[0093] The CRISPR repeat sequence fragment shown in SEQ ID NO.8 can be processed into a segment of crRNA. The processed crRNA sequence is SEQ ID NO.9: 5'-AATTTCTACTATTGTAGAT-3', and the formed secondary structure is shown in Figure 1.
[0094] Amino acid sequence alignment was performed using the Align tool on the Uniprot official website. The highest similarity between EDG-01 of the present invention and the CRISPR / Cas gene editing component reported in the existing literature (Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. doi: 10.1016 / j.cell.2015.09.038.) is FnCas12a (also called FnCpf1), and the similarity of the entire sequence is 46.18%. The amino acid sequence alignment results are shown in Figures 2 and 3.
[0095] Aphafold3 was used to predict the three-dimensional structure of the CasEDG-01 protein. The amino acid sequence used was SEQ ID NO.1+SEQ ID NO.24, the nucleic acid sequence of the non-targeting chain was (UTS, SEQ ID NO.21): TTTCAAGTGATTGAAGCATACACAAGAA, and the nucleic acid sequence of the targeting chain was (TS, SEQ ID NO.22): TTCTTGTGTATGCTTCAATCACTTGAAA; the nucleic acid sequence of sgRNA-1 (SEQ ID NO.23) was: AAUUUCUACUAUUGUAGAUAAGUGAUUGAAGCAUACACAAGAA. We then used Foldseek (URL: https: / / search.foldseek.com / search) to search for similar protein structures based on this predicted structure. In the PDB database of the Foldseek prediction results, we found that the proteins with the highest structural similarity to CasEDG-01 include FnCas12a and AsCas12a (also called AsCpf1). We then used the Uniprot database to annotate the domains of CasEDG-01 based on the domain information of proteins AsCas12a (uniprot id: U2UMQ6) and FnCas12a (uniprot id: A0Q7Q2). By aligning the domains of CasEDG-01 with those of FnCas12a and AsCas12a, the domain names and corresponding amino acid ranges of CasEDG-01 were finally obtained: wedge domain I (WEDⅠ; 1-33); recognition domain I (RECⅠ; 34-303); recognition domain II (RECⅡ; 304-572); wedge domain II (WEDⅡ; 573-640); PAM interaction domain (PI domain; 641-730); wedge domain III (WEDⅢ; 731-905); RuvC endonuclease domain I (RuvCⅠ; 906-965); (Bridge Helix; 966-982); RuvC endonuclease domain II (RuvCⅡ; 983-1090); nuclease domain (Nuc; 1091-1274); RuvC endonuclease domain III (RuvCⅢ; 1275-1318). Different domains were then assigned different colors, and the final result is shown in Figure 4.
[0096] To further identify the PAM sequence of the target site recognized by the EDG-01 and sgRNA complex, the present invention uses a PAM in vitro cleavage method to identify the PAM characteristics of the component cleaving the target site sequence, as follows:
[0097] ① By designing primers, a 1.3kb sequence containing PAM and target was amplified as a substrate. The target PAM used for PAM identification was set to 6 consecutive degenerate bases (N, N is selected from A, T, C, G), and the specific sequence was: 5'-NNNNNNAAGTGAAGATTGACATACACAGAA-3'.
[0098] ② Mix CasEDG-01 protein and sgRNA at a molar ratio of 1:2, then mix with 1 μg of PAM library substrate for cutting, perform in a PCR instrument at a temperature of 37°C, and then add proteinase K to digest at 55°C for 45 minutes (to terminate the reaction).
[0099] ③ After the reaction is complete, run a 2% agarose gel and purify the 1.3 kb residual substrate band. Amplify the approximately 250 bp fragment containing the 6N PAM region of the recovered product using designed primers. Then, through a second round of PCR, connect the product to an adapter with a second-generation high-throughput sequencing (Illumina).
[0100] ④ After PCR amplification, the DNA product was purified by 2% agarose gel electrophoresis. The DNA sample concentration was determined using Nanodrop, and 200 ng of the purified product was subjected to PE150 sequencing.
[0101] ⑤ After sequencing is completed, perform data analysis to determine the recognition of the PAM of the CasEDG-01 protein (i.e., the results in Figure 5).
[0102] As shown in Figure 5, after repeated experimental verification, the PAMs recognized by the EDG-01 and sgRNA complex are NTTV, TTCN, and TTTT, and weakly recognize TCTV, where V represents A, C, or G; and N represents A, C, G, or T. Therefore, the EDG-01 identified in the present invention has a relatively broad recognition PAM.
[0103] From Figure 5, we can see that the PAM sequence recognized by CasEDG-01 is close to NNNTTN. At the same time, the 4th and 5th positions of PAM have activity with not only T base but also C base. Therefore, we list all the data with PAM sequence NYYN (N is selected from A, T, C, G; Y is selected from T, C) in the second-generation sequencing data separately, specifically ATTA, ATTG, ATTC, ATTT, GTTA, GTTG, GTTC, GTTT, CTTA, CTTG, CTTC, CTTT, TTTA, TTTG, TTTC, TTTT, ATCA, ATCG, ATCC, ATC The 64 PAMs were T, GTCA, GTCG, GTCC, GTCT, CTCA, CTCG, CTCC, CTCT, TTCA, TTCG, TTCC, TTCT, ACTA, ACTG, ACTC, ACTT, GCTA, GCTG, GCTC, GCTT, CCTA, CCTG, CCTC, CCTT, TCTA, TCTG, TCTC, TCTT, ACCA, ACCG, ACCC, ACCT, GCCA, GCCG, GCCC, GCCT, CCCA, CCCG, CCCC, CCCT, TCCA, TCCG, TCCC, and TCCT. We calculated the proportions of these 64 PAMs in the total sequencing data in EDG-01 and the control group (the number of reads for each of the 64 PAMs was divided by the total number of reads).
[0104] To more intuitively visualize the sequencing data, we divided the proportion of these 64 PAMs in the control group by the corresponding PAM proportions in CasEDG-01. This ratio is called the PAM depletion index. A higher ratio indicates a greater decrease in the proportion of that PAM in the CasEDG-01 sequencing data compared to the control group, indicating a greater ability of CasEDG-01 to cleave that PAM. The final results are shown in Figure 6.
[0105] In order to further verify the ability of EDG-01 to cleave substrates with different PAMs, the present invention further determined its PAM preference through in vitro cleavage experiments. EDG-01 constructed on the pET28a vector was expressed in the prokaryotic expression strain BL21, and sgRNA-1 was obtained by in vitro transcription and purification using T7 transcriptase. In order to purify the EDG-01 protein expressed in the cell, a 6×His tag amino acid sequence was added to the N-terminus of its sequence. The final amino acid sequence is SEQ ID NO.6+SEQ ID NO.28 (where "HHHHHH" is a 6×His tag):
[0106] The CasEDG-01 recombinant protein was purified by nickel column affinity chromatography, heparin column chromatography, and Superdex 200Increase 10 / 300GL molecular sieve, and finally stored in buffer (20mM Tris-HCl, pH 7.5, 150mM KCl, 1mM MgCl2, 50% glycerol, 0.5mM TCEP). The protein concentration was determined by Nanodrop, and the protein supernatant was obtained by high-speed centrifugation and then quick-frozen in liquid nitrogen. Finally, the CasEDG-01 protein solution was stored at -80°C for future use.
[0107] Based on the PAM consumption index of CasEDG-01, the present invention further verified in vitro whether CasEDG-01 has NTTN and TTCN activity and weak TCTV activity. The CasEDG-01 protein and sgRNA were mixed in a molar ratio of 1:2 to obtain a CasEDG-01-sgRNA complex. The complex was mixed with cleavage buffer (10*buffer composition is 500mM tris pH 8.0 1M NaCl 100mM MgCl2) and nuclease-free water. After incubation on ice for 30 minutes, 100ng DNA substrate (containing 24 PAMs: NTTN, TTCN and TCTN) was added to make the final CasEDG-01:sgRNA:substrate = 10:20:1. After thorough mixing, incubation was carried out at 37°C for 45 minutes, followed by addition of 1μL of proteinase K, digestion was carried out at 55°C for 45 minutes, and a 2% agarose gel was run. The final gel image is shown in Figure 7, where M(DNA Markers indicate bands of 200bp, 1000bp, 750bp, 500bp, 250bp, and 100bp in size; P in the figure indicates two bands after in vitro fragment cleavage. Compared to the control (Ctrl) without the EDG-01 and sgRNA complex, cleavage of the double-stranded DNA by the EDG-01 and sgRNA complex produced two short DNA strands. Furthermore, the in vitro cleavage results for these 24 PAMs showed similar results to the PAM depletion index of CasEDG-01, indicating that CasEDG-01 has strong activities against NTTV, TTTT, and TTCN, and weak TCTV. TTTN has the strongest activity.
[0108] To compare the cleavage activity of CasEDG-01, the present invention compared CasEDG-01 with its most similar and well-studied target, Fncpf1. First, the differences in their PAM consumption indices were compared. The results are shown in Figure 8. As can be seen from the figure, both CasEDG-01 and Fncpf1 can recognize the PAM of NTTV, but CasEDG-01 has cleavage activity against TTCT and TTTT compared to Fncpf1, while Fncpf1 has no or very weak cleavage activity. Further in vitro cleavage verification was then performed, and the results are shown in Figure 9. It can be seen that the in vitro cleavage results are consistent with the PAM consumption index results.
[0109] At the same time, the present invention also verified the inactivation site of CasEDG-01. By mutating the aspartic acid (D929) at position 929 or the glutamic acid (E1018) at position 1018 in SEQ ID NO.1+SEQ ID NO.24 into alanine separately or simultaneously (D929A, E1018A and D929A / E1018A, the three mutations are represented by 929; 1018 and DD in the figure), CasEDG-01 can be completely inactivated. The results are shown in Figure 10. The PAM of the cleavage substrate position TTTC on the left and the PAM of the cleavage substrate position TTTN on the right, M (DNA Marker) represents bands of 200 bp, 1000 bp, 750 bp, 500 bp, 250 bp, and 100 bp in size.
[0110] The present invention provides a gene editing system for eukaryotes. A plurality of gene editing vectors in eukaryotes are further constructed (as described in the examples) to verify the gene editing activity of the EDG-01 and crRNA complex. It is found that the complex has high gene editing activity in eukaryotes. Taking plant cells as an example, its editing efficiency is 74% to 79%.
[0111] The target recognition sequence of the present invention is preferably 24 nt. sgRNA is formed by designing a guide sequence consistent with the target sequence and connecting it to the crRNA nucleotide sequence. The specific sequence is shown as follows:
[0112] 5'-AATTTCTACTATTGTAGATNNNNNNNNNNNNNNNNNNNNNNNN-3', wherein the 24nt N is the target sequence, and N is selected from A, T, C, and G.
[0113] After transcription of this sequence, the generated sgRNA forms a ribonucleoprotein complex with EDG-01, which can specifically cut the target sequence with a PAM type of TTTV. The mutations produced after cutting are mainly base deletions.
[0114] This study identified a class II, type V CRISPR / Cas effector protein with gene-editing properties in the Flavobacterium branchiophilum strain. This Cas protein, designated EDG-01 or CasEDG-01, is capable of site-directed gene editing of prokaryotic and eukaryotic genomic DNA under the guidance of a single guide RNA (sgRNA). This study provides a novel, efficient, and stable gene editing system that uses EDG-01 to edit target sequences, resulting in insertions or deletions of DNA sequences.
[0115] The present invention will be described in detail below through examples. In the following examples, unless otherwise specified, the raw materials are all commercially available.
[0116] Example 1
[0117] The rice variety "Zhonghua 11" used in the present invention was provided by Wuhan Aidijing Biotechnology Co., Ltd. Based on the research on the EDG-01 gene editing tool enzyme, the crRNA of EDG-01 is the nucleotide sequence shown in SEQ ID NO.9, the target site is designed to be about 24 nt in length, and the PAM site is "TTTV" with cleavage activity. A control gene editing tool enzyme, LbCpf1, was established. This enzyme has been used as a mature gene editing tool enzyme and has relatively efficient editing activity in animals, plants, and microorganisms. The nucleotide sequence of LbCpf1's crRNA is SEQ ID NO.10: 5'-AATTTCTACTAAGTGTAGAT-3', the target site is designed to be about 24 nt in length, and the PAM site is "TTTN" with cleavage activity. Gene editing experiments were carried out using EDG-01 in eukaryotic organisms (especially plants), and rice OsBEL (LOC_Os03g55240) was selected as the target gene.
[0118] The detailed process is as follows:
[0119] 1. Target design;
[0120] Based on the gene sequence of the test gene OsBEL retrieved from the National Center for Biotechnology Information (NCBI) database, the target sequence of EDG-01 was designed using the CRISPR-GE website (http: / / skl.scau.edu.cn / ). The present invention tested the editing efficiency of the EDG-01 nuclease, and the test PAM sequence was "TTTC".
[0121] The target sequence (SEQ ID NO.11) is as follows:
[0122] >OsBEL
[0123] TTTCATCTCCTTCTAGAAGCACAAGCGC
[0124] 2. Design and assembly of EDG-01 expression vector:
[0125] Based on the amino acid sequence of EDG-01 (as shown in SEQ ID NO.1 + SEQ ID NO.24), the nucleotide codons of rice were optimized, and a nucleotide sequence encoding a nuclear localization signal peptide was added to each end of the optimized nucleotide sequence. The resulting nucleotide sequence is shown in SEQ ID NO.2. The Golden Gate method was used for enzyme digestion, ligation and cloning into a plant expression vector to assemble into the pEGEDG-01Pubi-H expression vector (nucleotide sequence shown in SEQ ID NO.3 + SEQ ID NO.25) for the next step of target assembly. The present invention uses the pEGLbCpf1Pubi-H expression vector (provided by Wuhan Aidijing Biotechnology Co., Ltd.) as a control vector. The structures of the two expression vectors are shown in Figure 11.
[0126] 3. Construction of EDG-01 editing vector:
[0127] The pEGEDG-01Pubi-H vector designed and constructed in the present invention features the use of the maize ubiquitin constitutive promoter UBI promoter to drive the EDG-01 gene, and the target assembly method adopts the rice OsU6a small RNA promoter commonly used in the plant multi-gene editing system to drive crRNA. At the same time, the editing efficiency of EDG-01 and LbCpf1 in the rice protoplast gene editing system is compared.
[0128] The target primer synthesis is shown in Table 2:
[0129] Table 2
[0130] Using the above primers, overlapping PCR was used to amplify the crRNA expression cassettes OsU6a-crRNA-BEL-T of EDG-01 and LbCpf1. The crRNA expression cassette amplification system is as follows:
[0131] (1) The first round of PCR amplification of the OsU6a promoter and crRNA was performed using the amplification systems shown in Tables 3 and 4 using the reaction conditions described in Table 5.
[0132] Table 3: pOsU6a amplification system
[0133] Table 4: crRNA amplification system
[0134] Table 5: PCR reaction parameters
[0135] (2) The second round of PCR amplification of the crRNA expression cassette was performed using the amplification system shown in Table 6 using the reaction conditions described in Table 7.
[0136] Table 6: crRNA expression cassette amplification system
[0137] Table 7: PCR reaction parameters
[0138] The expression cassette fragments were recovered using a gel recovery kit and ligated using the Golden Gate method: 15 μL of a mixture of 100 ng of the enzyme ligation system pEGEDG-01Pubi-H or pEGLbCpf1Pubi-H vector, 30 ng of the first group of crRNA expression cassette fragments or the second group of crRNA expression cassette fragments, 1.5 μl of 10× CutSmart Buffer, 35 U of T4 DNA ligase, 10 U of Bsa I restriction endonuclease, and sterile water was reacted in a PCR amplification instrument at 37°C for 5 min / 20°C for 5 min for 10 cycles. The ligated mixture was transformed into DH5α Escherichia coli competent peptide cells, and then subjected to bacterial testing and sequencing confirmation to finally form the pEGEDG-01Pubi-H-BEL-T (nucleotide sequence shown in SEQ ID NO.4+SEQ ID NO.26) and pEGLbCpf1Pubi-H-BEL-T (nucleotide sequence shown in SEQ ID NO.5+SEQ ID NO.27) editing vector, and the map is shown in Figure 12.
[0139] 4. Rice protoplast transformation:
[0140] The two gene-editing vectors, pEGEDG-01Pubi-H-BEL-T and pEGLbCpf1Pubi-H-BEL-T, were introduced into protoplasts of Zhonghua 11 using PEG-mediated rice protoplast transformation, enabling their functional development. The rice variety "Zhonghua 11" was provided by Wuhan Aidigen Biotechnology Co., Ltd., and the rice protoplast transformation process was performed by Wuhan Aidigen Biotechnology Co., Ltd.
[0141] The detailed process of rice protoplast transformation is as follows:
[0142] (1) Preparation of protoplasts;
[0143] Select etiolated rice seedlings that have been cultured in the dark at 28°C for about one week, take the stems, and remove the outermost leaf sheaths; cut the stems into segments approximately 0.5 mm in length on a clean plastic board and place them in a 50 ml conical flask; add 5-10 ml of an enzymatic hydrolyzate consisting of cellulase and cleavage enzyme to completely immerse the tissue, evacuate with a vacuum pump, wrap with sealing film, and incubate at 28°C in the dark for 4 hours with slow shaking for enzymatic hydrolysis.
[0144] The protoplasts after the above enzymatic hydrolysis were filtered through a 40 μm filter into a 2 ml EP tube, and the middle part of the filtrate was retained and centrifuged at 600 rpm in a 4 ° C refrigerated centrifuge for 5 min; then the supernatant was aspirated with a pipette tip, discarded, and washed twice with 1 ml of pre-cooled W5 solution, each time centrifuged at 500 rpm in a 4 ° C refrigerated centrifuge for 5 min; finally, 1 ml of W5 solution was added to gently suspend the suspension and allowed to stand on ice for 30 min; the suspension was centrifuged at 500 rpm for 5 min, the supernatant was discarded, and MMG solution was added to resuspend it as needed, and it was allowed to stand for 8-10 min. After the protoplast quality was tested, it was used in the next step.
[0145] (2) Introduction of vector:
[0146] The above-mentioned pEGEDG-01Pubi-H-BEL-T and pEGLbCpf1Pubi-H-BEL-T test plasmids (10 μg each) were added to each group in triplicate for a total of six tubes. 100 μL of protoplasts and an equal volume of 40% PEG4000 solution were added to each tube. After slowly inverting to mix, the tubes were incubated in a 24.5°C water bath for 15 min. 1 ml of W5 solution was added to dilute the protoplasts and mixed to terminate the reaction. The tubes were centrifuged at 500 rpm in a 4°C refrigerated centrifuge for 5 min, and the supernatant was removed with a pipette tip. 1 ml of W5 solution was then added and the tubes were washed and centrifuged twice. Finally, 1 ml of W5 solution was added, the tubes were slowly mixed, and the tubes were incubated in the dark at 28°C for 48 h.
[0147] (3) Extraction of protoplast DNA:
[0148] After culturing the protoplasts in step (2) for 48 h, centrifuge at 500 rpm for 5 min in a 4°C refrigerated centrifuge. Remove the supernatant with a pipette tip and add 500 μL of 2X CTAB buffer. Vortex the mixture thoroughly to break the protoplasts and allow the DNA to fully precipitate and dissolve in the buffer. Extract the protoplast genomic DNA from the two groups of experiments using the CTAB genomic DNA extraction method.
[0149] 5. Detection of mutation types of rice BEL gene:
[0150] The present invention detects the efficiency and type of gene editing mutations with reference to the higher throughput, more flexible and easy-to-use detection means and analysis software HiDecode system based on second-generation sequencing in the prior art CN112322704A. The detection method is based on the construction of a PCR library by Illumina sequencing. First, a pair of specific primers are designed to amplify the 200-250bp band before and after the target, and a bridge sequence is added to the 5' end of the primer to amplify the first round and amplify the target sequence; in the second round, the bridge sequence of the first round is used, and the universal library construction adapter forward / reverse primer pos-N / plate-N (N=digital number) is used: pos-N contains a 6-nt position-barcode (hole barcode sequence) indicating the 96-well position (A1~H12) on the PCR plate corresponding to each sample, a total of 96; plate-N contains a 6-nt plate-barcode (plate barcode sequence) indicating the PCR plate number; universal amplification primers 2P-F / 2P-R are used: 2P-R contains an 8-nt marker for the sample of the own library. barcode (library marker barcode sequence); to distinguish the library you built, inform the sequencing company, and let them split their own library data according to the barcode; finally, through HiDecode software analysis, you can obtain the mutation type of each target sequence of each sample.
[0151] The first round amplification primers are as follows:
[0152] BEL-seqF (SEQ ID NO.19): CTCGGAGTGATCGCACGCCTGGATTGTCTCCGCAA
[0153] BEL-seqR (SEQ ID NO.20): CTGAGAGGCTGGATGGCCGAGGAGGTAGTAGTGGA
[0154] The second round of PCR and subsequent library construction and analysis were performed according to the HiDecode protocol. After sequencing, the library was filtered with a threshold of 0.01 to reduce nonspecific amplification and highly differential fragments. The results of the editing efficiency analysis are shown in Table 8 and Figure 13 , and the types of base mutations generated after editing are shown in Figure 14 .
[0155] Table 8: Number of reads sequenced by NGS and their editing efficiency:
[0156] The number of reads refers to the number of DNA sequence fragments read during high-throughput sequencing.
[0157] It can be seen from the above results that the editing efficiency of the Cas protein EDG-01 provided by the present invention has a comparable editing efficiency to the known LbCpf1 nuclease, both of which are around 76%. Cas12a nucleases shear genomic DNA mostly with sticky ends, and the mutation type is mainly fragment deletion. Therefore, the mutation types of all the above-mentioned mutant samples of EDG-01 are analyzed. As shown in Figure 15, deletions within 20bp account for 10.27%, deletions within 21-40bp account for 40.30%, deletions within 41-60bp account for 23.62%, and deletions within 61-80bp account for 25.81%. Deletions of more than 80bp did not occur. It may be because of the small number of samples or the limited length of PCR library amplification, which is filtered out during the library construction process. The results show that in plant gene editing applications, the most common mutation type of EDG-01 is deletion within 50bp.
[0158] In the present invention, the nucleotide sequence of SEQ ID NO.2 is as follows (the underlined bold mark is the nuclear localization signal sequence):
[0159] The preferred embodiments of the present invention have been described in detail above, but the present invention is not limited thereto. Within the technical concept of the present invention, various simple variations of the technical solution of the present invention may be made, including combining the various technical features in any other appropriate manner. These simple variations and combinations should also be regarded as disclosed in the present invention and fall within the scope of protection of the present invention.
Claims
1. A Cas protein, characterized in that, The amino acid sequence of the Cas protein is an amino acid sequence selected from at least one of the following: (1) The amino acid sequence shown as SEQ ID NO.1 + SEQ ID NO.24; (2) An amino acid sequence derived from SEQ ID NO.1 + SEQ ID NO.24 that has at least 90% identity, preferably at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity with SEQ ID NO.1 + SEQ ID NO.24 and whose enzyme activity remains unchanged.
2. The Cas protein according to claim 1, wherein The Cas protein contains an RuvC nuclease domain.
3. The Cas protein according to claim 1 or 2, wherein The Cas protein is derived from the Flavobacterium branchiophilum strain.
4. A combined protein, characterized in that, The combined protein contains the Cas protein described in any one of claims 1 - 3 and fragment a linked to the Cas protein; fragment a is provided by at least one selected from a His tag and a nuclear localization signal peptide; Preferably, the His tag and / or the nuclear localization signal peptide are linked to the N-terminus and / or C-terminus of the Cas protein.
5. The combined protein according to claim 4, wherein The combined protein has an amino acid sequence selected from at least one of the following: (1) The amino acid sequence shown as SEQ ID NO.6 + SEQ ID NO.28; (2) An amino acid sequence derived from SEQ ID NO.6 + SEQ ID NO.28 that has at least 90% identity, preferably at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity with SEQ ID NO.6 + SEQ ID NO.28 and whose enzyme activity remains unchanged.
6. A gene I encoding the Cas protein described in any one of claims 1 - 3; Preferably, the nucleotide sequence of gene I is a nucleotide sequence selected from at least one of the following: (1) The nucleotide sequence shown as SEQ ID NO.7; (2) A nucleotide sequence that has at least 90% identity, preferably at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity with SEQ ID NO.
7.
7. A gene II encoding the combined protein described in claim 4 or 5; Preferably, the nucleotide sequence of gene II is a nucleotide sequence selected from at least one of the following: (1) The nucleotide sequence shown as SEQ ID NO.2; (2) A nucleotide sequence having at least 90% identity, preferably at least 91% identity, at least 92% identity, at least 93% identity, at least 94% identity, at least 95% identity, at least 96% identity, at least 97% identity, at least 98% identity, or at least 99% identity with SEQ ID NO.
2.
8. A nucleic acid molecule, characterized in that, The nucleic acid molecule contains the nucleotide sequence shown in SEQ ID NO.8 and / or SEQ ID NO.9; Preferably, the nucleic acid molecule is RNA.
9. A ribonucleoprotein complex, characterized in that, The ribonucleoprotein complex contains a protein component and a nucleic acid component; The protein component is selected from at least one of the Cas proteins described in any one of claims 1-3 and the combined proteins described in claim 4 or 5; The nucleic acid component contains a guide sequence identical to the target sequence and a nucleic acid molecule; the nucleic acid molecule is the nucleic acid molecule described in claim 8.
10. The ribonucleoprotein complex according to claim 9, wherein The guide sequence is linked to the 3' end of the nucleic acid molecule; Preferably, the length of the guide sequence is 20-28 nt.
11. The ribonucleoprotein complex according to claim 9 or 10, characterized in that, The PAM sequence of the target site recognized by the ribonucleoprotein complex is selected from at least one of NTTV, TTCN, TTTT, TCTV, TGTA; wherein N is selected from A, C, G or T, and V is selected from A, C or G; Preferably, the PAM sequence of the target site recognized by the ribonucleoprotein complex is TTTV, wherein V is selected from A, C or G.
12. A recombinant vector, characterized in that, The recombinant vector contains the nucleic acid molecule described in claim 8.
13. A transgenic cell, characterized in that, The transgenic cell contains the recombinant vector described in claim 12; Preferably, the transgenic cell is a eukaryotic cell.
14. Use of the Cas protein described in any one of claims 1-3, the combined protein described in claim 4 or 5, the gene I described in claim 6, the gene II described in claim 7, the nucleic acid molecule described in claim 8, the ribonucleoprotein complex described in any one of claims 9-11, or the recombinant vector described in claim 12 in gene editing; Preferably, the gene editing is gene editing in eukaryotes; Preferably, the mode of gene editing is gene deletion.
15. A method for gene editing, characterized in that, Delivering the ribonucleoprotein complex described in any one of claims 9-11 and / or the recombinant vector described in claim 12 to a cell containing a target gene; the target gene contains a target sequence; Preferably, the cell is a eukaryotic cell.
Citation Information
Patent Citations
Methods for generating barcoded combinatorial libraries
CN109688820A
A modular universal plasmid design strategy for the assembly and editing of multiple DNA constructs for multiple hosts
CN110312797A
Nucleic acid-guided nucleases
CN111511906A
APPLICATIONS OF CRISPRi IN HIGH THROUGHPUT METABOLIC ENGINEERING
CN112703250A
Improved CRISPR-Cas9 Genome Editing Tool
US20200291369A1