Novel cas12a nuclease best5 and use thereof
By developing the new Cas12a protein BEST5, the limitations of the existing CRISPR/Cas system on PAM dependence have been solved, and a wider PAM recognition capability and more target site selection have been achieved, providing a richer tool library for gene editing and nucleic acid detection.
Patent Information
- Application Number
- PCT/CN2023/128157
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
The existing CRISPR/Cas system relies heavily on the existence of PAM sequences in targeted sequence recognition, limiting its application scope in cell editing and nucleic acid detection, and commercial Cas12a has strict PAM requirements, limiting the design of target sites.
A new Cas12a protein is developed called BEST5. It has a wider PAM recognition capability, can recognize 5’-YYN-3’ type PAM sequences, provide more target site selection, and enrich the library of nucleic acid detection and gene editing tools by screening out new candidate Cas12a proteases in human intestinal microbiota and existing databases.
The genome editing activity of BEST5 protein in vivo has been verified, providing more tools and effective site selection for in vivo gene editing application direction, and providing a wider application possibility in the field of nucleic acid detection.
Smart Images

Figure PCTCN2023128157-FTAPPB-I100001 
Figure PCTCN2023128157-FTAPPB-I100002 
Figure PCTCN2023128157-FTAPPB-I100003
Abstract
Description
Novel Cas12a nuclease BEST5 and its applications Technical Field
[0001] The present invention relates to the field of genetic engineering technology, and in particular to a novel Cas12a nuclease BEST5 and its application in gene editing and nucleic acid detection. Background Art
[0002] The CRISPR (clustered regularly interspaced short palindromic repeats) system can be divided into class 1 and class 2 based on homology. Class 1 includes type I, type III, and type IV, and class 2 includes type II, type V, and type VI. The most obvious feature of Class 2 is that it forms a complex with a single Cas (CRISPR-associated protein) protein and crRNA (CRISPR RNA), which performs targeted cutting and is simpler to operate. Although the Class 2 system only has a single effector protein, the different types of effector proteins that have been discovered have obvious differences in protein molecular weight, structural domains, crRNA, PAM (Protospacer-adjacent motif) preference, and nucleic acid cleavage methods, providing more flexible options for gene editing. For example, the well-known Cas9, Cas12, and Cas13a all belong to class 2 (Makarova, KS, Wolf, YI, Iranzo, J., Shmakov, SA, Alkhnbashi, OS, Brouns, SJJ, Charpentier, E., Cheng, D., Haft, DH, Horvath, P., et al. (2020). Evolutionary classification of CRISPR-Cas systems: a burst of class 2 and derived variants. Nat Rev Microbiol 18, 67-83.). Compared with gene editing technologies such as ZFN (zinc-finger nucleases) and TALEN (transcription activator-like effector nucleases), the CRISPR / Cas system has obvious advantages such as good specificity, wider targeting, simple steps, ability to edit multiple sites simultaneously, and low cost (Shan Qiwei, Gao Caixia. (2015). Latest research progress in plant genome editing and derivative technologies. Heredity 37, 953-973.), 2015).
[0003] In 2015, it was first reported that the new nuclease Cas12a can bind to and cut specific sites of target DNA under the guidance of single-stranded guide RNA, and verified the effectiveness of 8 Cas12a family proteins in genome editing in mammalian cells HEK293FT (Zetsche, B., Gootenberg, JS, Abudayyeh, OO, Slaymaker, IM, Makarova, KS, Essletzbichler, P., Volz, SE, Joung, J., van der Oost, J., Regev, A., et al. (2015). Cpf1 is a single RNA-guided endonuclease of a class 2 CRISPR-Cas system. Cell 163, 759-771.). The shorter guide RNA backbone portion (crRNA only) than Cas9 is easier to identify and predict, making the assembly of Cas12a gene editing tools more convenient and efficient. Its own preference for T-rich PAM recognition properties greatly enriches the site selection of editing tools in genome editing.
[0004] In addition, studies have found that after the Cas proteins of the Cas12 and Cas13 families exert their specific cleavage effects, they activate their nonspecific incidental cleavage activities (Chen, JS, Ma, E., Harrington, LB, Da Costa, M., Tian, X., Palefsky, JM, and Doudna, JA (2018). CRISPR-Cas12a target binding unleashes indiscriminate single-stranded DNase activity. Science (New York, NY) 360, 436-439; Liu, L., Li, X., Ma, J., Li, Z., You, L., Wang, J., Wang, M., Zhang, X., and Wang, Y. (2017). If fluorescence-quenching short nucleotides are added to the reaction system, the activated Cas protein will non-specifically cleave these short nucleotides, releasing fluorescent signals, thereby detecting the targeted sequence. The CRISPR / Cas system has been successfully applied in nucleic acid detection due to its unique DNA or RNA-specific cleavage and non-specific incidental cleavage activities (Zhu, C.-s., Liu, C.-y., Qiu, X.-y., Xie, S.-s., Li, W.-y., Zhu, L., and Zhu, L.-y. (2020). Novel nucleic acid detection strategies based on CRISPR-Cas systems: From construction to application. Biotechnology and Bioengineering 117, 2279-2294.).
[0005] As an acquired immune mechanism of prokaryotes, the CRISPR / Cas system has RNA-mediated nuclease activity. The earliest discovery that can be used for gene editing is the Cas9 system mediated by crRNA and tracrRNA (Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, JA, and Charpentier, E. (2012a). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science (New York, NY) 337, 816-821.). The system captures phage viral DNA or exogenous plasmid DNA through certain Cas proteins, mainly Cas1 and Cas2, and inserts it into its own direct repeat sequence to form a CRISPR sequence. The CRISPR sequence is transcribed into pre-crRNA, which is processed and modified into crRNA to form RNP with Cas9, which has RNA-guided DNA endonuclease activity. When the bacteria are infected with the virus again, the invading DNA can be targeted for cleavage. This process also requires the participation of tracrRNA (Trans-activating crRNA) and the presence of a specifically recognized PAM (Barrangou, R., Fremaux, C., Deveau, H., Richards, M., Boyavall, P., Moineau, S., Romero, DA, and Horvath, P. (2007). CRISPR Provides Acquired Resistance Against Viruses in Prokaryotes. Science (New York, NY) 315, 1709-1712.). The CRISPR / Cas system is widely used in gene editing because of its RNA-mediated nuclease activity.
[0006] Since the effectiveness of Cas12a in genome editing in vivo was first verified in 2015, more types of Cas12a tools have been discovered and used. Researchers have found that the VA Cas12a system has conservative characteristics, and systems from different sources have certain similarities in the sequence, secondary structure and PAM screening of mature crRNA (Teng, F., Li, J., Cui, T., Xu, K., Guo, L., Gao, Q., Feng, G., Chen, C., Han, D., Zhou, Q., et al. (2019). Enhanced mammalian genome editing by new Cas12a orthologs with optimized crRNA scaffolds. Genome Biol 20,15.) This feature gives most Cas12a genome editing tools an "innate advantage" in the selection of T-rich nucleic acid sequence target sites. Of course, in order to adapt to different application scenarios, a series of new Cas12a proteins and variants have emerged, such as being more relaxed in PAM recognition and having fewer off-target events (Huang, H., Huang, G., Tan, Z., Hu, Y., Shan, L., Zhou, J., Zhang, X., Ma, S., Lv, W., Huang, T., et al. (2022). Engineered Cas12a-Plus nuclease enables gene editing with enhanced activity and specificity. BMC Biol 20, 91.;Kleinstiver, BP, Sousa, AA, Walton, RT, Tak, YE, Hsu, JY, Clement, K., Welch, MM, Horng, JE, Malagon-Lopez, J., Scarfo, I., et al. (2019). Engineered CRISPR-Cas12a variants with increased activities and improved targeting ranges for gene, epigenetic and base editing. Nature biotechnology 37, 276-282.).
[0007] From the first demonstration of the gene editing activity of Cas9 by Jennifer Anna Doudna's laboratory, to its application in mammalian cell gene editing by Zhang Feng's laboratory, to the nucleic acid detection methods developed using the accessory cleavage activity of Cas13a / Cas12a: SHERLOCK (Specific High-Sensitivity Enzymatic Reporter UnLOCKing) and DETECTR (Endonuclease Targeted CRISPR Trans Reporter) (Cong, L., Ran, FA, Cox, D., Lin, S., Barretto, R., Habib, N., Hsu, PD, Wu, X., Jiang, W., Marraffini, LA, et al. (2013). Multiplex Genome Engineering Using CRISPR / Cas Systems. Science (New York, NY) 339, 819-823.; Jinek, M., Chylinski, K., Fonfara, I., Hauer, M., Doudna, JA, and Charpentier, E. (2012b). A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity. Science (New York, NY) 337, 816-821.; Joung, J., Ladha, A., Saito, M., Kim, NG, Woolley, AE, Segel, M., Barretto, RPJ, Ranu, A., Macrae, RK, Faure, G., et al. (2020). Detection of SARS-CoV-2with SHERLOCK One-Pot Testing.N Engl J Med 383,1492-1494.;Patchsung,M.,Jantarug,K.,Pattama,A.,Aphicho,K.,Suraritdechachai,S.,Meesawat,P.,Sappakhaw,K.,Leelahakorn,N.,Ruenkam,T.,Wongsatit,T.,et al.(2020).Clinical validation of a Cas13-based assay for the detection of SARS-CoV-2 RNA. Nature biomedical engineering 4, 1140-1149.). The CRISPR / Cas system has shown great potential for commercial application. Many research institutions and companies have begun to apply for patent protection in gene editing technologies, including CRISPR / Cas, and have established a monopoly in the manufacturing industry (Wang Huiyuan, Fan Yuelei, Chu Xin, Yu Jianrong (2018). Analysis of the Development Trend of CRISPR Gene Editing Technology. Life Sciences 30, 1019-1029.).
[0008] Commercial Cas12a (such as LbaCas12a) has strict PAM requirements, which limits the design of target sites. Therefore, it is urgent to develop a Cas12a protein that is not strict with PAM requirements and is suitable for a variety of PAM sequences.
[0009] Summary of the Invention
[0010] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.To this end, an object of the present invention is to provide a novel Cas12a protein, i.e. BEST5, which can be used for nucleic acid detection system and can be used as genome editing tool enzyme, for in vivo gene editing application direction provides more tools and effective site selection.
[0011] To this end, the first aspect of the present invention provides a Cas12a protein. According to an embodiment of the present invention, the Cas12a protein has:
[0012] (1) a protein having the amino acid sequence shown in SEQ ID NO: 1;
[0013] (2) A protein having at least 80% sequence identity with SEQ ID NO: 1 and having the same or similar biological function as the Cas12a protein.
[0014] CRISPR technology has developed rapidly, enabling applications in gene editing in bacteria, archaea, and eukaryotic cells. However, it has also exposed numerous challenges. CRISPR / Cas systems rely heavily on the presence of PAM sequences for target sequence recognition, and these limitations significantly restrict their applications in cell editing and nucleic acid detection. Furthermore, concerns about off-target effects, including low specificity and the resulting large deletions and complex gene rearrangements, require further attention. Although engineered CRISPR / Cas systems have significantly improved their specificity and have been applied to gene therapy for conditions such as β-thalassemia and sickle cell anemia, their PAM limitations continue to constrain their applications. Currently, commercially available Cas proteins include SpCas9, which is restricted by the 3' PAM NGG, and LbaCas12a / AsCas12a, which is restricted by the 5' PAM TTTN. The Cas12a system's dependence on a T-rich PAM significantly limits its application in gene editing and nucleic acid detection.
[0015] Commercial Cas12a (such as LbaCas12a) has strict PAM requirements, which limits the design of target sites, and its editing activity needs to be further improved. On the other hand, in the application scenario of in vitro nucleic acid detection, according to the kinetics of enzymatic reactions, in order to increase the reaction speed, it is necessary to improve the enzyme with a higher suitable temperature for the enzymatic reaction. It is necessary to expand the existing Cas12a nuclease tool library to meet the diverse needs of gene editing and nucleic acid detection.
[0016] The inventors identified new candidate Cas12a proteases by screening human intestinal flora and existing databases, exploring novel Cas12a proteases and enriching the library of nucleic acid detection and gene editing tools. These novel Cas12a proteases were characterized and evaluated using bioinformatics, biochemistry, and cytology experiments, focusing on cleavage efficiency, PAM ubiquity, accessory cleavage activity, and stability.
[0017] At the same time, in order to break the current patent barriers of commercial Cas12a in vivo genome editing, enrich the resources of genome editing tools, and enhance the potential of subsequent engineering transformation, it is particularly important to discover systems with in vivo genome editing activity. The newly discovered and successfully verified Cas12a system with in vivo editing activity makes the tool selection more diverse, providing more directions for the later application of gene therapy, agricultural breeding, and bioengineering transformation, and providing more effective site options for gene editing.
[0018] According to an embodiment of the present invention, the having the same or similar biological function refers to having at least one of the following activities:
[0019] The activity of binding to crRNA, the activity of binding to a specific site of the target sequence under the guidance of crRNA, the activity of endonuclease, the activity of binding to a specific site of the target sequence and cutting nucleic acid under the guidance of crRNA, or the recognition of PAM site.
[0020] According to an embodiment of the present invention, the PAM site is characterized by 5'-YYN-3', wherein Y is C or T, and N is any one of A, G, C and T.
[0021] The Cas12a-BEST5 provided by the present invention has more recognizable PAM sequences and relatively low PAM requirements. Compared with the existing LbaCas12a / AsCas12a restricted by the 5' end PAM TTTN, the BEST5 of the present invention provides more options for the design of target sites.
[0022] The second aspect of the present invention provides a fusion protein. According to an embodiment of the present invention, the fusion protein includes the Cas12a protein described in the first aspect and other modified parts.
[0023] The third aspect of the present invention provides an isolated polynucleotide. According to an embodiment of the present invention, the polynucleotide is a polynucleotide encoding the Cas12a protein described in the first aspect or a polynucleotide encoding the fusion protein described in the second aspect.
[0024] According to an embodiment of the present invention, the polynucleotide is selected from the polynucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO: 31.
[0025] The fourth aspect of the present invention provides an expression vector. According to an embodiment of the present invention, the expression vector comprises the isolated polynucleotide described in the third aspect.
[0026] According to an embodiment of the present invention, the expression vector further comprises a crRNA designed according to the target sequence and the PAM sequence,
[0027] Wherein, in the expression vector, the Cas12a protein described in the first aspect, encoded by the isolated polynucleotide described in the third aspect, is co-expressed with the crRNA,
[0028] Or the Cas12a protein described in the first aspect encoded by the isolated polynucleotide of the third aspect and the crRNA are expressed separately in different vectors.
[0029] The fifth aspect of the present invention provides a CRISPR-Cas system. According to an embodiment of the present invention, the CRISPR-Cas system comprises the Cas12a protein described in the first aspect and at least one crRNA,
[0030] The crRNA includes a backbone region capable of binding to the Cas12a protein described in the first aspect and a guide sequence capable of targeting a target sequence.
[0031] According to an embodiment of the present invention, the backbone region sequence in crRNA that binds to the Cas12a protein is UAAUUUCUACUGUUGUAGAU, as shown in SEQ ID NO: 49. A sixth aspect of the present invention provides a composition. According to an embodiment of the present invention, the composition comprises:
[0032] (i) a protein component selected from the group consisting of: the Cas12a protein described in the first aspect or the fusion protein described in the second aspect;
[0033] (ii) a nucleic acid component selected from: crRNA, or a nucleic acid encoding the crRNA; or a precursor crRNA of the crRNA, or a nucleic acid encoding the precursor crRNA;
[0034] The protein component and the nucleic acid component combine with each other to form a complex,
[0035] Wherein, the crRNA includes a backbone region capable of binding to the Cas12a protein described in the first aspect and a guide sequence capable of targeting a target sequence.
[0036] The seventh aspect of the present invention provides a host cell. According to an embodiment of the present invention, the host cell comprises the Cas12a protein described in the first aspect, or the fusion protein described in the second aspect, or the isolated polynucleotide described in the third aspect, or the expression vector described in the fourth aspect, or the CRISPR-Cas system described in the fifth aspect, or at least one of the composition described in the sixth aspect.
[0037] The eighth aspect of the present invention provides the Cas12a protein described in the first aspect, the fusion protein described in the second aspect, the isolated polynucleotide described in the third aspect, the expression vector described in the fourth aspect, the CRISPR-Cas system described in the fifth aspect, the composition described in the sixth aspect, and the host cell described in the seventh aspect for use in gene targeting, gene editing, gene modification, gene transcription and expression regulation.
[0038] The ninth aspect of the present invention provides the Cas12a protein of the first aspect, the fusion protein of the second aspect, the isolated polynucleotide of the third aspect, the expression vector of the fourth aspect, the CRISPR-Cas system of the fifth aspect, the composition of the sixth aspect, and the host cell of the seventh aspect selected from any one or more of the following:
[0039] Targeting and / or editing target nucleic acids; cleaving double-stranded DNA, single-stranded DNA, or single-stranded RNA; non-specific cleavage and / or degradation of collateral nucleic acids; non-specific cleavage of single-stranded nucleic acids; nucleic acid detection; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids.
[0040] The tenth aspect of the present invention provides the Cas12a protein described in the first aspect, the fusion protein described in the second aspect, the isolated polynucleotide described in the third aspect, the expression vector described in the fourth aspect, the CRISPR-Cas system described in the fifth aspect, the composition described in the sixth aspect, and the host cell described in the seventh aspect for constructing a CRISPR-Cas12a gene editing system.
[0041] The eleventh aspect of the present invention provides a method for editing a target nucleic acid, targeting a target nucleic acid, or cutting a target nucleic acid. According to an embodiment of the present invention, the method includes contacting the target nucleic acid with at least one of the Cas12a protein described in the first aspect, the fusion protein described in the second aspect, the isolated polynucleotide described in the third aspect, the expression vector described in the fourth aspect, the CRISPR-Cas system described in the fifth aspect, the composition described in the sixth aspect, and the host cell described in the seventh aspect.
[0042] A twelfth aspect of the present invention provides a method for cutting a single-stranded nucleic acid. According to an embodiment of the present invention, the method comprises:
[0043] The nucleic acid population is contacted with the Cas12a protein and crRNA described in the first aspect,
[0044] wherein the nucleic acid population comprises a target nucleic acid and at least one non-target single-stranded nucleic acid,
[0045] The crRNA is capable of targeting the target nucleic acid, and the Cas12a protein cuts the non-target single-stranded nucleic acid.
[0046] The crRNA includes a backbone region capable of binding to the Cas12a protein described in the first aspect and a guide sequence capable of targeting a target sequence.
[0047] The thirteenth aspect of the present invention provides a kit for gene editing, gene targeting or gene cutting. According to an embodiment of the present invention, the kit includes the Cas12a protein described in the first aspect, the fusion protein described in the second aspect, the isolated polynucleotide described in the third aspect, the expression vector described in the fourth aspect, the CRISPR-Cas system described in the fifth aspect, the composition described in the sixth aspect, and at least one of the host cell described in the seventh aspect.
[0048] A fourteenth aspect of the present invention provides a kit for detecting a target nucleic acid in a sample. According to an embodiment of the present invention, the kit comprises:
[0049] (a) the Cas12a protein described in the first aspect, or a nucleic acid encoding the Cas12a protein;
[0050] (b) crRNA, or a nucleic acid encoding the crRNA, or a precursor crRNA of the crRNA, or a nucleic acid encoding the precursor crRNA;
[0051] (c) a single-stranded nucleic acid detector that is single-stranded and does not hybridize with the crRNA,
[0052] The crRNA includes a backbone region capable of binding to the Cas12a protein described in the first aspect and a guide sequence capable of targeting a target sequence.
[0053] The fifteenth aspect of the present invention provides the Cas12a protein of the first aspect, the fusion protein of the second aspect, the isolated polynucleotide of the third aspect, the expression vector of the fourth aspect, the CRISPR-Cas system of the fifth aspect, the composition of the sixth aspect, and the host cell of the seventh aspect for preparing a preparation or a kit, wherein the preparation or kit is used for:
[0054] i) gene or genome editing;
[0055] ii) target nucleic acid detection;
[0056] iii) editing a target sequence in a target locus to modify the nucleic acid of an organism;
[0057] iv) diagnosis and / or treatment of disease;
[0058] v) Construction of disease models and drug screening.
[0059] A sixteenth aspect of the present invention provides a method for detecting a target nucleic acid in a sample. According to an embodiment of the present invention, the method comprises contacting the sample with the Cas12a protein, crRNA, and a single-stranded nucleic acid detector described in the first aspect, detecting a detectable signal generated by the Cas12a protein cleaving the single-stranded nucleic acid detector, thereby detecting the target nucleic acid,
[0060] Wherein the crRNA comprises a backbone region capable of binding to the Cas12a protein described in the first aspect and a guide sequence capable of targeting a target sequence,
[0061] The single-stranded nucleic acid detector does not hybridize with the crRNA.
[0062] CRISPR-Cas12a protein can realize accurate recognition and cutting of different DNAs by the design of crRNA. After binding to the target DNA, the Cas12a protein cuts and forms double-strand breaks in the target DNA chain. In this process, once the target sequence is combined, the RuvC domain of the Cas12a protein is activated, and Cas12a has the activity of an endonuclease that degrades any sequence single-stranded DNA (trans single-stranded DNA), which can non-specifically cut ssDNA. Therefore, by carrying the ssDNA molecule of the reporter group, efficient recognition and amplification of the targeted DNA can be achieved. This principle enables the Cas12a protein to be used as a recognition element for detecting nucleic acids, while also being able to amplify the output signal, and has important application value in the field of nucleic acid detection. The inventors discovered a new Cas12a protein in the Clostridium sp.AM42-36 strain and named it BEST5. After optimizing the new protein based on human codons, a genome editing plasmid was constructed, and effective system delivery was achieved by liposome transfection. The new system's genome editing activity was verified at the "Safe harbor gene" AAVS1 site in the human genome, and effective editing sites were screened, laying the foundation for later tools or system modifications based on the new system. The present invention provides a new nucleic acid detection system and genome editing tool enzyme. The genome editing activity of BEST5 in mammalian cells has been verified, providing more tools and effective site selection for in vivo gene editing applications.
[0063] The principle of specific signal amplification is achieved by utilizing the sequence recognition of the Cas12a effector protein and the cutting of trans single-stranded DNA. The sample containing the double-stranded DNA to be tested is mixed with the RNP formed by Cas12a and specific crRNA and the trans single-stranded DNA probe for incubation. The target nucleic acid sequence carried by the sample will stimulate the trans-cleavage activity of Cas12a, continuously cut the single-stranded DNA probe, and amplify the output signal while detecting the recognition element of the nucleic acid. Further combined with in vitro nucleic acid extraction and amplification, reverse transcription and fluorescent biosensing technologies, it can be applied to biological detection of multiple targets such as pathogens, DNA viruses, and RNA viruses.
[0064] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:
[0066] Figures 1A-1B show the plasmid map (Figure 1A) after the Cas12a protein BEST5 nucleic acid sequence is inserted into the pET28a (+) vector and the SDS-PAGE electrophoresis diagram (Figure 1B) after purification in Example 1 of the present invention;
[0067] Figure 2 shows the results of detecting the PAM sequence of Cas12a protein BEST5 using DocMF in Example 2 of the present invention;
[0068] Figure 3 shows the results of in vitro dsDNA cleavage activity detection of the Cas12a protein BEST5 in Example 3 of the present invention, wherein A. the results of in vitro dsDNA cleavage activity detection of BEST5 by agarose gel electrophoresis; B. the ratio of the remaining amount of reaction substrate to the input amount by band intensity analysis (the ordinate is percentage, unit: %), wherein LbaCas12a is a positive control commercial protein;
[0069] Figure 4 shows the results of the comparison of the accessory cleavage activities of BEST5 and LbaCas12a in Example 4 of the present invention;
[0070] Figure 5 shows a comparison of the differences in BEST5 cleavage activity of different crRNA lengths in Example 5 of the present invention;
[0071] FIG6 shows the comparative results of BEST5 accessory cleavage activity at different reaction temperatures in Example 6 of the present invention;
[0072] FIG7 shows the results of in vitro PAM region validation of BEST5 in Example 7 of the present invention;
[0073] Figure 8 shows the electrophoresis diagram of the editing plasmid and editing product T7E1 of the BEST5 system in mammalian cell genome editing (AAVS1) in Example 8 of the present invention, wherein A. plasmid map, B. T7E1 electrophoresis, wherein M: DL2000 DNA marker; +: T7EI enzyme added; -: T7EI enzyme not added.
[0074] Figures 9A-9B show the editing plasmid maps of the BEST5 and SpCas9 systems in Example 9 of the present invention in the HBG promoter and BCL11a targeting region;
[0075] FIG10 shows the editing efficiency of BEST5 and SpCas9 proteins in the HBG promoter and BCL11a targeting region in Example 9 of the present invention.
[0076] Detailed Description of the Invention
[0077] The embodiments of the present invention are described in detail below. The embodiments described below are exemplary and are only used to explain the present invention, and should not be understood as limiting the present invention.
[0078] Conventional techniques such as immunology, biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and recombinant DNA used in the present invention can be found in Sambrook, Fritsch and Maniatis, MOLECULAR CLONING: A LABORATORY MANUAL, 2nd ed. (1989); CURRENT PROTOCOLS IN MOLECULAR BIOLOGY (F M Ausubel et al., eds., (1987)); METHODS IN ENZYMOLOGY series (Academic Press, Inc.): PCR 2: A PRACTICAL METHOD. APPROACH) (MJ MacPherson, BD Hames and GR Taylor, eds. (1995)), Harlow and Lane, eds. (1988) ANTIBODIES, A LABORATORY MANUAL, and ANIMAL CELL CULTURE (RI Freshney, ed. (1987)).
[0079] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include one or more of such features. Furthermore, in the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0080] The endpoints of the ranges and any values disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoints of each range, the endpoints of each range and individual point values, and the individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered to be specifically disclosed herein.
[0081] In order to make the present invention more easily understood, certain technical and scientific terms are specifically defined below. Unless otherwise clearly defined elsewhere in this document, all other technical and scientific terms used herein have the meaning commonly understood by those skilled in the art to which the present invention belongs.
[0082] In this document, the terms “include” or “comprising” are open expressions, that is, including the contents specified in the present invention, but not excluding other contents.
[0083] As used herein, the terms "optionally," "optional," or "optionally" generally mean that the subsequently described event or circumstance may but need not occur, and that the description includes instances where the event or circumstance occurs and instances where it does not.
[0084] According to a specific embodiment of the present invention, the present invention provides a Cas12a protein, wherein the Cas12a protein has:
[0085] (1) a protein having the amino acid sequence shown in SEQ ID NO: 1;
[0086] (2) A protein having at least 80% sequence identity with SEQ ID NO: 1 and having the same or similar biological function as the Cas12a protein.
[0087] According to some specific embodiments of the present invention, "at least 80% sequence identity" refers to an amino acid sequence that has at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity with SEQ ID NO: 1, and substantially retains the biological function of the Cas12a protein shown in SEQ ID NO: 1.
[0088] According to some specific embodiments of the present invention, the Cas12a protein provided by the present invention has one or more amino acid substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 amino acid substitutions, deletions or additions) compared to the sequence shown in SEQ ID NO: 1, and has the same or similar biological function as the Cas12a protein.
[0089] According to a specific embodiment of the present invention, the having the same or similar biological function refers to having at least one of the following activities:
[0090] The activity of binding to crRNA, the activity of binding to a specific site of the target sequence under the guidance of crRNA, the activity of endonuclease, the activity of binding to a specific site of the target sequence and cutting nucleic acid under the guidance of crRNA, or the activity of recognizing PAM sites.
[0091] According to an embodiment of the present invention, the PAM site is characterized by 5'-YYN-3', wherein Y is C or T, and N is any one of A, G, C and T.
[0092] The present invention further provides a fusion protein, which includes the Cas12a protein described above and other modified parts.
[0093] According to a specific embodiment of the present invention, the modified portion is selected from another protein or polypeptide, a detectable label or any combination thereof. For example, the fusion protein of the present invention comprises a detectable label, such as a fluorescent dye.
[0094] The present invention also provides an isolated polynucleotide, which is a polynucleotide encoding the aforementioned Cas12a protein or a polynucleotide encoding the aforementioned fusion protein.
[0095] The isolated polynucleotide can be based on codon degeneracy, can encode the polynucleotide of the Cas12a protein or all polynucleotide sequences encoding the aforementioned fusion protein (including sequences codon-optimized when expressed in prokaryotic cells), which are all encompassed within the scope of protection of the present invention. As a preferred embodiment, the polynucleotide is selected from the polynucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO: 31.
[0096] In some embodiments of the present invention, the polynucleotide is preferably a single-stranded or double-stranded nucleotide.
[0097] The present invention also provides an expression vector comprising the isolated polynucleotide described above. Preferably, the expression vector further comprises a regulatory element operably linked to the isolated polynucleotide.
[0098] In one embodiment, the regulatory element is selected from one or more of the following groups: enhancer, transposon, promoter, terminator, polyadenylation sequence, marker gene.
[0099] In some embodiments of the present invention, the vector includes a cloning vector, an expression vector, a shuttle vector, and an integration vector.
[0100] In some embodiments of the present invention, the vector is a viral vector (e.g., a retroviral vector, a lentiviral vector, an adenoviral vector, an adeno-associated vector, and a herpes simplex vector), and can also be a plasmid, a virus, a cosmid, a phage, etc., which are well known to those skilled in the art.
[0101] In some embodiments of the present invention, the expression vector further comprises a crRNA designed according to the target sequence and the PAM sequence.
[0102] Wherein, in the expression vector, the aforementioned Cas12a protein encoded by the aforementioned isolated polynucleotide is co-expressed with the crRNA,
[0103] Or the aforementioned Cas12a protein and the crRNA encoded by the aforementioned separated polynucleotide are expressed respectively in different vectors.
[0104] The present invention further provides a CRISPR-Cas system, which comprises the Cas12a protein described above and at least one crRNA.
[0105] The crRNA includes a backbone region capable of binding to the aforementioned Cas12a protein and a guide sequence capable of targeting a target sequence.
[0106] According to an embodiment of the present invention, the backbone region sequence that binds to the Cas12a protein in crRNA is UAAUUUCUACUGUUGUAGAU, as shown in SEQ ID NO: 49.
[0107] The present invention provides an engineered non-naturally occurring vector system, or a CRISPR-Cas system, which includes a Cas12a protein or a nucleic acid sequence encoding the Cas12a protein and a nucleic acid encoding one or more crRNAs.
[0108] In some embodiments of the present invention, the nucleic acid sequence encoding the Cas12a protein and the nucleic acid encoding one or more crRNAs are artificially synthesized.
[0109] In some embodiments of the present invention, the nucleic acid sequence encoding the Cas12a protein and the nucleic acid encoding one or more crRNAs do not naturally exist at the same time.
[0110] The one or more crRNAs target one or more target sequences in the cell. The one or more target sequences are hybridized with the genome of the DNA molecule encoding one or more gene products, and the Cas12a protein is guided to the genomic position of the DNA molecule of the one or more gene products. After the Cas12a protein reaches the target sequence position, the target sequence is modified, edited or cut, thereby changing or modifying the expression of the one or more gene products.
[0111] The present invention provides a composition comprising:
[0112] (i) a protein component selected from the group consisting of: the aforementioned Cas12a protein or the aforementioned fusion protein;
[0113] (ii) a nucleic acid component selected from: crRNA, or a nucleic acid encoding the crRNA; or a precursor crRNA of the crRNA, or a nucleic acid encoding the precursor crRNA;
[0114] The protein component and the nucleic acid component combine with each other to form a complex,
[0115] Wherein, the crRNA includes a backbone region capable of binding to the aforementioned Cas12a protein and a guide sequence capable of targeting a target sequence.
[0116] In some embodiments of the present invention, the composition is non-naturally occurring or modified. In one embodiment, at least one component of the composition is non-naturally occurring or modified. In one embodiment, the protein component is non-naturally occurring or modified; and / or the nucleic acid component is non-naturally occurring or modified.
[0117] In some embodiments of the present invention, the Cas12a of the present invention, crRNA, fusion protein, nucleic acid molecules, carriers, systems, compositions, can be delivered by any method known in the art.Such methods include but are not limited to electroporation, lipofection, nuclear transfection, microinjection, sonoporation, gene gun, calcium phosphate-mediated transfection, cationic transfection, liposome transfection, dendritic transfection, heat shock transfection, nuclear transfection, magnetofection, lipofection, puncture transfection, optical transfection, reagent enhancement nucleic acid uptake and via liposomes, immunoliposomes, viral particles, artificial virions, etc. delivery.In one embodiment, the delivery vector is selected from lipid particles, sugar particles, metal particles, protein particles, liposomes, exosomes, microvesicles, gene guns or viral vectors (for example, replication-defective retrovirus, slow virus, adenovirus or adeno-associated virus).
[0118] The present invention further provides a host cell, comprising at least one of the aforementioned Cas12a protein, or the aforementioned fusion protein, or the aforementioned isolated polynucleotide, or the aforementioned expression vector, or the aforementioned CRISPR-Cas system, or the aforementioned composition.
[0119] In some embodiments of the invention, the cell is a prokaryotic cell.
[0120] In some embodiments of the present invention, the cell is a eukaryotic cell. In some embodiments of the present invention, the cell is a mammalian cell. In some embodiments of the present invention, the cell is a human cell. In some embodiments of the present invention, the cell is a non-human mammalian cell, such as a cell of a non-human primate, a cow, a sheep, a pig, a dog, a monkey, a rabbit, a rodent (such as a rat or a mouse). In some embodiments of the present invention, the cell is a non-mammalian eukaryotic cell, such as a cell of a poultry bird (such as a chicken), a fish, or a crustacean (such as a clam, a shrimp).
[0121] In some embodiments of the invention, the cell is a stem cell or a stem cell line.
[0122] In some embodiments of the invention, the host cells of the invention comprise genetic or genomic modifications that are not present in their wild type.
[0123] The present invention further provides the aforementioned Cas12a protein, the aforementioned fusion protein, the aforementioned isolated polynucleotide, the aforementioned expression vector, the aforementioned CRISPR-Cas system, the aforementioned composition, and the aforementioned host cell for use in gene targeting, gene editing, gene modification, gene transcription and expression regulation.
[0124] In some embodiments of the present invention, the gene targeting, gene editing, and gene modification are performed inside and / or outside cells.
[0125] In some embodiments of the present invention, the gene editing or editing of the target nucleic acid includes modifying a gene, knocking out a gene, changing the expression of a gene product, repairing a mutation, and / or inserting a polynucleotide, a gene mutation. The editing can be performed in prokaryotic cells and / or eukaryotic cells.
[0126] The present invention further provides the aforementioned Cas12a protein, the aforementioned fusion protein, the aforementioned isolated polynucleotide, the aforementioned expression vector, the aforementioned CRISPR-Cas system, the aforementioned composition, and the aforementioned host cell selected from any one or more of the following:
[0127] Targeting and / or editing target nucleic acids; cleaving double-stranded DNA, single-stranded DNA, or single-stranded RNA; non-specific cleavage and / or degradation of collateral nucleic acids; non-specific cleavage of single-stranded nucleic acids; nucleic acid detection; specifically editing double-stranded nucleic acids; base editing double-stranded nucleic acids; base editing single-stranded nucleic acids.
[0128] The present invention further provides the aforementioned Cas12a protein, the aforementioned fusion protein, the aforementioned isolated polynucleotide, the aforementioned expression vector, the aforementioned CRISPR-Cas system, the aforementioned composition, and the aforementioned host cell for use in constructing a CRISPR-Cas12a gene editing system.
[0129] The present invention further provides a method for editing a target nucleic acid, targeting a target nucleic acid, or cutting a target nucleic acid, the method comprising contacting the target nucleic acid with at least one of the aforementioned Cas12a protein, the aforementioned fusion protein, the aforementioned isolated polynucleotide, the aforementioned expression vector, the aforementioned CRISPR-Cas system, the aforementioned composition, and the aforementioned host cell.
[0130] In some embodiments of the invention, the contacting can be inside a cell in vitro, ex vivo, or in vivo.
[0131] In some embodiments of the present invention, cleavage of the single-stranded nucleic acid is non-specific cleavage.
[0132] The present invention further provides a method for cleaving a single-stranded nucleic acid, the method comprising:
[0133] The nucleic acid population is contacted with the Cas12a protein and crRNA as described above,
[0134] wherein the nucleic acid population comprises a target nucleic acid and at least one non-target single-stranded nucleic acid,
[0135] The crRNA is capable of targeting the target nucleic acid, and the Cas12a protein cuts the non-target single-stranded nucleic acid.
[0136] The crRNA includes a backbone region capable of binding to the aforementioned Cas12a protein and a guide sequence capable of targeting a target sequence.
[0137] The present invention further provides a kit for gene editing, gene targeting or gene cutting, the kit comprising at least one of the aforementioned Cas12a protein, the aforementioned fusion protein, the aforementioned isolated polynucleotide, the aforementioned expression vector, the aforementioned CRISPR-Cas system, the aforementioned composition, and the aforementioned host cell.
[0138] The present invention further provides a kit for detecting a target nucleic acid in a sample, the kit comprising:
[0139] (a) the Cas12a protein described above, or a nucleic acid encoding the Cas12a protein;
[0140] (b) crRNA, or a nucleic acid encoding the crRNA, or a precursor crRNA of the crRNA, or a nucleic acid encoding the precursor crRNA;
[0141] (c) a single-stranded nucleic acid detector that is single-stranded and does not hybridize with the crRNA,
[0142] The crRNA includes a backbone region capable of binding to the aforementioned Cas12a protein and a guide sequence capable of targeting a target sequence.
[0143] The present invention further provides the aforementioned Cas12a protein, the aforementioned fusion protein, the aforementioned isolated polynucleotide, the aforementioned expression vector, the aforementioned CRISPR-Cas system, the aforementioned composition, and the aforementioned host cell for preparing a preparation or a kit, wherein the preparation or kit is used for:
[0144] i) gene or genome editing;
[0145] ii) target nucleic acid detection;
[0146] iii) editing a target sequence in a target locus to modify the nucleic acid of an organism;
[0147] iv) diagnosis and / or treatment of disease;
[0148] v) Construction of disease models and drug screening.
[0149] In some embodiments of the present invention, the above-mentioned gene or genome editing is performed inside or outside the cell.
[0150] In some embodiments of the present invention, the target nucleic acid detection is performed in vitro.
[0151] In some embodiments of the invention, the treatment of the disease is the treatment of a condition caused by a defect in the target sequence in the target locus.
[0152] The present invention further provides a method for detecting a target nucleic acid in a sample, the method comprising contacting the sample with the aforementioned Cas12a protein, crRNA, and a single-stranded nucleic acid detector, detecting a detectable signal generated by the Cas12a protein cleaving the single-stranded nucleic acid detector, thereby detecting the target nucleic acid.
[0153] The crRNA comprises a backbone region capable of binding to the aforementioned Cas12a protein and a guide sequence capable of targeting a target sequence.
[0154] The single-stranded nucleic acid detector does not hybridize with the crRNA.
[0155] The principle of specific signal amplification is achieved by utilizing the sequence recognition of the Cas12a effector protein and the cutting of trans single-stranded DNA. The sample containing double-stranded DNA to be tested is mixed with the RNP formed by Cas12a and specific crRNA and the trans single-stranded DNA probe to be incubated. The target nucleic acid sequence carried by the sample will excite the trans-cleavage activity of Cas12a, continuously cut the single-stranded DNA probe, and amplify the output signal while detecting the recognition element of the nucleic acid. Further combined with in vitro nucleic acid extraction and amplification, reverse transcription and fluorescent biosensor technologies, it can be applied to biological detection of multiple targets such as pathogens, DNA viruses, RNA viruses. Utilizing the sequence recognition of Cas12a effector protein and the cutting of trans single-stranded DNA, the target nucleic acid sequence carried by the positive sample excites the trans-cleavage activity of the Cas12a effector protein, and the trans single-stranded DNA probes with FAM and BHQ1 at both ends are cut to achieve specific signal amplification, collect fluorescence intensity, and perform nucleic acid detection.
[0156] According to a specific embodiment of the present invention, a sample containing double-stranded DNA to be tested is mixed with the RNP (a complex of Cas effector protein and crRNA) formed by Cas12a and specific crRNA, and a trans single-stranded DNA probe with FAM and BHQ1 at both ends. The target nucleic acid sequence carried by the sample excites the trans-cleavage activity of Cas12a, cuts the single-stranded DNA probe, releases the fluorescent group, and collects the fluorescence intensity after 10-30 minutes. The scheme of the present disclosure will be explained below with reference to the examples. Those skilled in the art will understand that the following examples are only used to illustrate the present disclosure and should not be regarded as limiting the scope of the present disclosure. In the examples, if specific techniques or conditions are not indicated, they are carried out according to the techniques or conditions described in the literature in this area or according to the product specifications. The reagents or instruments used are not indicated by the manufacturer and are all conventional products that can be obtained commercially.
[0157] Example 1: Screening and purification of Cas12a nuclease BEST5
[0158] Using a bioinformatics prediction process (patent application number: 201610741844.0, "A method and apparatus for screening novel CRISPR-Cas systems"), a new Cas12a protein was discovered in the human gut bacterial reference genome (Zou Y., Xue W., Luo G., Deng Z., Qin P., Guo R., et al. (2019). 1,520 reference genomes from cultivated human gut bacteria enable functional microbiome analyses. Nature Biotechnology). This protein was named BEST5. Combining previous and uniquely optimized molecular experimental techniques, several new Cas12a proteases with potential gene editing capabilities were identified. The selected Cas12a protease, BEST5, was selected for further validation.
[0159] 1. Construction of prokaryotic expression vector
[0160] The amino acid sequence of the Cas12a protein BEST5 identified in the present invention is shown in SEQ ID NO: 1 in Table 1. The Cas12a protein BEST5 editing gene (as shown in SEQ ID NO: 2 in Table 1) and its protein purification and expression-related tag sequence (DNA fragment inserted into the plasmid as shown in SEQ ID NO: 3 in Table 1) were integrated into the pET28a (+) vector (Figure 1A).
[0161] 2. Expression and purification of Cas12a protein BEST5
[0162] Before protein expression and purification, the physicochemical properties of the protein, including isoelectric point, relative molecular mass, and extinction coefficient, were analyzed based on the protein sequence using the ProtParam tool provided by ExPasy (https: / / web.expasy.org / protparam / ) to adjust the purification process and buffer.
[0163] The plasmid was introduced into competent cells BL21 (DE3) (Takara) using the heat shock transformation method, and 300 μL of antibiotic-free culture medium was added and cultured for 60 minutes. The plate was spread (LB plate, kanamycin resistance) and cultured overnight at 37°C, and a single colony was selected for expansion culture. The protein-expressing BL21 (DE3) cells were cultured in LB medium (supplemented with 50 mg / L kanamycin) at 37°C until the OD600 reached 0.6, and protein expression was induced by adding 0.5 mM isopropyl β-D-thiogalactopyranoside (IPTG). The BL21 (DE3) cells were further cultured overnight at 16°C (low temperature induction). The cells were collected by centrifugation at 6000 rpm and 4°C for 10 min. The collected cells were resuspended in a 1 g:20 mL binding buffer (50 mM Tris-HCl, pH 7.8, 500 mM NaCl, 5 mM imidazole) and lysed by sonication. Lysozyme (10 mg / mL) and PMSF (0.1 M) were added at a volume ratio of 1:100 before sonication. The sonicated cells were centrifuged at 12000 rpm and 4°C for 60 min, and the supernatant was collected.
[0164] (1) Affinity chromatography. Because the target protein carries a His tag, we first use a Ni-NTA gravity column for affinity chromatography to purify the protein. Before use, the filler should be washed three times with water and once with binding buffer. The filler is combined with the bacterial supernatant for 30 minutes, and is fully shaken every 5 minutes to allow as much target protein as possible to bind to the Ni on the filler. Collect the flow-through. Wash with 5% elution buffer (50mM Tris-HCl, pH7.8, 500mM NaCl, 500mM imidazole) and collect the washed components. Elute the target protein with 50% elution buffer. Collect the target protein. Rinse the filler with elution buffer and collect the components. All collected components are sampled for SDS-PAGE to confirm the purification efficiency and recovery efficiency of the target protein.
[0165] (2) Molecular sieve chromatography. Use a 50K ultrafiltration tube to concentrate the target protein component to a volume of <2 mL and filter with a 0.56 μm filter membrane. Use AKTA (Cytiva) to separate proteins of different molecular weights through HiLoad 16 / 600 Superdex 200 pg (Cytiva). Load the sample using a 2 mL sample loop and pass it through the chromatography column in low buffer (30 mM phosphate, pH 7.0, 150 mM NaCl, 0.4 mM DTT) at a flow rate of 0.5 mL / min. The collection plate is continuously collected. Samples are collected from the collection tubes at all UV peaks for SDS-PAGE electrophoresis to confirm the target protein component.
[0166] (3) Ion exchange chromatography. The target protein fraction was supplemented with low buffer to 20 mL. Filtered with a 0.56 μm filter membrane, the protein was separated and purified using HiTrap Capto SP ImpRes (Cytiva) using AKTA. The target protein was loaded onto the column and gradient eluted with 50% high buffer (30 mM phosphate, pH 7.0, 1 M NaCl, 0.4 mM DTT). The collection plate was continuously collected. Samples from each tube were collected and subjected to SDS-PAGE electrophoresis to confirm the purification efficiency and recovery efficiency. For systems with poor purity (<90%), cation exchange was used for further purification. The main peak collection tube fraction was concentrated using a 50 kDa ultrafiltration tube to a volume of <3 mL, and then supplemented with low buffer to 20 mL. The above process was repeated to reduce the NaCl content in the target protein fraction. The protein was separated and purified using HiTrap Capto Q ImpRes (Cytiva) using AKTA. The target protein was loaded onto the column and gradient eluted with 50% high buffer. The collection plate was continuously collected. Samples from each tube were collected and subjected to SDS-PAGE electrophoresis. The electrophoresis results are shown in Figure 1B, confirming that the purification efficiency and recovery efficiency meet the requirements of subsequent activity determination.
[0167] Concentrate the protein using a 50K ultrafiltration tube and measure the absorbance A280 (1 Abs = 1 mg / mL) with a microplate reader. Divide by the extinction coefficient to obtain the actual protein concentration. Add 70% sterile glycerol at a volume ratio of 1:1 and store at -20°C.
[0168] Table 1 BEST5 protein, encoding nucleic acid and prokaryotic expression vector insertion sequence involved in this example
[0169] Example 2: Identification of the PAM sequence of Cas12a protein BEST5
[0170] The PAM of Cas12a nuclease BEST5 was identified using the DocMF method in a published article (Li, Z., Wang, X., Xu, D., Zhang, D., Wang, D., Dai, X., Wang, Q., Li, Z., Gu, Y., Ouyang, W., et al. (2020). DNB-based on-chip motif finding: A high-throughput method to profile different types of protein-DNA interactions. Science Advances 6, eabb3350.). The specific DocMF experimental process includes dsDNA library and DNB preparation, on-chip sequencing, and protein cleavage and imaging. The crRNA sequence information for mediating BEST5 protein targeted cutting of dsDNA library is shown in Table 2 (shown in SEQ ID NO: 4 in Table 2); the dsDNA library consists of a 23nt fixed sequence region and a 15nt random base region on both sides, wherein the fixed sequence serves as the gRNA target sequence and the random base region serves as the PAM recognition site. The sequence information is shown in Table 2 (shown in SEQ ID NO: 5 in Table 2). By statistically analyzing the base distribution of random N in the DNA sequence that presents signal differences before and after dsDNA library cutting, the PAM sequence of BEST5 was identified to be 5'-YYN, wherein Y is C or T, and N is A or T or G or C (as shown in Figure 2).
[0171] Table 2 Nucleic acid sequences used in PAM identification
[0172] Example 3: In vitro cleavage ability detection of Cas12a protein BEST5
[0173] 1. Preparation of cleavage substrate
[0174] (1) HBG gene DNA fragment was amplified from 293T cell genomic DNA by PCR as a dsDNA cleavage substrate (SEQ ID NO: 8) for experimental use. The primers were synthesized by Beijing Liuhe BGI Genomics Co., Ltd. The amplified products were identified by 1.5% agarose gel electrophoresis, and the corresponding bands were cut out of the gel according to their size and analyzed by Qubit TM Gel recovery and purification using the dsDNA HS Assay Kit were performed to obtain a higher concentration of the cleaved substrate. The substrate was diluted with 10× NEBuffer 2.1 (New England Biolabs, B7202) and enzyme-free water to a final concentration of 50 ng / μL, with a final concentration of 1× NEBuffer 2.1.
[0175] Table 3 below shows the primers used in amplification.
[0176] Table 3 dsDNA cleavage substrate PCR amplification primer information
[0177] The sequence information of the dsDNA cleavage substrate HBG is as follows (the underlined bold sequence is the target site, the same below):
[0178] 2. Preparation of crRNA
[0179] Three crRNAs were designed for the HBG gene dsDNA cleavage substrate, each targeting three sites of the HBG gene, and were named HBG-1, HBG-2, and HBG-3, as shown in Table 4. HBG-1, HBG-2, and HBG-3 crRNAs were expressed by MEGAshortscript. TM The RNA was transcribed using the T7 Transcription Kit. The corresponding single-stranded DNA templates corresponded to HBG-T1, HBG-T2, and HBG-T3 (Table 5), respectively, and were synthesized by BGI Liuhe. 2 pmol of double-stranded DNA template and T7 primer (5'-TAATACGACTCACTATAGGG-3', SEQ ID NO: 9) were added and incubated for 12 hours at 37°C using a Bio-rad S1000TM polymerase chain reaction (PCR) instrument. RNA was purified using saturated phenol, chloroform, and isopropanol solutions. Afterwards, the RNA was purified using Qubit TM The RNA HS Assay Kit was used for quantification. The final crRNA sequence information is shown in Table 4.
[0180] Table 4 crRNA sequences used in this example
[0181] Table 5: Sequence information of the crRNA transcription DNA template used in this example
[0182] 3. In vitro cleavage experiment
[0183] To compare the enzymatic activity of BEST5 in vitro, Lba Cas12a (Cpf1) (New England Biolabs, M0653T) was used as a positive control. The protein was diluted with 10× NEBuffer 2.1 and enzyme-free water to a final concentration of 500 nM. The purified crRNA was diluted with 10× NEBuffer 2.1 and enzyme-free water to a final concentration of 500 nM. The final concentration of NEBuffer 2.1 was 1×. 1 μL of diluted Cas protein was mixed with 1 μL of crRNA, 1 μL of 10× NEBuffer 2.1, and 7 μL of enzyme-free water for a total of 10 μL and reacted at 37°C for 15 minutes to form RNPs. Subsequently, 1 μL of 50 ng / μL HBG cleavage substrate was added to a final volume of 10 μL and reacted at 45°C for 10 minutes. 1 μL of 10 mg / ml RNase A was added, and the mixture was incubated at 37°C for 10 min. Product bands were detected by 1.5% agarose gel electrophoresis (Figure 3A). The DNA marker was a 200 bp DNA ladder (Tian Gen, MD115-01). The substrate band intensity was analyzed using the ImageLab software provided with the Bio-lab gel imager (Figure 3B). The reference formula was: quantitative percentage of DNA cleavage = 100 × (1-a) / (a+b+c), where a is the integrated intensity of the undigested product band, and b and c are the integrated intensities of each product band produced by cleavage.
[0184] The results in Figure 3 show that the BEST5 cleavage product fragments have clear bands and the size is in line with expectations. In addition, the editing efficiency at target sites 1 and 2 is higher than that of LbaCas12a, showing an effective activity advantage.
[0185] Example 4: BEST5 system accessory cleavage activity
[0186] The transcription recovery and substrate preparation methods of crRNA are the same as those in Example 3. The purified crRNA and 40U / μL RNase inhibitor were diluted with 10×NEBuffer2.1 and enzyme-free water to a final concentration of 2μM, wherein the final concentration of RNase inhibitor was 4U / μL, and the final concentration of NEBuffer2.1 was 1×. Then, Cas12a protein was diluted with 10×NEBuffer2.1, 50mM DTT and enzyme-free water to a final concentration of 1μM, wherein the final concentrations of NEBuffer2.1 and DTT were 1× and 0.5mM, respectively. Take 1μL of diluted Cas protein and 1μL of crRNA and react in 1×NEBuffer2.1 at 37°C for 15 minutes to form 8μL RNP. Afterwards, 1 μL of 50 ng / μL targeted double-stranded DNA and 1 μL of CRISPR reporter buffer 2.1 (containing the single-stranded DNA reporter sequence FAM-reporter, as prepared in Table 6 below) were added to a final volume of 10 μL. The system was incubated at 45°C for 30 minutes in a qPCR instrument. The detection method was set to FAM fluorescence in the qPCR software. The fluorescence value at the set wavelength was recorded every minute, and a time curve was plotted to reflect the progress of the enzyme digestion reaction (Figure 4). NC is a blank control containing only Cas protein but no crRNA in the system.
[0187] Table 6 CRISPR reporter buffer 2.1 preparation
[0188] Among them, the FAM-reporter sequence is: 5'-FAM-AAAAAA-BHQ1-3'.
[0189] As shown in the figure, BEST5 protein has a certain degree of accessory cleavage ability in the presence of crRNA and target double-stranded DNA, and can exhibit trans-cleavage activity against FAM reporter. After activation by HBG-T2 crRNA and target DNA, the reaction can reach a signal intensity comparable to LbaCas12a in 30 minutes.
[0190] Example 5: Effect of crRNA sequence length on BEST5 cleavage activity
[0191] To determine the optimal sequence length of crRNA, HBG gene dsDNA cleavage substrate was used, targeting three sites, and four crRNAs were designed for each site. The lengths of the target sequences were 18 bp, 20 bp, 23 bp, and 25 bp, respectively. The specific sequences are shown in Table 7 below. The crRNA and substrate dsDNA were prepared as in Example 4, and the enzyme digestion reaction system and fluorescence detection were prepared.
[0192] Table 7
[0193] The FAM fluorescence value after 20 minutes of reaction was detected (Figure 5). The results showed that among the three sites, the enzyme activity was best when the crRNA-DNA target chain binding sequence was 23 bp, followed by 20 bp, 25 bp, and 18 bp.
[0194] Example 6: Activity of the BEST5 system at different temperatures
[0195] A 10 μL enzyme digestion reaction system was prepared in the same manner as in Example 4, and incubated at 37 ° C, 45 ° C, and 60 ° C for 30 minutes, respectively. The Cas12a reaction activity at each temperature was compared by fluorescence value. The results shown in Figure 5 show that BEST5 performed better at 45 ° C than at 37 ° C and basically lost its activity at 60 ° C.
[0196] Example 7: In vitro PAM sequence recognition and verification assay of BEST5
[0197] Since DocMF analysis of the PAM of BEST5 revealed a 5'YYN sequence, a validation experiment was designed. For the HBG gene, gRNAs with PAM sequences of TTN, YTN, YYN, and TYN were designed, as shown in Table 8.
[0198] Table 8 crRNA sequences used in this example
[0199] 10 μL of enzyme digestion reaction system was prepared in the same manner as in Example 3, incubated at 45 ° C for 10 minutes, and the products were detected by 1.5% agarose gel electrophoresis. As shown in Figure 7, BEST5 has the highest cutting efficiency when the PAM sequence is TTN, but at the same time, we found that the target site when PAM is YTN, YYN, and TYN also showed corresponding enzyme digestion product bands, proving that it has a wider PAM recognition site than the conventional LbaCas12a that recognizes 5'TTTN.
[0200] Example 8: Cas12a system mammalian cell genome editing experiment
[0201] 1. Human cell culture
[0202] The human embryonic kidney cell-derived cell line HEK293T was selected for in vivo editing activity testing.
[0203] The culture conditions were: DMEM medium (high glucose, Gibco) containing 10% fetal bovine serum (FBS, Gibco), 1% non-essential amino acids (NEAA, Gibco) and 1% glutamine (GlutaMAX, Gibco), 37°C, 5% CO2 concentration.
[0204] 2. Preparation of targeted editing plasmids in eukaryotic cells
[0205] For editing of HEK293T cells, the endogenous gene AAVS1 (gene bank ID: AC005782.1) was selected for targeted cleavage verification.
[0206] The nucleotide sequence of the targeted region of AAVS1 is as follows (SEQ ID NO: 30):
[0207] For the above genes, based on the T-rich PAM characteristics of the Cas12a family, three targeting sites were designed as shown in Table 9. The gene editing plasmids of the corresponding proteins were designed and synthesized, and the map is shown in Figure 8 A (taking AAVS1-g1 as an example).
[0208] Codon-optimized BEST5 coding sequence (SEQ ID NO: 31):
[0209] The nucleotide sequences were synthesized by Beijing Liuhe BGI Genomics Co., Ltd.
[0210] Table 9 crRNA and target site information
[0211] The following steps are performed to complete intracellular plasmid delivery (transfection), genome editing, and activity verification experiments:
[0212] (1) Extraction of endotoxin-free targeted gene editing plasmid
[0213] A. Take 15 mL of LB liquid medium (sterilized in advance at high temperature and high pressure, room temperature), add 15 μL of 1000X Amp antibiotic (100 mg / mL), use a 10 μL pipette tip to pick up the punctured strain containing the target plasmid (Beijing Liuhe BGI), place it in the medium, and culture at 37°C, 200 rpm, for 12-16 hours;
[0214] B. Centrifuge the cultured and expanded bacterial solution at 8000 rpm for 3 minutes and discard the culture medium.
[0215] C. Use the endotoxin-free plasmid miniprep kit (Tian Gen, DP118) to extract the target plasmid according to the instructions;
[0216] D. After extraction, the DNA concentration was quantified using a NanoDrop2000 / 2000C ultramicro spectrophotometer (Thermofisher) and stored at -20°C.
[0217] (2) Plasmid transfection
[0218] A. One day before transfection, use a pipette to remove the original culture medium from the HEK293T cells (~90% confluence) in the seed plate. Slowly add about 2 mL of 37°C preheated DPBS (Gibco) along the wall to wash the cell surface. Then add 1 mL of preheated digestion solution (TrypLE Express, Gibco) to digest. After about 3 minutes, add an appropriate amount of preheated DMEM medium containing serum to terminate the digestion. Resuspend by pipetting, take a small amount of cell suspension and gently pipette to mix with an equal proportion of trypan blue dye (Solarbio). Take about 20 μL of the mixture and add it to a cell counting plate (Countstar) and use a cell analyzer (Countstar Rigel S2) to count the viable cells. Finally, use a 12-well cell culture plate for plating and culture, with approximately 0.5 to 1×10 cells per well. 6 cells;
[0219] B. Once the cells reach 50% to 70% confluency, transfect the cells with the target plasmid using the Lipofectamine 3000 kit (Invitrogen) according to the manufacturer's instructions (2 μg of plasmid and 2.4 μL of Lipofectamine 3000 Reagent per well). Replace the culture medium as needed 6 hours after transfection.
[0220] C. After transfection, cells need to be cultured for 2-3 days to allow for sufficient gene editing.
[0221] D. After cell culture is complete, calculate the transfection efficiency and harvest the cells. Follow the same cell digestion and resuspension steps as above (using the same reagent amounts as those calculated for the cell culture area). Transfer approximately 20 μL of the cell suspension to a cell counting plate and analyze the plasmid transfection efficiency using the green fluorescence (GFP) channel on a cell analyzer. Transfer the remaining cells to a 1.5 mL centrifuge tube and centrifuge at 12,000 rpm for 1 minute. Remove the supernatant and harvest the cells.
[0222] (3) Identification of genome editing activity
[0223] After harvesting the cells, perform genome extraction and T7E1 enzyme digestion assay to initially detect editing. The steps are as follows:
[0224] A. Genomic DNA extraction: Genomic DNA was extracted using a blood / cell / tissue genomic DNA extraction kit (Tiangen, DP304). The genomic DNA concentration was quantified using Nanodrop and stored at -20°C.
[0225] B. Targeted region PCR: A high-fidelity amplification enzyme (PrimeSTAR GXL DNA Polymerase, Takara) was used to amplify the target site region from genomic DNA. The amplification primers are shown in Table 10 below. All deoxynucleotide sequences used were synthesized at the Shenzhen National Gene Bank Synthesis and Editing Platform. After a clear and single target band was observed under a gel imager by 1% agarose (TAE) gel electrophoresis (7.5 V / cm, 30 min), the gel was cut and the target band was purified and recovered using a PCR purification and gel extraction kit (NucleoSpin Extract, MN). The concentration was measured using a Nanodrop;
[0226] Table 10 PCR amplification primers for the targeted region of the AAVS1 gene
[0227] C. Denaturation and Annealing: Mix the mutant DNA and control reaction system as shown in Table 11 and perform heat denaturation and annealing. The PCR instrument (Bio-rad) settings are shown in Table 12.
[0228] Table 11 Annealing reaction system
[0229] Table 12 Denaturation and annealing conditions
[0230] D. T7E1 digestion: Add 0.3 μL of T7E1 nuclease I (NEB, M0302S) to the reaction mixture from step C, for a total of 20 μL. Incubate at 37°C for 20 min.
[0231] E. Activity detection: After the reaction is completed, add 4 μL 6× Gel Loading Dye (NEB) and detect the bands on agarose gel.
[0232] The agarose gel detection results are shown in Figure 8B, indicating that the new system for gene editing using BEST5 has human cell editing activity, and it has genome editing activity at the three sites of AAVS1g1, g2, and g3.
[0233] Example 9: Amplicon library construction and sequencing to evaluate target site editing efficiency
[0234] For editing of HEK293T cells, two nucleic acid sequences, HBG prompter and BCL11a, were used for targeted cleavage verification. The nucleotide sequence of the HBG prompter target region is as follows (SEQ ID NO: 37):
[0235] The nucleotide sequence of the BCL11a targeting region is as follows (SEQ ID NO: 38)
[0236] For the above sequences, editing plasmids were designed based on PAM features, the cell genome was edited, and the target site region was amplified. The editing efficiency of BEST5 was tested by amplicon library sequencing and compared with the SpCas9 system. Among them, SpCas9 (SpCas9-B1 / B2 / H3) and BEST5 target sites (BEST5-B5.1 / B6.1 / H3.1) are shown in Table 13, and the plasmid maps are shown in Figures 9A-9B. Plasmid construction and sequence synthesis are all from Beijing Liuhe BGI Genomics Co., Ltd., and the sequences are shown in Table 13. Primer3 (https: / / primer3.ut.ee / ) was used to prime the above sites. The primer sequences are shown in Table 14. The primers in the above steps were synthesized by Beijing Liuhe BGI Genomics Co., Ltd.
[0237] Table 13 HBG promoter and BCL11a sequence targeting sites
[0238] (1) DNB library construction and sequencing: The purified amplification products obtained above were used for PCR-free library construction. For detailed procedures, refer to the MGIEasy PCR-Free DNA Library Preparation Reagent Kit instruction manual (1000013453). Sequencing was performed using MGISEQ-2000RS (PE100).
[0239] (2) The above amplicon sequencing data were analyzed using Crispresso2 (https: / / github.com / pinellolab / CRISPResso2), see Figure 10.
[0240] The results in Figure 10 show that the editing efficiency of SpCas9 B1 (excluding transfection efficiency) was 52.58%, SpCas9 B2 was 65.40%, and SpCas9 H3 was 67.41%. The editing efficiency of BEST5-B5.1 (excluding transfection efficiency) was 66.22%, BEST5-B6.1 was 75.14%, and BEST5-H3.1 was 83.38%. BEST5 showed higher editing efficiency than SpCas9.
[0241] Table 14 Primers for amplification of BCL11a and HBG promoter target sites
[0242] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", "some implementation plans" or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0243] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A Cas12a protein, wherein The Cas12a protein has: (1) a protein having the amino acid sequence shown in SEQ ID NO: 1; (2) A protein having at least 80% sequence identity with SEQ ID NO: 1 and having the same or similar biological function as the Cas12a protein.
2. Cas12a protein according to claim 1, wherein, The said having the same or similar biological function refers to having at least one of the following activities: The activity of binding to crRNA, the activity of binding to a specific site of the target sequence under the guidance of crRNA, the activity of endonuclease, the activity of binding to a specific site of the target sequence and cutting nucleic acid under the guidance of crRNA, or the activity of recognizing PAM site.
3. Cas12a protein according to claim 2, wherein, The PAM site is characterized by 5'-YYN-3', wherein Y is C or T, and N is any one of A, G, C and T.
4. A fusion protein, wherein: The fusion protein includes the Cas12a protein according to any one of claims 1 to 3 and other modified parts.
5. An isolated polynucleotide, wherein: The polynucleotide is a polynucleotide encoding the Cas12a protein according to any one of claims 1 to 3 or a polynucleotide encoding the fusion protein according to claim 4.
6. The polynucleotide according to claim 5, wherein The polynucleotide is selected from the polynucleotide sequence shown in SEQ ID NO: 2 or SEQ ID NO:
31.
7. An expression vector, wherein: The expression vector comprises the isolated polynucleotide of claim 5 or 6.
8. The expression vector according to claim 7, wherein The expression vector further comprises a crRNA designed according to the target sequence and the PAM sequence, Wherein, in the expression vector, the Cas12a protein according to any one of claims 1 to 3 encoded by the isolated polynucleotide according to claim 5 or 6 is co-expressed with the crRNA.
9. A CRISPR-Cas system, wherein: The CRISPR-Cas system comprises the Cas12a protein of any one of claims 1 to 3 and at least one crRNA, The crRNA includes a backbone region capable of binding to the Cas12a protein of any one of claims 1 to 3 and a guide sequence capable of targeting a target sequence.
10. A composition, wherein The composition comprises: (i) a protein component selected from: the Cas12a protein of any one of claims 1 to 3 or the fusion protein of claim 4; (ii) a nucleic acid component selected from: crRNA, or a nucleic acid encoding the crRNA; or a precursor crRNA of the crRNA, or a nucleic acid encoding the precursor crRNA; The protein component and the nucleic acid component combine with each other to form a complex, Wherein, the crRNA includes a backbone region capable of binding to the Cas12a protein described in any one of claims 1 to 3 and a guide sequence capable of targeting a target sequence.
11. A host cell, wherein The host cell comprises at least one of the Cas12a protein of any one of claims 1 to 3, or the fusion protein of claim 4, or the isolated polynucleotide of claim 5 or 6, or the expression vector of claim 7 or 8, or the CRISPR-Cas system of claim 9, or the composition of claim 10.
12. The Cas12a protein according to any one of claims 1 to 3, the fusion protein according to claim 4, the isolated polynucleotide according to claim 5 or 6, the expression vector according to claim 7 or 8, the CRISPR-Cas system according to claim 9, the composition according to claim 10, the host cell according to claim 11 in gene targeting, gene editing, gene modification, gene transcription and expression regulation.
13. The Cas12a protein according to any one of claims 1 to 3, the fusion protein according to claim 4, the isolated polynucleotide according to claim 5 or 6, the expression vector according to claim 7 or 8, the CRISPR-Cas system according to claim 9, the composition according to claim 10, and the host cell according to claim 11 are selected from any one or more of the following: Targeting and / or editing target nucleic acid; cleavage of double-stranded DNA, single-stranded DNA or single-stranded RNA; non-specific cleavage and / or degradation of collateral nucleic acid; non-specific cleavage of single-stranded nucleic acid; nucleic acid detection; specific editing of double-stranded nucleic acid; base editing of double-stranded nucleic acid; base editing of single-stranded nucleic acid.
14. Use of the Cas12a protein according to any one of claims 1 to 3, the fusion protein according to claim 4, the isolated polynucleotide according to claim 5 or 6, the expression vector according to claim 7 or 8, the CRISPR-Cas system according to claim 9, the composition according to claim 10, and the host cell according to claim 11 in constructing a CRISPR-Cas12a gene editing system.
15. A method for editing a target nucleic acid, targeting a target nucleic acid or cleaving a target nucleic acid, wherein: The method comprises contacting the target nucleic acid with at least one of the Cas12a protein of any one of claims 1 to 3, the fusion protein of claim 4, the isolated polynucleotide of claim 5 or 6, the expression vector of claim 7 or 8, the CRISPR-Cas system of claim 9, the composition of claim 10, and the host cell of claim 11.
16. A method for cleaving a single-stranded nucleic acid, wherein: The method comprises: The nucleic acid population is contacted with the Cas12a protein and crRNA of any one of claims 1 to 3, wherein the nucleic acid population comprises a target nucleic acid and at least one non-target single-stranded nucleic acid, The crRNA can target the target nucleic acid, and the Cas12a protein cuts the non-target single-stranded nucleic acid. The crRNA includes a backbone region capable of binding to the Cas12a protein of any one of claims 1 to 3 and a guide sequence capable of targeting a target sequence.
17. A kit for gene editing, gene targeting or gene cutting, the kit comprising at least one of the Cas12a protein according to any one of claims 1 to 3, the fusion protein according to claim 4, the isolated polynucleotide according to claim 5 or 6, the expression vector according to claim 7 or 8, the CRISPR-Cas system according to claim 9, the composition according to claim 10, and the host cell according to claim 11.
18. A kit for detecting a target nucleic acid in a sample, the kit comprising: (a) the Cas12a protein of any one of claims 1 to 3, or a nucleic acid encoding the Cas12a protein; (b) crRNA, or a nucleic acid encoding the crRNA, or a precursor crRNA of the crRNA, or a nucleic acid encoding the precursor crRNA; (c) a single-stranded nucleic acid detector that is single-stranded and does not hybridize with the crRNA, The crRNA includes a backbone region capable of binding to the Cas12a protein of any one of claims 1 to 3 and a guide sequence capable of targeting a target sequence.
19. Use of the Cas12a protein according to any one of claims 1 to 3, the fusion protein according to claim 4, the isolated polynucleotide according to claim 5 or 6, the expression vector according to claim 7 or 8, the CRISPR-Cas system according to claim 9, the composition according to claim 10, the host cell according to claim 11 in the preparation of a preparation or a kit, wherein, The preparation or kit is used for: i) Gene or genome editing; ii) target nucleic acid detection; iii) editing a target sequence in a target locus to modify the nucleic acid of an organism; iv) diagnosis and / or treatment of disease; v) Construction of disease models and drug screening.
20. A method for detecting a target nucleic acid in a sample, wherein: The method comprises contacting a sample with a Cas12a protein, crRNA, and a single-stranded nucleic acid detector according to any one of claims 1 to 3, detecting a detectable signal generated by cleavage of the single-stranded nucleic acid detector by the Cas12a protein, thereby detecting a target nucleic acid. wherein the crRNA comprises a backbone region capable of binding to the Cas12a protein of any one of claims 1 to 3 and a guide sequence capable of targeting a target sequence, The single-stranded nucleic acid detector does not hybridize with the crRNA.
Citation Information
Patent Citations
Nucleic acid detection method
CN109837328A
Novel crispr-associated protein and use thereof
CN112567031A
Compositions and methods for genome engineering with cas12a proteins
CN113227367A
Applications of Streptococcus-derived Cas9 nucleases on minimal Adenine-rich PAM targets
US20220162620A1
Applications of recombined streptococcus canis cas9 enzymes for PAM-free DNA modification
WO2022266268A2