CRISPR-Cas12b2 gene editing system and application thereof
By optimizing the Cas12b2 protein and sgRNA in the CRISPR-Cas12b2 system, the problems of excessively long protein length and low editing efficiency in existing systems have been solved, enabling efficient gene editing of bacterial genomes, which is suitable for gene editing applications in non-disease diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-03
AI Technical Summary
Existing CRISPR-Cas9 and CRISPR-Cas12 systems suffer from problems in gene editing, such as excessively long protein lengths leading to difficulties in vector construction and low cell transformation efficiency. Furthermore, the gene editing efficiency of the Cas12 system is relatively low, limiting its widespread application.
A CRISPR-Cas12b2 gene editing system is provided, comprising Cas12b2 protein and sgRNA. The PAM site recognized by Cas12b2 protein is GTA and/or GTG. By optimizing the amino acid sequence and nucleotide sequence, a highly efficient CRISPR-Cas12b2 complex is formed to achieve efficient gene editing.
It achieves highly efficient gene knockout, insertion, and point mutation in bacterial genomes, with an editing efficiency of up to 100%, balancing the needs of small size and high efficiency, and is suitable for gene editing for non-disease diagnosis and treatment purposes.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This application relates to the field of gene editing technology, specifically to the CRISPR-Cas12b2 gene editing system and its applications. Background Technology
[0002] The CRISPR-Cas system, an adaptive immune system of prokaryotes, has been artificially modified into an important gene-editing tool. When applied to gene editing, the CRISPR-Cas system comprises two key components: the Cas protein and sgRNA. The Cas protein is an RNA-guided endonuclease; sgRNA is a strand of ribonucleic acid that forms a complementary base pair with the target site on genomic DNA, guiding the Cas protein to cleave the DNA strand near that site. This facilitates gene editing through in vivo DNA repair or exogenous DNA modification.
[0003] Two major categories and six subcategories of CRISPR-Cas systems have been discovered in nature. The first major category consists of a complex of multiple proteins as the effector for defensive cleavage, while the second major category consists of a single protein as the effector, making them more convenient to use and widely applied in gene editing. The second major category includes three subcategories: Type II, Type V, and Type VI. Type II, namely the CRISPR-Cas9 system, is the most classic gene editing tool. Type V, namely the CRISPR-Cas12 system, uses Cas12 as its effector. In recent years, CRISPR-Cas12 systems such as Cas12a, Cas12b, Cas12f, Cas12d, Cas12e, and Cas12f have been discovered, validating their gene editing capabilities.
[0004] However, existing gene editing tools still have problems. The effector Cas9 of the CRISPR-Cas9 system is approximately 1400 amino acids long, which is not conducive to expression vector construction and significantly reduces cell transformation and delivery efficiency. Although the effector Cas12 of the CRISPR-Cas12 system is generally smaller in size than Cas9, its wild-type gene editing efficiency is generally significantly lower than that of Cas9, limiting its widespread application. Summary of the Invention
[0005] Based on this, this application provides a CRISPR-Cas12b2 gene editing system. The CRISPR-Cas12b2 system can achieve highly efficient gene editing in cells.
[0006] In a first aspect of this application, a Cas12b2 protein is provided, wherein the protospacer adjacent sequence (PAM site) it recognizes includes GTA and / or GTG, and is an effector of the CRISPR-Cas12b2 gene editing system, selected from the following group:
[0007] (a) A polypeptide having an amino acid fragment with the sequence shown in SEQ ID NO:1;
[0008] (b) A polypeptide having the biological function of (a) with ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% homology with (a);
[0009] (c) A derivative polypeptide formed by substituting, deleting or adding one or more amino acid residues of (a) and retaining the biological function of (a); the derivative polypeptide comprises an amino acid fragment with an amino acid sequence as shown in SEQ ID NO:2.
[0010] In a second aspect of this application, an sgRNA is provided, obtained through transcriptome sequencing and bioinformatics analysis, which is capable of specifically binding to the Cas12b2 protein as described in the first aspect. The sgRNA is selected from the following group:
[0011] (a) RNA having the nucleotide sequence shown in SEQ ID NO:3;
[0012] (b) RNA having a biological function of (a) with ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% homology to (a);
[0013] (c) A derivative RNA formed by substituting, deleting or adding one or more nucleotides of (a) and retaining the biological function of (a).
[0014] A third aspect of this application provides protein variants, including at least one of truncated proteins and fusion proteins:
[0015] (a) The truncated protein is a polypeptide composed of one or more functional domains of the Cas12b2 protein described in the first aspect;
[0016] (b) The fusion protein comprises the Cas12b2 protein described in the first aspect, and one or more functional domains.
[0017] The fourth aspect of this application provides isolated polynucleotides encoding the Cas12b2 protein of the first aspect, the sgRNA of the second aspect, or the protein variant of the third aspect.
[0018] The fifth aspect of this application provides a CRISPR-Cas12b2 complex comprising:
[0019] (i) Protein components selected from the group consisting of: the Cas12b2 protein described in the first aspect, the protein variants described in the third aspect, and combinations thereof; and
[0020] (ii) Nucleic acid components selected from the group consisting of: sgRNA as described in the second aspect, nucleic acids encoding sgRNA as described in the second aspect, precursor RNA of sgRNA as described in the second aspect, nucleic acids encoding precursor RNA of sgRNA as described in the second aspect, and combinations thereof;
[0021] The protein component and the nucleic acid component combine to form a complex.
[0022] The sixth aspect of this application provides an activated CRISPR-Cas12b2 complex comprising:
[0023] (i) Protein components selected from the group consisting of the Cas12b2 protein described in the first aspect, protein variants as described in the third aspect, and combinations thereof;
[0024] (ii) Nucleic acid components selected from the group consisting of: sgRNA as described in the second aspect, nucleic acids encoding sgRNA as described in the second aspect, precursor RNA of sgRNA as described in the second aspect, nucleic acids encoding precursor RNA of sgRNA as described in the second aspect, and combinations thereof; and,
[0025] (iii) The target sequence that binds to the sgRNA.
[0026] A seventh aspect of this application provides a CRISPR-Cas12b2 system comprising one or more vectors, said one or more vectors comprising:
[0027] (i) a first nucleic acid comprising the isolated polynucleotides described in the fourth aspect; and
[0028] (ii) a second nucleic acid, which contains a nucleotide sequence encoding the sgRNA described in the second aspect;
[0029] in:
[0030] The first nucleic acid and the second nucleic acid may exist on the same or different vectors.
[0031] The eighth aspect of this application provides a recombinant vector comprising:
[0032] The isolated polynucleotides as described in the fourth aspect;
[0033] Alternatively, polynucleotides encoding sgRNA as described in the second aspect.
[0034] In some embodiments, the backbone of the recombinant vector is selected from pUC19 and pBR322.
[0035] In some embodiments, the recombinant vector further includes a T7 promoter, a pJ23119 promoter, or a PvanP promoter.
[0036] In some embodiments, the backbone of the recombinant vector includes, but is not limited to, pACYC, pBAD, pET28a, and pHT08.
[0037] A ninth aspect of this application provides an engineered host cell that is not an animal or plant species, the host cell comprising one or more of the following: the Cas12b2 protein as described in the first aspect, the protein variant as described in the third aspect, the sgRNA as described in the second aspect, the isolated polynucleotide as described in the fourth aspect, the CRISPR-Cas12b2 complex as described in the fifth aspect, the activated CRISPR-Cas12b2 complex as described in the sixth aspect, the CRISPR-Cas12b2 system as described in the seventh aspect, and the recombinant vector as described in the eighth aspect.
[0038] A tenth aspect of this application provides the use of the CRISPR-Cas12b2 complex as described in the sixth aspect, the activated CRISPR-Cas12b2 complex as described in the seventh aspect, the CRISPR-Cas12b2 system as described in the eighth aspect, or the engineered host cell as described in the ninth aspect, said use including:
[0039] (1) Gene editing, gene targeting, or gene cutting for purposes other than disease diagnosis and treatment;
[0040] (2) Use in the preparation of reagents or kits; wherein the preparation or kit is used for:
[0041] (i) Gene or genome editing, including at least one of gene knockout, insertion, point mutation and substitution;
[0042] (ii) Target nucleic acid detection and / or diagnosis;
[0043] (iii) Editing target sequences in target loci to modify organisms;
[0044] (iv) Treatment of the disease;
[0045] (v) Targeting the target gene;
[0046] (vi) Cut the target gene.
[0047] The eleventh aspect of this application provides a method for editing or cleaving a target nucleic acid for purposes other than disease diagnosis and treatment, the method comprising contacting the target nucleic acid with a CRISPR-Cas12b2 complex as described in the fifth aspect, an activated CRISPR-Cas12b2 complex as described in the sixth aspect, or an engineered host cell as described in the ninth aspect.
[0048] The twelfth aspect of this application provides a kit for gene editing, gene targeting, or gene cutting, the kit comprising one or more of the following: the Cas12b2 protein as described in the first aspect, the protein variant as described in the third aspect, the isolated polynucleotide as described in the fourth aspect, the sgRNA as described in the second aspect, the recombinant vector as described in the eighth aspect, the CRISPR-Cas12b2 complex as described in the fifth aspect, the activated CRISPR-Cas12b2 complex as described in the sixth aspect, the CRISPR-Cas12b2 system as described in the seventh aspect, and the engineered host cell as described in the ninth aspect.
[0049] This application provides a novel CRISPR-Cas12b2 gene editing system. It is a newly discovered and characterized CRISPR-Cas system, differing from previously reported CRISPR-Cas systems in sequence characteristics and evolutionary relationships. Cas12b2 belongs to the same Cas12b subclass as the previously reported Cas12b1. However, compared to Cas12b1, the amino acid sequence similarity is generally less than 30%, and the lengths of each domain are significantly shorter, classifying it as a distinct subtype of Cas12b. The Cas12b2 sequence length is 692 aa, nearly 50% shorter than SpCas9 (1368 aa) and 39% shorter than AacCas12b1 (1129 aa).
[0050] The beneficial effects of this application include that, compared with existing technologies, the amino acid length of this Cas12b2 protein is only 692 aa, which is nearly 50% shorter than other common Type II Cas proteins (such as SpCas9 protein, 1368 aa), making vector construction easier and improving cell transformation efficiency. This CRISPR-Cas12b2 system can achieve one or more types of gene editing near the target DNA sequence in the cell genome, including gene knockout, insertion, point mutation, and substitution. For gene knockout, it can achieve gene knockout of hundreds to 15,000 base pairs in the bacterial genome, with editing efficiency reaching 100%, and point mutation efficiency reaching 90%, balancing the requirements of small size and high efficiency. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments and examples of this application, and to more completely understand this application and its beneficial effects, the drawings used in the description of the embodiments or examples will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of this application. Those skilled in the art can obtain other drawings based on these drawings without creative effort. It should also be noted that the drawings are all drawn in a simplified form and are only used to conveniently and clearly assist in illustrating this application.
[0052] Figure 1 This is a phylogenetic tree of the Cas12b2 protein in this application.
[0053] Figure 2 This application compares the component composition and arrangement of the Cas12b2 protein with those of other CRISPR-Cas systems.
[0054] Figure 3 This is a comparison of the domains of the Cas12b2 protein and the Cas12b1 protein in this application.
[0055] Figure 4 The PAM site recognized by the Cas12b2 protein in this application.
[0056] Figure 5 This application compares the PAM sites of the Cas12b2 protein with those of other Cas proteins.
[0057] Figure 6 RNA-seq was used to determine the tracRNA and crRNA sequences of the Cas12b2 protein in this application.
[0058] Figure 7 This is an in vitro enzyme cleavage activity test for the CRISPR-Cas12b2-sgRNA complex in this application.
[0059] Figure 8 This is a schematic diagram illustrating the principle of applying the Cas12b2-sgRNA complex to gene editing in one embodiment of this application.
[0060] Figure 9 This is a schematic diagram of the gene map of plasmid pEcas12b2 in one embodiment of this application.
[0061] Figure 10 This is a schematic diagram of the gene map of plasmid pEcgb2 in one embodiment of this application.
[0062] Figure 11 This is a schematic diagram of the gene map of the Escherichia coli MG1655 envc gene editing plasmid pEcgb2-envc in one embodiment of this application.
[0063] Figure 12 This is a schematic diagram of the gene map of the Escherichia coli MG1655 envc gene editing plasmid pEcgb2-HA-envc in one embodiment of this application.
[0064] Figure 13 This is an agarose gel electrophoresis pattern of Escherichia coli MG1655 envc gene knockout verified by colony PCR in one embodiment of this application.
[0065] Figure 14 This is an agarose gel electrophoresis pattern of Escherichia coli BL21(DE3) dgka gene knockout verified by colony PCR in one embodiment of this application.
[0066] Figure 15 This is a sequencing result of the Escherichia coli BL21(DE3) kpLE2 gene fragment knocked out in one embodiment of this application.
[0067] Figure 16 This is a sequencing result of the point-mutated Escherichia coli BL21(DE3) cadA gene in one embodiment of this application.
[0068] Figure 17 This is a sequencing result of the point-mutated Escherichia coli BL21(DE3) gadc gene in one embodiment of this application. Detailed Implementation
[0069] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings. Preferred embodiments of this application are shown in the drawings. However, this application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of this application.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0071] In this application, unless otherwise specified, "one or more" means any one of the listed items or any combination of the listed items. Similarly, "one or more" and other instances that otherwise indicate "one or more" shall be understood in the same way unless otherwise specified.
[0072] The terms “combinations thereof,” “any combination thereof,” and “any combination thereof” as used in this application include all suitable combinations of any two or more of the listed items.
[0073] In this application, the word "suitable" in "suitable combination", "suitable method", "any suitable method" etc., shall be defined as being able to implement the technical solution of this application, solve the technical problem of this application, and achieve the expected technical effect of this application.
[0074] In this application, terms such as "further," "even more," "particularly," "for example," "like," "example," and "exemplary" are used for descriptive purposes to indicate that different technical solutions preceding and following each other are related in terms of their coverage, but should not be construed as limiting the preceding technical solution or restricting the scope of protection of this application. In this application, unless otherwise specified, A (e.g., B) indicates that B is a non-limiting example of A, and it can be understood that A is not limited to B.
[0075] In this application, "optionally," "optionally," and "optional" mean that something is optional, that is, it refers to either "with" or "without" a parallel solution. If multiple "options" appear in a technical solution, unless otherwise specified and there are no contradictions or mutual constraints, each "option" is independent. Unless otherwise specified, the descriptions such as "optionally include" and "optionally contain" in this application, taking "optionally include" as an example, mean "may include or not include."
[0076] The terms “containing,” “comprising,” and “including” as used in this application are synonyms and are inclusive or open-ended, not excluding additional, uncited members or features. Members or features include, for example, materials or components, structures, elements, instruments, etc.; non-limiting examples of members or features include actions, conditions under which actions occur, timing, states, etc.
[0077] In this application, the technical features or solutions described in open-ended language include both closed-ended technical features or solutions consisting of the listed contents and open-ended technical features or solutions that include the listed contents.
[0078] In this application, the exemplary descriptions such as "in some implementations (or embodiments)" and "in one implementation (or embodiment)" may cover, but are not limited to, the following meanings: these solutions can be combined with other solutions in a suitable manner to form new technical solutions.
[0079] In this application, the terms "first aspect," "second aspect," "third aspect," "fourth aspect," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or quantity, nor should they be construed as implicitly indicating the importance or quantity of the indicated technical features. Moreover, "first," "second," "third," "fourth," etc., serve only a non-exhaustive enumeration purpose and should be understood not to constitute a closed limitation on quantity.
[0080] In this application, when numerical intervals (i.e., numerical ranges) are involved, unless otherwise specified, the distribution of selectable numerical values within the numerical interval is considered continuous, and includes the two endpoints of the numerical interval (i.e., the minimum and maximum values), as well as every numerical value between these two endpoints. Unless otherwise specified, when a numerical interval refers only to integers within that numerical interval, it includes the two endpoint integers of the numerical range, as well as every integer between the two endpoints, which is equivalent to directly listing every integer. When multiple numerical ranges are provided to describe features or characteristics, these numerical ranges can be merged. In other words, unless otherwise specified, the numerical ranges disclosed herein should be understood to include any and all subranges included therein. The "numerical value" in the numerical interval can be any quantitative value, such as a number, percentage, ratio, etc. The term "numerical interval" can be broadly included to include numerical interval types such as percentage intervals, ratio intervals, and proportion intervals.
[0081] Unless otherwise specified, the term "selected from the group below" in this application may include one or more of them.
[0082] In this application, where the method flow involves multiple steps, unless otherwise explicitly stated herein, there is no strict order restriction on the execution of these steps; they can be executed in any order other than those described. Moreover, any step may include multiple sub-steps or multiple stages, which are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or simultaneously with other steps or parts of the sub-steps or stages of other steps.
[0083] Unless otherwise specified, the term "Cas protein" in this application may be used interchangeably with Cas enzyme and Cas effector protein. Cas protein is used in its broadest sense, including wild-type Cas protein, its derivatives or variants, analogs, and its functional fragments such as oligonucleotide-binding fragments.
[0084] Unless otherwise specified, the term "variant" in this application may be used interchangeably with "derivative" or "analyte" and refers to a polypeptide that substantially retains the function or activity of a Cas protein (e.g., Cas12b2 protein).
[0085] Unless otherwise specified, the term "CRISPR-Cas12b2 complex" in this application refers to a complex formed by the binding of sgRNA and Cas12b2 protein, which contains a guide sequence that hybridizes to the target sequence and binds to the Cas12b2 protein, and the complex is capable of recognizing and cleaving target nucleotides that can hybridize with the guide RNA or mature crRNA.
[0086] Unless otherwise specified, the term "target nucleic acid" in this application may be used interchangeably with "target sequence," "target nucleic acid sequence," or "target nucleic acid molecule," referring to a specific nucleic acid containing a nucleic acid sequence that is wholly or partially complementary to the spacer sequence in the guide RNA. "Target sequence" refers to a polynucleotide targeted by the spacer sequence in the guide RNA, such as a sequence complementary to that spacer sequence, wherein hybridization between the target sequence and the spacer sequence will promote the formation of a CRISPR-Cas12b2 complex (including the Cas12b2 protein and sgRNA). Complete complementarity is not required, as long as sufficient complementarity exists to induce hybridization and promote the formation of a CRISPR-Cas12b2 complex. In some embodiments, the target nucleic acid contains a non-coding region (e.g., a promoter or terminator). In some embodiments, the target nucleic acid is single-stranded or double-stranded.
[0087] Unless otherwise specified, the term "vector" in this application refers to a nucleic acid molecule capable of delivering another nucleic acid molecule linked thereto. Vectors include, but are not limited to, single-stranded, double-stranded, or partially double-stranded nucleic acid molecules; nucleic acid molecules including one or more free ends, or without free ends (e.g., circular); nucleic acid molecules including DNA, RNA, or both; and a wide variety of other polynucleotides known in the art. A vector can be introduced into a host cell through transformation, transduction, or transfection, thereby enabling the expression of its carried genetic material elements in the host cell. A vector can be introduced into a host cell to produce transcripts, proteins, or peptides, including proteins, fusion proteins, isolated nucleic acid molecules, etc., as described herein (e.g., CRISPR transcripts, such as nucleic acid transcripts, proteins, or enzymes). A vector may contain a variety of elements controlling expression, including, but not limited to, promoter sequences, transcription initiation sequences, enhancer sequences, selection elements, and reporter genes. The vector may also contain a replication initiation site.
[0088] Vectors include plasmids and viral vectors. A plasmid is a circular double-stranded DNA loop in which another DNA fragment can be inserted, for example, using standard molecular cloning techniques. A viral vector contains a virus-derived DNA or RNA sequence within a vector used to package the virus; viruses include, for example, retroviruses, replication-defective retroviruses, adenoviruses, replication-defective adenoviruses, and adeno-associated viruses. Viral vectors also contain polynucleotides carried by a virus intended for transfection into a host cell. Some vectors (e.g., bacterial vectors with bacterial origins of replication and augmented mammalian vectors) are capable of autonomous replication in the host cells into which they are introduced.
[0089] Other vectors (e.g., non-attachment mammalian vectors) integrate into the host cell's genome after introduction and thereby replicate along with the host genome. Furthermore, some vectors can direct the expression of genes they are operatively linked to. Such vectors are called "expression vectors."
[0090] In some embodiments, the vector (e.g., a viral vector or a non-viral vector, such as a lentiviral vector or plasmid) can be delivered to the target tissue via, for example, intramuscular injection, intravenous administration, percutaneous administration, intranasal administration, oral administration, or mucosal administration. The delivery can be performed via a single dose or multiple doses. Those skilled in the art will understand that the actual dose to be delivered herein can vary considerably depending on a variety of factors, including but not limited to the choice of vector, target cells, organism, tissue, general condition of the subject to be treated, the degree of transformation / modification sought, the route of administration, the manner of administration, and the type of transformation / modification sought.
[0091] Unless otherwise specified, the term "promoter" in this application refers to a non-coding nucleotide sequence located upstream of a gene that initiates the expression of a downstream gene. A constitutive promoter is a nucleotide sequence that, when operatively linked to a polynucleotide encoding or defining a gene product, will result in the production of that gene product in the cell under most or all physiological conditions. An inducible promoter is a promoter that selectively expresses a coding sequence or functional RNA in response to the presence of endogenous or exogenous stimuli, such as by a chemical compound (chemical inducer), or in response to environmental, hormone, chemical, and / or developmental signals. Inducible or regulatory promoters include promoters induced or regulated, for example, by light, heat, stress, flooding or drought, salt stress, osmotic stress, plant hormones, wounds, or chemicals (such as ethanol, abscisic acid (ABA), jasmonic acid esters, salicylic acid, or safeners).
[0092] Unless otherwise specified, the term "host cell" in this application refers to eukaryotic cells (e.g., animal cells, plant cells, fungal cells, etc.), prokaryotic cells (e.g., some microbial cells, Escherichia coli, Bacillus subtilis, etc.), or cells derived from multicellular organisms (e.g., cell lines) cultured in the form of single-celled entities, said cells serving as recipients of nucleic acids (e.g., expression vectors), and includes the offspring of the original cells that have been genetically modified with nucleic acids.
[0093] Unless otherwise specified, the term "delivery" in this application refers to a destination-providing entity. For example, components of the CRISPR-Cas12b2 system / complex of this application may be delivered in various forms, such as DNA / RNA or RNA / RNA or a combination of protein and RNA. For example, the Cas12b2 protein may be delivered as a polynucleotide encoding DNA or a polynucleotide encoding RNA or as a protein.
[0094] Unless otherwise specified, the term "complementarity" in this application refers to the ability of one nucleic acid sequence to form one or more hydrogen bonds with another nucleic acid sequence via conventional Watson-Crick or other non-conventional types. The complementarity percentage represents the percentage of residues in a nucleic acid molecule that can form hydrogen bonds (e.g., Watson-Crick base pairing) with another nucleic acid sequence (e.g., if 5, 6, 7, 8, 9, or 10 out of 10 are complementary, the complementarity percentages are 50%, 60%, 70%, 80%, 90%, and 100%). "Complete complementarity" means that all consecutive residues in one nucleic acid sequence form hydrogen bonds with the same number of consecutive residues in another nucleic acid sequence. "Substantially complementary" refers to a complementarity of at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, 99%, or 100% in a region of 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, 50 or more nucleotides, or to two nucleic acids that hybridize under strict conditions.
[0095] Unless otherwise specified, the term "strict condition" in this application refers to conditions under which a nucleic acid complementary to the target sequence hybridizes primarily with the target sequence and substantially does not hybridize to non-target sequences. Strict conditions are typically sequence-dependent and depend on many factors. Generally, the longer the sequence, the higher the temperature at which it specifically hybridizes to its target sequence.
[0096] Unless otherwise specified, the term "hybridization" in this application refers to a reaction in which one or more polynucleotides react to form a complex that is stabilized by hydrogen bonding between the bases of those nucleotide residues. The complex may comprise two strands forming a duplex, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination thereof. Hybridization reactions can constitute a step in a broader process, such as the initiation of PCR or the cleavage of a polynucleotide by an enzyme. A sequence capable of hybridizing with a given sequence is referred to as the "complement" of that given sequence.
[0097] This application discloses a newly discovered CRISPR-Cas12b2 gene editing system that enables precise gene editing in cells. The CRISPR-Cas12b2 system differs from other reported CRISPR-Cas systems in sequence characteristics and evolutionary relationships, representing a novel type of CRISPR-Cas system. This gene editing system comprises a cluster of regularly spaced short palindromic repeats (CRISPR)-Cas12b2 complexes and guide RNA, namely the Cas12b2 protein and sgRNA. This application also provides a gene editing method based on the CRISPR-Cas12b2 complex. This method can achieve one or more types of gene editing, such as gene knockout, insertion, point mutation, and substitution, near the target DNA sequence in the cellular genome, achieving a gene editing efficiency of up to 100% in *E. coli*. The technical solutions in this application will be further described in detail below.
[0098] This application provides an exemplary Cas12b2 protein.
[0099] The PAM sites recognized by the Cas12b2 protein include GTA and / or GTG (e.g., GTA or GTG), which are effectors of the CRISPR-Cas12b2 gene editing system.
[0100] In some embodiments, the Cas12b2 protein is selected from the group consisting of:
[0101] (a) A polypeptide having an amino acid fragment with the sequence shown in SEQ ID NO:1;
[0102] (b) A polypeptide having the biological function of (a) with ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% homology with (a);
[0103] (c) A derivative polypeptide formed by substituting, deleting or adding one or more amino acid residues of (a) and retaining the biological function of (a).
[0104] (d)dCas12b2, namely a polypeptide with an amino acid fragment as shown in SEQ ID NO:2, is a derivative polypeptide with expression regulation function formed by the substitution of one or more amino acid residues at the key catalytic site of Cas12b2.
[0105] SEQ ID NO: 1 (692aa):
[0106] MSVKSFQAKVVCDTPEKREYLWLTHRLFNQGVCELLPYLFKMRRGDLGHEFKTIYEAIRNSQNSFAKLEPITTTKAWSSKHVGTGDPKNQWAVLCAKLNASGRILFDRDKPPFAYASEFWRKICEMAVQLMHSHDGLFKDWREERAAWKERKADWESNHQLYMKAKPTLDAFQEEAGRLSGSRKRWLLYLDFLAKHPELAAWHGGKAHVDPLTAQERKGCRRPGDHFDIFWDKNPELAALNALDRTWRREFANFKRRPTWTNPSPEKHPAWYSFKRGATYKDLDLATGTLRLRVLTGDDGKGRRGEWHAYTFQADNRLRRLRPAPESVKIGRNSFSWLYADPFLGIDRPAEVRGIKLVFRAKRPYLLFSVDIADEQLARLTRPGFKDTESGKTATPADIPDGTRFLAVDFGQRTLGACSVCAFKDGEPLPPESVFLLRLPGLSFSDIGRHESTIRRRTSKMYRSGPRRRRTHHAPRGGTTFADQRHHVLKMKEDRYKKAANLLIKAALRHRASVILVENLRNYRPDLERPARENRARMQWNVQRIIEFLDKTAKPLGIRLWRVSPWYSSQFCSACGHPGKRFSIPRKTRWEHFYARRHGPVRKPVIEPGGQFFVCSNPDCPRPTGIIHADVNASLNLHRILAETFERPQGKGKQKTWQGQPLNWKPITDQCRARIETHFLANKADLAAQTPW。
[0107] SEQ ID NO: 2 (692aa):
[0108] MSVKSFQAKVVCDTPEKREYLWLTHRLFNQGVCELLPYLFKMRRGDLGHEFKTIYEAIRNSQNSFAKLEPITTTKAWSSKHVGTGDPKNQWAVLCAKLNASG RILFDRDKPPFAYASEFWRKICEMAVQLMHSHDGLFKDWREERAAWKERKADWESNHQLYMKAKPTLDAFQEEAGRLSGSRKRWLLYLDFLAKHPELAAWHG GKAHVDPLTAQERKGCRRPGDHFDIFWDKNPELAALNALDRTWRREFANFKRRPTWTNPSPEKHPAWYSFKRGATYKDLDLATGTLRLRVLTGDDGKGRRGEWHAYTFQADNRLRRLRPAPESVKIGRNSFSWLYADPFLGIDRPAEVRGIKLVFRAKRPYLLFSVDIADEQLARLTRPGFKDTESGKTATPADIPDGTRFLAV X FGQRTLGACSVCAFKDGEPLPPESVFLLRLPGLSFSDIGRHESTIRRRTSKMYRSGPRRRRTHHAPRGGTTTFADQRHHVLKMKEDRYKKAANLLIKAALRHRASVILV X NLRNYRPDLERPARENRARMQWNVQRIIEFLDKTAKPLGIRLWRVSPWYSSQFCSACGHPGKRFSIPRKTRWEHFYARRHGPVRKPVIEPGGQFFVCSNPDCPRPTGIIHA X VNASLNLHRILAETFERPQGKGKQKTWQGQPLNWKPITDQCRARIETHFLANKADLAAQTPW.
[0109] The Cas12b2 protein in this application belongs to the newly discovered CRISPR-Cas12b2 system. This CRISPR-Cas12b2 system differs from other reported CRISPR-Cas systems in sequence characteristics and evolutionary relationships. Cas12b2 and Cas12b1 both belong to the Cas12b subclass. However, compared to Cas12b1, their amino acid sequence similarity is generally less than 20%, classifying them as a different subtype of Cas12b. The Cas12b2 sequence is 692 aa in length, nearly 50% shorter than SpCas9 (1368 aa).
[0110] Cas12b2, a polypeptide having an amino acid fragment as shown in SEQ ID NO: 1, can form a complex with the sgRNA of this application, target specific DNA sequences on the genome, perform DNA endonuclease activity, and exert gene editing function.
[0111] dCas12b2, a polypeptide with the amino acid sequence shown in SEQ ID NO: 2, is a derivative polypeptide with expression regulatory function formed by substituting one or more amino acid residues at the key catalytic site of Cas12b2. The mutation sites of dCas12b2 are the active sites D409, E518, D630, or combinations thereof in the RuvC domain of Cas12b2, and the mutation mode is to mutate to any other arbitrary amino acid. The mutated dCas12b2 can still form a complex with the sgRNA of this application, targeting specific DNA sequences on the genome, but it no longer has DNA endonuclease activity. Instead, it regulates gene expression intensity by binding to DNA and inhibiting gene transcription and translation.
[0112] In some embodiments, the Cas12b2 protein is derived from Thioglobus bacterium. When applying it to a specific host for gene editing, the nucleotide sequence should be codon-optimized based on the host codon usage frequency.
[0113] In some embodiments, the Cas12b2 protein is generated by an endogenous constitutive promoter upstream of the gene in its original host, the T7 promoter, and the tetracycline-induced promoter P. tet Xylose-induced promoter P xylA mannose-induced promoter P manP One of the control expressions.
[0114] In some embodiments, the sgRNA in this application has the nucleotide sequence shown in SEQ ID NO: 3.
[0115] The sgRNA consists of CRISPR RNA (crRNA), trans-activating CRISPR RNA (tracrRNA), and a designed linker sequence (loop). The crRNA contains a repeat sequence (DR) and a spacer sequence. More specifically, the sgRNA sequence, from 5' to 3', is tracrRNA, loop, DR, and spacer. The tracrRNA binds to the Cas12b2 protein and is linked to the crRNA via the loop sequence. The DR sequence in the crRNA forms a base pair with the anti-DR sequence of the tracrRNA; the spacer sequence forms a base pair with the target DNA sequence on the genome, guiding the CRISPR-Cas12 complex to locate the target DNA region on the genome.
[0116] The sgRNAs used in this application include both primitive sgRNAs and engineered sgRNAs.
[0117] The crRNA and tracrRNA sequences of the original sgRNA were identified from the upstream and downstream regions of the Cas12b2 gene on the original host genome of CRISPR-Cas12b2. The engineered sgRNA was engineered based on the sequence of the original sgRNA.
[0118] In some embodiments, the stem-loop region of the tracrRNA portion of the engineered sgRNA sequence backbone and the base complement sequence of the repeat-inverted repeat sequence are deleted for a length of 7nt, 10nt, 17nt, 18nt, 22nt, 26nt, 33nt, or any range or value between any two of these values.
[0119] Using the engineered sgRNA described in the above embodiments, the gene editing efficiency of the CRISPR / Cas12b2 complex is comparable to that using the original sgRNA.
[0120] In some embodiments, the nucleotide sequence of the artificially designed linker sequence loop is at least one of AAGG, AACC, GGAA, and CCAA.
[0121] In some implementations, the length of the interval sequence is 18nt, 20nt, 22nt, 24nt, 26nt, 28nt, 30nt, or a range or value between any two of these values.
[0122] The specific spacer sequence of the sgRNA needs to be designed and modified based on the gene editing site sequence.
[0123] SEQ ID NO:3:
[0124] GGCTACGGCCCACCGGGCGCTATAGGCCCCCGGTGCTGCGTCGGGCGAGCTGAGTGTCGCCGACCGGTTGCCGAAAGGCTTTGCCGAGTAGGGTCGCCCCATCCCCTTGCGGGATGACGGCCTCCGGTTAACCCCGACGCAGAGAGATCCTTTTCTGAT (tracrRNA) AAGG (loop) ATCAGAAAAGGACCTCTCTGGACAC (DR)NNNNNNNNNNNNNNNN (spacer).
[0125] The gene editing method based on the CRISPR-Cas12b2 complex in this application includes:
[0126] (1) Design sgRNA sequences that can target DNA at the gene editing target site.
[0127] (2) The Cas12b2 protein and the sgRNA were constructed into a gene editing vector.
[0128] (3) Deliver the gene editing vector into the cells.
[0129] (4) Verify the correctness of gene editing by polymerase chain reaction (PCR) and gene sequencing.
[0130] In some embodiments, the host for performing gene editing includes at least one of Escherichia coli, Bacillus subtilis, Corynebacterium glutamicum, and yeast.
[0131] In this application, the types of gene editing include one or more of gene knockout, gene insertion, point mutation, and gene replacement.
[0132] In some implementations, the gene knockout length is 100bp to 15000bp.
[0133] In some embodiments, the length of the gene insertion is 10bp to 5000bp.
[0134] In some embodiments, the point mutation is a mutation between any two of the four bases A, T, G, and C.
[0135] In some embodiments, the length of the gene replacement is 10bp to 15000bp.
[0136] In this application, designing the sgRNA sequence refers to designing the spacer sequence in the sgRNA.
[0137] In some implementations, the genomic DNA sequence that is complementary to the sgRNA spacer sequence, i.e., the protospacer adjacent sequence (PAM sequence), is one or more of NNAGTA, NNCGTA, NNGGTA, NNTTTG, and NNTTTC.
[0138] In some embodiments, after determining the gene editing target site, the sgRNA sequence is designed using CHOPCHOP (https: / / chopchop.cbu.uib.no). The sgRNA length is set to the length of the spacer sequence in the aforementioned embodiments. The 5'-PAM is set to the PAM sequence in the aforementioned embodiments.
[0139] In some embodiments, the expression vectors for the Cas12b2 protein and the sgRNA include, but are not limited to, plasmids or viruses.
[0140] In some embodiments, the plasmid expression vector for Cas12b2 is pACYC, pBAD, pET28a, or pHT08.
[0141] In some embodiments, the plasmid expression vector for the sgRNA is pUC19 or pBR322.
[0142] In some embodiments, the sgRNA is expressed under the control of the T7 promoter, pJ23119 promoter, and PvanP promoter.
[0143] In some embodiments, the method of delivering the gene-editing vector into cells includes chemical transformation or electroporation transformation.
[0144] In some implementations, achieving precise gene editing in cells requires the provision of a homology repair template.
[0145] In some implementations, the homology repair template is upstream and downstream DNA of the gene editing target site.
[0146] Optionally, the homology repair template can be inserted into the expression vector of the sgRNA or delivered to the cell in the form of linear DNA.
[0147] In some implementations, the success rate of gene editing can reach 80% to 100% by verifying the correctness of gene editing through polymerase chain reaction (PCR) and gene sequencing.
[0148] The following are some examples.
[0149] The embodiments of this application will be described in detail below with reference to examples. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of this application. For experimental methods in the following embodiments where conditions are not specified, reference should be made to the guidelines given in this application, or to experimental manuals or conventional conditions in the art, or to the conditions recommended by the manufacturer, or to experimental methods known in the art.
[0150] Example 1: Mining and Identification of the CRISPR-Cas12b2 System
[0151] 1. Using the Cas12b amino acid sequences T0D7A2, A0A939A6Q9, A0A971PHS7, A0A971QJY0, A0A7T5ELD8, A0A936LT08, A0A975EKD8, and A0A1G3LDQ4 from the UniProt database as seed sequences, the RuvC domains of the above sequences were annotated using PF18516 (RuvC nuclease domain) from the InterProPfam database, and a new implicit Markov model was generated through multiple sequence alignment.
[0152] 2. CRISPR arrays on genomes were identified using CRISPRCasFinder in databases such as the National Center for Biotechnology Information (NCBI), the Marine Microbial Gene Database (GOMC), and the Earth Microbiome (GEM).
[0153] 3. Extract protein sequence information from the genome within a 20,000 bp region near the aforementioned CRISPR array. Based on the constructed hidden Markov model, perform protein conserved sequence alignment using HMMER software. If the e-value of the alignment result is less than or equal to 10... -5 If so, it is determined to be the hypothetical Cas12 enzyme, and the system consisting of the CRISPR array and the Cas12 enzyme is the hypothetical CRISPR-Cas12 system.
[0154] 4. A phylogenetic tree was constructed using reported and putative Cas12 enzymes. The specific procedure was as follows: the RuvC domain of the aforementioned Cas12 enzymes was annotated using HMMER; multiple sequence alignment and domain annotation of the RuvC domain were performed using MAFFT; and a phylogenetic tree was constructed using FastTree based on the multiple sequence alignment results. The protein evolutionary relationships between the CRISPR-Cas12b2 system and other CRISPR-Cas systems were compared (e.g.,...). Figure 1 As shown), the arrangement of components (such as...) Figure 2 As shown), protein structural features (such as...) Figure 3 As shown), the CRISPR-Cas12b2 system has the following characteristics:
[0155] (1) Based on the evolutionary relationship of proteins, the Cas12b2 protein is on a different evolutionary branch than other Cas proteins.
[0156] (2) Based on the component arrangement characteristics of the CRISPR-Cas system, the CRISPR-Cas12b2 system has Cas1, Cas2, Cas4 proteins and tracrRNA. Among them, Cas1, Cas2, and Cas4 are located upstream of the Cas12b2 protein, and there is a unique Cas4-Cas1 fusion protein.
[0157] (3) The Cas12b2 protein is relatively closely related to the Cas12b1 protein, and the two have similar domain compositions. However, the Cas12b2 protein does not have a Bridge Helix (BH) domain, and the lengths of its other domains are shorter than those of the Cas12b1 protein.
[0158] In summary, the CRISPR-Cas12b2 system is a new subtype of CRISPR-Cas system.
[0159] Example 2: Characterizing the PAM site recognized by the Cas12b2 protein
[0160] 1. Using CRISPRCasFinder software, predicted repeat sequences and spacers were identified in the host genome region containing the Cas12b2 protein.
[0161] 2. DNA was synthesized, encompassing the Cas12b2 gene on the host genome containing the Cas12b2 protein, the Repeat 1-Spacer 1-Repeat 2-Spacer 2-Repeat 3-Spacer 3 array, and the DNA sequence between them. Genentech Biotechnology Co., Ltd. was commissioned to perform the gene synthesis.
[0162] 3. Construct the Cas12b2 loci expression vector. The above DNA was amplified by PCR, digested and ligated, and then inserted into the pBAD-Ptet plasmid vector to obtain the plasmid pBAD-Ptet-Cas12b2_loci.
[0163] 4. Construction of PAM library plasmid. Primer PAM-6N was synthesized. This primer has the following structure: it contains a Spacer1 sequence with 6 random bases at the 5' end. Using primer PAM-6N-R, primer PAM-6N was completed into double-stranded DNA under the catalysis of DNA polymerase I large fragment (Klenow, NEB). Subsequently, the double-stranded DNA was ligated to plasmid pUC19 via a Gibson ligation reaction, yielding the PAM library plasmid pUC19-PAM6N.
[0164] 5. Plasmids pBAD-Ptet-Cas12b2_loci and pUC19-PAM6N were sequentially transformed into Escherichia coli BL21(DE3) cells via electroporation. The recovered bacterial culture was spread onto LB agar plates containing 50 μg / mL kanamycin, 50 μg / mL spectinomycin, and 200 ng / mL tetracycline hydrochloride, and incubated overnight at 37°C.
[0165] 6. Take 1 mL of culture medium, gently scrape off all cells from the plate, and extract the plasmid library pUC19-PAM6N. Using primers PAM-NGS-F / R, perform PCR on the pUC19-PAM6N plasmid libraries before and after transformation. The PCR products were used by Suzhou Genewiz Biotechnology Co., Ltd. for library construction and next-generation sequencing (Illumina Novaseq 2x150bp sequencing platform). The sequencing results were analyzed using a Python script to calculate the sequence deletion status of the PAM library before and after Cas12b2 expression, such as... Figure 4 As shown. The results indicate that Cas12b2 recognizes either GTA or GTG as the PAM site. Compared with other CRISPR-Cas systems ( Figure 5 The Cas12b2 protein described in this application has a unique PAM site.
[0166] The primer sequences (5'-3') are as follows:
[0167] PAM-6N (SEQ ID NO: 4): cacacaggaaacagctatgaccatgattacgccaagcttgNNNNNNGCGCATCGTGTACCACAGCAgcacccacctgcttcgcgaatt.
[0168] PAM-6N-R (SEQ ID NO: 5): aattcgCGAAGCAGGTGGGTG.
[0169] PAM-NGS-F (SEQ ID NO: 6): AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATCTTGTTGTGTGGAATTGTGAGCGG.
[0170] PAM-NGS-R (SEQ ID NO: 7): CAAGCAGAAGACGGCATACGAGATACATCGGTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCCAGGGTTTTCCCAGTCACGAC.
[0171] Example 3: Analysis and acquisition of sgRNA sequence of CRISPR-Cas12b2 system
[0172] Fresh E. coli culture containing plasmid pBAD-Ptet-Cas12b2_loci, constructed in Example 2, was centrifuged at 10000×g for 5 min, the supernatant was discarded, and total RNA was extracted from the bacterial samples. Library construction and next-generation sequencing (Illumina Novaseq 2x150bp sequencing platform) were performed by Suzhou Genewiz Biotechnology Co., Ltd. The sequencing results were processed as follows:
[0173] 1. In the command line, use the BWA program to match the sequencing results to the plasmid DNA sequence.
[0174] bwa index plasmid.fasta
[0175] bwa mem plasmid.fasta f1.fastq f2.fastq >RNA.sam.
[0176] Among them, plasmid.fasta is the sequence of plasmid pBAD-Ptet-Cas12b2_loci. f1.fastq and f2.fastq are the results of paired-end next-generation sequencing.
[0177] 2. In the command line, use the samtools program to further process the data.
[0178] samtools sort -@ 4 RNA.sam -o . / RNA.bam
[0179] samtools depth RNA.bam > result.depth.txt.
[0180] The result.depth.txt file contains sequencing depth information for each site, as shown in the following example. Figure 6 As shown. Combined with analysis of the promoter and terminator sequences within the region, the specific sequence information of tracrRNA and crRNA was obtained. The tracrRNA sequence is shown in SEQ ID NO: 8, and the crRNA sequence is shown in SEQ ID NO: 9. The tracrRNA and crRNA sequences were ligated through a short artificial AAGG loop to obtain the sgRNA sequence shown in SEQ ID NO: 3.
[0181] SEQ ID NO: 8:
[0182] GGCTACGGCCCACCGGGCGCTATAGGCCCCCGGTGCTGCGTCGGGCGAGCTGAGTGTCGCCGACCGGTTGCCGAAAGGCTTTGCCGAGTAGGGTCGCCCCATCCCCTTGCGGGATGACGGCCTCCGGTTAACCCCGACGCAGAGAGATCCTTTTCTGAT;
[0183] SEQ ID NO: 9:
[0184] ATCAGAAAAGGACCTCTCTGGACACNNNNNNNNNNNNNNNNNN.
[0185] Example 4: Testing the in vitro enzyme activity of the CRISPR-Cas12b2 complex
[0186] 1. Expression and purification of Cas12b2 protein. The Cas12b2 gene sequence was cloned into the pET-28a(+) plasmid with a 6×His tag at the N-terminus. Single colonies were picked and inoculated into LB medium containing 50 μg / mL kanamycin and cultured overnight at 37°C and 200 rpm. The next day, 10 mL of seed culture was inoculated into 1 L of LB medium containing 50 μg / mL kanamycin and cultured at 37°C and 200 rpm until OD600 = 0.8. 0.2 mM IPTG was added, and protein expression was induced at 16°C and 200 rpm for 20 hours. Cell pellet was obtained by centrifugation. Cells were resuspended in lysis buffer (20 mM HEPES-Na pH 8.0, 800 mM NaCl, 20 mM imidazole, 10% glycerol, 1 mM DTT) and homogenized at 1000 bar using an autoclave. The lysis buffer was centrifuged at 15000×g for 60 minutes at 4°C. The supernatant was purified by HisTrap, Heparin, and SD200 increase 10 / 300 column chromatography to obtain Cas12b2 protein, which was then flash-frozen in liquid nitrogen and stored at -80°C.
[0187] 2. Purify sgRNA by in vitro transcription.
[0188] (a) sgRNA was transcribed in vitro using the HiScribe T7 High Yield RNA Synthesis Kit (NEB). Following the manufacturer's instructions, 500 ng of DNA substrate and 2 μL of T7 RNA polymerase were added, and the mixture was incubated at 37°C for 16 hours.
[0189] (b) After the reaction, add 160 μL of nuclease-free water and 20 μL of 3M sodium acetate (pH 5.2) to each 20 μL system, and mix thoroughly. Add an equal volume of a 1:1 phenol:chloroform mixture, centrifuge at 13000×g for 5 minutes, collect the aqueous phase (supernatant) and transfer it to a new tube. Add an equal volume of chloroform, mix thoroughly for extraction, collect the supernatant and transfer it to a new tube, and repeat once. Add 2 volumes of ethanol to precipitate the RNA and incubate at -20°C for at least 30 minutes.
[0190] (c) Centrifuge at 13000×g for 15 min at 4℃, and discard the supernatant. Wash the precipitate with 500 μL of ice-cold 70% ethanol, centrifuge at 13000×g for 15 min, discard the supernatant, and leave the cap open for 2 min to dry the sample. Finally, resuspend the RNA in 50 μL of DEPC water, flash freeze in liquid nitrogen, and store at -80℃.
[0191] 3. In vitro assembly and cutting experiments.
[0192] (a) In a 50 μL reaction system, add 2×Assembly Buffer (40 mM HEPES-Na pH 7.5, 300 mM NaCl, 10 mM MgCl2, 1 mM DTT), 1 μM Cas12b2 protein, and 1.2 μM sgRNA, and incubate at room temperature for 30 min to complete the assembly of the CRISPR-Cas12b2-sgRNA complex.
[0193] (b) Add 20 nM double-stranded DNA substrate and incubate at 25℃, 30℃, 37℃, 45℃, 50℃, and 60℃ for 2 h each. Add 10 mM EDTA and 0.5 μL RNase A to terminate the reaction.
[0194] (c) The reaction products were analyzed by 2% agarose gel electrophoresis. The gel was imaged using a UV gel imaging system, and the substrate DNA cleavage ratio was analyzed using ImageJ software. The results are as follows: Figure 7 As shown. The results indicate that the CRISPR-Cas12b2-sgRNA complex in this application possesses in vitro enzymatic cleavage activity and the ability to be applied to in vivo gene editing.
[0195] Example 5: Construction of recombinant plasmid pEcas12b2 carrying Cas12b2 expression cassette
[0196] The principle behind the application of the Cas12b2-sgRNA complex in gene editing is as follows: Figure 8 As shown. In this embodiment, Cas12b2 is expressed in vivo by a plasmid vector, and the specific construction method is as follows.
[0197] 1. The Cas12b2 enzyme was selected, and its amino acid sequence is shown in SEQ ID NO: 1. Codon optimization was performed for the *E. coli* host, and the optimized nucleotide sequence of the gene is shown in SEQ ID NO: 16. Genesynthetic work was commissioned to Genscript Biotech Inc.
[0198] 2. Using Cas12b2-1-F / R primers and the gene sequence synthesized in the above steps as a template, the Cas12b2 gene was amplified by PCR. Using pEcCas9 plasmid as a template and pEcCas9-F / R primers, the arabinose-induced promoter pBAD and the *E. coli* λ-Red recombination system gene were amplified by PCR. Using pBAD plasmid as a template and ori-F / R primers, the pSC101 replicon and kanamycin resistance gene were amplified by PCR. Using pEvolvR-enCas9-PolI3m-TBD plasmid as a template and promoter-F / R primers, the tetracycline-induced promoter Ptet gene was amplified by PCR. The PCR products were subjected to agarose gel electrophoresis and purified. The PCR products were ligated using a Gibson Assembly kit (Clonesmarter brand, purchased from Sino-American Taihe Biotechnology (Beijing) Co., Ltd.).
[0199] Primers were synthesized by Suzhou Genewiz Biotechnology Co., Ltd. PCR amplification was performed using 2×Phanta MaxMaster Mix (Dye Plus) high-fidelity DNA polymerase from Nanjing Novizan Biotechnology Co., Ltd.
[0200] The PCR reaction system consisted of: 25 μL of 2×Phanta Max Master Mix (Dye Plus), 2 μL of upstream primer, 2 μL of downstream primer, 20 μL of sterile water, 1 μL of template, and a total volume of 50 μL.
[0201] The PCR reaction conditions were: 95℃ for 5 min; 95℃ for 15 s, 55℃ for 15 s, 72℃ for 30 s-3 min, 32 cycles; 72℃ for 5 min; 4℃ forever.
[0202] The PCR product ligation system consisted of 1 μL of each of the four recovered PCR products, 1 μL of sterile water, and 5 μL of 2×Gibson Mix.
[0203] The PCR product ligation reaction conditions were: 50℃ for 60 min.
[0204] The gene and primer sequences (5'-3') are as follows:
[0205] SEQ ID NO: 16 (2076 bp):
[0206]
[0207] Cas12b2-1-F (SEQ ID NO: 17): TCTAAAGAGGAGAAAGGCTCGATGAGCGTTAAATCTTTTCAGGCC;
[0208] Cas12b2-1-R (SEQ ID NO: 18):CATTCAAATATGTATCCGCTCATGTAAACCACGGGGTTTGTGCAGC;
[0209] pEcCas9-F (SEQ ID NO: 19):TTTACATGAGCGGATACATATTTGAATG;
[0210] pEcCas9-R (SEQ ID NO: 20): GATTCTTCGTCTGTTTCTACTGGT;
[0211] ori-F (SEQ ID NO: 21): CCAGTAGAAACAGACGAAGAATCCATGGGTATGGACAGTTTTCCC;
[0212] ori-R (SEQ ID NO: 22): TGGCTTTCCCTGCAGCTG;
[0213] promoter-F (SEQ ID NO: 23): GCTTTCCCTGCAGCTGATGCGAGAGTAGGGAACTGCC;
[0214] promoter-R (SEQ ID NO: 24): CGAGCCTTTCTCCTCTTTAGATCTTTTGAATTC.
[0215] 3. Take 10 μL of the above ligation product and chemically transform it into *E. coli* TOP10 competent cells (purchased from Beijing Bomed Gene Technology Co., Ltd.). The specific transformation process is as follows: Place 50 μL of TOP10 competent cells on ice, add the ligation product to the competent cells, incubate on ice for 30 min, heat shock at 42℃ for 30 s, and then place on ice for 2 min. Then add 500 μL of SOC resuscitation solution and incubate at 37℃ and 200 rpm for 1 h. Spread the bacterial culture on LB agar plates containing 50 μg / mL kanamycin and incubate overnight at 37℃. The next day, select single clones for sequencing verification.
[0216] Selected single clones with correct sequencing were cultured overnight in LB liquid medium containing 50 μg / mL kanamycin. The plasmid was extracted the next day to obtain the recombinant vector pEcas12b2, and its plasmid map is shown below. Figure 9 As shown.
[0217] The culture medium components involved in this embodiment are as follows:
[0218] SOC medium: Sodium chloride 0.5 g / L, yeast extract 5 g / L, peptone 20 g / L, magnesium chloride 10 mM, potassium chloride 2.5 mM, glucose 20 mM, water.
[0219] LB medium: sodium chloride 10 g / L, yeast extract 5 g / L, peptone 10 g / L, water; solid plate medium contains agar 16 g / L.
[0220] Example 6: Construction of recombinant plasmid pEcgb2 carrying sgRNA expression cassette
[0221] 1. Using the DNA sequence SEQ ID NO: 3 as a template and sgRNA-F / R as primers, the sgRNA expression cassette was amplified by PCR. sgRNA expression was controlled by the PJ23119 promoter. The plasmid backbone was amplified using pEcgRNA as a template and pEcgRNA-vec-F / R as primers. The PCR products were purified by agarose gel electrophoresis. The PCR products were ligated using a Gibson Assembly kit (Clonesmarter brand, purchased from Sino-American Taihe Biotechnology (Beijing) Co., Ltd.).
[0222] 2. Take 10 μL of the above ligation product and chemically transform it into *E. coli* TOP10 competent cells (purchased from Beijing Bomed Gene Technology Co., Ltd.). See Example 5 for the specific operation procedure. Select the correctly sequenced single clones and culture them overnight in LB liquid medium containing 50 μg / mL spectinomycin. The plasmid is extracted the next day to obtain the recombinant vector pEcgb2, whose plasmid map is shown below. Figure 10 As shown.
[0223] The primer sequences (5'-3') are as follows:
[0224] sgRNA-F (SEQ ID NO: 25): TAGCTCAGTCCTAGGTATAATGCTAGCGGCTACGGCCCACCGGGC;
[0225] sgRNA-R (SEQ ID NO: 26): GTCGACTCTAGAGAATTCAAAAAAAGTGTCCAGAGAGGTCCTTTTCTGA;
[0226] pEcgRNA-vec-F (SEQ ID NO: 27): TTTTTTTGAATTCTCTAGAGTCGACCTGC;
[0227] pEcgRNA-vec-R (SEQ ID NO: 28): GCTAGCATTATACCTAGGACTGAGC.
[0228] Example 7: Construction of E. coli MG1655 envc gene knockout gene editing plasmids pEcgb2-envc and pEcgb2-HA-envc
[0229] 1. Design the sgRNA spacer sequence. Using the CHOPCHOP online website (https: / / chopchop.cbu.uib.no), paste the E. coli MG1655 envc gene sequence, select Escherichia colistr. MG1655 as the host, and choose CRISPR / Cpf1, Cas12, or CasX as the Cas enzyme. Modify the options to sgRNA length to 20 nt, 5'-PAM to Non-stardard GTA, and leave other options unchanged. After completing the settings, the website automatically searches for the optimal spacer sequence. The spacer sequence with the highest score is selected, as shown below.
[0230] spacer-envc-1 (SEQ ID NO: 29): GCCTTGCAGCCAGTCAGCCA.
[0231] 2. Construction of the gene-editing plasmid pEcgb2-envc. Using b21-sgRNA-F and b21-envc-sg-R as primers and plasmid pEcgb2 as a template, PCR was performed to amplify the sgRNA containing the SEQ ID NO: 3 spacer sequence; using pEcgb2-vec-F / R as primers and plasmid pEcgb2 as a template, the plasmid backbone was amplified. The PCR products were subjected to agarose gel electrophoresis and purified. The PCR products were ligated using a Gibson Assembly kit. 10 μL of the above products were chemically transformed into E. coli TOP10 competent cells (purchased from Beijing Bomed Gene Technology Co., Ltd.). The specific operation procedure is described in Example 5. Selected single clones with correct sequencing were cultured overnight in LB liquid medium containing 50 μg / mL spectinomycin. The plasmid was extracted the next day to obtain the envc gene-editing plasmid pEcgb2-envc, the plasmid map of which is shown below. Figure 11 As shown.
[0232] 3. During gene editing, the homologous recombination repair template can be provided in the form of linear double-stranded DNA or in the form of plasmid circular double-stranded DNA. When the latter is provided as a template, the gene editing plasmid pEcgb2-sgRNA-HA is constructed. Using b21-sgRNA-F and b21-envc-sg-HR as primers, PCR reaction is used to amplify the sgRNA containing the SEQ ID NO: 3 spacer sequence (fragment envc-sgRNA); using envc-UP-F / R as primers and E. coli MG1655 genome as template, the upstream 500 bp homologous arm of the envc gene is amplified (fragment envc-UP); using envc-DOWN-F / R as primers, the downstream 500 bp homologous arm of the envc gene is amplified (fragment envc-DOWN); using pEcgb2-vec-F / R as primers and plasmid pEcgb2 as template, the plasmid backbone (fragment pEcgb2-vec) is amplified. The PCR products were subjected to agarose gel electrophoresis and purified. The PCR products were then ligated using a Gibson Assembly kit (Clonesmarter brand, purchased from Sino-American Taihe Biotechnology (Beijing) Co., Ltd.).
[0233] The ligation reaction mixture consisted of: 5 μL of 2×Gibson Mix, 2 μL of envc-sgRNA, 1 μL of envc-UP, 1 μL of envc-DOWN, and 1 μL of pEcgb2-vec.
[0234] The connection reaction conditions were: 50℃, 60 min.
[0235] Take 10 μL of the above ligation product and chemically transform it into E. coli TOP10 competent cells (purchased from Beijing Bomed Gene Technology Co., Ltd.). See Example 5 for the specific operation procedure. Selected single clones with correct sequencing were cultured overnight in LB liquid medium containing 50 μg / mL spectinomycin. The plasmid was extracted the next day to obtain the envc gene editing plasmid pEcgb2-HA-envc, the plasmid map of which is shown below. Figure 12 As shown.
[0236] 4. When using linear double-stranded DNA as a template for homology repair, perform overlap extension PCR on the recovered and purified DNA fragments envc-UP and envc-DOWN. The steps for overlap extension PCR are as follows.
[0237] Step 1: The system consists of 25 μL of 2×Phanta Max Master Mix (Dye Plus), 2 μL of envc-UP, 2 μL of envc-DOWN (total amount of each DNA fragment is approximately 0.2 pmol), and 17 μL of sterile water, for a total volume of 46 μL.
[0238] The PCR reaction conditions were: 95℃ for 5 min; 95℃ for 15 s, 55℃ for 15 s, 72℃ for 1 min, 5 cycles; 72℃ for 5 min; 4℃ forever.
[0239] Step 2: Based on the system in Step 1, add 2 μL of primer envc-DOWN-R and 2 μL of primer envc-UP-F, for a total volume of 50 μL.
[0240] The PCR reaction conditions were: 95℃ for 5 min; 95℃ for 15 s, 55℃ for 15 s, 72℃ for 1 min, 27 cycles; 72℃ for 5 min; 4℃ forever.
[0241] The PCR product was subjected to agarose gel electrophoresis and purified to obtain the linear double-stranded DNA fragment envc-HA.
[0242] The primer sequences (5'-3') are as follows:
[0243] b21-sgRNA-F (SEQ ID NO: 30): ACCGATATGCTGATCCTTGACAGCTAGCTCAGTC;
[0244] b21-envc-sg-R (SEQ ID NO: 31):GCAGGTCGACTCTAGAGAATTCAAAAAAATGGCTGACTGGCTGCAAGGCGTGTCCAGAGAGGTCCTTTTCTG;
[0245] pEcgb2-vec-F (SEQ ID NO: 32): TTTTTTTGAATTCTCTAGAGTCGACCTGC;
[0246] pEcgb2-vec-R (SEQ ID NO: 33):GGATCCAGCATATGCGGTGT;
[0247] b21-envc-sg-HR (SEQ ID NO: 34): CTGCAGGCAGGTCGACTCTAGAGAATTCAAAAAAATGGCTGACTGGCTGCAAGGCGTGTCCAGAGAGGTCCTTTTCTG;
[0248] envc-UP-F (SEQ ID NO: 35): AGAGTCGACCTGCAGAAGATCAACTCACCGAAAGTGGC;
[0249] envc-UP-R (SEQ ID NO: 36): GGTATTAATCGCCTTTCCCCTC;
[0250] envc-DOWN-F (SEQ ID NO: 37): GAGGGGAAAGGCGATTAATACCGTTTTGTTTCCATTTCGTCGTAACG;
[0251] envc-DOWN-R (SEQ ID NO: 38): CTCGAGTAGGGATAACAGGGCATCGCCTGGGTATTACCGA.
[0252] Example 8: Gene editing to knock out the envc gene in E. coli MG1655
[0253] 1. Seed culture. Inoculate activated single colonies of E. coli MG1655 from plates into LB medium and incubate at 37℃ and 200 rpm for 12-16 h to obtain seed culture.
[0254] 2. Transfer culture. Transfer the seed culture to 100 mL LB medium at a ratio of 1:100, and incubate at 37℃ and 200 rpm for 1.5-2 h until the OD600 reaches about 0.6. Stop the culture and place the bacterial culture on ice to pre-cool.
[0255] 3. Preparation of competent cells. Centrifuge the bacterial culture at 4000 rpm for 10 min, discard the supernatant, resuspend the cells in an equal volume of 10% glycerol, and centrifuge at 4000 rpm for 10 min. Discard the supernatant, resuspend the cells in 50% of the volume of 10% glycerol, and centrifuge at 4000 rpm for 15 min. Discard the supernatant, resuspend the cells in 10% of the volume of 10% glycerol, and centrifuge at 4000 rpm for 15 min. Discard the supernatant, resuspend the cells in 0.5% of the volume of 10% glycerol to obtain competent cells. Aliquot 50 μL into each tube and use directly or flash-freeze in liquid nitrogen and store in an ultra-low temperature freezer.
[0256] 4. Electroporation transformation of plasmid pEcas12b2. Place one tube of E. coli MG1655 competent cells on ice and add 1 ng of pEcas12b2 plasmid to the cells. Transfer the mixture to a 0.1 cm bio-Rad electroporation cuvette, set the electroporation field strength to 18 kV / cm, and immediately add 950 μL of SOC medium after electroporation. Incubate at 37°C and 200 rpm for 1 h. Spread an appropriate amount of bacterial culture onto LB agar plates containing 50 μg / mL kanamycin and incubate overnight at 37°C. The next day, perform colony PCR on single clones on the plate using primers Cas12b2-1-F / R to verify successful transformation of plasmid pEcas12b2.
[0257] 5. Transformation of plasmid pEcgb2-envc-HA. Pick a single colony of the overnight culture seed culture transformed with plasmid pEcas12b2 and transfer it 1:100 to LB medium containing 50 μg / mL kanamycin. Incubate at 37°C and 200 rpm for 0.5 h until OD600 = approximately 0.2. Add 0.2% L-arabinose to induce λ-Red recombinant system expression. Incubate for another 1.5-2 h until OD600 = approximately 0.6. Subsequently, prepare competent cells and electroporate to transform plasmid pEcgb2-envc-HA, following the procedure described above. Spread the recovered bacterial culture onto LB agar plates containing 50 μg / mL kanamycin, 50 μg / mL spectinomycin, and 200 ng / mL tetracycline hydrochloride, and incubate overnight at 37°C.
[0258] 6. Verify gene editing results. Using envc-test-F / R as primers, colony PCR was performed on single colonies on the plate. The PCR products were then subjected to agarose gel electrophoresis. Without knocking out the envc gene, the PCR product size was 2341 bp; after knocking out the 1260 bp envc gene, the PCR product size was 1081 bp. The agarose gel electrophoresis results are shown below. Figure 13 As shown, the gene knockout ratio is 12 / 12, and the gene editing success rate is 100%.
[0259] The primer sequences (5'-3') are as follows:
[0260] envc-test-F (SEQ ID NO: 39): CGTTCAAAGGCGAAGATCGC;
[0261] envc-test-R (SEQ ID NO: 40): TCGTTCTGCGAATCGTCG.
[0262] Example 9: Gene editing to knock out the E. coli BL21(DE3) dgka gene fragment
[0263] 1. Design of the sgRNA spacer sequence. The dgka gene is 369 bp in length. The process for designing the sgRNA spacer sequence is as described in Example 7. The spacer sequence selected is:
[0264] spacer-kpLE2-1 (SEQ ID NO: 41): GCGGTATTGTTGGCGGTGGT.
[0265] 2. Construct the gene-editing plasmid pEcgb2-HA-dgka, as described in Example 7.
[0266] 3. Prepare E. coli BL21(DE3) competent cells, and sequentially electroporate plasmids pEcas12b2 and pEcgb2-HA-dgka into the cells, as described in Example 8.
[0267] 4. Verify gene editing results. Colony PCR was performed on 17 single colonies on the plate using dgka-test-F / R primers. The PCR products were subjected to agarose gel electrophoresis, and the results are as follows: Figure 14 As shown, this indicates that gene knockout was achieved in 16 out of 17 single colonies, with a gene editing success rate of 94%.
[0268] The primer sequences (5'-3') are as follows:
[0269] dgka-test-F (SEQ ID NO: 42):TCTGCACATCCAGATTTGGG;
[0270] dgka-test-R (SEQ ID NO: 43): GACCGTTACGTACATCCTGAG.
[0271] Example 10: Gene editing to knock out the E. coli BL21(DE3) kpLE2 gene fragment
[0272] 1. Design of the sgRNA spacer sequence. kpLE2 is a prophage sequence from the E. coli BL21(DE3) genome, 36.98 kb in length, from which a 15 kb segment was selected for knockout. The process for designing the sgRNA spacer sequence is as described in Example 7. The selected spacer sequence is:
[0273] spacer-kpLE2-1 (SEQ ID NO: 10):ACCCACCACCGGCATCCATT.
[0274] 2. Construct the gene editing plasmid pEcgb2-HA-kpLE2, as described in Example 7.
[0275] 3. Prepare E. coli BL21(DE3) competent cells, and sequentially electroporate plasmids pEcas12b2 and pEcgb2-HA-kpLE2 into the cells, as described in Example 8.
[0276] 4. Verify gene editing results. Colony PCR was performed on six single colonies on the plate using kpLE2-test-F / R primers. The PCR products were subjected to agarose gel electrophoresis and sequencing. The sequencing results are as follows: Figure 15 This indicates that the 15 kb kpLE2 gene fragment was successfully knocked out in all six single colonies, with a gene editing success rate of 100%.
[0277] The primer sequences (5'-3') are as follows:
[0278] kpLE2-test-F (SEQ ID NO: 11): TGGTCCTTCTCTGGTGTTGTT;
[0279] kpLE2-test-R (SEQ ID NO: 12): CGGCTTTCAAACCAGTCGTGA.
[0280] Example 11: Gene editing point mutation of E. coli BL21(DE3) cadA gene
[0281] 1. Design of sgRNA spacer sequence. The 1060th base pair of the cadA gene on the E. coli BL21(DE3) genome was mutated from G:C to C:G. The 500 bp upstream and downstream of this site were selected as the target fragment for designing the sgRNA spacer sequence, as described in Example 7. The selected spacer sequence is:
[0282] spacer-cadA-1 (SEQ ID NO: 15):GAAGGGAAAGTGATTTACGA.
[0283] 2. Construct the gene editing plasmid pEcgb2-HA-cadA, as described in Example 7.
[0284] 3. Prepare E. coli BL21(DE3) competent cells, and sequentially electroporate plasmids pEcas12b2 and pEcgb2-HA-cadA into the cells, as described in Example 8.
[0285] 4. Verify gene editing results. Using cadA-test-F / R primers, colony PCR was performed on 10 single colonies on the plate. PCR products were subjected to agarose gel electrophoresis and sequencing. The sequencing result of one PCR product is shown below. Figure 16 Nine out of ten single colonies successfully achieved the G:C > C:G point mutation shown in the figure, resulting in a gene editing success rate of 90%.
[0286] The primer sequences (5'-3') are as follows:
[0287] cadA-test-F (SEQ ID NO: 13): TGCTGAACAGGAACAGCAGG;
[0288] cadA-test-R (SEQ ID NO: 14): GCCTGTTCTATGATTTCTTTGGTCCG.
[0289] Example 12: Gene editing point mutation of E. coli BL21(DE3) gadc gene
[0290] 1. Design of sgRNA spacer sequence. The 488th base pair of the gadc gene on the E. coli BL21(DE3) genome was mutated from C:G to G:C. The 500 bp upstream and downstream of this site were selected as the target fragment for designing the sgRNA spacer sequence, as described in Example 7. The selected spacer sequence is:
[0291] spacer-cadA-1 (SEQ ID NO: 44): TCCTGTTACCTGCATTTATT.
[0292] 2. Construct the gene editing plasmid pEcgb2-HA-gadc, as described in Example 7.
[0293] 3. Prepare E. coli BL21(DE3) competent cells, and sequentially electroporate plasmids pEcas12b2 and pEcgb2-HA-gadc into the cells as described in Example 8.
[0294] 4. Verify gene editing results. Using gadc-test-F / R primers, colony PCR was performed on 10 single colonies on the plate. The PCR products were subjected to agarose gel electrophoresis and sequencing. The sequencing results are as follows: Figure 17 As shown in the figure, 8 out of 10 single colonies successfully achieved the C:G > G:C point mutation, resulting in a gene editing success rate of 80%.
[0295] The primer sequences (5'-3') are as follows:
[0296] gadc-test-F (SEQ ID NO: 45):AGCTGCGAAATGACCAGCG;
[0297] gadc-test-R (SEQ ID NO: 46): TCACTGGCATTAGCAACGG.
[0298] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0299] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims, and the specification and drawings can be used to interpret the content of the claims.
Claims
1. A Cas12b2 protein, characterized in that, The protospacer sequence it identifies includes GTA and / or GTG adjacent sequences, and the Cas12b2 protein is selected from the following group: (a) A polypeptide having an amino acid fragment with the sequence shown in SEQ ID NO:1; (b) A polypeptide having the biological function of (a) with ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 99.5% homology with (a); (c) A derivative polypeptide formed by substituting, deleting or adding one or more amino acid residues of (a) and retaining the biological function of (a); the derivative polypeptide comprises an amino acid fragment with an amino acid sequence as shown in SEQ ID NO:
2.
2. An sgRNA, characterized in that, It can specifically bind to the RNA of the Cas12b2 protein as described in claim 1; Optionally, the sgRNA is selected from the group consisting of: (a) RNA having the nucleotide sequence shown in SEQ ID NO:3; (b) RNA having a biological function of (a) with ≥80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% homology to (a); (c) A derivative RNA formed by substituting, deleting or adding one or more nucleotides of (a) and retaining the biological function of (a).
3. A protein variant, including at least one of truncated proteins and fusion proteins, characterized in that: (a) The truncated protein is a polypeptide composed of one or more functional domains of the Cas12b2 protein as described in claim 1; (b) The fusion protein comprises the Cas12b2 protein as described in claim 1, and one or more functional domains.
4. An isolated polynucleotide, characterized in that, The polynucleotide encodes the Cas12b2 protein as described in claim 1, or the sgRNA as described in claim 2, or the protein variant as described in claim 3.
5. A CRISPR-Cas12b2 complex, characterized in that, It includes: (i) Protein components selected from the group consisting of: the Cas12b2 protein of claim 1, the protein variants of claim 3, and combinations thereof; and (ii) A nucleic acid component selected from the group consisting of: the sgRNA as described in claim 2, a nucleic acid encoding the sgRNA as described in claim 2, a precursor RNA of the sgRNA as described in claim 2, a nucleic acid encoding the precursor RNA of the sgRNA as described in claim 2, and combinations thereof; The protein component and the nucleic acid component combine to form a complex.
6. An activated CRISPR-Cas12b2 complex, characterized in that, Include: (i) Protein components selected from the group consisting of: the Cas12b2 protein of claim 1, the protein variants of claim 3, and combinations thereof; (ii) A nucleic acid component selected from the group consisting of: the sgRNA of claim 2, a nucleic acid encoding the sgRNA of claim 2, a precursor RNA of the sgRNA of claim 2, a nucleic acid encoding the precursor RNA of the sgRNA of claim 2, and combinations thereof; and, (iii) The target sequence that binds to the sgRNA.
7. The CRISPR-Cas12b2 system, characterized in that, It comprises one or more carriers, wherein the one or more carriers comprise: (i) a first nucleic acid comprising the isolated polynucleotide as described in claim 4; and (ii) a second nucleic acid comprising a nucleotide sequence encoding the sgRNA as described in claim 2; in: The first nucleic acid and the second nucleic acid may exist on the same or different vectors.
8. A recombinant vector, characterized in that, It includes: The isolated polynucleotide as described in claim 4, wherein the backbone of the recombinant vector is optionally selected from pACYC, pBAD, pET28a and pHT08; Alternatively, the recombinant vector may encode a polynucleotide of the sgRNA as described in claim 2, wherein the backbone of the recombinant vector is optionally selected from pUC19 and pBR322, and the recombinant vector may optionally further comprise a T7 promoter, a pJ23119 promoter, or a PvanP promoter.
9. An engineered host cell, characterized in that, The host cell is a non-plant or animal variety, and the host cell contains one or more of the following: the Cas12b2 protein as described in claim 1, the protein variant as described in claim 3, the sgRNA as described in claim 2, the isolated polynucleotide as described in claim 4, the CRISPR-Cas12b2 complex as described in claim 5, the activated CRISPR-Cas12b2 complex as described in claim 6, the CRISPR-Cas12b2 system as described in claim 7, and the recombinant vector as described in claim 8.
10. Use of the CRISPR-Cas12b2 complex of claim 6, the activated CRISPR-Cas12b2 complex of claim 7, the CRISPR-Cas12b2 system of claim 8, or the engineered host cell of claim 9, characterized in that, The uses include: (1) Gene editing, gene targeting, or gene cutting for purposes other than disease diagnosis and treatment; (2) The intended use in the preparation of the reagent or kit; wherein the preparation or kit is used for: (i) Gene or genome editing, including at least one of gene knockout, insertion, point mutation and substitution; (ii) Target nucleic acid detection and / or diagnosis; (iii) Editing target sequences in target loci to modify organisms; (iv) Treatment of the disease; (v) Targeting the target gene; (vi) Cut the target gene.
11. A method for editing and cleaving target nucleic acids for purposes other than disease diagnosis and treatment, characterized in that, The method includes contacting the target nucleic acid with the CRISPR-Cas12b2 complex as described in claim 5, the activated CRISPR-Cas12b2 complex as described in claim 6, or the engineered host cell as described in claim 9.
12. A kit for gene editing, gene targeting, or gene cutting, characterized in that, The kit comprises one or more of the following: the Cas12b2 protein as claimed in claim 1, the protein variant as claimed in claim 3, the isolated polynucleotide as claimed in claim 4, the sgRNA as claimed in claim 2, the recombinant vector as claimed in claim 8, the CRISPR-Cas12b2 complex as claimed in claim 5, the activated CRISPR-Cas12b2 complex as claimed in claim 6, the CRISPR-Cas12b2 system as claimed in claim 7, and the engineered host cell as claimed in claim 9.