Novel effector nuclease and use thereof
By developing novel effector nucleases Cas-kf63 and Cas-kf66, combined with a gRNA-based CRISPR/Cas system, the problem of frequent insertion or deletion mutations in gene editing in existing technologies has been solved, enabling efficient and precise gene editing in eukaryotic cells.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIDAHUANG KENFENG SEED CO LTD
- Filing Date
- 2025-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
In existing CRISPR/Cas gene editing technologies, the high efficiency of non-homologous end linkage repair leads to frequent insertion or deletion mutations, while homology-dependent repair has low efficiency, making it difficult to achieve precise gene editing and limiting its widespread application.
Novel effector nucleases Cas-kf63 and Cas-kf66 were developed, and their functions were validated by designing and screening amino acid sequences. They were then combined with gRNA to form a CRISPR/Cas system for gene editing in eukaryotic cells.
This technology enables efficient gene editing in eukaryotic cells, reduces insertion or deletion mutations, and improves the precision and efficiency of gene editing.
Smart Images

Figure CN2025073830_30072026_PF_FP_ABST
Abstract
Description
Novel effector nucleases and their applications Technical Field
[0001] This invention belongs to the field of biotechnology and relates to novel effector nucleases Cas-kf63 or Cas-kf66 associated with regularly clustered short palindromic repeats (CRISPR), and their applications in gene editing. Background Technology
[0002] The rapidly developing gene editing technology of recent years utilizes Cas (CRISPR-associated) effector nucleases related to the clustered regularly interspersed short palindromic repeats (CRISPR) in the bacterial acquired immune system to cut specific genes in cells at pre-selected sites in the genome. The breaks are then repaired through two inherent DNA repair mechanisms: non-homologous end joining (NHEJ) and homology-dependent repair (HDR), ensuring cell survival. NHEJ is highly efficient, requires no template, and introduces insertion or deletion mutations during the rejoining process at the cut site, thus inactivating the gene; it is the dominant mechanism. HDR requires a repair template and achieves precise repair according to the template; it accounts for a very small proportion. If a pre-designed DNA fragment is provided as a template, the sequence can be inserted at a pre-selected site to modify the target gene, but its efficiency is extremely low, severely limiting its widespread application.
[0003] Acquired immune systems containing CRISPR structures are widely found in bacteria and archaea. Based on the number of proteins involved in the immune interference response, they are divided into two main classes: Class I and Class II. Class I systems require multiple proteins to form a complex to perform immune interference, while Class II systems can achieve interference with just a single protease containing a multifunctional group (Makarova et al., 2020, Nat Rev Microbiol 18, 67-83; Nidhi et al., 2021, Int J Mol Sci 22, 3327). Each class is further divided into several types based on structure and composition. The commonly used SpCas9 is a typical representative of Class II CRISPR systems, requiring a conserved NGG sequence at the 3' end of the target sequence as a PAM sequence, and blunt-end cleavage of the target sequence (Jinek et al., 2012, Science 337, 816-821). Other effector nucleases of the other two types of CRISPR systems, such as SaCas9 derived from Staphylococcus aureus and NmeCas9 derived from Neisseria meningitidis, have also been successfully used for editing (Ran et al., 2015, Nature 520, 186-191; Amrani et al., 2018, Genome Biol 19, 214).
[0004] Cas12 effector nucleases belonging to the Class 2 type V CRISPR system, such as FnCpfI and LbCpf1 (later also known as FnCas12a and LbCas12a) derived from bacterial strains Francisella tularensis subsp. Novida U112 and Lachnospiraceae bacterium ND2006 respectively, require the target sequence to have a conserved TTTN as a PAM sequence at the 5' end, and then cleave the target sequence with sticky ends (Zetsche et al., 2015, Cell 163, 1-13). Several other Class 2 type V CRISPR system Cas12a endonucleases have also been used for gene editing to varying degrees. Most of these primarily recognize target DNA sequences containing a conserved TTTV (V being A, C, or G) PAM sequence at the 5' end, and cleave the DNA double strand at the distal end of the PAM, producing sticky ends. In addition, the Cas12a effector nuclease also has RNase function, which can autonomously process and cleave the pre-crRNA of the CRISPR repeat array to form a single mature crRNA, which guides the recognition of the target DNA sequence.
[0005] Although Cas12a effector nucleases have been identified in many bacteria based on the principle of homology, the homology between protease sequences from different sources is low. Enzymes with large genetic distances often share only about 30% amino acid sequence similarity. Besides the difficulty in directly determining the biological function of a specific Cas12a enzyme based on sequence due to low homology, not every enzyme protein is an active protein with immune function. Some Cas12a immune systems are incomplete, lacking other necessary components such as Cas1, Cas2, Cas4, or CRISPR arrays (Yan et al., 2019, Science 363, 88-91). Even Cas12a enzyme proteins with complete composition and sequence may not necessarily have enzymatic activity; only a very few may retain activity after expression in eukaryotes, thus requiring specific experimental screening for verification. Currently discovered CRISPR / Cas gene editing systems each have different characteristics. Based on various reasons and the complex and diverse research and application needs, the continuous development of new gene editing tool enzymes is of great significance to the development and application research of biotechnology. Summary of the Invention
[0006] Based on the discovered and studied Cas12a effector nuclease sequences, this invention designs, screens, and verifies the functions of the novel effector nucleases Cas-kf63 and Cas-kf66, which have gene editing functions in eukaryotic cells. Based on this discovery, this invention develops a new CRISPR / Cas system and a gene editing method based on this system.
[0007] Cas protein
[0008] On the one hand, the present invention provides a Cas protein (or, referred to as "effect nuclease", "Cas enzyme" or "effect protein"), which is an effector protein in the CRISPR / Cas system. The Cas protein is Cas-kf63 and Cas-kf66, and the amino acid sequence of Cas-kf63 is shown in SEQ ID No 1, and the amino acid sequence of Cas-kf66 is shown in SEQ ID No 4.
[0009] Compared to the sequence of SEQ ID No. 1 or 4, in some embodiments, the Cas protein is a sequence having one or more amino acid substitutions, deletions, or additions, and substantially retains the biological function of its derived sequence; the one or more amino acids include substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more amino acids; in some embodiments, the amino acid sequence has at least 50% sequence identity with the sequence shown in SEQ ID No. 1 or 4, for example, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% identity, and substantially retains the biological function of its derived sequence.
[0010] Those skilled in the art will understand that the primary structure of a protein can be altered without adversely affecting its activity and function. For example, one or more conserved amino acid substitutions can be introduced into the protein's amino acid sequence without adversely affecting the protein molecule's activity and / or three-dimensional structure. Examples and implementations of conserved amino acid substitutions are familiar to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted; that is, a nonpolar amino acid residue can replace another nonpolar amino acid residue, a polar, uncharged amino acid residue can replace another polar, uncharged amino acid residue, a basic amino acid residue can replace another basic amino acid residue, and an acidic amino acid residue can replace another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitution, where an amino acid is replaced by another amino acid belonging to the same group, falls within the scope of this invention, provided that the substitution does not lead to inactivation of the protein's biological activity. Furthermore, this invention also covers proteins containing one or more other nonconserved amino acid substitutions, provided that such nonconservative substitution does not significantly affect the desired function and biological activity of the protein of this invention.
[0011] The functions and activities include, but are not limited to, gRNA binding activity, endonuclease activity, and gRNA-guided binding and cleavage of a specific site of a target sequence (or, referred to as the “target sequence” or “target nucleic acid”), including but not limited to Cis cleavage activity and Trans cleavage activity.
[0012] As is well known in the art, one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N- and / or C-termini of a protein while retaining its functional activity. Therefore, proteins whose N- and / or C-termini of the Cas protein of this invention have been altered, while retaining their desired functional activity, are also within the scope of this invention. These alterations may include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.
[0013] It should be recognized that proteins can also be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known in the art. For example, amino acid sequence variants of Cas proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.
[0014] Those skilled in the art will understand that these minor amino acid changes in the Cas protein of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using rDNA technology, recombinant DNA) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.
[0015] Fusion protein
[0016] The present invention also provides a fusion protein comprising a Cas protein with amino acid sequences SEQ ID No. 1 (Cas-kf63) and SEQ ID No. 4 (Cas-kf66) and other modified portions.
[0017] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or any combination thereof.
[0018] In one embodiment, the modified portion is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting portions, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminase, methyltransferase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.
[0019] The NLS (Nuclear Localization Signal) nuclear localization sequence is well known to those skilled in the art, and examples include, but are not limited to, SV40 large T antigen, Nucleoplasmin, EGL-13, c-Myc, and TUS protein.
[0020] In one embodiment, the NLS sequence is located at, near, or close to the end (e.g., N-terminus, C-terminus, or both ends) of the Cas protein of the present invention.
[0021] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can choose other suitable epitope tags (e.g., for purification, detection, or tracing).
[0022] The reporter gene sequence is well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0023] In one embodiment, the fusion protein of the present invention contains a detectable marker, such as a fluorescent dye, such as FITC or DAPI.
[0024] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.
[0025] In one embodiment, the modified portion is attached to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0026] The Cas protein, protein derivative, or fusion protein of the present invention is not limited by the manner of its production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0027] The nucleic acid (polynucleotide) of the Cas protein.
[0028] On the other hand, the present invention provides an isolated polynucleotide comprising a polynucleotide sequence encoding the Cas protein or fusion protein of the present invention.
[0029] In one embodiment, the polynucleotide sequence is codon-optimized for expression in prokaryotic cells. In another embodiment, the polynucleotide sequence is codon-optimized for expression in eukaryotic cells.
[0030] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0031] In one specific embodiment, the present invention provides a sequence encoding Cas-kf63 as shown in SEQ ID No 2 and a sequence encoding Cas-kf66 as shown in SEQ ID No 5.
[0032] In some embodiments, the polynucleotide sequence has one or more base substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) compared to the sequence shown in SEQ ID No. 2 or 5.
[0033] Those skilled in the art will understand that, based on the degeneracy of the genetic code, most amino acids are encoded by multiple different codons. The genetic code is the process by which genetic information on DNA or mRNA is translated into proteins in an organism. Each amino acid is encoded by one or more specific trinucleotide sequences (codons). Based on this characteristic, different nucleotide sequences (codons) can be transcribed and translated into the same amino acid. As long as the change of the polynucleotide sequence does not lead to the translation into different amino acids or the inactivation of the protein's biological activity, the polynucleotide sequence falls within the scope of this invention.
[0034] carrier
[0035] The present invention also provides a carrier comprising polynucleotides isolated from Cas protein as described above; preferably, it further comprises a regulatory element operatively linked thereto.
[0036] In one embodiment, the regulatory element is selected from one or more of the following: enhancers, transposons, promoters, terminators, leader sequences, polyadenylation sequences, and marker genes.
[0037] In one embodiment, the vector includes cloning vectors, expression vectors, shuttle vectors, and integration vectors, which are vector structures that can be naturally occurring or artificially synthesized and are known to those skilled in the art.
[0038] carrier system
[0039] The present invention provides an engineered, non-naturally occurring vector system, or CRISPR-Cas system, comprising a Cas protein or a nucleic acid sequence encoding the Cas protein and a nucleic acid encoding one or more gRNAs, wherein the gRNAs comprise a repeating sequence for binding the Cas protein and a spacer sequence for targeting a target sequence, the spacer sequence being complementary to the target sequence.
[0040] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more gRNAs are artificially synthesized.
[0041] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more gRNAs do not coexist naturally.
[0042] The one or more gRNAs target one or more target sequences in the cell. The one or more gRNAs guide the Cas protein to the genomic locus of the DNA molecule encoding one or more gene products by hybridizing with the target sequence. Once at the target sequence location, the Cas protein modifies, edits, or cuts the target sequence, thereby altering or modifying the expression of the one or more gene products.
[0043] The cells of this invention include one or more of animals, plants, or microorganisms.
[0044] In some embodiments, the Cas protein is codon-optimized for expression in cells.
[0045] In some embodiments, the Cas protein directs the cleavage of one or both nucleic acid strands at the target sequence location.
[0046] In one specific embodiment, the repeating sequence of the gRNA binding Cas-kf63 protein designed in this invention is shown in SEQ ID No. 3, and the repeating sequence of the gRNA binding Cas-kf66 protein is shown in SEQ ID No. 6.
[0047] The present invention also provides an engineered, non-naturally occurring vector system, which may include one or more vectors, the vectors comprising: a gRNA expression unit and the Cas protein expression unit, wherein the gRNA expression unit and the Cas protein expression unit are located on the same or different vectors of the system.
[0048] The expression unit connects to a variety of regulatory elements, including promoters (e.g., constitutive or inducible promoters), enhancers (e.g., 35S promoter or 35S enhanced promoter), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylate (PolyA) signals and polyU sequences).
[0049] In some implementations, the vector in the system is a viral vector (e.g., a retroviral vector, lentiviral vector, adenovirus vector, and herpes simplex vector), or it can be a plasmid, virus, granule, bacteriophage, or other type known to those skilled in the art.
[0050] In one embodiment, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence represented by TTTV, wherein V is selected from A, G, or C.
[0051] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell.
[0052] In one implementation, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0053] In one embodiment, the target sequence is present within a cell. In another embodiment, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0054] In one embodiment, the Cas protein is linked to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.
[0055] Protein-nucleic acid complexes / compositions
[0056] On the other hand, the present invention provides a composition (or complex) comprising:
[0057] (i) a protein component selected from the Cas protein or the fusion protein; (ii) a nucleic acid component selected from gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a nucleic acid encoding the precursor RNA of the gRNA; wherein the gRNA comprises a repeating sequence that binds to the Cas protein and a spacer sequence for targeting the target sequence, the spacer sequence being complementary to the target sequence.
[0058] The Cas protein and gRNA of this invention can form a binary complex, which is activated upon binding to a target sequence substrate to form an activated CRISPR complex. The target sequence substrate is complementary to the spacer sequence (or guide sequence for hybridization with the target nucleic acid) in the gRNA. In some embodiments, the spacer sequence of the gRNA perfectly matches the target sequence substrate. In other embodiments, the spacer sequence of the gRNA partially (continuously or discontinuously) matches the target sequence substrate, and the activated complex can exhibit trans-cleavage activity. This trans-cleavage activity refers to the non-specific or random cleavage activity of the activated CRISPR complex against other nearby single-stranded nucleic acids, also known in the art as trans-cleavage activity.
[0059] host cells
[0060] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: the Cas protein, fusion protein, polynucleotide, protein-nucleic acid composition, vector or vector system described in the present invention.
[0061] In some implementations, the cell is a prokaryotic cell.
[0062] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (e.g., rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (e.g., chickens), fish, or crustaceans (e.g., clams, shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).
[0063] In some implementations, the cell is a stem cell or stem cell line.
[0064] In some implementations, the cells are microbial cells, such as Saccharomyces cerevisiae.
[0065] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.
[0066] application
[0067] The Cas protein, fusion protein, polynucleotide, vector, CRISPR-Cas system, vector system, composition, or host cell provided by this invention have, in some embodiments, any of the following applications:
[0068] (1) Applications in gene editing or gene targeting for the purpose of diagnosis and treatment of non-human diseases;
[0069] (2) Applications in the preparation of gene editing or gene targeting reagents or kits;
[0070] (3) Applications of gene cutting for purposes other than disease diagnosis and treatment;
[0071] (4) Applications in the preparation of gene cutting reagents or kits;
[0072] (5) Application in nonspecific nucleic acid side branch cleavage;
[0073] (6) Application in nonspecific nucleic acid side branch degradation.
[0074] Gene cleavage as described herein includes: DNA or RNA breaks in a target sequence by the Cas protein described herein (Cis cleavage), and breaks in non-target DNA or RNA (single-stranded nucleic acid substrates) via side branch cleavage (i.e., nonspecific or non-targeted, trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.
[0075] In one embodiment, the present invention also provides a method for editing, targeting, or cutting a target sequence, the method being a method not for disease diagnosis and treatment purposes, the object of the method not being human, the method comprising contacting the target sequence with the CRISPR-Cas system, or the vector system, or the composition, or the host cell.
[0076] In one embodiment, the present invention also provides a kit for gene editing or gene targeting, the kit comprising the CRISPR-Cas system, or the vector system, or the composition, or the host cell. The gene editing or editing target nucleic acid includes gene modification, gene knockout, alteration of gene product expression, mutation repair, insertion of polynucleotides, and gene mutation.
[0077] In one embodiment, the present invention provides a kit for detecting a target sequence in a sample, the kit comprising:
[0078] (a) the Cas protein described herein, or the nucleic acid encoding the Cas protein; and / or
[0079] (b) gRNA, or nucleic acid encoding said gRNA, or precursor RNA containing said gRNA, or nucleic acid encoding said precursor RNA; said gRNA comprising a repeating sequence that binds to said Cas protein and a spacer sequence that targets and is complementary to the target sequence; and / or
[0080] (c) A single-stranded nucleic acid detector, wherein the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.
[0081] In one embodiment, the present invention provides a method for detecting a target sequence in a sample, the method being a method for non-disease diagnosis and treatment purposes, the method comprising contacting the sample with the Cas protein, gRNA, and a single-stranded nucleic acid detector, detecting a detectable signal generated by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting the target gene; the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.
[0082] In this invention, the single-stranded nucleic acid detector includes, but is not limited to, single-stranded DNA, single-stranded RNA, DNA-RNA hybrids, nucleic acid analogs, base modifiers, and other single-stranded nucleic acid detectors; "nucleic acid analogs" include, but are not limited to: locked nucleic acid, bridging nucleic acid, morpholine nucleic acid, ethylene glycol nucleic acid, hexitol nucleic acid, threonine nucleic acid, arabinose nucleic acid, 2'-oxymethyl RNA, 2'-methoxyacetyl RNA, 2'-fluoroRNA, 2'-aminoRNA, 4'-sulfur RNA, and combinations thereof, including optional ribonucleotide or deoxyribonucleotide residues.
[0083] In this invention, the detectable signal is achieved through the following methods: vision-based detection, sensor-based detection, color detection, fluorescence signal-based detection, gold nanoparticle-based detection, fluorescence polarization, fluorescence detection, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based detection.
[0084] In this invention, preferably, a fluorescent group and a quenching group are respectively disposed at both ends of the single-stranded nucleic acid detector, so that the single-stranded nucleic acid detector can exhibit a detectable fluorescent signal after being cleaved. The fluorescent group is selected from one or any combination of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red, or LC RED460; the quenching group is selected from one or any combination of BHQ1, BHQ2, BHQ3, Dabcy1, or Tamra.
[0085] Plant editing vector
[0086] In the specific field of plant gene editing breeding, this invention further constructs a vector for plant gene editing and gene targeting.
[0087] In one embodiment, the editing vector includes a gRNA expression unit, a Cas protein expression unit, and a resistance selection gene expression unit. The gRNA includes a crRNA sequence capable of targeting a target sequence. The Cas protein expression unit expresses the Cas protein. In some embodiments, the resistance selection gene in the resistance selection gene expression unit is used during genetic transformation to screen for successfully transformed cells or organisms. Examples of such genes include, but are not limited to, marker genes such as hygromycin resistance gene (HYG), neomycin resistance gene, tetracycline resistance gene, penicillin resistance gene, streptomycin resistance gene, herbicide resistance gene, heavy metal resistance gene, and antibiotic resistance gene.
[0088] In one specific embodiment, the gRNA expression unit is ZmUBI1 pro:gRNA:Nos term, and / or the Cas protein expression unit is ZmUBI1 pro:Cas-kf63 / Cas-kf66:PsE9 term, and / or the resistance selection gene expression unit is 35S pro:HYG:35S term. In one embodiment, the three expression units are set on the same vector. Specifically, the editing vector uses pCambia3301 as the vector backbone, and the three expression units are set starting from the right boundary of the T-DNA. Based on the expression characteristics of different vectors, the three expression units can be set on different vectors or the same vector. In some embodiments, the optional vectors include pCAMBIA series vectors (such as pCAMBIA1300, pCAMBIA3301, or pCAMBIA3300, etc.), pBI series vectors (such as pBI121), pGreen series vectors (such as pGreenII), and other vectors such as pCHF3, pCambia1380, and other commonly used gene editing vectors in plants.
[0089] In one specific implementation, a BsaI restriction site is left between ZmUBI pro and Nos term, allowing the insertion of one or more gRNA sequences. The gRNAs are tandemly linked using HH and HDV ribozyme technology to form a gene editing vector targeting any site of the target sequence. Specifically, gene editing sites are selected for different regions of the target sequence to be edited, and a recognition sequence that can be paired with the crRNA repeat sequence is designed to form a gRNA sequence. One or more gRNA sequences are cross-linked using HH and HDV ribozyme technology. After linearizing the basic editing vector, one or more gRNA sequences are inserted using DNA ligase or Gibson cloning method to form the target sequence editing vector.
[0090] As a specific embodiment, the gRNA sequences for the target sequences constructed in this invention are shown in SEQ ID NO 7, 8, 9, and 10.
[0091] As a more specific embodiment, the Cas-kf63 target sequence editing vector constructed in this invention has a gRNA expression sequence as shown in SEQ ID NO 13.
[0092] In one embodiment, the present invention provides an engineered bacterium containing any of the editing vectors described above. The editing vector is transferred into plant cells via Agrobacterium-mediated transformation to edit the target gene. Selectable engineered bacteria include Agrobacterium, Bacillus subtilis, and Escherichia coli.
[0093] Furthermore, the present invention provides the application of any of the described editing vectors, or the described engineered bacteria, in plant gene editing breeding. In one embodiment, the present invention provides a method for genetic transformation of plants, wherein the obtained target gene editing vector is introduced into recipient plant cells, and then screened to obtain heritable, non-transgenic, and stably inherited progeny plants.
[0094] Furthermore, the genetic transformation method includes the following steps:
[0095] (1) Genetic transformation of recipient plant cells to obtain regenerated plants (including: 1) induction of callus tissue; 2) subculture of callus tissue; 3) infection of callus tissue; 4) screening and identification of resistant callus tissue; 5) differentiation of callus tissue into seedlings; 6) rooting culture; 7) domestication and transplanting to obtain regenerated plants);
[0096] (2) To detect whether gene editing of the regenerated plants was successful;
[0097] (3) Screening to obtain homozygous edited progeny plants with the target gene trait, non-transgenic, and stably inherited.
[0098] Terminology Definition
[0099] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.
[0100] The term "gRNA (guide RNA)" refers to any polynucleotide sequence that is sufficiently complementary to a target sequence to hybridize with it and guide the CRISPR / Cas system to specifically bind to the target sequence. The structure of a gRNA typically includes a crRNA (CRISPR RNA) sequence derived from a CRISPR array. This sequence contains repeat sequences and spacer sequences. The spacer sequence is derived from foreign genetic material and is complementary to the target sequence. Typically, gRNAs may also have a non-complementary single-stranded region at the 5' end, called a "5' overhang," which helps distinguish the gRNA from the target sequence. In some cases, the 3' end of the gRNA may contain an inversely complementary tail, which helps improve the binding stability of the gRNA to the Cas protein. In some applications, gRNAs may also contain additional modifications, such as phosphorylation or chemical modifications, to improve their stability or functionality.
[0101] The terms “regularly clustered short palindromic repeats (CRISPR), CRISPR-associated (Cas) (CRISPR-Cas) system” or “CRISPR system” are used interchangeably and have the meaning commonly understood by those skilled in the art, which typically includes transcripts or other elements associated with the expression of CRISPR-associated (“Cas”) genes, or transcripts or other elements capable of directing the activity of said Cas genes.
[0102] The term "target sequence" is synonymous with "target sequence" and "target nucleic acid" in this invention, referring to a polynucleotide targeted by a spacer sequence in crRNA. The target sequence can contain any polynucleotide, such as DNA or RNA. Hybridization between the target sequence and crRNA will promote the expression of a CRISPR / Cas system. Perfect complementarity is not required, as long as there is sufficient complementarity to induce hybridization and promote the expression of a CRISPR / Cas system.
[0103] The term "non-natural" as used herein, and the terms "non-natural" or "engineered," are used interchangeably and indicate artificial involvement. When these terms are used to describe nucleic acid molecules or peptides, they refer to those that do not exist in nature or are not synthesized naturally by organisms.
[0104] As used herein, the term “wildtype” has the meaning commonly understood by those skilled in the art as referring to the typical form of an organism, strain, or gene, or the characteristic that distinguishes it from mutant or variant forms when it exists in nature, is separable from its natural source and has not been intentionally modified by humans.
[0105] The term "engineered host cell" refers to cells that can be used to introduce vectors, including but not limited to prokaryotic cells such as Agrobacterium, Escherichia coli, or Bacillus subtilis, as well as eukaryotic cells such as microbial cells, fungal cells, animal cells, or plant cells.
[0106] The term "identity," as used herein, refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are considered identical at that position when a position is occupied by the same base or amino acid monomeric subunit (e.g., a position in each of two DNA molecules occupied by adenine, or a position in each of two polypeptides occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT have 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum possible identity. Such comparisons can be performed, for example, by computer programs, such as the Needleman and Wunsch (J Mol Biol 48:444-453 (1970)) algorithm in the GAP program, using a Blossum 62 matrix or a PAM250 matrix and gap weights of 16, 14, 12, 10, 8, 6 or 4 and length weights of 1, 2, 3, 4, 5 or 6 to determine the percentage identity between two amino acid sequences.
[0107] The term "plant" should be understood as any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development, particularly monocotyledonous or dicotyledonous plants; vegetable crops, including artichokes, kohlrabi, arugula, leeks, asparagus, lettuce, cabbage, cauliflower, broccoli, kale, taro, and melons (e.g., cantaloupe, watermelon); and fruit crops such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, blackberries, grapes, avocados, and bananas. Kiwifruit, persimmon, pomegranate, pineapple, mango, papaya, and lychee, etc.; field crops such as alfalfa, corn / maize (feed corn, sweet corn, popcorn), peanuts, small grain crops (barley, oats, rye, rice, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas, soybeans), oil plants (rapeseed, mustard, poppy, olive, sunflower, coconut, castor oil plants, cocoa beans, peanuts), fiber plants (cotton, flax, jute), etc.
[0108] The technical effects achieved by this invention are as follows:
[0109] Based on the discovered and studied Cas12a effector nuclease sequences, this invention designs, screens, and verifies the functions of novel effector nucleases Cas-kf63 and Cas-kf66, which exhibit nuclease activity and have broad application prospects.
[0110] This invention yielded two novel Cas12a-type effector nuclease proteins, Cas-kf63 and Cas-kf66. Through codon optimization and plant gene editing transformation experiments, these proteins demonstrated the ability to recognize and cleave target sites with TTTV sequences at the 5' end, resulting in novel effector nucleases for gene editing. Furthermore, this invention provides a highly efficient plant gene editing vector for targeted genome modification in plants. Attached Figure Description
[0111] Figure 1 shows the effector nucleases Cas-kf63 and Cas-kf66 and their CRISPR array structures. (In the figure: (a) Cas-kf63 and Cas-kf66 are two types of CRISPR-related Cas12a effector nucleases derived from the bacteria Oscillospiraceae bacterium and Ruminococcus bacterium, respectively. The genome of Cas-kf63 does not contain Cas4, Cas1, or Cas2 enzyme genes, but only contains 5 repeating CRISPR arrays; the genome of Cas-kf66 contains Cas4 and Cas2 enzyme genes and 3 repeating CRISPR arrays, but does not contain Cas1 enzyme gene; (b) The crRNA repeat sequences of Cas-kf63 and Cas-kf66 both contain conserved sequences UCUACU and GUAGAU at their 3' ends, which can form hairpin structures.)
[0112] Figure 2 shows a schematic diagram of the structure of the rice starch synthase gene and the selected editing sites (in the figure: the starch synthase gene Waxy (LOC_Os06g04200) consists of 14 exons and 13 introns. One site each in intron 5, exons 6, 7 and 8 was selected to design OsWx-T3, OsWx-T1, OsWx-T4 and OsWx-T2 gRNAs, and PCR-specific primers OsWx-F5 and OsWx-R1 were designed upstream and downstream of these sites for amplification and sequencing analysis of the edited DNA fragments).
[0113] Figure 3 is a schematic diagram of the gene editing transformation vector KF374 (in the figure: T-DNA RB is the right boundary of T-DNA; ZmUBI1 pro and ZmUBI1 intron are the promoters of the maize ubiquitin gene including introns; HH and HDV are the hammerhead ribozyme and hepatitis delta virus ribozyme, respectively; crRNA is the mature CRISPR RNA repeat sequence of FnCpf1; OsWx-T1, OsWx-T2, OsWx-T3 and OsWx-T4 are gRNAs; Nos term is the terminator of the Agrobacterium tumefaciens nopaline synthase gene; ZmUBI1 pro and ZmUBI1 intron are the promoters of the maize ubiquitin gene including introns; Cas-kf63 is a novel Cas12a effector nuclease belonging to the class 2 type V CRISPR system; PsE9 term is the terminator of the pea ribulose-1,5-bisphosphate carboxylase / oxygenase small subunit E9 protein gene; CaMV 35S pro is the cauliflower mosaic virus 35S promoter; HygR is hygromycin phosphotransferase; CaMV 35S term is the cauliflower mosaic virus 35S terminator; T-DNA LB is the left boundary of T-DNA.
[0114] Figure 4 shows the DNA sequences near the OsWx-T3 and OsWx-T1 editing sites in the target region of partial gene editing events in Cas-kf63.
[0115] Figure 5 shows the DNA sequences near the OsWx-T4 and OsWx-T2 editing sites in the target region of Cas-kf63 partial gene editing events. The 120bp intron7 and exon8 sequences are not shown in the figure.
[0116] Figure 6 shows the DNA sequences near the OsWx-T3 and OsWx-T1 editing sites in the target region of partial gene editing events in Cas-kf66.
[0117] Figure 7 shows the DNA sequences near the OsWx-T4 and OsWx-T2 editing sites in the target region of partial gene editing events in Cas-kf66. The 120bp intron7 and exon8 sequences are not shown in the figure. Detailed Implementation
[0118] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments, but this does not limit the invention to the scope of the embodiments described. Unless otherwise specified, the experiments and methods described in the embodiments are generally carried out according to conventional methods well known in the art and described in various references. The reagents and raw materials used in the present invention are all commercially available. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are followed. Where the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be obtained commercially. Those skilled in the art will understand that the embodiments describe the present invention by way of example and should not be construed as limiting the scope of protection claimed by the present invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.
[0119] In this invention, the molecular biology experiments involving DNA sequence analysis and alignment, DNA vector cloning and design, PCR primers and CRISPR gRNA design were all performed using the DNA analysis software Geneious Prime (GraphPad Software LLC).
[0120] The concept of this invention includes:
[0121] 1. Obtaining novel effector nuclease genes: Based on the amino acid sequences of the Cas12a effector nucleases that have been discovered and studied, we searched microbial genome and metagenomic databases such as NCBI (www.ncbi.nlm.nih.gov), JGI (img.jgi.doe.gov), and UniProt (www.uniprot.org) multiple times using different query programs such as Blastp and Tblastn. We screened for recently sequenced genes with certain homology and containing the RuvC nuclease domain at the carboxyl terminus of the protein. We retrieved the corresponding genomic nucleic acid sequences and used DNA analysis software such as Geneious to find CRISPR repeat sequence arrays and protein-coding regions. Based on their conserved Cas12a-Cas4-Cas1-Cas2-CRISPR gene arrangement, we identified Cas12a effector nucleases. We selected novel Cas12a effector nucleases with low similarity to those reported in the literature and designed experiments for functional verification and screening.
[0122] 2. Establishing a plant gene editing system using novel effector nucleases: A novel Cas12a effector nuclease with low similarity to previously reported nucleases was selected. Suitable plant expression promoters and terminators were chosen to construct gene expression units, and a basic plant transformation expression vector was built. The rice starch synthase gene (Waxy or OsWx) LOC_Os06g04200|Chr6:1765621..1770656forward was selected as the target gene. A Waxy gene-specific gRNA was designed and cloned into the aforementioned plant transformation vector to simultaneously express the novel effector nuclease and its gRNA. This gRNA was then introduced into Agrobacterium EHA105 for rice transformation.
[0123] 3. Novel effector nuclease-mediated plant gene editing: Rice was transformed with the above-mentioned Agrobacterium, resistant callus tissue was selected, genomic DNA was extracted, and the transformation vector-specific gene fragment was amplified by PCR to confirm that the resistant callus was a positive transgenic event. The edited region fragment of the target gene was amplified using the same genomic DNA, and the gene editing effect was analyzed by sequencing. The editing efficiency at different sites was statistically calculated.
[0124] Example 1
[0125] I. Obtaining novel effector nuclease genes and their crRNA repeat sequences
[0126] Microbial genome search: Published Cas12a effector nuclease sequences were collected, and a search library was constructed. Sequences identical to known Cas12 effector nucleases were excluded, and sequences were sorted by similarity. Sequences with low similarity were selected for detailed sequence analysis and refinement. The genomic nucleic acid sequences corresponding to the selected Cas12a effector nucleases were retrieved, and CRISPR repeat sequence arrays were searched using DNA analysis software such as Geneious. Based on the conserved Cas12a-Cas4-Cas1-Cas2-CRISPR gene arrangement, it was determined whether the sequence was a Cas12a effector nuclease. Cas12a effector nucleases with intact and conserved CRISPR structures, appropriate lengths, and recent sequencing, along with their corresponding crRNA repeat sequences, were selected and experimentally designed for functional verification and further screening.
[0127] Through the above-mentioned multi-step screening and subsequent multiple experimental verifications, the inventors obtained several novel type V CRISPR Cas12a effector nucleases. The two that were verified to have cleavage activity were named Cas-kf63 and Cas-kf66, respectively. The genome of Cas-kf63 only has the Cas12a gene and CRISPR repeat sequence array, while the genome of Cas-kf66 has a typical Cas12a-Cas4-Cas1-Cas2-CRISPR gene structure and CRISPR repeat sequence array, but the Cas1 gene is missing (Figure 1a). Cas-kf63 was derived from shotgun sequencing of the Oscillospiraceae genome, and its protein, nucleotide, and CRISPR repeat sequences are shown in SEQ ID NO 1-3; Cas-kf66 was derived from shotgun sequencing of the Ruminococcus genome, and its protein, nucleotide, and CRISPR repeat sequences are shown in SEQ ID NO 4-6; the protein contains artificially added SV40 and Nucleoplasmin nuclear localization signal (NLS) sequences at both ends, and the encoding nucleotides have been optimized to codon sequences suitable for eukaryotic cell expression; the 3' end of the crRNA repeat sequence has conserved UCUACU and GUAGAU sequences that can form hairpin structures (Figure 1b).
[0128] Example 2: Establishment of a plant gene editing vector using a novel effector nuclease
[0129] Beidahuang Kenfeng Seed Industry has independently developed a rice-specific gene-editing Agrobacterium-mediated transformation DNA vector. Using the pCambia3301 vector as its backbone, it contains three gene expression units: ZmU6 pro:gRNA:AtU6-26 term, ZmUBI1 pro:SpCas9:PsE9 term, and 35S pro:HYG:35S term. A BsaI restriction site is left between ZmU6 pro and gRNA. After linearization, one or more gRNA sequences can be inserted using DNA ligase or Gibson cloning methods to form a basic gene-editing vector targeting any site. Specific sequence information can be found in patent CN113604501B, "Gene Editing Method and Application of Aroma Control Gene in Improved Indica Rice Lines," specifically the KF80 transformation vector. All gene-editing transformation vectors described in this invention are modified and constructed based on this disclosed basic gene-editing vector.
[0130] Using the KF80 transformation vector as the basic backbone, ZmU6 pro was replaced by ZmUBI1 pro, AtU6-26 term was replaced by Nos term, SpCas9 was replaced by LbCpf1 gene, and gRNA was replaced by artificially synthesized HH-OsWx-T1-HDV-HH-OsWx-T2-HDV-HH-OsWx-T3-HDV-HH-OsWx-T4-HDV (SEQ ID NO 13). Among them, OsWx-T1, OsWx-T2, OsWx-T3 and OsWx-T4 are gRNAs that edit the corresponding sites of the rice starch synthase gene (OsWx) (SEQ ID NO 7, 8, 9, 10), and HH and HDV are the ribozymes of hammerhead and hepatitis delta virus, respectively (SEQ ID NO 11, 12). The final gene editing vector KF365, based on the LbCpf1 effector nuclease, was developed, containing three gene expression units: ZmUBI1 pro:gRNA (SEQ ID NO 13):Nos term, ZmUBI1 pro:LbCpf1:PsE9 term, and 35S pro:HYG:35S term. These served as positive control gene editing vectors to edit the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites in rice. Subsequently, the LbCpf1 gene in KF365 was replaced with the Cas-kf63 and Cas-kf66 (SEQ ID NO 2, 5) genes, respectively, to form gene editing vectors KF374 and KF385 based on the Cas-kf63 and Cas-kf66 effector nucleases (Table 2). These vectors were then used to edit the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites in rice. The composition and sequence of the KF374 vector are shown in Figure 3 and SEQ ID NO 14, respectively. Other vectors, except for the different effector nucleases used, have the same composition and sequence as KF374, and therefore are not listed further. All the above gene editing vectors were transformed into Agrobacterium EHA105 and stored at -80℃ for later use.
[0131] III. Novel effector nuclease-mediated plant gene editing
[0132] Gene editing transformation in rice: Gene editing vectors based on novel effector nucleases must be transformed into plant cells to verify their function. Using Agrobacterium-mediated transformation, healthy mature seeds of the rice variety Kendao 32 were used as explants to induce callus tissue. Gene editing vectors KF374 and KF385 were transformed into cells, respectively. Resistance transgenic events were screened, and the DNA sequences of the target genes OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites were examined for specific changes to evaluate the gene editing effects of each effector nuclease in plants. The formulations of various basic culture media used in the genetic transformation experiments are listed in Table 1, and the specific operating procedures are as follows:
[0133] Induction of rice callus: For each transformation experiment, select approximately 200 mature, plump, healthy, and clean rice seeds, remove the husks, and place them in a 100 ml sterile glass bottle. Wash the seeds three times with 75% ethanol, add an equal volume of sterile water and 10% sodium hypochlorite (NaClO), and two drops of Tween 20. Sterilize by shaking for 20 minutes, and wash with sterile water 3-5 times until the foam is removed. Blot the seeds dry with sterile filter paper, place them flat on the induction medium, and incubate in the dark at 33℃ for 3-4 weeks until callus tissue grows from the embryo.
[0134] Subculture of rice callus: Select vigorous healthy callus tissue, place it on a new induction medium and subculture for 1-2 weeks to proliferate and maintain the vitality of the callus tissue.
[0135] Preparation of Agrobacterium: Remove the Agrobacterium glycerol storage tube carrying the gene-editing vector from the -80℃ freezer. Streak a small amount of Agrobacterium onto a solid YEP medium plate containing 50 mg / L kanamycin and incubate in the dark at 28℃ for 2-3 days. Pick a single colony and inoculate it into 5 mL of liquid YEP medium, incubating overnight in the dark at 28℃. Extract plasmid DNA from a small amount of bacterial culture for vector-specific PCR to confirm that the strain carries the correct gene-editing vector. Simultaneously, plate a portion of the bacterial culture onto a YEP medium plate for observation. The Agrobacterium used for infection must be free of contamination. The entire plate should be smooth and free of granular or other colored fungi or molds to ensure the purity of the obtained Agrobacterium. Strict aseptic technique must be maintained to ensure no contamination in subsequent work. After confirming the Agrobacterium plate is free of contamination, directly scrape the bacterial cells into a 100 mL Erlenmeyer flask containing 25 mL of suspension infection medium and incubate at 100-120 rpm on a rotating shaker at 25℃ for 2-3 hours. Take samples to measure OD. 600 The OD value of Agrobacterium was adjusted using suspension infection culture medium. 600 A value of 0.1-0.2 is used for callus infection.
[0136] Infection of callus tissue: Select small granular, healthy callus tissue into a 250 ml sterile glass bottle, add Agrobacterium and shake to infect for 3-5 minutes, then use sterile filter paper to absorb the bacterial solution on the surface of the callus tissue, place it on sterile filter paper that has been placed on the surface of the co-culture medium (the filter paper prevents the growth of Agrobacterium), and co-culture at 28°C in the dark for 3 days.
[0137] Screening and identification of resistant callus: Infected rice callus was transferred to a sterile 250 mL glass bottle and washed several times with sterile water until the water was clear. Finally, it was soaked and washed for 0.5 hours on a shaker at 100 rpm with carbenicillin at a final concentration of approximately 250 mg / L. The callus was then transferred to a blank culture dish, and the surface moisture was blotted dry with sterile filter paper. After air-drying in a sterile laminar flow hood for 2.5 hours, the callus was aliquoted and placed on screening medium. The medium was then incubated in the dark at 33°C for 3-4 weeks. The screening medium contained 400 mg / L carbenicillin to inhibit Agrobacterium growth and 50 mg / L hygromycin to screen for transformed cells. Non-transformed cells stopped growing and gradually died on the screening medium. After 3-4 weeks of culture, successfully transformed cells grew resistant callus. After the resistant callus continued to grow, a suitable amount was sampled, genomic DNA was extracted, and PCR identification was performed using primers specific to the transformation vector. Each PCR-positive callus was considered an independent transformation event.
[0138] Table 1. Composition and preparation method of basic culture medium used in Agrobacterium genetic transformation experiments.
[0139] 2. Analysis of gene editing effects:
[0140] Resistance transformation events need to be detected using molecular biology methods to confirm the success of genetic transformation. A commonly used, simple, and reliable detection method is PCR amplification. Rice resistance callus tissue was sampled, and genomic DNA was extracted using a rapid DNA extraction method based on SDS (Sodium dodecyl sulfate). First, vector-specific PCR detection of the HygR gene was performed to identify the transformation event. The PCR reaction system used was the Quick Taq HSDyeMix (DTM-101) kit (TOYOBO Life Science). A 20 μL PCR reaction system contained 10 μL 2x Quick Taq HSDyeMix, 1.0 μL 10 pmol / μL primer Hyg-F1 (SEQ ID NO 15) and 1.0 μL 10 pmol / μL primer Hyg-R1 (SEQ ID NO 16), 6.0 μL sterile water, and finally 2.0 μL 50 ng / μL of the sample's genomic DNA. The PCR reaction conditions were as follows: denaturation at 95°C for 5 minutes, followed by denaturation at 95°C for 30 seconds, annealing at 60°C for 1 minute, extension at 72°C for 1 minute, for a total of 30 cycles, and a final extension at 72°C for 7 minutes, followed by holding at 4°C. The PCR amplification products were separated by 1% agarose gel electrophoresis. Samples that successfully amplified a 402bp long fragment specific to the HygR gene (SEQ ID NO 17) were identified as transgenic positive.
[0141] Using a similar PCR amplification method and specific primers OsWx-F5 and OsWx-R1 (SEQ ID NO 18, 19) for OsWx(LOC_Os06g04200), target fragments (SEQ ID NO 20) near the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 gene editing sites can be amplified from HygR-positive samples. After purification, these fragments are directly sequenced using the same upstream or downstream primers and compared with the same gene fragment sequence from the wild-type Kendao 32. If the DNA fragment sequence of the transgenic event is indistinguishable from the wild-type sequence, gene editing has not occurred. If overlapping peaks appear near the expected effector nuclease cleavage sites, the sequenced PCR fragment is a mixed template containing both edited and unedited DNA fragments, indicating heterozygous gene editing. If there is a definite sequence difference from the wild-type, it indicates homozygous gene editing. The proportion of fragments with edited sites detected in the sequenced target gene PCR fragment represents the overall gene editing efficiency. Gene editing experiments confirmed that the novel effector nucleases Cas-kf63 and Cas-kf66 have gene editing functions, with overall gene editing efficiencies of 35.4% and 89.3%, respectively (Table 2).
[0142] Table 2 Gene editing efficiency of novel effector nucleases
[0143] *The length of an enzyme protein includes the amino acid sequence of short peptides located at both ends of the cell nucleus.
[0144] **Editing efficiency = Number of editing events at any site / Total number of sequencing events for the target gene PCR fragment.**
[0145] The TOPO PCR fragment of the target gene from the Cas-kf63 editing event was cloned into the pCR2.1 vector (Thermo Fisher Scientific). Four clones from each sample were sequenced to accurately analyze the specific editing changes in the DNA sequence at the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites. Different clones from the same callus editing event often exhibited multiple editing sequence changes. Some clones only edited at the OsWx-T3 or OsWx-T2 sites, while others edited at both sites simultaneously. Four clones showed editing at all three sites (OsWx-T3, OsWx-T1, and OsWx-T2), but no editing occurred at the OsWx-T4 site. The editing efficiency among the four sites showed a significant difference: OsWx-T3 > OsWx-T2 > OsWx-T1 > OsWx-T4 (Table 3). Cloning and sequencing revealed that the novel effector nuclease Cas-kf63 selectively edits sites with different PAM sequences. It exhibits high editing efficiency for TTTA (OsWx-T3) and TTTC (OsWx-T2), very low efficiency for TTTG (OsWx-T1), and no ability to edit sites lacking typical TTTV (V = A, C, or G) PAMs (OsWx-T4). The specific editing changes of the DNA sequences near the OsWx-T3 and OsWx-T1 sites, and the OsWx-T4 and OsWx-T2 sites in the 20 clones marked with an asterisk in Table 3 are shown in Figures 4 and 5, respectively.
[0146] Table 3. Gene site editing efficiency of the effector nuclease Cas-kf63
[0147] Note: *The aligned sequences of the 20 selected clones are shown in Figures 4 and 5; NE (not edited) is unedited.
[0148] The TOPO PCR fragment of the target gene from the Cas-kf66 editing event was cloned into the pCR2.1 vector (Thermo Fisher Scientific). Four clones from each sample were sequenced to accurately analyze the specific editing changes in the DNA sequence at the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites. Different clones from the same callus editing event often exhibit multiple editing sequence changes. Some clones only edited at the OsWx-T3 or OsWx-T2 sites, while others edited at both sites simultaneously. No clones showing editing at all three sites (OsWx-T3, OsWx-T1, and OsWx-T2) were detected, while no editing occurred at the OsWx-T4 site. The editing efficiency among the four sites showed a significant difference: OsWx-T3 > OsWx-T2 > OsWx-T1 > OsWx-T4 (Table 4). Cloning and sequencing revealed that the novel effector nuclease Cas-kf66 selectively edits sites with different PAM sequences. It exhibits high editing efficiency for TTTA (OsWx-T3) and TTTC (OsWx-T2), very low efficiency for TTTG (OsWx-T1), and no ability to edit sites lacking typical TTTV (V = A, C, or G) PAMs (OsWx-T4). The specific editing changes of the DNA sequences near the OsWx-T3 and OsWx-T1 sites, and the OsWx-T4 and OsWx-T2 sites in the 20 clones marked with an asterisk in Table 4 are shown in Figures 6 and 7, respectively.
[0149] Table 4. Gene site editing efficiency of the effector nuclease Cas-kf66
[0150] Note: *The aligned sequences of the 20 selected clones are shown in Figures 6 and 7; NE (not edited) is unedited.
[0151] The above target gene sequence analysis results revealed that gene editing mediated by Cas-kf63 and Cas-kf66 effector nucleases is mostly small fragment deletion, which mostly occurs at the expected target cleavage site in the distal part of the target gene PAM, consistent with other known Cas12a effector nucleases such as LbCpf1 and FnCpf1.
[0152] This invention identified several novel Cas12a-class effector nuclease proteins. Through codon optimization and plant gene editing transformation experiments, it was found that Cas-kf63 and Cas-kf66 possess the ability to recognize and cleave target sites with TTTV sequences at the 5' end, resulting in two novel effector nucleases suitable for plant gene editing. Furthermore, a highly efficient plant gene editing vector is provided, enabling effective targeted genome modification in plants.
[0153] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Based on all the teachings that have been published, various modifications and changes can be made to the details. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention should be considered equivalent substitutions and are included within the protection scope of the present invention.
[0154] DNA sequence description in this invention
[0155] SEQ ID NO 1:Cas-kf63 protein 1227aa
[0156] SEQ ID NO 2: Cas-kf63 coding DNA 3684bp
[0157] SEQ ID NO 3:Cas-kf63 crRNA repeat DNA 36bp
[0158] SEQ ID NO 4:Cas-kf66 protein 1258aa
[0159] SEQ ID NO 5: Cas-kf66 coding DNA 3777bp
[0160] SEQ ID NO 6:Cas-kf66 crRNA repeat DNA 36bp
[0161] SEQ ID NO 7:synthesized OsWx-T1 fragment 27bp
[0162] SEQ ID NO 8:synthesized OsWx-T2 fragment 27bp
[0163] SEQ ID NO 9:synthesized OsWx-T3 fragment 27 bp
[0164] SEQ ID NO 10:synthesized OsWx-T4 fragment 27 bp
[0165] SEQ ID NO 11:synthesized hammerhead ribozyme 37 bp
[0166] SEQ ID NO 12:synthesized hepatitis delta virus ribozyme 68 bp
[0167] SEQ ID NO 13:synthesized HH-OsWx-T1-HDV-HH-OsWx-T2-HDV-HH-OsWx-T3-HDV-HH-OsWx-T4-HDV DNA fragment 628 bp
[0168] SEQ ID NO 14:KF374 T-DNA 11678 bp
[0169] SEQ ID NO 15:Hyg-F1 oligo 22 bp
[0170] SEQ ID NO 16:Hyg-R1 oligo 22 bp
[0171] SEQ ID NO 17:HygR gene PCR fragment 402 bp
[0172] SEQ ID NO 18:OsWx-F5 oligo 22 bp
[0173] SEQ ID NO 19:OsWx-R1 oligo 22 bp
[0174] SEQ ID NO 20:Oryza sativa Kendao32 Waxy gene OsWx(LOC_Os06g04200)PCR fragment 916 bp
Claims
1. A Cas protein, characterized in that, The Cas protein is Cas-kf63 or Cas-kf66, the Cas-kf63 amino acid sequence is shown as SEQ ID No 1, and the Cas-kf66 amino acid sequence is shown as SEQ ID No 4.
2. A fusion protein comprising the Cas protein of claim 1 and other modified parts.
3. An isolated polynucleotide, comprising, The polynucleotide is a polynucleotide sequence encoding the Cas protein of claim 1 or a polynucleotide sequence encoding the fusion protein of claim 2; preferably, the nucleotide sequence encoding Cas-kf63 is shown as SEQ ID No 2, and the nucleotide sequence encoding Cas-kf66 is shown as SEQ ID No 5.
4. A vector, characterized by, The vector comprises the polynucleotide of claim 3 and regulatory elements operably linked thereto.
5. A CRISPR-Cas system, characterized in that, The system comprises the Cas protein of claim 1 and at least one gRNA, the gRNA comprising a repeat sequence binding to the Cas protein and a spacer sequence for targeting a target sequence; preferably, the repeat sequence binding to the Cas-kf63 gRNA is shown as SEQ ID No 3, and the repeat sequence binding to the Cas-kf66 gRNA is shown as SEQ ID No 6.
6. A vector system characterized by comprising The vector system comprises one or more vectors comprising: a gRNA expression unit and a Cas protein expression unit of claim 1, the gRNA comprising a crRNA sequence capable of targeting a target sequence, the gRNA expression unit and the Cas protein expression unit being located on the same or different vectors of the system.
7. A composition characterized in that, The composition comprises: (i) a protein component selected from: the Cas protein of claim 1 or the fusion protein of claim 2; (ii) a nucleic acid component selected from: a gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a nucleic acid encoding the precursor RNA of the gRNA; the gRNA comprising a repeat sequence binding to the Cas protein and a spacer sequence for targeting a target sequence.
8. An engineered host cell, characterized in that, The host cell comprises the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7.
9. Use of the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4 or the CRISPR-Cas system of claim 5, or the vector system of claim6, or the composition of claim 7, or the host cell of claim 8 in any one of: (1) gene editing or gene targeting for non-human disease diagnosis and treatment purposes; (2) preparation of gene editing or gene targeting reagents or kits; (3) gene cleavage for non-disease diagnosis and treatment purposes; (4) other purposes. (4) use in preparing a gene cleavage reagent or kit; (5) use in non-specifically cleaving a collateral nucleic acid; (6) use in non-specifically degrading a collateral nucleic acid.
10. A kit for gene editing or gene targeting, characterized in that, The kit comprises the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7, or the host cell of claim 8.
11. A kit for detecting a target sequence in a sample, characterized in that, The kit comprises: (a) the Cas protein of claim 1, or a nucleic acid encoding the Cas protein; and / or (b) a gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA comprising the gRNA, or a nucleic acid encoding the precursor RNA; the gRNA comprising a crRNA sequence capable of targeting a target sequence; and / or (c) a single-stranded nucleic acid detector, which is a single-stranded nucleic acid that does not hybridize to the gRNA.
12. A method of editing or targeting a target sequence, comprising, The method is a method for non-disease diagnosis and treatment purposes, and the subject to which the method is applied does not include humans, and the method comprises contacting a target sequence with the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7, or the host cell of claim 8.
13. A method of cleaving a target sequence, comprising, The method is a method for non-disease diagnosis and treatment purposes, and the subject to which the method is applied does not include humans, and the method comprises contacting a target sequence with the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7, or the host cell of claim 8.
14. A method of detecting a target sequence in a sample, characterized in that, The method is a method for non-disease diagnosis and treatment purposes, and the method comprises contacting a sample with the Cas protein, gRNA and single-stranded nucleic acid detector of claim 11, detecting a detectable signal generated by the cleavage of the single-stranded nucleic acid detector by the Cas protein, thereby detecting a target sequence; the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize to the gRNA.
15. An editing vector, comprising, The vector comprises a gRNA expression unit, a Cas protein expression unit and a resistance screening gene expression unit, the gRNA comprising a crRNA sequence capable of targeting a target sequence, and the Cas protein expression unit expressing the Cas protein of claim 1.
16. The editing vector of claim 15, wherein, The gRNA expression unit is ZmUBI1 pro:gRNA:Nos term, and / or the Cas protein expression unit is ZmUBI1 pro:Cas-kf63 / Cas-kf66:PsE9 term, and / or the resistance screening gene expression unit is 35S pro:HYG:35S term; preferably, the editing vector takes pCambia3301 as the vector backbone, and three expression units are arranged starting from the T-DNA right border.
17. The editing vector of claim 16, wherein, A BsaI enzyme cleavage site is left between ZmUBI pro and Nos term, and one or more gRNA sequences can be inserted, and each gRNA is connected in series by HH and HDV ribozyme technology to form a gene editing vector for any target site of a target sequence.
18. An engineered bacterium, characterized in that, The editing vector according to any one of claims 15-17.
19. Use of the editing vector according to any one of claims 15-17 or the engineering bacteria according to claim 18 in plant gene editing breeding.
20. A method for genetic transformation of a plant, comprising: The editing vector according to any one of claims 15-17 is transformed into a recipient plant cell, and a heritable, non-transgenic and stably inherited progeny plant is screened; preferably, the process comprises the following steps: (1) genetic transformation of the recipient plant cell to obtain a regenerated plant; (2) detection of whether the gene editing of the regenerated plant is successful; (3) screening of a progeny plant with homozygous editing of the target sequence, non-transgenic and stable inheritance.