Novel effector nuclease and application thereof
By designing and screening new effector nucleases Cas-kf63 and Cas-kf66, the activity and stability problems of the existing CRISPR/Cas system when applied in eukaryotes are solved, and efficient gene editing function in eukaryotic cells is achieved.
Patent Information
- Application Number
- CN202580000088.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-03
AI Technical Summary
The existing CRISPR/Cas gene editing system has problems with activity and stability when applied in eukaryotes, and the homology between Cas12a enzyme proteins of different sources is low, making it difficult to directly judge their biological functions.
Design, screen and perform functional verification to obtain novel effector nucleases Cas-kf63 and Cas-kf66, which have the gene editing function of eukaryotic cells, and develop new CRISPR/Cas system and gene editing methods.
It realizes effective gene editing functions in eukaryotic cells, provides a new CRISPR/Cas system and gene editing method, and improves the accuracy and efficiency of gene editing.
Smart Images

Figure CN120092084A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biotechnology and relates to a novel effector nuclease Cas-kf63 or Cas-kf66 associated with clustered regularly interspaced short palindromic repeats (CRISPR), and an application thereof in the field of gene editing. Background Art
[0002] Gene editing technology has developed rapidly in recent years. It uses the effector nuclease Cas (CRISPR-associated) associated with the regularly interspaced short palindromic repeats (CRISPR) in the bacterial acquired immune system to cut specific genes in cells at pre-selected sites in the genome. Through two DNA repair mechanisms inherent to cells, Non-Homologous End Joining (NHEJ) and Homology Dependent Repair (HDR), the breaks are repaired to ensure cell survival. NHEJ has high repair efficiency and does not require a template. The restoration of the linking process at the cut point will introduce insertion or deletion mutations, thereby inactivating the gene, which is dominant. HDR requires a repair template and achieves precise repair according to the template. It accounts for a very small proportion. If a pre-designed DNA fragment is provided as a template, the fragment sequence can be inserted into the pre-selected site to modify the target gene, but the efficiency is extremely low, which seriously restricts its widespread application.
[0003] The acquired immune system containing CRISPR structures widely exists in bacteria and archaea. It is divided into two major categories, 1 and 2, according to the number of proteins involved in the immune interference reaction. Class 1 systems require multiple proteins to form a complex to perform immune interference functions, while class 2 only requires a single protease with multifunctional groups to achieve interference (Makarova et al., 2020, Nat Rev Microbiol 18, 67 - 83; Nidhi et al., 2021, Int J Mol Sci 22, 3327). Each class of system is further divided into multiple types according to structure, composition, etc. The commonly used SpCas9 is a typical representative of the class 2 type II CRISPR system, requiring a conserved NGG at the 3' end of the target sequence as the PAM sequence and performing blunt-end cleavage on the target sequence (Jinek et al., 2012, Science 337, 816 - 821). Other effector nucleases of class 2 type II CRISPR systems, such as SaCas9 derived from Staphylococcus aureus and NmeCas9 derived from Neisseria meningitidis, have also been successfully used for gene editing (Ran et al., 2015, Nature 520, 186 - 191; Amrani et al., 2018, Genome Biol 19, 214).
[0004] The Cas12 effector nucleases of the type V CRISPR system, which also belongs to class 2, such as FnCpfI and LbCpf1 (later also known as FnCas12a and LbCas12a) derived from the bacterial strains Francisella tularensis subsp. Novicida U112 and Lachnospiraceae bacterium ND2006 respectively, require a conserved TTTN at the 5' end of the target sequence as the PAM sequence and perform sticky-end cleavage on the target sequence (Zetsche et al., 2015, Cell 163, 1 - 13). Other Cas12a endonucleases of multiple class 2 type V CRISPR systems have also been used for gene editing to varying degrees. Most of them mainly recognize target DNA sequences containing a conserved PAM sequence with TTTV (V is A, C, or G) at the 5' end and cleave the DNA double strand at the distal end of the PAM, generating sticky ends. In addition, the Cas12a effector nuclease also has RNase function and can autonomously process and cleave the precursor pre-crRNA of the CRISPR repeat array to form a single mature crRNA, guiding the recognition of the target DNA sequence.
[0005] Although Cas12a effector nucleases have been identified in many bacteria based on the principle of homology, the homology between protease sequences from different sources is not high, and there is often only about 30% amino acid sequence similarity between enzyme proteins with a relatively large genetic distance. In addition to the difficulty in directly judging the biological function of a specific Cas12a enzyme based on the sequence due to low homology, not every enzyme protein is an active protein with immune function. Some Cas12a immune systems are incomplete, lacking other necessary components such as auxiliary proteins like Cas1, Cas2, Cas4, etc., or the CRISPR array is missing (Yan et al., 2019, Science 363, 88 - 91). Even for Cas12a enzyme proteins with complete composition and sequence, they may not necessarily have nuclease activity. Only a very small number may still maintain activity after expression in eukaryotes. Therefore, it must be verified through specific experiments. Currently, the discovered CRISPR / Cas gene editing systems each have different characteristics. Based on various reasons and complex and diverse research and application requirements, continuously developing new gene editing tool enzymes is of great significance for the development and application research of biotechnology. Summary of the Invention
[0006] Based on the discovered and studied Cas12a effector nuclease sequences, the present invention designed, screened, and verified the functions, and obtained the novel effector nucleases Cas-kf63 and Cas-kf66 with gene editing functions in eukaryotic cells. Based on this discovery, the present invention developed a new CRISPR / Cas system and a gene editing method based on this system.
[0007] Cas protein
[0008] On the one hand, the present invention provides a Cas protein (or, referred to as "effector nuclease", "Cas enzyme", "effector protein"). The Cas protein is an effector protein in the CRISPR / Cas system. The Cas protein is Cas-kf63 and Cas-kf66. The amino acid sequence of Cas-kf63 is shown in SEQ ID No 1, and the amino acid sequence of Cas-kf66 is shown in SEQ ID No 4.
[0009] Compared with the amino acid sequence of the Cas protein and the sequence of SEQ ID No 1 or 4, in some embodiments, the Cas protein is a sequence with one or more amino acid substitutions, deletions or additions, and substantially retains its biological function derived from the sequence; the one or more amino acids include substitutions, deletions or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids; in some embodiments, the amino acid sequence has at least 50% sequence identity with the sequence shown in SEQ ID No 1 or 4, for example, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% identity, and substantially retains its biological function derived from the sequence.
[0010] Those skilled in the art are aware that the primary structure of a protein can be altered without adversely affecting its activity and functionality. For example, one or more conservative amino acid substitutions can be introduced into the amino acid sequence of a protein without adversely affecting the activity and / or three-dimensional structure of the protein molecule. Those skilled in the art are aware of examples and embodiments of conservative amino acid substitutions. Specifically, an amino acid residue can be replaced with another amino acid residue belonging to the same group as the site to be replaced, i.e., a non-polar amino acid residue is replaced with another non-polar amino acid residue, a polar uncharged amino acid residue is replaced with another polar uncharged amino acid residue, a basic amino acid residue is replaced with another basic amino acid residue, and an acidic amino acid residue is replaced with another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. As long as the substitution does not result in the inactivation of the biological activity of the protein, conservative substitutions in which one amino acid is replaced by another amino acid belonging to the same group fall within the scope of the present invention. Additionally, the present invention also encompasses proteins containing one or more other non-conservative amino acid substitutions, as long as such non-conservative substitutions do not significantly affect the desired functions and biological activities of the proteins of the present invention.
[0011] The functions and activities include, but are not limited to, the activity of binding to gRNA, endonuclease activity, and the activity of binding to and cleaving specific sites of a target sequence (or, referred to as "target sequence", "target nucleic acid") under the guidance of gRNA, including but not limited to Cis cleavage activity and Trans cleavage activity.
[0012] As is well known in the art, one or more amino acid residues can be altered (substituted, deleted, truncated or inserted) from the amino N- and / or carboxyl C-terminus of a protein while still retaining its functional activity. Accordingly, proteins in which one or more amino acid residues have been altered from the N- and / or C-terminus of the Cas proteins of the present invention while retaining their desired functional activity are also within the scope of the present invention. These alterations can include those introduced by modern molecular methods such as PCR, which include PCR amplification that modifies or extends a protein coding sequence by virtue of including an amino acid coding sequence among the oligonucleotides used in the PCR amplification.
[0013] It should be appreciated that proteins can also be altered in a variety of ways, including amino acid substitutions, deletions, truncations and insertions, and methods for such manipulations are generally known in the art. For example, amino acid sequence variants of Cas proteins can be prepared by mutagenesis of the DNA. It can also be accomplished by other forms of mutagenesis and / or by directed evolution, for example, using known mutagenesis, recombination and / or shuffling methods, in combination with relevant screening methods, to effect single or multiple amino acid substitutions, deletions and / or insertions.
[0014] Those skilled in the art will appreciate that these minor amino acid changes in the Cas proteins of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using rDNA technology, recombinant DNA) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site or other functional domain of the protein, the properties of the polypeptide can be altered, but the polypeptide can retain its activity. If the mutations present are not close to the catalytic domain, active site or other functional domain, a lesser effect can be expected.
[0015] Fusion protein
[0016] The present invention also provides a fusion protein comprising a Cas protein having the amino acid sequences of SEQ ID No 1 (Cas-kf63) and SEQ ID No 4 (Cas-kf66) and other modifying moieties.
[0017] In one embodiment, the modifying moiety is selected from additional proteins or polypeptides, detectable labels or any combination thereof.
[0018] In one embodiment, the modifying moiety is selected from an epitope tag, a reporter gene sequence, a nuclear localization signal (NLS) sequence, a targeting moiety, a transcriptional activation domain (e.g., VP64), a transcriptional repression domain (e.g., KRAB domain or SID domain), a nuclease domain (e.g., Fok1), and a domain having an activity selected from the following: nucleotide deaminase, methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.
[0019] The NLS (Nuclear Localization Signal) nuclear localization sequence is well known to those skilled in the art, and its examples include, but are not limited to, SV40 large T antigen, Nucleoplasmin, EGL-13, c-Myc, and TUS protein.
[0020] In one embodiment, the NLS sequence is located at, near, or close to the end (e.g., N-terminus, C-terminus, or both ends) of the Cas protein of the present invention.
[0021] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can select other suitable epitope tags (e.g., for purification, detection, or tracing).
[0022] The reporter gene sequence is well known to those skilled in the art, and its examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.
[0023] In one embodiment, the fusion protein of the present invention comprises a detectable label, such as a fluorescent dye, such as FITC or DAPI.
[0024] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated, or fused to the modifying moiety via a linker.
[0025] In one embodiment, the modifying moiety is linked to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and their examples include, but are not limited to, linkers comprising one or more (e.g., 1, 2, 3, 4, or 5) amino acids (such as Glu or Ser) or amino acid derivatives (such as Ahx, β-Ala, GABA, or Ava), or PEG, etc.
[0026] The Cas protein, protein derivative or fusion protein of the present invention is not limited by its production method. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.
[0027] Nucleic acid (polynucleotide) of Cas protein
[0028] On the other hand, the present invention provides an isolated polynucleotide comprising a polynucleotide sequence encoding the Cas protein or fusion protein of the present invention.
[0029] In one embodiment, the polynucleotide sequence is codon-optimized for expression in prokaryotic cells. In one embodiment, the polynucleotide sequence is codon-optimized for expression in eukaryotic cells.
[0030] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.
[0031] In a specific embodiment, the present invention provides that the sequence encoding Cas-kf63 is as shown in SEQ ID No 2, and the sequence encoding Cas-kf66 is as shown in SEQ ID No 5.
[0032] In some embodiments, the polynucleotide sequence has one or more base substitutions, deletions or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more base substitutions, deletions or additions) compared to the sequence shown in SEQ ID No 2 or 5.
[0033] Those skilled in the art are aware that due to the degeneracy in the genetic code, most amino acids are encoded by multiple different codons. The genetic code is the process in living organisms of translating genetic information on DNA or mRNA into proteins, where each amino acid is encoded by one or more specific trinucleotide sequences (codons). Based on this property, different nucleotide sequences (codons) can be transcribed and translated into the same amino acid. As long as the change in the polynucleotide sequence does not result in translation into a different amino acid or inactivation of the biological activity of the protein, the polynucleotide sequence falls within the scope of the present invention.
[0034] Vector
[0035] The present invention also provides a vector comprising the isolated polynucleotide of the Cas protein as described above; preferably, it further includes regulatory elements operably linked thereto.
[0036] In one embodiment, the regulatory element is selected from one or more of the following groups: enhancer, transposon, promoter, terminator, leader sequence, polyadenylation sequence, marker gene.
[0037] In one embodiment, the vector includes a cloning vector, an expression vector, a shuttle vector, an integration vector, which are vector structures that are known to those skilled in the art and can be naturally occurring or synthetic.
[0038] Vector system
[0039] The present invention provides an engineered non-naturally occurring vector system, or a CRISPR-Cas system, which includes a Cas protein or a nucleic acid sequence encoding the Cas protein and a nucleic acid encoding one or more gRNAs. The gRNA includes a repeat sequence that binds to the Cas protein and a spacer sequence for targeting a target sequence, and the spacer sequence is complementary to the target sequence.
[0040] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more gRNAs are synthetic.
[0041] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more gRNAs do not co-exist naturally.
[0042] The one or more gRNAs target one or more target sequences in a cell. The one or more gRNAs hybridize with the target sequences at the genomic loci of DNA molecules encoding one or more gene products, guiding the Cas protein to the genomic loci of the DNA molecules of the one or more gene products. After the Cas protein reaches the target sequence position, it modifies, edits, or cleaves the target sequence, thereby altering or modifying the expression of the one or more gene products.
[0043] The cells of the present invention include one or more of animals, plants, or microorganisms.
[0044] In some embodiments, the Cas protein is codon-optimized for expression in cells.
[0045] In some embodiments, the Cas protein directs the cleavage of one or both nucleic acid strands at the target sequence position.
[0046] In a specific embodiment, the repeat sequence of the gRNA that binds to the Cas-kf63 protein designed by the present invention is as shown in SEQ ID No 3, and the repeat sequence of the gRNA that binds to the Cas-kf66 protein is as shown in SEQ ID No 6.
[0047] The present invention also provides an engineered non-naturally occurring vector system, which may include one or more vectors, and the one or more vectors include: a gRNA expression unit and the Cas protein expression unit, and the gRNA expression unit and the Cas protein expression unit are located on the same or different vectors of the system.
[0048] The expression unit is connected to a variety of regulatory elements including promoters (e.g., constitutive promoters or inducible promoters), enhancers (e.g., 35S promoter or 35S enhanced promoter), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation PolyA signals and poly PolyU sequences).
[0049] In some embodiments, the vectors in the system are viral vectors (e.g., retroviral vectors, lentiviral vectors, adenoviral vectors, and herpes simplex vectors), and can also be of types such as plasmids, viruses, cosmids, phages, etc., which are well known to those skilled in the art.
[0050] In one embodiment, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has the sequence shown by TTTV, where V is selected from A, G, or C.
[0051] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic cell or a eukaryotic cell.
[0052] In one embodiment, the target sequence is a non-naturally occurring DNA or RNA sequence.
[0053] In one embodiment, the target sequence is present in a cell. In one embodiment, the target sequence is present in the nucleus or cytoplasm (e.g., organelles) of the cell. In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.
[0054] In one embodiment, the Cas protein is linked to one or more NLS sequences. In one embodiment, the fusion protein contains one or more NLS sequences. In one embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.
[0055] Protein-nucleic acid complex / composition
[0056] On the other hand, the present invention provides a composition (or complex) comprising:
[0057] (i) A protein component selected from the Cas protein or the fusion protein described above; (ii) A nucleic acid component selected from: the gRNA, or a nucleic acid encoding the gRNA, or the precursor RNA of the gRNA, or a nucleic acid encoding the precursor RNA of the gRNA; the gRNA includes a repeat sequence that binds to the Cas protein and a spacer sequence for targeting the recognition of the target sequence, and the spacer sequence is complementary to the target sequence.
[0058] The Cas protein and gRNA of the present invention can form a binary complex, which is activated when binding to the target sequence substrate to form an activated CRISPR complex. The target sequence substrate is complementary to the spacer sequence in the gRNA (or, the guiding sequence hybridized to the target nucleic acid). In some embodiments, the spacer sequence of the gRNA is completely matched with the target sequence substrate. In other embodiments, the spacer sequence of the gRNA is partially (continuous or discontinuous) matched with the target sequence substrate, and the activated complex can exhibit collateral nuclease activity, which refers to the non-specific cleavage activity or random cleavage activity of the activated CRISPR complex on other nearby single-stranded nucleic acids, also known as trans cleavage activity in the art.
[0059] Host cell
[0060] The present invention also relates to a cell or cell line or their progeny in vitro, ex vivo or in vivo, which contains: the Cas protein, fusion protein, polynucleotide, protein-nucleic acid composition, vector or vector system of the present invention.
[0061] In certain embodiments, the cell is a prokaryotic cell.
[0062] In certain embodiments, the cell is a eukaryotic cell. In certain embodiments, the cell is a mammalian cell. In certain embodiments, the cell is a human cell. In certain embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, rodents (such as rats or mice). In certain embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry birds (such as chickens), fish or crustaceans (such as clams, shrimps). In certain embodiments, the cell is a plant cell, such as cells of monocotyledonous plants or dicotyledonous plants, or cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats or rice, such as algae, trees or production plants, fruits or vegetables (for example, tree species such as citrus trees, nut trees; solanaceous plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).
[0063] In certain embodiments, the cell is a stem cell or a stem cell line.
[0064] In certain embodiments, the cell is a microbial cell, such as Saccharomyces cerevisiae and the like.
[0065] In certain cases, the host cell of the present invention comprises a modification of a gene or genome that is not present in its wild type.
[0066] Application
[0067] The Cas protein, or the fusion protein, or the polynucleotide, or the vector, or the CRISPR-Cas system, or the vector system, or the composition, or the host cell provided by the present invention, in some embodiments, has the following applications:
[0068] (1) Application in gene editing or gene targeting for non-human disease diagnosis and treatment purposes;
[0069] (2) Application in the preparation of gene editing or gene targeting reagents or kits;
[0070] (3) Application in gene cleavage for non-disease diagnosis and treatment purposes;
[0071] (4) Application in the preparation of gene cleavage reagents or kits;
[0072] (5) Application in non-specific nucleic acid collateral cleavage;
[0073] (6) Application in non-specific nucleic acid collateral degradation.
[0074] Gene cleavage herein includes: DNA or RNA cleavage in the target sequence by the Cas protein described herein (Cis cleavage), and cleavage of non-target DNA or RNA (single-stranded nucleic acid substrate) in collateral cleavage (i.e., non-specific or non-targeted, Trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.
[0075] In one embodiment, the present invention further provides a method for editing a target sequence, targeting a target sequence or cleaving a target sequence, the method being for non-disease diagnosis and treatment purposes, the method being applied to an object that does not include humans, and the method comprising contacting the target sequence with the CRISPR-Cas system, or the vector system, or the composition, or the host cell.
[0076] In one embodiment, the present invention also provides a kit for gene editing or gene targeting, which kit comprises the CRISPR-Cas system as described above, or the vector system as described above, or the composition as described above, or the host cell as described above. The gene editing or editing of the target nucleic acid includes modifying a gene, knocking out a gene, altering the expression of a gene product, repairing a mutation, inserting a polynucleotide, and gene mutation.
[0077] In one embodiment, the present invention provides a kit for detecting a target sequence in a sample, which kit comprises:
[0078] (a) the Cas protein as described above, or a nucleic acid encoding the Cas protein; and / or
[0079] (b) a gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA comprising the gRNA, or a nucleic acid encoding the precursor RNA; the gRNA comprises a repeat sequence that binds to the Cas protein and a spacer sequence that targets and is complementary to the target sequence; and / or
[0080] (c) a single-stranded nucleic acid detector, which single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.
[0081] In one embodiment, the present invention provides a method for detecting a target sequence in a sample, which method is a method for non-disease diagnosis and treatment purposes, and which method comprises contacting the sample with the Cas protein, gRNA and single-stranded nucleic acid detector as described above, and detecting a detectable signal generated by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting the target gene; the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.
[0082] In the present invention, the single-stranded nucleic acid detector includes, but is not limited to, single-stranded DNA, single-stranded RNA, DNA-RNA hybrids, nucleic acid analogs, base modifiers and other single-stranded nucleic acid detectors; "nucleic acid analogs" include, but are not limited to: locked nucleic acid, bridged nucleic acid, morpholino nucleic acid, ethylene glycol nucleic acid, hexitol nucleic acid, threose nucleic acid, arabinonucleic acid, 2'-O-methyl RNA, 2'-methoxyacetyl RNA, 2'-fluoro RNA, 2'-amino RNA, 4'-thio RNA and combinations thereof, including optionally ribonucleotide or deoxyribonucleotide residues.
[0083] In the present invention, the detectable signal is achieved by the following means: visual-based detection, sensor-based detection, color detection, fluorescence signal-based detection, gold nanoparticle-based detection, fluorescence polarization, fluorescence detection, colloid phase transition / dispersion, electrochemical detection and semiconductor-based detection.
[0084] In the present invention, preferably, a fluorescent group and a quenching group are respectively provided at both ends of the single-stranded nucleic acid detector. When the single-stranded nucleic acid detector is cleaved, a detectable fluorescent signal can be exhibited. The fluorescent group is selected from one or any several of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red or LC RED460; the quenching group is selected from one or any several of BHQ1, BHQ2, BHQ3, Dabcy1 or Tamra.
[0085] Plant editing vector
[0086] In the specific field of plant gene editing and breeding of the present invention, a vector for plant gene editing and gene targeting is further constructed.
[0087] In one embodiment, the editing vector includes a gRNA expression unit, a Cas protein expression unit and a resistance screening gene expression unit. The gRNA includes a crRNA sequence capable of targeting a target sequence. The Cas protein expression unit expresses the Cas protein. In some embodiments, the resistance screening gene in the resistance screening gene expression unit is used in the genetic transformation process to facilitate the screening of successfully transformed cells or organisms. Examples thereof include, but are not limited to, marker genes such as hygromycin resistance gene (HYG), neomycin resistance gene, tetracycline resistance gene, penicillin resistance gene, streptomycin resistance gene, herbicide resistance gene, heavy metal resistance gene, antibiotic resistance gene, etc.
[0088] In a specific embodiment, the gRNA expression unit is ZmUBI1 pro:gRNA:Nos term, and / or the Cas protein expression unit is ZmUBI1 pro:Cas-kf63 / Cas-kf66:PsE9 term, and / or the resistance screening gene expression unit is 35S pro:HYG:35S term; in one embodiment, the three expression units are arranged on the same vector. Specifically, the editing vector uses pCambia3301 as the vector backbone, and the three expression units are arranged starting from the right border of T-DNA. Based on the expression characteristics of different vectors, the three expression units can be arranged on different vectors or the same vector. In certain embodiments, optional vectors include pCAMBIA series vectors (such as pCAMBIA1300, pCAMBIA3301, or pCAMBIA3300, etc.), pBI series vectors (such as pBI121), pGreen series vectors (such as pGreenII), and other vectors such as pCHF3, pCambia1380 and other commonly used plant gene editing vectors.
[0089] In a specific embodiment, a BsaI cleavage site is left between ZmUBI pro and Nos term, into which one or more gRNA sequences can be inserted. The gRNAs are concatenated by HH and HDV ribozyme technology to form a gene editing vector targeting any target site of the target sequence. Specifically, gene editing sites are selected for different regions of the target sequence to be edited, and recognition sequences that can pair with the editing sites are designed to form a gRNA sequence with the crRNA repeat sequence. One or more gRNA sequences are interconnected by HH and HDV ribozyme technology. After linearizing the basic editing vector, one or more gRNA sequences are inserted using DNA ligase or Gibson cloning method to form a target sequence editing vector.
[0090] As a specific example, the gRNA sequences targeting the target sequence constructed in the present invention are shown in SEQ ID NO 7, 8, 9, and 10.
[0091] As a more specific example, for the Cas-kf63 target sequence editing vector constructed in the present invention, its gRNA expression sequence is shown in SEQ ID NO 13.
[0092] In one embodiment, the present invention provides an engineered bacterium containing any of the above-mentioned editing vectors. The editing vector is transferred into plant cells through Agrobacterium-mediated transformation to edit the target gene. Optional engineered bacteria include Agrobacterium, Bacillus subtilis, Escherichia coli, etc.
[0093] Furthermore, the present invention provides the application of any of the above-mentioned editing vectors or the above-mentioned engineered bacterium in plant gene editing and breeding. In one embodiment, the present invention provides a method for genetic transformation of plants, in which the obtained target gene editing vector is introduced into recipient plant cells, and then heritable, non-transgenic and stably inherited offspring plants are screened.
[0094] Furthermore, the genetic transformation method includes the following steps:
[0095] (1) Genetic transformation of recipient plant cells to obtain regenerated plants (including: 1) induction of callus; 2) subculture of callus; 3) infection of callus; 4) screening and identification of resistant callus; 5) differentiation of callus into seedlings; 6) rooting culture; 7) acclimatization and transplantation to obtain regenerated plants);
[0096] (2) Detect whether the gene editing of the regenerated plants is successful;
[0097] (3) Screen to obtain homozygous edited offspring plants with the target gene trait, non-transgenic and stable inheritance.
[0098] Term definition
[0099] In the present invention, unless otherwise specified, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Moreover, the operating steps such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA used herein are all conventional steps widely used in the corresponding fields. Meanwhile, to better understand the present invention, the definitions and explanations of relevant terms are provided below.
[0100] The term "gRNA (guide RNA)" refers to any polynucleotide sequence that has sufficient complementarity with a target sequence to hybridize with the target sequence and guide the specific binding of the CRISPR / Cas system to the target sequence. The structure of gRNA generally includes a crRNA (CRISPR RNA) sequence, which is derived from the CRISPR array and contains a repeat sequence and a spacer sequence. The spacer sequence is a part derived from foreign genetic material and is complementary to the target sequence for base pairing. Usually, a non-complementary single-stranded region called "5' overhang" can be designed at the 5' end of gRNA, which helps to distinguish gRNA from the target sequence. In some cases, the 3' end of gRNA may contain a reverse complementary tail, which helps to improve the binding stability of gRNA to the Cas protein. In some applications, gRNA may also contain additional modifications such as phosphorylation, chemical modification, etc. to improve its stability or functionality.
[0101] The terms "clustered regularly interspaced short palindromic repeats (CRISPR), CRISPR-associated (Cas) (CRISPR-Cas) system" or "CRISPR system" are used interchangeably and have the meanings commonly understood by those skilled in the art, and generally include transcripts or other elements related to the expression of CRISPR-associated ("Cas") genes, or transcripts or other elements capable of guiding the activity of the Cas genes.
[0102] The terms "target sequence", "target nucleic acid" and "target nucleic acid" have the same meaning. In the present invention, it refers to the polynucleotide targeted by the spacer sequence in crRNA. The target sequence can include any polynucleotide, such as DNA or RNA. The hybridization between the target sequence and crRNA will promote the expression of the CRISPR / Cas system. Complete complementarity is not required, as long as there is sufficient complementarity to cause hybridization and promote the expression of a CRISPR / Cas system.
[0103] As used herein, the term "non-naturally occurring" or "engineered" are used interchangeably and denote human intervention. When these terms are used to describe nucleic acid molecules or polypeptides, they denote nucleic acid molecules or polypeptides that do not exist in nature or are not naturally synthesized by an organism.
[0104] As used herein, the term "wild-type" has the meaning commonly understood by those skilled in the art, which denotes the typical form of an organism, strain, gene, or characteristics that distinguish it from mutant or variant forms when it exists in nature, which can be isolated from natural sources and has not been deliberately modified by humans.
[0105] The term "engineered host cell" refers to a cell that can be used to introduce a vector, including but not limited to prokaryotic cells such as Agrobacterium, Escherichia coli, or Bacillus subtilis, and eukaryotic cells such as microbial cells, fungal cells, animal cells, or plant cells.
[0106] As used herein, the term "identity" is used to refer to the sequence matching between two polypeptides or two nucleic acids. When a position in two sequences being compared is occupied by the same base or amino acid monomer subunit (e.g., a position in each of two DNA molecules is occupied by adenine, or a position in each of two polypeptides is occupied by lysine), then the molecules are identical at that position. The "percent identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions being compared × 100. For example, if 6 out of 10 positions of two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT have 50% identity (3 out of a total of 6 positions match). Generally, the comparison is made when the two sequences are aligned to yield maximum identity. Such alignment can be calculated using, for example, a computer program such as the Needleman and Wunsch algorithm in the GAP program (J Mol Biol 48:444-453 (1970)), using the Blossum 62 matrix or PAM250 matrix and a gap weight of 16, 14, 12, 10, 8, 6, or 4 and a length weight of 1, 2, 3, 4, 5, or 6 to determine the percent identity between two amino acid sequences.
[0107] The term "plant" should be understood as any differentiated multicellular organism capable of photosynthesis, including crop plants at any mature or developmental stage, especially monocotyledonous or dicotyledonous plants, vegetable crops, including artichokes, kohlrabi, arugula, leeks, asparagus, lettuce, cabbage, cauliflower, broccoli, kale, taro, melons (e.g., cantaloupe, watermelon); fruit crops such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, blackberries, grapes, avocados, bananas, kiwis, persimmons, pomegranates, pineapples, mangoes, papayas, and lychees, etc.; field crops such as alfalfa, corn / maize (forage corn, sweet corn, popcorn), peanuts, small grain cereals (barley, oats, rye, rice, wheat, etc.), sorghum, tobacco, kapok, leguminous plants (beans, lentils, peas, soybeans), oil plants (rapeseed, mustard, poppy, olive, sunflower, coconut, castor oil plant, cocoa bean, groundnut), fiber plants (cotton, flax, jute), etc.
[0108] Technical effects achieved by the present invention:
[0109] Based on the discovered and studied Cas12a effector nuclease sequences, the present invention designed, screened, and verified the functions, and obtained novel effector nucleases Cas-kf63 and Cas-kf66, which showed nuclease activity in experiments and have broad application prospects.
[0110] The present invention obtained two new Cas12a-like effector nuclease proteins, Cas-kf63 and Cas-kf66. After codon optimization and experimental verification of plant gene editing transformation, they have the ability to recognize and cleave target sites with a TTTV sequence at the 5' end, and novel effector nucleases for gene editing are obtained. On this basis, an efficient basic vector for plant gene editing is also provided, which can effectively perform targeted genomic modification on plants. Description of the Drawings
[0111] Figure 1Schematic diagrams of the Cas-kf63 and Cas-kf66 effector nucleases and their CRISPR array structures (in the figure: (a) Cas-kf63 and Cas-kf66 are type II-V CRISPR-associated Cas12a effector nucleases derived from the bacteria Oscillospiraceae bacterium and Ruminococcus bacterium, respectively. The genome of Cas-kf63 does not contain the Cas4, Cas1, and Cas2 enzyme genes and only contains 5 repeated CRISPR arrays; the genome of Cas-kf66 contains the Cas4 and Cas2 enzyme genes and 3 repeated CRISPR arrays and does not contain the Cas1 enzyme gene; (b) The crRNA repeat sequences of Cas-kf63 and Cas-kf66 both contain the conserved sequences UCUACU and GUAGAU that can form hairpin structures at their 3' ends).
[0112] Figure 2 Schematic diagram of the rice starch synthase gene structure and the selected editing sites (in the figure: The starch synthase gene Waxy (LOC_Os06g04200) consists of 14 exons and 13 introns. One site each in intron 5, exons 6, 7, and 8 was selected to design OsWx-T3, OsWx-T1, OsWx-T4, and OsWx-T2 gRNAs, and the PCR specific primers OsWx-F5 and OsWx-R1 were designed upstream and downstream respectively for amplifying and sequencing the DNA fragments in the editing region).
[0113] Figure 3Schematic diagram of the gene editing transformation vector KF374 (in the figure: T-DNA RB is the right border of T-DNA; ZmUBI1pro and ZmUBI1 intron are the promoters of the maize ubiquitin gene including the intron; HH and HDV are the hammerhead ribozyme and hepatitis delta virus ribozyme respectively; crRNA is the mature CRISPR RNA repeat sequence of FnCpf1; OsWx-T1, OsWx-T2, OsWx-T3 and OsWx-T4 are gRNAs; Nosterm is the terminator of the Agrobacterium tumefaciens nopaline synthase gene; ZmUBI1 pro and ZmUBI1 intron are the promoters of the maize ubiquitin gene including the intron; Cas-kf63 is a new Cas12a effector nuclease belonging to the class 2 type V CRISPR system; PsE9 term is the terminator of the pea ribulose-1,5-bisphosphate carboxylase / oxygenase small subunit E9 protein gene; CaMV 35S pro is the cauliflower mosaic virus 35S promoter; HygR is the hygromycin phosphotransferase; CaMV 35S term is the cauliflower mosaic virus 35S terminator; T-DNA LB is the left border of T-DNA).
[0114] Figure 4 DNA sequences near the OsWx-T3 and OsWx-T1 editing sites in the target region for some gene editing events of Cas-kf63.
[0115] Figure 5 DNA sequences near the OsWx-T4 and OsWx-T2 editing sites in the target region for some gene editing events of Cas-kf63. The 120bp-long partial intron7 and exon8 sequences are not shown in the figure.
[0116] Figure 6 DNA sequences near the OsWx-T3 and OsWx-T1 editing sites in the target region for some gene editing events of Cas-kf66.
[0117] Figure 7 DNA sequences near the OsWx-T4 and OsWx-T2 editing sites in the target region for some gene editing events of Cas-kf66. The 120bp-long partial intron7 and exon8 sequences are not shown in the figure. Detailed implementation methods
[0118] The technical solutions of the present invention will be described in detail below in conjunction with the accompanying drawings and embodiments, but the present invention is not limited to the scope of the described embodiments. Unless otherwise specified, the experiments and methods described in the embodiments are basically carried out according to the conventional methods well-known in the art and described in various reference documents. The reagents and raw materials used in the present invention are all commercially available. For those conditions not specified in the embodiments, they are carried out according to the conventional conditions or the conditions recommended by the manufacturer. For those reagents or instruments without indicating the manufacturer, they are all conventional products that can be obtained through commercial purchase. Those skilled in the art know that the embodiments describe the present invention by way of example and should not limit the scope claimed by the present invention. All the published cases and other reference materials mentioned herein are incorporated herein by reference in their entirety.
[0119] In the present invention, DNA sequence analysis and alignment, DNA vector cloning construction design, PCR primer and CRISPR gRNA design involved in molecular biology experiments are all completed using the DNA analysis software Geneious Prime (GraphPad Software LLC).
[0120] The concept of the present invention includes:
[0121] 1. Obtaining a novel effector nuclease gene: According to the amino acid sequence of the discovered and studied Cas12a effector nuclease, use different query programs such as Blastp and Tblastn to search multiple times in microbial genome and metagenome databases such as NCBI ( www.ncbi.nlm.nih.gov ), JGI (img.jgi.doe.gov) and UniProt (www.uniprot.org), screen newly sequenced genes with a certain degree of homology and containing the RuvC nuclease domain at the carboxyl terminus of the protein, retrieve the corresponding genomic nucleic acid sequences, use DNA analysis software such as Geneious to find CRISPR repeat sequence arrays and protein-coding regions, and identify Cas12a effector nucleases according to their conserved Cas12a-Cas4-Cas1-Cas2-CRISPR gene arrangement. Preferably, novel Cas12a effector nucleases with a relatively low similarity to those already reported are designed for functional verification and screening experiments.
[0122] 2. Establish a plant gene editing system using novel effector nucleases: Select novel Cas12a effector nucleases with low similarity to those reported publicly, and use applicable plant expression promoters, terminators, etc. to construct gene expression units, and build a basic vector for plant transformation and expression. Select the rice starch synthase gene (Waxy or OsWx) LOC_Os06g04200|Chr6:1765621..1770656 forward as the target gene, design a gRNA specific to the Waxy gene, clone it into the above plant transformation vector to co-express the novel effector nuclease and its gRNA, and introduce it into Agrobacterium tumefaciens EHA105 for rice transformation.
[0123] 3. Plant gene editing mediated by novel effector nucleases: Transform rice with the above Agrobacterium tumefaciens, select resistant calli, extract genomic DNA, PCR amplify the specific gene fragment of the transformation vector, and confirm that the resistant calli are positive transgenic events. Amplify the edited region fragment of the target gene with the same genomic DNA, sequence and analyze the gene editing effect, and statistically calculate the editing efficiency at different sites.
[0124] Example 1
[0125] I. Obtaining of novel effector nuclease genes and their crRNA repeat sequences
[0126] Query of microbial genomes: Collect the published Cas12a effector nuclease sequences and construct a query library. Exclude the sequences that are exactly the same as the known Cas12 effector nucleases, sort them according to the similarity degree, select those with low similarity for detailed sequence analysis and screening, retrieve the genomic nucleic acid sequences corresponding to the selected Cas12a effector nucleases, use DNA analysis software such as Geneious to find the CRISPR repeat sequence array, and identify whether it is a Cas12a effector nuclease according to its conserved Cas12a-Cas4-Cas1-Cas2-CRISPR gene arrangement. Preferably, select Cas12a effector nucleases with a complete and conserved CRISPR structure, moderate length, and newly sequenced, and their corresponding crRNA repeat sequences, and design experiments for functional verification and further screening.
[0127] Through the above multi-step screening and subsequent multiple experimental verifications, the inventors obtained multiple novel type II-V CRISPR Cas12a-like effector nucleases. Two of them with verified cleavage activity were named Cas-kf63 and Cas-kf66 respectively. The genome of Cas-kf63 only has the Cas12a gene and the CRISPR repeat sequence array, and the genome of Cas-kf66 has the typical Cas12a-Cas4-Cas1-Cas2-CRISPR gene structure arrangement and the CRISPR repeat sequence array, but the Cas1 gene is missing ( Figure 1a). Cas-kf63 is derived from shotgun sequencing of the Oscillospiraceae genome. Its protein, nucleotide, and CRISPR repeat sequences are shown in SEQ ID NOs 1-3. Cas-kf66 is derived from shotgun sequencing of the Ruminococcus genome. Its protein, nucleotide, and CRISPR repeat sequences are shown in SEQ ID NOs 4-6. The two ends of the protein contain artificially added SV40 and Nucleoplasmin nuclear localization signal (NLS) sequences respectively. The coding nucleotides have been optimized to codon sequences suitable for eukaryotic cell expression. The 3' end of the crRNA repeat sequence has conserved sequences UCUACU and GUAGAU that can form a hairpin structure ( Figure 1 b).
[0128] Example 2 Establishment of a plant gene editing vector using a novel effector nuclease
[0129] Beidahuang Kenfeng Seed Industry independently developed a DNA vector for Agrobacterium-mediated transformation for rice gene editing. Based on the pCambia3301 vector, it contains three gene expression units: ZmU6 pro:gRNA:AtU6-26 term, ZmUBI1pro:SpCas9:PsE9term, and 35S pro:HYG:35S term. There is a BsaI restriction site between ZmU6 pro and gRNA. After linearizing the vector, one or more gRNA sequences can be inserted by methods such as DNA ligase or Gibson cloning to form a basic gene editing vector targeting any target site. The specific sequence information can be found in the KF80 transformation vector in the patent CN113604501B "Gene Editing Method for Controlling the Fragrance Gene of Improved Indica Rice Lines and Its Application". Each gene editing transformation vector described in the present invention is constructed by modifying the disclosed basic gene editing vector.
[0130] Using the above KF80 conversion vector as the basic backbone, replace ZmU6 pro with ZmUBI1 pro, replace AtU6-26 term with Nosterm, replace SpCas9 with the LbCpf1 gene, and replace gRNA with the synthetic HH-OsWx-T1-HDV-HH-OsWx-T2-HDV-HH-OsWx-T3-HDV-HH-OsWx-T4-HDV (SEQ ID NO 13), where OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 are gRNAs (SEQ ID NO 7, 8, 9, 10) for editing corresponding sites of the rice starch synthase gene (OsWx), and HH and HDV are the hammerhead and hepatitis delta virus ribozymes (SEQ ID NO 11, 12), respectively. Finally, the gene editing vector KF365 based on the LbCpf1 effector nuclease was formed, which contains three gene expression units ZmUBI1 pro:gRNA (SEQ ID NO 13):Nos term, ZmUBI1 pro:LbCpf1:PsE9 term, and 35S pro:HYG:35S term. As a positive control gene editing vector, it was used to edit the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites of rice. Subsequently, the LbCpf1 gene in KF365 was replaced with the Cas-kf63 and Cas-kf66 (SEQ ID NO 2, 5) genes respectively to form gene editing vectors KF374 and KF385 based on the Cas-kf63 and Cas-kf66 effector nucleases (Table 2), and also tried to edit the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites of rice. The composition structure and sequence of the KF374 vector are shown in Figure 3 and SEQ ID NO 14. Except for the different effector nucleases used, the composition structures and other sequences of other vectors are exactly the same as those of KF374, so they will not be listed again. The above-mentioned various gene editing vectors were respectively transferred into Agrobacterium tumefaciens EHA105 and stored at -80 °C for future use.
[0131] III. Plant Gene Editing Mediated by Novel Effector Nucleases
[0132] Gene editing transformation of rice: Gene editing vectors based on novel effector nucleases must be transformed into plant cells to verify their functions. Using the Agrobacterium-mediated transformation method, healthy mature seeds of the rice variety Kendao 32 were used as explants to induce callus. The gene editing vectors KF374 and KF385 were respectively transformed into cells, resistant transgenic events were screened, and the DNA sequences of the target genes OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 loci were detected for specific changes to evaluate the gene editing effects of each effector nuclease in plants. The formulations of various basic media used in the genetic transformation experiment are listed in Table 1, and the specific operation procedures are as follows:
[0133] Induction of rice callus: For each transformation experiment, about 200 mature, plump, healthy and clean rice seeds were selected, the glumes were removed, placed in a 100-ml sterile glass bottle, washed 3 times with 75% ethanol, equal volume of sterile water and 10% sodium hypochlorite (NaClO), two drops of Tween 20 were added, and shaken for sterilization for 20 minutes. Then washed 3-5 times with sterile water until the foam was washed away. The surface moisture of the seeds was blotted dry with sterile filter paper and placed flat on the induction medium, and incubated in the dark at 33°C in a constant temperature incubator for 3-4 weeks until callus grew from the embryo.
[0134] Subculture of rice callus: Select and peel healthy callus with strong growth, place it on a new induction medium for subculture for 1-2 weeks to expand and maintain the vitality of the callus.
[0135] Preparation of Agrobacterium: Take out the Agrobacterium glycerol preservation tube carrying the gene editing vector from the -80°C low-temperature refrigerator. Take a small amount of Agrobacterium and streak it on a solid YEP medium plate containing 50 mg / L kanamycin, and incubate it in the dark at 28°C for 2-3 days. Pick a single colony and inoculate it into 5 ml of YEP liquid medium and incubate it in the dark at 28°C overnight. Take a small amount of the bacterial liquid to extract plasmid DNA for vector-specific PCR detection to confirm that the strain carries the correct gene editing vector. At the same time, take a part of the bacterial liquid and spread it on a YEP medium plate for cultivation and observation. The Agrobacterium used for infection must be free of contamination. The whole plate is required to be smooth without granular or other colored fungi or molds and other contaminants to ensure the purity of the obtained Agrobacterium, and strict aseptic operation is required to ensure no contamination in the subsequent work. After determining that the Agrobacterium plate is free of contamination, directly scrape the bacteria into a 100-ml triangular flask containing 25 ml of suspension infection medium, and incubate it on a rotary shaker at a constant temperature of 25°C at 100-120 rpm for 2-3 hours, and take samples to measure the OD 600 value, and adjust the OD 600 value of Agrobacterium to 0.1-0.2 with suspension infection medium for callus infection.
[0136] Infection of callus: Pick small granular and healthy callus and place it in a 250 - milliliter sterile glass bottle. After gently shaking and infecting with Agrobacterium for 3 - 5 minutes, blot the bacterial liquid on the surface of the callus with a sterile filter paper, and place it on a sterile filter paper previously placed on the surface of the co - culture medium (the filter paper prevents the growth of Agrobacterium), and co - culture in the dark at 28°C for 3 days.
[0137] Screening and identification of resistant callus: Transfer the infected rice callus to a sterilized 250 - milliliter glass bottle, wash it several times with sterile water until the water is clear and transparent, and finally soak and wash it for 0.5 hour at 100 revolutions per minute on a shaker with carbenicillin at a final concentration of about 250 mg / L. Transfer the callus to a blank petri dish, blot the moisture on the surface of the callus with a sterile filter paper, air - dry it in a sterile operation laminar flow hood for 2.5 hours, then divide and place the callus on the screening medium, and culture it in the dark in a 33°C constant - temperature incubator for 3 - 4 weeks. The screening medium contains 400 mg / L carbenicillin to inhibit the growth of Agrobacterium and 50 mg / L hygromycin to screen for transformed cells. Non - transformed cells stop growing and gradually die on the screening medium. After 3 - 4 weeks of culture, the successfully transformed cells will grow resistant callus. Appropriately sample the resistant callus when it continues to grow, extract genomic DNA, and perform PCR identification with primers specific to the transformation vector. Each PCR - positive callus represents an independent transformation event.
[0138] Table 1 Composition and preparation method of the basic medium used in Agrobacterium - mediated genetic transformation experiment
[0139]
[0140]
[0141] 2. Analysis of gene editing effect:
[0142] Resistant transformation events need to be detected by molecular biology methods to confirm whether genetic transformation is successful. A commonly used simple and reliable detection method is PCR amplification. Sample the rice resistant callus and extract genomic DNA using a rapid DNA extraction method based on SDS (Sodium dodecylsulphate). First, perform PCR detection of the vector-specific HygR gene to identify the transformation event. The PCR reaction system uses the Quick Taq HS DyeMix (DTM-101) kit (TOYOBO LifeScience). A 20-μl PCR reaction system contains 10 μl of 2x Quick Taq HS DyeMix, 1.0 μl of 10 pmol / μl primer Hyg-F1 (SEQ ID NO 15) and 1.0 μl of 10 pmol / μl primer Hyg-R1 (SEQ ID NO 16), 6.0 μl of sterile water, and finally 2.0 μl of 50 ng / μl genomic DNA of the sample. The PCR reaction conditions are denaturation at 95°C for 5 minutes, then denaturation at 95°C for 30 seconds, annealing at 60°C for 1 minute, extension at 72°C for 1 minute for a total of 30 cycles, and finally extension at 72°C for 7 minutes and hold at 4°C. The PCR amplification products are separated by 1% agarose gel electrophoresis. Samples that successfully amplify the 402-bp long fragment (SEQ ID NO17) specific to the HygR gene are determined to be transgenic positive.
[0143] By a similar PCR amplification method and using the OsWx (LOC_Os06g04200)-specific primers OsWx-F5 and OsWx-R1 (SEQ ID NO 18, 19), a 916-bp long target fragment (SEQ ID NO 20) near the gene editing sites of OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 can be amplified from HygR-positive samples. After purification, it is directly sequenced with the same upstream or downstream primers and compared with the sequence of the same gene fragment of Kendao 32 wild type. If the DNA fragment sequence of the transgenic event has no difference from the wild-type sequence, then gene editing has not occurred; if overlapping peaks appear near the expected cleavage site of the effector nuclease, it indicates that the sequenced PCR fragment is a mixed template, containing DNA fragments with and without editing, which is heterozygous gene editing; if there are definite sequence differences from the wild type, it is homozygous gene editing. The proportion of fragments with editing detected at any site in the PCR fragment of the target gene is the overall gene editing efficiency. Gene editing experiments confirmed that the novel effector nucleases Cas-kf63 and Cas-kf66 have gene editing functions, and the overall gene editing efficiencies are 35.4% and 89.3% respectively (Table 2).
[0144] Table 2 Gene editing efficiencies of novel effector nucleases
[0145]
[0146] *The number of amino acids in the length of the enzyme protein includes sequences such as the nuclear localization short peptide at both ends.
[0147] **Editing efficiency = Number of editing events occurring at any site / Total number of sequencing events of the PCR fragment of the target gene.
[0148] TOPO clone the PCR fragment of the target gene of the Cas-kf63 editing event into the pCR2.1 vector (Thermo Fisher Scientific). Send 4 clones for sequencing for each sample, and the specific editing changes in the DNA sequences at the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 sites can be accurately analyzed. Different clones derived from the same edited callus event often have various editing sequence changes. Some clones only have editing at the OsWx-T3 or OsWx-T2 site, some clones achieve editing at both sites simultaneously, and four clones have editing at all three sites of OsWx-T3, OsWx-T1, and OsWx-T2, while no editing occurs at the OsWx-T4 site. There are obvious differences in the editing efficiency among the four sites: OsWx-T3 > OsWx-T2 > OsWx-T1 > OsWx-T4 (Table 3). Clone sequencing shows that the novel effector nuclease Cas-kf63 has selectivity for editing sites with different PAM sequences, with higher editing efficiency for TTTA (OsWx-T3) and TTTC (OsWx-T2), very low editing efficiency for TTTG (OsWx-T1), and no ability to edit (OsWx-T4) without a typical TTTV (V is A, C, or G) PAM. The specific editing changes in the DNA sequences near the OsWx-T3 and OsWx-T1 sites, and the OsWx-T4 and OsWx-T2 sites of the 20 clones marked with asterisks in Table 3 are shown in Figure 4 and Figure 5 .
[0149] Table 3 Editing efficiency of gene sites of the effector nuclease Cas-kf63
[0150]
[0151]
[0152] Note: *The alignment sequences of the 20 selected clones are shown in Figure 4 and Figure 5 ; NE (not edited) means not edited.
[0153] The PCR fragments of the target genes of the Cas-kf66 editing events were TOPO cloned into the pCR2.1 vector (Thermo Fisher Scientific). Four clones of each sample were sent for sequencing, and the specific editing changes in the DNA sequences at the OsWx-T1, OsWx-T2, OsWx-T3, and OsWx-T4 loci could be accurately analyzed. Different clones derived from the same edited callus event often had multiple editing sequence changes. Some clones only had editing at the OsWx-T3 or OsWx-T2 locus, some clones achieved editing at both loci simultaneously, and no clones were detected with editing at all three loci of OsWx-T3, OsWx-T1, and OsWx-T2. No editing occurred at the OsWx-T4 locus, and there were significant differences in the editing efficiencies among the four loci: OsWx-T3 > OsWx-T2 > OsWx-T1 > OsWx-T4 (Table 4). Clone sequencing showed that the novel effector nuclease Cas-kf66 had selectivity for editing at sites with different PAM sequences, with higher editing efficiencies for TTTA (OsWx-T3) and TTTC (OsWx-T2), a very low editing efficiency for TTTG (OsWx-T1), and no ability to edit at (OsWx-T4) without a typical TTTV (V is A, C, or G) PAM. The specific editing changes in the DNA sequences near the OsWx-T3 and OsWx-T1 loci, and the OsWx-T4 and OsWx-T2 loci of the 20 clones marked with asterisks in Table 4 are shown in Figure 6 and Figure 7 .
[0154] Table 4 Editing efficiencies of gene loci of the effector nuclease Cas-kf66
[0155]
[0156]
[0157]
[0158] Note: *The alignment sequences of the 20 selected clones are shown in Figure 6 and Figure 7 ; NE (not edited) not edited.
[0159] The above results of the target gene sequence analysis revealed that gene editing mediated by the Cas-kf63 and Cas-kf66 effector nucleases was mostly small fragment deletions, which mostly occurred at the expected target cleavage sites in the distal part of the PAM of the target gene, consistent with other known Cas12a-like effector nucleases such as LbCpf1 and FnCpf1.
[0160] The present invention has identified multiple novel Cas12a-like effector nuclease proteins. After codon optimization and verification through plant gene editing transformation experiments, it is found that Cas-kf63 and Cas-kf66 among them have the ability to recognize target sites with TTTV sequences at the 5' end and cleave them, resulting in two novel effector nucleases suitable for plant gene editing. On this basis, an efficient basic vector for plant gene editing is also provided, which can perform effective genomic directed modification on plants.
[0161] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. According to all the teachings that have been published, various modifications and changes can be made to the details. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
[0162] DNA sequence description in the present invention
[0163] SEQ ID NO 1: Cas-kf63 protein 1227aa
[0164] SEQ ID NO 2: Cas-kf63 coding DNA 3684bp
[0165] SEQ ID NO 3: Cas-kf63 crRNA repeat DNA 36bp
[0166] SEQ ID NO 4: Cas-kf66 protein 1258aa
[0167] SEQ ID NO 5: Cas-kf66 coding DNA 3777bp
[0168] SEQ ID NO 6: Cas-kf66 crRNA repeat DNA 36bp
[0169] SEQ ID NO 7: synthesized OsWx-T1 fragment 27bp
[0170] SEQ ID NO 8: synthesized OsWx-T2 fragment 27bp
[0171] SEQ ID NO 9: synthesized OsWx-T3 fragment 27 bp
[0172] SEQ ID NO 10: synthesized OsWx-T4 fragment 27 bp
[0173] SEQ ID NO 11: synthesized hammerhead ribozyme 37 bp
[0174] SEQ ID NO 12: synthesized hepatitis delta virus ribozyme 68 bp
[0175] SEQ ID NO 13: synthesized HH-OsWx-T1-HDV-HH-OsWx-T2-HDV-HH-OsWx-T3-HDV-HH-OsWx-T4-HDV DNA fragment 628 bp
[0176] SEQ ID NO 14: KF374 T-DNA 11678 bp
[0177] SEQ ID NO 15: Hyg-F1 oligo 22 bp
[0178] SEQ ID NO 16: Hyg-R1 oligo 22 bp
[0179] SEQ ID NO 17: HygR gene PCR fragment 402 bp
[0180] SEQ ID NO 18: OsWx-F5 oligo 22 bp
[0181] SEQ ID NO 19: OsWx-R1 oligo 22 bp
[0182] SEQ ID NO 20: Oryza sativa Kendao32 Waxy gene OsWx(LOC_Os06g04200)PCRfragment 916bp
Claims
1. A Cas protein, characterized in that The Cas protein is Cas-kf63 or Cas-kf66, the Cas-kf63 amino acid sequence is shown in SEQ ID No 1, and the Cas-kf66 amino acid sequence is shown in SEQ ID No 4.
2. A fusion protein comprising the Cas protein according to claim 1 and other modified parts.
3. An isolated polynucleotide, characterized in that The polynucleotide is a polynucleotide sequence encoding the Cas protein according to claim 1, or a polynucleotide sequence encoding the fusion protein according to claim 2; preferably, the nucleotide sequence encoding Cas-kf63 is shown in SEQ ID No 2, and the nucleotide sequence encoding Cas-kf66 is shown in SEQ ID No 5.
4. A carrier, characterized in that The vector comprises the polynucleotide according to claim 3 and a regulatory element operably linked thereto.
5. A CRISPR-Cas system, characterized in that: The system comprises the Cas protein according to claim 1 and at least one gRNA, wherein the gRNA comprises a repetitive sequence binding to the Cas protein and a spacer sequence for targeting a target sequence; preferably, the repetitive sequence binding to the Cas-kf63 gRNA is as shown in SEQ ID No 3, and the repetitive sequence binding to the Cas-kf66 gRNA is as shown in SEQ ID No 6.
6. A vector system, characterized in that The vector system includes one or more vectors, which include: a gRNA expression unit and the Cas protein expression unit according to claim 1, the gRNA includes a crRNA sequence capable of targeting a target sequence, and the gRNA expression unit and the Cas protein expression unit are located on the same or different vectors of the system.
7. A composition, characterized in that The composition comprises: (i) a protein component selected from: the Cas protein of claim 1 or the fusion protein of claim 2; (ii) a nucleic acid component selected from: gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a nucleic acid encoding the precursor RNA of the gRNA; the gRNA comprises a repeat sequence that binds to the Cas protein and a spacer sequence for targeting a target sequence.
8. An engineered host cell, characterized in that The host cell comprises the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3, or the vector of claim 4, or the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7.
9. Use of the Cas protein according to claim 1, or the fusion protein according to claim 2, or the polynucleotide according to claim 3, or the vector according to claim 4, or the CRISPR-Cas system according to claim 5, or the vector system according to claim 6, or the composition according to claim 7, or the host cell according to claim 8 in any of the following: (1) Applications in gene editing or gene targeting for the diagnosis and treatment of non-human diseases; (2) Application in the preparation of gene editing or gene targeting reagents or kits; (3) Application in gene cutting for purposes other than disease diagnosis and treatment; (4) Application in the preparation of gene cleavage reagents or kits; (5) Application in non-specific cleavage of side branch nucleic acids; (6) Application in non-specific degradation of collateral nucleic acids.
10. A kit for gene editing or gene targeting, characterized in that: The kit comprises the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7, or the host cell of claim 8.
11. A kit for detecting a target sequence in a sample, characterized in that: The kit comprises: (a) the Cas protein of claim 1, or a nucleic acid encoding the Cas protein; and / or (b) gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA comprising the gRNA, or a nucleic acid encoding the precursor RNA; the gRNA comprises a crRNA sequence capable of targeting a target sequence; and / or (c) a single-stranded nucleic acid detector, wherein the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.
12. A method for editing a target sequence or targeting a target sequence, characterized in that: The method is for non-disease diagnosis and treatment purposes, the subject to which the method is applied does not include humans, and the method comprises contacting the target sequence with the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7, or the host cell of claim 8.
13. A method for cutting a target sequence, characterized in that: The method is for non-disease diagnosis and treatment purposes, the subject to which the method is applied does not include humans, and the method comprises contacting the target sequence with the CRISPR-Cas system of claim 5, or the vector system of claim 6, or the composition of claim 7, or the host cell of claim 8.
14. A method for detecting a target sequence in a sample, characterized in that: The method is for non-disease diagnosis and treatment purposes, and comprises contacting a sample with the Cas protein, gRNA and single-stranded nucleic acid detector described in claim 11, detecting a detectable signal generated by the Cas protein cutting the single-stranded nucleic acid detector, thereby detecting a target sequence; the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.
15. An editing vector, characterized in that: The vector comprises a gRNA expression unit, a Cas protein expression unit and a resistance screening gene expression unit, the gRNA comprises a crRNA sequence capable of targeting a target sequence, and the Cas protein expression unit expresses the Cas protein according to claim 1.
16. The editing vector according to claim 15, characterized in that The gRNA expression unit is ZmUBI1pro:gRNA:Nos term, and / or the Cas protein expression unit is ZmUBI1 pro:Cas-kf63 / Cas-kf66:PsE9term, and / or the resistance screening gene expression unit is 35S pro:HYG:35S term; preferably, the editing vector uses pCambia3301 as the vector backbone, and three expression units are set starting from the right border of T-DNA.
17. The editing vector according to claim 16, characterized in that: A BsaI restriction site is left between ZmUBI pro and Nos term, into which one or more gRNA sequences can be inserted. Each gRNA is connected in series using HH and HDV ribozyme technology to form a gene editing vector targeting any target site of the target sequence.
18. An engineered bacterium, characterized in that: Containing the editing vector described in any one of claims 15-17.
19. Use of the editing vector described in any one of claims 15 to 17, or the engineered bacteria described in claim 18 in plant gene editing breeding.
20. A method for genetic transformation of a plant, characterized in that: The editing vector described in any one of claims 15 to 17 is transferred into a recipient plant cell, and then screened to obtain heritable, non-transgenic and stably inherited offspring plants; preferably, the process comprises the following steps: (1) Genetically transforming recipient plant cells to obtain regenerated plants; (2) Detecting whether gene editing of the regenerated plants is successful; (3) Screening to obtain offspring plants that are homozygous for the target sequence traits, non-transgenic, and stably inherited.
Citation Information
Patent Citations
Gene editing methods for aroma-controlling genes in improved indica rice lines and their applications
CN113604501B