Effector nuclease and use thereof

By developing novel effector nucleases Cas-kf8 and Cas-kf9, combined with gRNA and an optimized vector system, the problems of low repair efficiency and high mutation rate in CRISPR/Cas9 technology have been solved, enabling more efficient and precise gene editing.

WO2025260207A1PCT designated stage Publication Date: 2025-12-26BEIDAHUANG KENFENG SEED CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099527
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing CRISPR/Cas9 gene editing technology is prone to introducing insertion or deletion mutations during non-homologous end linking repair, and its homology-dependent repair efficiency is low, which limits its widespread application in gene editing.

Method used

Novel effector nucleases Cas-kf8 and Cas-kf9 were developed. The amino acid sequences of these enzymes were obtained through design and screening, and they were combined with gRNA to form a riboprotein complex, enabling specific cleavage of target DNA sequences. This expanded the number of editable gene sites, and the vector system was optimized to improve editing efficiency.

Benefits of technology

It has improved the precision and efficiency of gene editing, expanded the editable gene loci, and enhanced the flexibility and application potential of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099527_26122025_PF_FP_ABST
    Figure CN2024099527_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of biotechnology. Provided are a new effector nuclease and the use thereof. On the basis of discovered and studied Cas12a effector nuclease sequences, two new clustered regularly interspersed short palindromic repeats (CRISPR)-associated effector nucleases are obtained by means of design, screening, and functional verification. Specifically provided is a Cas protein. The Cas protein is Cas-kf8 or Cas-kf9, wherein the Cas-kf8 has an amino acid sequence as shown in SEQ ID No. 1, and the Cas-kf9 has an amino acid sequence as shown in SEQ ID No. 4. The provided Cas protein demonstrates nuclease activity in experiments and has broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Effector nucleases and their applications Technical Field

[0001] This invention belongs to the field of biotechnology and relates to novel effector nucleases Cas-kf8 and Cas-kf9 associated with regularly clustered short palindromic repeats (CRISPR), and their applications in gene editing. Background Technology

[0002] The rapidly developing gene editing technology of recent years utilizes Cas (CRISPR-associated) effector nucleases related to the clustered regularly interspersed short palindromic repeats (CRISPR) in the bacterial acquired immune system to cut specific genes in cells at pre-selected sites in the genome. The breaks are then repaired through two inherent DNA repair mechanisms: non-homologous end joining (NHEJ) and homology-dependent repair (HDR), ensuring cell survival. NHEJ is highly efficient, requires no template, and introduces insertion or deletion mutations during the rejoining process at the cut site, thus inactivating the gene; it is the dominant mechanism. HDR requires a repair template and achieves precise repair according to the template; it accounts for a very small proportion. If a pre-designed DNA fragment is provided as a template, the sequence can be inserted at a pre-selected site to modify the target gene, but its efficiency is extremely low, severely limiting its widespread application.

[0003] Gene editing technology based on the regularly clustered, interspaced short palindromic repeats of the CRISPR effector nuclease SpCas9 from *Streptococcus pyogenes* was first successfully implemented in vivo in 2013 (Cong et al., 2013, Science 339, 819-823; Mali et al., 2013, Science 339, 823-826). The Cas9 endonuclease forms a nucleoprotein complex with the guide gRNA, recognizing the target DNA sequence through base pairing between the 20 bp RNA sequence and the edited DNA sequence. This method exhibits high specificity, strong enzyme activity, and a simple experimental design, leading to the rapid and widespread successful application of the CRISPR / Cas9 gene editing system in various plant and animal systems. It is also a preferred method for plant breeding using gene editing technology, enabling knockout editing primarily via the NHEJ pathway in various plants, as well as precise HDR-mediated editing in a few cases (Li et al., 2015, Plant Physiol 169, 960-970; Ma et al.). Chen et al., 2016, Mol Plant 9, 961-974; Chen et al., 2019, Annu Rev Plant Biol 70, 667-97; Mao et al., 2019, Natl Sci Rev 6, 421-437. Combined with multiplex gRNA synchronous expression technology, CRISPR / Cas9 technology can simultaneously and efficiently edit multiple genes at multiple sites via the NHEJ pathway, achieving simultaneous improvement of multiple traits (Xie et al., 2015, Proc Natl Acad Sci USA 112, 3570-3575; Hsieh-Feng and Yang 2020, aBIOTECH, doi.org / 10.1007 / s42994-019-00014-w). By fusing Cas9 with other proteases such as deaminases or reverse transcriptases, precise editing of single bases or small fragments of DNA sequences can be achieved (Kim et al., 2017, Nat Biotechnol 35, 371-376; Komor et al., 2016, Nature 533, 420-424; Nishida et al., 2016, Science 353, aaf8729; Gaudelli et al., 2017, Nature 551, 464-471; Anzalone et al., 2019, Nature 576, 149-157).Mutating the SpCas9 protein sequence can yield mutants that recognize PAM (Protospacer Adjacent Motif) sequences other than NGG, expanding the number of editable gene sites (Hu et al., 2018, Nature 556, 57-63; Kleinstiver et al., 2015, Nature 523, 481-485; Nishimasu et al., 2018, Science 361, 1259-1262; Walton et al., 2020, Science 368, 290-296). By fusing Cas9 with functional groups of other transcription factors and other effector proteins, gene expression levels can be regulated (Pan et al., 2021, Curr Opinion Plant Biol 60, 101980).

[0004] Acquired immune systems containing CRISPR structures are widely found in bacteria and archaea. Based on the number of proteins involved in the immune interference response, they are divided into two main classes: Class I and Class II. Class I systems require multiple proteins to form a complex to perform immune interference, while Class II systems can achieve interference with just a single protease containing a multifunctional group (Makarova et al., 2020, Nat Rev Microbiol 18, 67-83; Nidhi et al., 2021, Int J Mol Sci 22, 3327). Each class is further divided into several types based on structure and composition. The commonly used SpCas9 is a typical representative of Class II CRISPR systems, requiring a conserved NGG sequence at the 3' end of the target sequence as a PAM sequence, and blunt-end cleavage of the target sequence (Jinek et al., 2012, Science 337, 816-821). Other effector nucleases of the other two types of CRISPR systems, such as SaCas9 derived from Staphylococcus aureus and NmeCas9 derived from Neisseria meningitidis, have also been successfully used for editing (Ran et al., 2015, Nature 520, 186-191; Amrani et al., 2018, Genome Biol 19, 214).

[0005] Cas12 effector nucleases belonging to the Class 2 type V CRISPR system, such as FnCpfI and LbCpf1 (later also known as FnCas12a and LbCas12a) derived from bacterial strains Francisella tularensis subsp. Novida U112 and Lachnospiraceae bacterium ND2006 respectively, require the target sequence to have a conserved TTTN as a PAM sequence at the 5' end, and then cleave the target sequence with sticky ends (Zetsche et al., 2015, Cell 163, 1-13). Several other Class 2 type V CRISPR system Cas12a endonucleases have also been used for gene editing to varying degrees. Most of these primarily recognize target DNA sequences containing a conserved TTTV (V being A, C, or G) PAM sequence at the 5' end, and cleave the DNA double strand at the distal end of the PAM, producing sticky ends. In addition, the Cas12a effector nuclease also has RNase function, which can autonomously process and cleave the pre-crRNA of the CRISPR repeat array to form a single mature crRNA, which guides the recognition of the target DNA sequence.

[0006] Although Cas12a effector nucleases have been identified in many bacteria based on the principle of homology, the homology between protease sequences from different sources is low. Enzymes with large genetic distances often share only about 30% amino acid sequence similarity. Besides the difficulty in directly determining the biological function of a specific Cas12a enzyme based on sequence due to low homology, not every enzyme is an active protein with immune function. Some Cas12a immune systems are incomplete, lacking other necessary components such as Cas1, Cas2, Cas4, or CRISPR arrays (Yan et al., 2019, Science 363, 88-91). Even Cas12a enzymes with complete composition and sequence do not necessarily possess enzymatic activity or even eukaryotic gene-editing function. Only a few may retain activity after expression in eukaryotes; therefore, specific experimental screening is necessary for verification. Zetsche et al. conducted detailed experimental studies on 16 Cas12a effector nucleases, including FnCpfI and LbCpf1. Only 8 of them showed effective activity in in vitro DNA cleavage experiments. The researchers then optimized the codons of these 8 effector nucleases and expressed them in human cell lines. Only 2 of them showed effective activity and achieved the editing of the target gene (Zetsche et al., 2015, Cell 163, 1-13).

[0007] The CRISPR / Cas gene editing systems discovered so far each have their own characteristics. Based on various reasons and complex and diverse research and application needs, the continuous development of new gene editing tool enzymes is of great significance to the development and application research of biotechnology.

[0008] Summary of the Invention

[0009] Based on the discovered and studied Cas12a effector nuclease sequences, this invention designs, screens, and verifies their functions, obtaining two novel effector nucleases, Cas-kf8 and Cas-kf9. Based on this discovery, this invention develops a new CRISPR / Cas system and a gene editing method based on this system.

[0010] Cas protein

[0011] On one hand, the present invention provides a Cas protein (or, referred to as "effect nuclease", "Cas enzyme" or "effect protein"), which is an effector protein in the CRISPR / Cas system. The Cas protein is Cas-kf8 or Cas-kf9, the amino acid sequence of Cas-kf8 is shown in SEQ ID No 1, and the amino acid sequence of Cas-kf9 is shown in SEQ ID No 4.

[0012] Compared to the sequence of SEQ ID No. 1 or 4, in some embodiments, the Cas protein is a sequence having one or more amino acid substitutions, deletions, or additions, and substantially retains the biological function of its derived sequence; the one or more amino acids include substitutions, deletions, or additions of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more amino acids; in some embodiments, the amino acid sequence has at least 50% sequence identity with the sequence shown in SEQ ID No. 1 or 4, for example, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% identity, and substantially retains the biological function of its derived sequence.

[0013] Those skilled in the art will understand that the primary structure of a protein can be altered without adversely affecting its activity and function. For example, one or more conserved amino acid substitutions can be introduced into the protein's amino acid sequence without adversely affecting the protein molecule's activity and / or three-dimensional structure. Examples and implementations of conserved amino acid substitutions are familiar to those skilled in the art. Specifically, an amino acid residue can be substituted with another amino acid residue belonging to the same group as the site to be substituted; that is, a nonpolar amino acid residue can replace another nonpolar amino acid residue, a polar, uncharged amino acid residue can replace another polar, uncharged amino acid residue, a basic amino acid residue can replace another basic amino acid residue, and an acidic amino acid residue can replace another acidic amino acid residue. Such substituted amino acid residues may or may not be encoded by the genetic code. Conservative substitution, where an amino acid is replaced by another amino acid belonging to the same group, falls within the scope of this invention, provided that the substitution does not lead to inactivation of the protein's biological activity. Furthermore, this invention also covers proteins containing one or more other nonconserved amino acid substitutions, provided that such nonconservative substitution does not significantly affect the desired function and biological activity of the protein of this invention.

[0014] The functions and activities include, but are not limited to, gRNA binding activity, endonuclease activity, and gRNA-guided binding to and cleavage of specific sites on target sequences, including but not limited to Cis cleavage activity and Trans cleavage activity.

[0015] As is well known in the art, one or more amino acid residues can be altered (replaced, deleted, truncated, or inserted) from the N- and / or C-termini of a protein while retaining its functional activity. Therefore, proteins whose N- and / or C-termini of the Cas protein of this invention have been altered, while retaining their desired functional activity, are also within the scope of this invention. These alterations may include those introduced by modern molecular methods such as PCR, which includes PCR amplification that alters or lengthens the protein-coding sequence by means of oligonucleotides containing amino acid-coding sequences used in the PCR amplification.

[0016] It should be recognized that proteins can also be altered in various ways, including amino acid substitutions, deletions, truncations, and insertions, and methods for such operations are generally known in the art. For example, amino acid sequence variants of Cas proteins can be prepared by mutating DNA. This can also be accomplished through other forms of mutagenesis and / or directed evolution, for example, using known mutagenesis, recombination, and / or shuffling methods, combined with relevant screening methods, to perform single or multiple amino acid substitutions, deletions, and / or insertions.

[0017] Those skilled in the art will understand that these minor amino acid changes in the Cas protein of the present invention can occur (e.g., naturally occurring mutations) or be generated (e.g., using rDNA technology, recombinant DNA) without loss of protein function or activity. If these mutations occur in the catalytic domain, active site, or other functional domains of the protein, the properties of the polypeptide may be altered, but the polypeptide may retain its activity. If the mutations are not located near the catalytic domain, active site, or other functional domains, a smaller impact can be expected.

[0018] The present invention also provides a fusion protein comprising a Cas protein with the amino acid sequence SEQ ID No 1 (Cas-kf8) or SEQ ID No 4 (Cas-kf9) and other modified portions.

[0019] In one embodiment, the modified portion is selected from other proteins or peptides, detectable markers, or any combination thereof.

[0020] In one embodiment, the modified portion is selected from epitope tags, reporter gene sequences, nuclear localization signal (NLS) sequences, targeting portions, transcriptional activation domains (e.g., VP64), transcriptional repression domains (e.g., KRAB or SID domains), nuclease domains (e.g., Fok1), and domains having activities selected from: nucleotide deaminase, methylase activity, demethylase, transcriptional activation activity, transcriptional repression activity, transcriptional release factor activity, histone modification activity, nuclease activity, single-stranded RNA cleavage activity, double-stranded RNA cleavage activity, single-stranded DNA cleavage activity, double-stranded DNA cleavage activity, and nucleic acid binding activity; and any combination thereof.

[0021] The NLS (Nuclear Localization Signal) nuclear localization sequence is well known to those skilled in the art, and examples include, but are not limited to, SV40 large T antigen, Nucleoplasmin, EGL-13, c-Myc, and TUS protein.

[0022] In one embodiment, the NLS sequence is located at, near, or close to the end (e.g., N-terminus, C-terminus, or both ends) of the Cas protein of the present invention.

[0023] The epitope tag is well known to those skilled in the art, including but not limited to His, V5, FLAG, HA, Myc, VSV-G, Trx, etc., and those skilled in the art can choose other suitable epitope tags (e.g., for purification, detection, or tracing).

[0024] The reporter gene sequence is well known to those skilled in the art, and examples include, but are not limited to, GST, HRP, CAT, GFP, HcRed, DsRed, CFP, YFP, BFP, etc.

[0025] In one embodiment, the fusion protein of the present invention contains a detectable marker, such as a fluorescent dye, such as FITC or DAPI.

[0026] In one embodiment, the Cas protein of the present invention is optionally coupled, conjugated, or fused to the modified portion via a linker.

[0027] In one embodiment, the modified portion is attached to the N-terminus or C-terminus of the Cas protein of the present invention via a linker. Such linkers are well known in the art, and examples include, but are not limited to, linkers containing one or more (e.g., 1, 2, 3, 4, or 5) amino acids (e.g., Glu or Ser) or amino acid derivatives (e.g., Ahx, β-Ala, GABA, or Ava), or PEG, etc.

[0028] The Cas protein, protein derivative, or fusion protein of the present invention is not limited by the manner of its production. For example, it can be produced by genetic engineering methods (recombinant technology) or by chemical synthesis methods.

[0029] Nucleic acid of Cas protein

[0030] On the other hand, the present invention provides an isolated polynucleotide comprising a polynucleotide sequence encoding the Cas protein or fusion protein of the present invention.

[0031] In one embodiment, the polynucleotide sequence is codon-optimized for expression in prokaryotic cells. In another embodiment, the polynucleotide sequence is codon-optimized for expression in eukaryotic cells.

[0032] In one embodiment, the polynucleotide is preferably single-stranded or double-stranded.

[0033] In one specific embodiment, the polynucleotide sequence provided by the present invention is shown as SEQ ID No. 2 or 5.

[0034] In some embodiments, the polynucleotide sequence has one or more base substitutions, deletions, or additions (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) compared to the sequence shown in SEQ ID No. 2 or 5.

[0035] Those skilled in the art will understand that, based on the degeneracy of the genetic code, most amino acids are encoded by multiple different codons. The genetic code is the process by which genetic information on DNA or mRNA is translated into proteins in an organism. Each amino acid is encoded by one or more specific trinucleotide sequences (codons). Based on this characteristic, different nucleic acid sequences (codons) can be transcribed and translated into the same amino acid. As long as the change of the polynucleotide sequence does not lead to the translation into different amino acids or the inactivation of the protein's biological activity, the polynucleotide sequence falls within the scope of this invention.

[0036] carrier

[0037] The present invention also provides a carrier comprising the Cas protein and isolated polynucleotides as described above; preferably, it further comprises a regulatory element operatively linked thereto.

[0038] In one embodiment, the regulatory element is selected from one or more of the following: enhancers, transposons, promoters, terminators, leader sequences, polyadenylation sequences, and marker genes.

[0039] In one embodiment, the vectors include cloning vectors, expression vectors, shuttle vectors, and integration vectors, which are vector structures that can be naturally occurring or artificially synthesized and are known to those skilled in the art. In a specific embodiment, the present invention provides commercially available vectors pACYC-Duet-1 and pUC19.

[0040] carrier system

[0041] The present invention provides an engineered, non-naturally occurring vector system, or a CRISPR-Cas system, comprising a Cas protein or a nucleic acid sequence encoding the Cas protein and a nucleic acid encoding one or more gRNAs, wherein the gRNAs comprise a repeating sequence for binding the Cas protein and a spacer sequence for targeting a target sequence, the spacer sequence being complementary to the target sequence.

[0042] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more gRNAs are artificially synthesized.

[0043] In one embodiment, the nucleic acid sequence encoding the Cas protein and the nucleic acid encoding one or more gRNAs do not coexist naturally.

[0044] The one or more gRNAs target one or more target sequences in the cell. The one or more gRNAs guide the Cas protein to the genomic locus of the DNA molecule encoding one or more gene products by hybridizing with the target sequence. Once at the target sequence location, the Cas protein modifies, edits, or cuts the target sequence, thereby altering or modifying the expression of the one or more gene products.

[0045] The cells of this invention include one or more of animals, plants, or microorganisms.

[0046] In some embodiments, the Cas protein is codon-optimized for expression in cells.

[0047] In some embodiments, the Cas protein directs the cleavage of one or both nucleic acid strands at the target sequence location.

[0048] In one specific embodiment, the repeating sequence designed by the present invention that binds to the Cas protein is shown in SEQ ID No. 3 or 6.

[0049] The present invention also provides an engineered, non-naturally occurring vector system, which may include one or more vectors, the one or more vectors including: a gRNA expression unit and the Cas protein expression unit, wherein the gRNA expression unit and the Cas protein expression unit are located on the same or different vectors in the system.

[0050] The expression unit connects to a variety of regulatory elements, including promoters (e.g., constitutive or inducible promoters), enhancers (e.g., 35S promoter or 35S enhanced promoter), internal ribosome entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylate (PolyA) signals and polyU sequences).

[0051] In one embodiment, the gRNA expression unit is ZmU6 pro (promoter):gRNA:AtU6-26 term. Furthermore, a BsaI restriction site is left between ZmU6 pro and gRNA, allowing one or more gRNA sequences to be inserted to form a gene editing vector system targeting any site of the target gene.

[0052] In one embodiment, the Cas protein expression unit is ZmUBI1 pro (promoter): Cas (Cas-kf8 or Cas-kf9): PsE9 term (terminator).

[0053] In one embodiment, the vector system further includes a resistance selection gene expression unit, wherein the resistance selection gene is used in the genetic transformation process to screen out successfully transformed cells or organisms. Examples of such genes include, but are not limited to, marker genes such as hygromycin resistance gene (HYG), neomycin resistance gene, tetracycline resistance gene, penicillin resistance gene, streptomycin resistance gene, herbicide resistance gene, heavy metal resistance gene, and antibiotic resistance gene.

[0054] In one specific embodiment, the resistance selection gene expression unit is 35S pro:HYG:35S term.

[0055] In some implementations, the vector in the system is a viral vector (e.g., a retroviral vector, lentiviral vector, adenovirus vector, and herpes simplex vector), or it can be a plasmid, virus, granule, bacteriophage, or other type known to those skilled in the art.

[0056] In one embodiment, when the target sequence is DNA, the target sequence is located at the 3' end of the protospacer adjacent motif (PAM), and the PAM has a sequence represented by TTN, wherein N is selected from A, G, T, and C.

[0057] In one embodiment, the target sequence is a DNA or RNA sequence from a prokaryotic or eukaryotic cell.

[0058] In one implementation, the target sequence is a non-naturally occurring DNA or RNA sequence.

[0059] In one embodiment, the target sequence is present within a cell. In another embodiment, the target sequence is present in the cell nucleus or cytoplasm (e.g., organelles). In one embodiment, the cell is a eukaryotic cell. In other embodiments, the cell is a prokaryotic cell.

[0060] In one embodiment, the Cas protein is linked to one or more NLS sequences. In one embodiment, the fusion protein comprises one or more NLS sequences. In one embodiment, the NLS sequence is linked to the N-terminus or C-terminus of the protein. In one embodiment, the NLS sequence is fused to the N-terminus or C-terminus of the protein.

[0061] Protein-nucleic acid complexes / compositions

[0062] On the other hand, the present invention provides a composition (or complex) comprising:

[0063] (i) a protein component selected from the Cas protein or the fusion protein; (ii) a nucleic acid component selected from gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a nucleic acid encoding the precursor RNA of the gRNA; wherein the gRNA comprises a repeating sequence that binds to the Cas protein and a spacer sequence for targeting the target sequence, the spacer sequence being complementary to the target sequence.

[0064] The Cas protein and gRNA of this invention can form a binary complex, which is activated upon binding to a target sequence substrate to form an activated CRISPR complex. The target sequence substrate is complementary to the spacer sequence (or guide sequence for hybridization with the target nucleic acid) in the gRNA. In some embodiments, the spacer sequence of the gRNA perfectly matches the target sequence substrate. In other embodiments, the spacer sequence of the gRNA partially (continuously or discontinuously) matches the target sequence substrate, and the activated complex can exhibit trans-cleavage activity. This trans-cleavage activity refers to the non-specific or random cleavage activity of the activated CRISPR complex against other nearby single-stranded nucleic acids, also known in the art as trans-cleavage activity.

[0065] host cells

[0066] The present invention also relates to an in vitro, ex vivo, or in vivo cell or cell line or its progeny, said cell or cell line or its progeny comprising: the Cas protein, fusion protein, polynucleotide, protein-nucleic acid composition, vector or vector system described in the present invention.

[0067] In some implementations, the cell is a prokaryotic cell.

[0068] In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is a non-human mammalian cell, such as cells of non-human primates, cattle, sheep, pigs, dogs, monkeys, rabbits, or rodents (e.g., rats or mice). In some embodiments, the cell is a non-mammalian eukaryotic cell, such as cells of poultry (e.g., chickens), fish, or crustaceans (e.g., clams, shrimp). In some embodiments, the cell is a plant cell, such as cells of monocotyledonous or dicotyledonous plants, or cells of cultivated plants or food crops such as cassava, corn, sorghum, soybeans, wheat, oats, or rice, such as algae, trees, or productive plants, fruits, or vegetables (e.g., trees such as citrus trees, nut trees; nightshade plants, cotton, tobacco, tomatoes, grapes, coffee, cocoa, etc.).

[0069] In some implementations, the cell is a stem cell or stem cell line.

[0070] In some cases, the host cells of the present invention contain genetic or genomic modifications that are not present in their wild type.

[0071] application

[0072] The Cas protein, fusion protein, polynucleotide, vector, CRISPR-Cas system, vector system, composition, or host cell provided by this invention have, in some embodiments, any of the following applications:

[0073] (1) Applications in gene editing or gene targeting for the purpose of diagnosis and treatment of non-human diseases;

[0074] (2) Applications in the preparation of gene editing or gene targeting reagents or kits;

[0075] (3) Applications of gene cutting for purposes other than disease diagnosis and treatment;

[0076] (4) Applications in the preparation of gene cutting reagents or kits;

[0077] (5) Application in nonspecific nucleic acid side branch cleavage;

[0078] (6) Application in nonspecific nucleic acid collateral degradation;

[0079] (7) Application in plant gene editing breeding.

[0080] Gene cleavage as described herein includes: DNA or RNA breaks in the target sequence of the Cas protein described herein (Cis cleavage), and breaks in non-target DNA or RNA (single-stranded nucleic acid substrates) during side branch cleavage (i.e., nonspecific or non-targeted, trans cleavage). In some embodiments, the cleavage is a double-stranded DNA break. In some embodiments, the cleavage is a single-stranded DNA break or a single-stranded RNA break.

[0081] Carrier system construction method

[0082] In the specific field of plant gene editing breeding, this invention further constructs vectors for plant gene editing and gene targeting.

[0083] In one embodiment, the present invention provides a basic editing vector with pCambia3301 as the vector backbone, and three expression units set from the right boundary of T-DNA: ZmU6 pro:gRNA:AtU6-26 term, ZmUBI1 pro:LbCpf1:PsE9 term, and 35S pro:HYG:35S term. The Cas-kf8 basic editing vector or the Cas-kf9 basic editing vector is obtained by replacing LbCpf1 with the Cas protein.

[0084] In one specific implementation, the present invention provides a target sequence editing vector, wherein gene editing sites are selected for different regions of the target sequence to be edited, a recognition sequence that can be paired with the editing site is designed to form a gRNA sequence with a crRNA repeat sequence, one or more gRNA sequences are cross-linked using tRNA technology, and one or more gRNA sequences are inserted after the basic editing vector is linearized using DNA ligase or Gibson cloning method to form the target sequence editing vector.

[0085] As a specific embodiment, the gRNA sequence for the target sequence constructed in this invention is shown in SEQ ID NO 13.

[0086] As a more specific embodiment, the Cas-kf8 target sequence editing vector constructed by the present invention has the sequence shown in SEQ ID NO 14.

[0087] Based on the above-mentioned vector, in one embodiment, the present invention provides a method for genetic transformation of plants, wherein the obtained target gene editing vector is introduced into recipient plant cells, and then screened to obtain heritable, non-transgenic and stably inherited progeny plants.

[0088] Furthermore, the genetic transformation method includes the following steps:

[0089] (1) Genetic transformation of recipient plant cells to obtain regenerated plants (including: 1) induction of callus tissue; 2) subculture of callus tissue; 3) infection of callus tissue; 4) screening and identification of resistant callus tissue; 5) differentiation of callus tissue into seedlings; 6) rooting culture; 7) domestication and transplanting to obtain regenerated plants);

[0090] (2) To detect whether gene editing of the regenerated plants was successful;

[0091] (3) Screening to obtain homozygous edited progeny plants with the target gene trait, non-transgenic, and stably inherited.

[0092] In one embodiment, the present invention also provides a method for editing, targeting, or cutting a target sequence, the method being a method not for disease diagnosis and treatment purposes, the object of the method not being human, the method comprising contacting the target sequence with the CRISPR-Cas system, or the vector system, or the composition, or the host cell.

[0093] In one embodiment, the present invention also provides a kit for gene editing or gene targeting, the kit comprising the CRISPR-Cas system, or the vector system, or the composition, or the host cell.

[0094] The gene editing or editing of target nucleic acids includes gene modification, gene knockout, alteration of gene product expression, mutation repair, insertion of polynucleotides, and gene mutation.

[0095] In one embodiment, the present invention provides a kit for detecting a target sequence in a sample, the kit comprising:

[0096] (a) the Cas protein described herein, or the nucleic acid encoding the Cas protein; and / or

[0097] (b) gRNA, or nucleic acid encoding said gRNA, or precursor RNA containing said gRNA, or nucleic acid encoding said precursor RNA; said gRNA comprising a repeating sequence that binds to said Cas protein and a spacer sequence that targets and is complementary to the target sequence; and / or

[0098] (c) A single-stranded nucleic acid detector, wherein the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.

[0099] In one embodiment, the present invention provides a method for detecting a target sequence in a sample, the method being a method for non-disease diagnosis and treatment purposes, the method comprising contacting the sample with the Cas protein, gRNA, and a single-stranded nucleic acid detector, detecting a detectable signal generated by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting the target gene; the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.

[0100] In this invention, the single-stranded nucleic acid detector includes, but is not limited to, single-stranded DNA, single-stranded RNA, DNA-RNA hybrids, nucleic acid analogs, base modifiers, and other single-stranded nucleic acid detectors; "nucleic acid analogs" include, but are not limited to: locked nucleic acid, bridging nucleic acid, morpholine nucleic acid, ethylene glycol nucleic acid, hexitol nucleic acid, threonine nucleic acid, arabinose nucleic acid, 2'-oxymethyl RNA, 2'-methoxyacetyl RNA, 2'-fluoroRNA, 2'-aminoRNA, 4'-sulfur RNA, and combinations thereof, including optional ribonucleotide or deoxyribonucleotide residues.

[0101] In this invention, the detectable signal is achieved through the following methods: vision-based detection, sensor-based detection, color detection, fluorescence signal-based detection, gold nanoparticle-based detection, fluorescence polarization, fluorescence detection, colloidal phase transition / dispersion, electrochemical detection, and semiconductor-based detection.

[0102] In this invention, preferably, a fluorescent group and a quenching group are respectively disposed at both ends of the single-stranded nucleic acid detector, so that the single-stranded nucleic acid detector can exhibit a detectable fluorescent signal after being cleaved. The fluorescent group is selected from one or any combination of FAM, FITC, VIC, JOE, TET, CY3, CY5, ROX, Texas Red, or LC RED460; the quenching group is selected from one or any combination of BHQ1, BHQ2, BHQ3, Dabcy1, or Tamra.

[0103] Terminology Definition

[0104] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the operational steps used herein, such as molecular genetics, nucleic acid chemistry, chemistry, molecular biology, biochemistry, cell culture, microbiology, cell biology, genomics, and recombinant DNA, are all conventional steps widely used in their respective fields. To better understand this invention, definitions and explanations of relevant terms are provided below.

[0105] The term "gRNA (guide RNA)" refers to any polynucleotide sequence that is sufficiently complementary to a target sequence to hybridize with it and guide the CRISPR / Cas system to specifically bind to the target sequence. The structure of a gRNA typically includes a crRNA (CRISPR RNA) sequence derived from a CRISPR array. This sequence contains repeat sequences and spacer sequences. The spacer sequence is derived from foreign genetic material and is complementary to the target sequence. Typically, gRNAs may also have a non-complementary single-stranded region at the 5' end, called a "5' overhang," which helps distinguish the gRNA from the target sequence. In some cases, the 3' end of the gRNA may contain an inversely complementary tail, which helps improve the binding stability of the gRNA to the Cas protein. In some applications, gRNAs may also contain additional modifications, such as phosphorylation or chemical modifications, to improve their stability or functionality.

[0106] The term "target sequence" is synonymous with "target sequence" and "target nucleic acid" in this invention, referring to a polynucleotide targeted by a spacer sequence in crRNA. The target sequence can contain any polynucleotide, such as DNA or RNA. Hybridization between the target sequence and crRNA will promote the expression of a CRISPR / Cas system. Perfect complementarity is not required, as long as there is sufficient complementarity to induce hybridization and promote the expression of a CRISPR / Cas system.

[0107] The term "non-natural" as used herein, and the terms "non-natural" or "engineered," are used interchangeably and indicate artificial involvement. When these terms are used to describe nucleic acid molecules or peptides, they refer to those that do not exist in nature or are not synthesized naturally by organisms.

[0108] The term "identity," as used herein, refers to the sequence matching between two polypeptides or two nucleic acids. Two compared sequences are considered identical at that position when a position is occupied by the same base or amino acid monomeric subunit (e.g., a position in each of two DNA molecules occupied by adenine, or a position in each of two polypeptides occupied by lysine). The "percentage identity" between two sequences is a function of the number of matching positions shared by the two sequences divided by the number of positions compared × 100. For example, if six out of ten positions in two sequences match, then the two sequences have 60% identity. For example, the DNA sequences CTGACT and CAGGTT have 50% identity (three out of six positions match). Typically, two sequences are compared to produce the maximum possible identity. Such comparisons can be performed, for example, by computer programs, such as the Needleman and Wunsch (J MoIBiol.48:444-453(1970)) algorithm in the GAP program, using a Blossum 62 matrix or a PAM250 matrix and gap weights of 16, 14, 12, 10, 8, 6 or 4 and length weights of 1, 2, 3, 4, 5 or 6 to determine the percentage identity between two amino acid sequences.

[0109] The term "plant" should be understood as any differentiated multicellular organism capable of photosynthesis, including crop plants at any stage of maturity or development, particularly monocotyledonous or dicotyledonous plants; vegetable crops, including artichokes, kohlrabi, arugula, leeks, asparagus, lettuce, cabbage, cauliflower, broccoli, kale, taro, and melons (e.g., cantaloupe, watermelon); and fruit crops such as apples, apricots, cherries, nectarines, peaches, pears, plums, prunes, cherries, blackberries, grapes, avocados, and bananas. Kiwifruit, persimmon, pomegranate, pineapple, mango, papaya, and lychee, etc.; field crops such as alfalfa, corn / maize (feed corn, sweet corn, popcorn), peanuts, small grain crops (barley, oats, rye, rice, wheat, etc.), sorghum, tobacco, kapok, legumes (beans, lentils, peas, soybeans), oil plants (rapeseed, mustard, poppy, olive, sunflower, coconut, castor oil plants, cocoa beans, peanuts), fiber plants (cotton, flax, jute), etc.

[0110] The technical effects achieved by this invention are as follows:

[0111] Based on the discovered and studied Cas12a effector nuclease sequences, this invention designs, screens, and verifies the functions of two novel effector nucleases, Cas-kf8 and Cas-kf9. Experiments show that they exhibit nuclease activity and have broad application prospects.

[0112] This invention yielded two novel Cas12a-type effector nuclease proteins, Cas-kf8 and Cas-kf9. Through codon optimization and plant gene editing transformation experiments, these proteins demonstrated the ability to recognize and cleave target sites with TTTV sequences at the 5' end, resulting in novel effector nucleases suitable for plant gene editing. Furthermore, this invention provides a highly efficient plant gene editing vector for targeted genome modification in plants. Attached Figure Description

[0113] Figure 1 shows a schematic diagram of the Cas-kf8 and Cas-kf9 effector nucleases and their CRISPR array structures (in the figure: (a) Cas-kf8 and Cas-kf9 are two types of CRISPR-related Cas12a effector nucleases derived from Peptostreptococcus equinus and Clostridium thermobutyricum bacteria, respectively, and their genomes contain characteristically arranged Cas4, Cas1, and Cas2 enzyme genes and CRISPR arrays; (b) the crRNA repeat sequences of Cas-kf8 and Cas-kf9, whose 3' ends both contain conserved sequences UCUACU and GUAGAU that can form hairpin structures).

[0114] Figure 2 shows a schematic diagram of the structure of the rice starch synthase gene and the selected editing sites (in the figure: the starch synthase gene Waxy (LOC_Os06g04200) consists of 14 exons and 13 introns. One site each in exons 6 and 8 was selected to design OsWx-T1 and OsWx-T2 gRNAs, and PCR-specific primers OsWx-F5 and OsWx-R1 were designed upstream and downstream of each for amplification and sequencing analysis of the edited DNA fragments).

[0115] Figure 3 is a schematic diagram of the gene editing transformation vector KF316 (in the figure: T-DNA RB is the right boundary of T-DNA; ZmU6 pro is the promoter of the maize nuclear small RNA (U6 snRNA) gene; tRNA and crRNA are combinations of tRNA and Cas12a-specific crRNA; AtU6 term is the terminator of the Arabidopsis U6 snRNA gene; ZmUBI1 pro and ZmUBI1 intron are promoters of the maize ubiquitin gene including introns; Cas-kf8 is a novel Cas12a effector nuclease belonging to the class 2 type V CRISPR system; PsE9 term is the terminator of the pea ribulose-1,5-bisphosphate carboxylase / oxygenase small subunit E9 protein gene; CaMV 35Spro is the cauliflower mosaic virus 35S promoter; HygR is hygromycin phosphotransferase; CaMV 35S term is the cauliflower mosaic virus 35S terminator; T-DNA LB is the left boundary of T-DNA).

[0116] Figure 4 shows the DNA sequence near the OsWx-T2 editing site in the target region during the Cas-kf8 gene editing event.

[0117] Figure 5 shows the DNA sequence of the Cas-kf9 gene editing event in the target region (in the figure: (a) DNA sequence near the OsWx-T1 editing site; (b) DNA sequence near the OsWx-T2 editing site). Detailed Implementation

[0118] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and embodiments, but this does not limit the invention to the scope of the embodiments described. Unless otherwise specified, the experiments and methods described in the embodiments are generally carried out according to conventional methods well known in the art and described in various references. The reagents and raw materials used in the present invention are all commercially available. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are followed. Where the manufacturer of the reagents or instruments used is not specified, they are all conventional products that can be obtained commercially. Those skilled in the art will understand that the embodiments describe the present invention by way of example and should not be construed as limiting the scope of protection claimed by the present invention. All disclosures and other references mentioned herein are incorporated herein by reference in their entirety.

[0119] In this invention, the molecular biology experiments involving DNA sequence analysis and alignment, DNA vector cloning and design, PCR primers and CRISPR gRNA design were all performed using the DNA analysis software Geneious Prime (GraphPad Software LLC).

[0120] The concept of this invention includes:

[0121] 1. Obtaining novel effector nuclease genes: Based on the amino acid sequences of the Cas12a effector nucleases that have been discovered and studied, we searched microbial genome and metagenomic databases such as NCBI (www.ncbi.nlm.nih.gov), JGI (img.jgi.doe.gov), and UniProt (www.uniprot.org) multiple times using different query programs such as Blastp and Tblastn. We screened for recently sequenced genes with certain homology and containing the RuvC nuclease domain at the carboxyl terminus of the protein. We retrieved the corresponding genomic nucleic acid sequences and used DNA analysis software such as Geneious to find CRISPR repeat sequence arrays and protein-coding regions. Based on their conserved Cas12a-Cas4-Cas1-Cas2-CRISPR gene arrangement, we identified Cas12a effector nucleases. We selected novel Cas12a effector nucleases with low similarity to those reported in the literature and designed experiments for functional verification and screening.

[0122] 2. Validation of the bacterial plasmid interference function of a novel effector nuclease: Most known Cas12a effector nucleases require the target DNA sequence to contain a conserved TTTV (V for A, C, or G) PAM sequence at the 5' end. Based on this, a target sequence containing TTTV at the 5' end was designed, cloned into the pUC19 vector, transformed into E. coli competent cells TOP10, and plasmid DNA was extracted as the target DNA. A bacterial expression unit Cas12a-gRNA containing a novel Cas12a effector nuclease and a corresponding gRNA recognizing the above target sequence was designed, cloned into the pACYC-Duet-1 vector, transformed into stable competent bacteria for expression, and strains containing this Cas12a-gRNA expression vector were prepared as new competent cells. The novel effector nuclease was transformed into competent cells containing Cas12a-gRNA using pUC19 plasmid DNA containing the target sequence. This verified the plasmid interference function of the novel effector nuclease. If the Cas12a-gRNA has effector nuclease function, it will recognize the target DNA sequence and cleave the plasmid at a specific site, interfering with the survival of the transformed pUC19 plasmid DNA and preventing normal colony formation. In contrast, the control transformation of competent cells without Cas12a-gRNA could form colonies normally. The difference in the number of colonies between the two reflects the strength of the plasmid interference function of the effector nuclease. Based on this, effector nucleases with strong plasmid interference function can be screened.

[0123] 3. Establish a plant gene editing system using novel effector nucleases: Select novel effector nucleases with high activity from bacterial plasmid interference experiments and their corresponding crRNA repeat sequences. Utilize suitable plant expression promoters and terminators to construct gene expression units and build basic vectors for plant transformation expression. Select the rice starch synthase gene (Waxy or OsWx) LOC_Os06g04200|Chr6:1765621..1770656forward as the target gene. Design two Waxy gene-specific gRNAs, clone them into the aforementioned plant transformation vector, and simultaneously express the novel effector nuclease and its gRNA. Introduce these gRNAs into Agrobacterium EHA105 for rice transformation.

[0124] 4. Novel effector nuclease-mediated plant gene editing: Rice was transformed with the above-mentioned Agrobacterium, resistant callus tissue was selected, genomic DNA was extracted, and the transformation vector-specific gene fragment was amplified by PCR to confirm that the resistant callus was a positive transgenic event. The edited region of the target gene was amplified using the same genomic DNA, and the gene editing effect was analyzed by sequencing. The editing efficiency at different sites was statistically calculated.

[0125] Example 1

[0126] I. Obtaining novel effector nuclease genes and their crRNA repeat sequences

[0127] Microbial genome search: Based on the amino acid sequences of known Cas12a effector nucleases such as LbCpf1 and FnCpf1 (Zetsche et al., 2015, Cell 163, 1-13), microbial genome and metagenomic databases such as NCBI, JGI, and UniProt were repeatedly searched using Blastp, Tblastn, etc., to collect genes with similar proteins, and repetitive sequences were removed. Based on the average length of known Cas12a proteins, proteins far exceeding this range were removed, retaining only genes with complete protein lengths between 900 and 1500 amino acids. Sequences completely identical to known Cas12a effector nucleases were excluded. Newly sequenced genes with low similarity but containing a RuvC nuclease domain at the carboxyl terminus were screened as initial Cas12a effector nucleases for further screening.

[0128] Patent application search: Using keywords such as Cas9, Cas12, and CRISPR, searches were conducted in the USPTO (www.uspto.gov) and Chinese patent applications (tysf.cponline.cnipa.gov.cn) to collect published effector nuclease sequences and construct a query library. The initially selected Cas12a effector nucleases were then used to search the collected effector nuclease library using Blastp, excluding sequences completely identical to known Cas12 effector nucleases. Sequences were then sorted by similarity, and those with low similarity were selected for detailed sequence analysis and further refinement.

[0129] CRISPR sequence analysis of effector nucleases: Genomic nucleic acid sequences corresponding to selected Cas12a effector nucleases were retrieved, and CRISPR repeat sequence arrays were searched using DNA analysis software such as Geneious. Based on the conserved Cas12a-Cas4-Cas1-Cas2-CRISPR gene arrangement, it was determined whether it was a Cas12a effector nuclease. Cas12a effector nucleases with intact and conserved CRISPR structures, appropriate lengths, and recent sequencing, along with their corresponding crRNA repeat sequences, were selected for functional verification and further screening experiments.

[0130] Through the above-mentioned multi-step screening and subsequent multiple experimental verifications, the inventors obtained two novel type V CRISPR Cas12a effector nucleases, named Cas-kf8 and Cas-kf9, respectively. Both of them have typical Cas12a-Cas4-Cas1-Cas2-CRISPR gene structure arrangement and CRISPR repeat sequence array (Figure 1a). Cas-kf8 was derived from shotgun sequencing of the Peptostreptococcus equinu genome, and its protein, nucleotide, and CRISPR repeat sequences are shown in SEQ ID NO 1-3; Cas-kf9 was derived from shotgun sequencing of the Clostridium thermobutyricum genome, and its protein, nucleotide, and CRISPR repeat sequences are shown in SEQ ID NO 4-6; the protein contains artificially added SV40 and Nucleoplasmin nuclear localization signal (NLS) sequences at both ends, and the nucleotides have been optimized to codon sequences suitable for eukaryotic cell expression; the 3' end of the crRNA repeat sequence has conserved UCUACU and GUAGAU sequences that can form hairpin structures (Figure 1b).

[0131] II. Validation of the bacterial plasmid interference function of novel effector nucleases

[0132] Prokaryotic expression vectors pACYC-Duet-1 and pUC19 were purchased from Beijing TransGen Biotech Co., Ltd., and *E. coli* competent cells TOP10 and Stable, and *Agrobacterium* EHA105 were purchased from Shanghai Weidi Biotechnology Co., Ltd. LB broth contained 10 g / L tryptone, 5 g / L yeast extract, and 10 g / L sodium chloride (NaCl), and was autoclaved at 121°C for 20 minutes. If antibiotics were added, they were added to a final concentration of 50 μg / mL after the medium had cooled appropriately. LB solid medium was prepared by adding 15 g / L agar to the LB broth before sterilization. Genes and primers requiring artificial synthesis were synthesized by Suzhou Genewise Biotech Co., Ltd., and Sanger sequencing of DNA was performed by Sangon Biotech (Shanghai) Co., Ltd.

[0133] 1. Design and construction of target sequence plasmids: The fragment of the rice starch synthase gene (Waxy or OsWx) LOC_Os06g04200|Chr6:1765621..1770656forward was selected as the target sequence for bacterial experiments (Figure 2). Two fragments located in exons 6 and 8, respectively, containing the TTTV sequence (PAM) at their 5' ends, were selected as the target sequences OsWx-T1 and OsWx-T2 (SEQ ID NO 7, 9). The same target sequences OsWx-T1 and OsWx-T2 will be used for subsequent plant gene editing experiments.

[0134] The target recognition sequences OsWx-T1 and OsWx-T2 (SEQ ID NO 7, 9) were artificially synthesized and cloned into the pUC19 vector to form KF280 and KF282, respectively, which were used as target genes for bacterial plasmid interference experiments. The OsWx-T1 and OsWx-T2 target sequences of the KF280 and KF282 plasmid DNA contain complete PAMs. After transformation into competent cells containing effector nuclease expression plasmids, they can be cleaved by effector nuclease complexes with specific recognition sequences. The fragmented plasmid DNA cannot form colonies, thus producing plasmid interference.

[0135] The target sequences of OsWx-T1 and OsWx-T2 (SEQ ID NO 8, 10) without the PAM sequence were artificially synthesized and cloned into the pUC19 vector to form KF281 and KF283, respectively. These were used as controls without effective target sequences for bacterial plasmid interference experiments. The incomplete OsWx-T1 and OsWx-T2 target sequences of KF281 and KF283, lacking PAM, could not be cleaved by the effector nuclease complex with the specific recognition sequence after transformation into competent cells containing effector nuclease expression plasmids. The intact plasmid DNA could form colonies normally, and plasmid interference did not occur.

[0136] 2. Construction of bacterial expression vectors for effector nucleases:

[0137] Design and construct bacterial expression vectors for effector nucleases: The prokaryotic expression vector pACYC-Duet-1 was modified as the basic backbone. Prokaryotic gene expression promoters proC and J23119 (Chavez et al., 2018, PNAS USA 115, 3669-3673; parts.igem.org / Part:BBa_J23100) were artificially synthesized to express effector nucleases and their gRNA recognition sequences, respectively, replacing elements such as the lac repressor gene in the pACYC-Duet-1 vector. Appropriate restriction endonuclease sites were reserved for subsequent insertion and cloning of effector nucleases and their corresponding crRNAs. Using LbCpf1 of Lachnospiraceae bacterium as a control, the LbCpf1 gene (SEQ ID NO 11) was artificially synthesized, and short peptides of SV40 (9 amino acids) and Nucleoplasmin (16 amino acids) nuclear localization signals were added to its amino and carboxyl terms, respectively, for subsequent nuclear localization after expression in eukaryotic cells. This was cloned downstream of the proC promoter of the modified pACYC-Duet-1 vector, and expression was terminated by the terminator sequence of the Streptococcus pyogenes Cas9 gene. The artificially synthesized LbCpf1 crRNA+OsWx-T1 recognition sequence (SEQ ID NO 12, 8) was cloned downstream of the J23119 promoter of the above vector, and gene transcription was terminated downstream by the T7 terminator, ultimately forming the bacterial expression vector KF278 containing the LbCpf1 effector nuclease and the OsWx-T1 recognition sequence. Replacing OsWx-T1 in KF278 with OsWx-T2 results in the bacterial expression vector KF279, which expresses the LbCpf1 effector nuclease and the OsWx-T2 recognition sequence.

[0138] Replacing the LbCpf1 gene in KF278 with the Cas-kf8 or Cas-kf9 (SEQ ID NO 2, 5) gene using conventional cloning methods yields bacterial expression vectors expressing Cas-kf8 or Cas-kf9 effector nucleases and the OsWx-T1 recognition sequence, respectively. Similarly, replacing the LbCpf1 gene in KF279 yields bacterial expression vectors expressing Cas-kf8 or Cas-kf9 effector nucleases and the OsWx-T2 recognition sequence, respectively. All of these bacterial expression vectors expressing effector nucleases and target recognition sequences were transformed into stable bacterial competent cells for later use.

[0139] 3. Bacterial plasmid interference experiment: Stable single colonies containing effector nuclease expression vectors such as KF278 and KF279 were selected and cultured to prepare competent cells. These cells were aliquoted and stored at -80°C. Transformation with pUC19 blank vector was performed to test the transformation efficiency, which reached 1 x 10⁻⁶. 7 Plasmid interference experiments can only be performed after obtaining cfu / μg DNA (colony-forming units per microgram of DNA). Using 10 nanograms of KF280, KF281, KF282, or KF283 plasmid DNA containing the target sequence, heat-shock transform competent cells containing appropriate effector nucleases and expression plasmids with the target recognition sequence. After recovery culture in liquid LB medium for 1 hour, the DNA is appropriately diluted and plated onto LB agar plates containing 50 μg / mL kanamycin. After overnight incubation at 37°C, colony counts are recorded. Each bacterial transformation experiment is repeated three times, and the average value is used to calculate the transformation efficiency.

[0140] Plasmid interference experiment of control effector nuclease LbCpf1: First, blank stable bacterial competent cells were transformed with plasmid DNA from KF280, KF281, KF282, and KF283, respectively. The number of colonies formed was compared, and the concentration of each plasmid DNA was adjusted to achieve the same transformation efficiency. Then, stable competent cells containing plasmid KF278 were transformed with plasmid DNA from KF280 and KF281. The number of colonies in the KF280 experiment and the KF281 control was counted. If the LbCpf1 effector nuclease expressed by KF278 and the OsWx-T1 target recognition sequence can recognize and cleave the OsWx-T1 target sequence in KF280, plasmid interference occurs, and fewer colonies will be found in the KF280 culture dish. Because the incomplete OsWx-T1 target sequence of the KF281 plasmid DNA lacks the 5' PAM sequence TTTG, it cannot be recognized and cleaved by the LbCpf1 effector nuclease expressed by KF278, thus avoiding plasmid interference. KF281 culture dishes will then show a normal number of colonies comparable to those of transformed blank Stable competent cells. This demonstrates the ability of LbCpf1 to cleave the OsWx-T1 site in bacteria. Transforming Stable competent cells containing the KF279 plasmid with KF282 and KF283 plasmid DNA, respectively, allows us to determine the ability of LbCpf1 to cleave the OsWx-T2 site in bacteria. Experiments show that LbCpf1 can effectively cleave both the OsWx-T1 and OsWx-T2 sites in bacteria.

[0141] Plasmid interference experiment of novel effector nucleases: Stable competent cells containing expression vectors of novel effector nucleases Cas-kf8 or Cas-kf9 and the OsWx-T1 target recognition sequence were transformed using the same KF280 and KF281 plasmid DNA as described above. The colony counts in the KF280 experiment and the KF281 control were counted to determine the cleavage capacity of Cas-kf8 or Cas-kf9 at the OsWx-T1 site in bacteria. Similarly, the cleavage capacity of Cas-kf8 or Cas-kf9 at the OsWx-T2 site in bacteria was determined by transforming Stable competent cells containing expression vectors of novel effector nucleases Cas-kf8 or Cas-kf9 and the OsWx-T2 target recognition sequence using KF282 and KF283 plasmid DNA. The experiments demonstrated that Cas-kf8 and Cas-kf9 have cleavage capacity at both the OsWx-T1 and OsWx-T2 sites in bacteria.

[0142] III. Establishing a plant gene editing system using novel effector nucleases

[0143] Beidahuang Kenfeng Seed Industry has independently developed a rice-specific gene-editing Agrobacterium-mediated transformation DNA vector. Using the pCambia3301 vector as its backbone, it contains three gene expression units: ZmU6 pro:gRNA:AtU6-26 term, ZmUBI1 pro:SpCas9:PsE9 term, and 35S pro:HYG:35S term. A BsaI restriction site is left between ZmU6 pro and gRNA. After linearization, one or more gRNA sequences can be inserted using DNA ligase or Gibson cloning methods to form a basic gene-editing vector targeting any site. Specific sequence information can be found in patent CN113604501B, "Gene Editing Method and Application of Aroma Control Gene in Improved Indica Rice Lines," specifically the KF80 transformation vector. All gene-editing transformation vectors described in this invention are modified and constructed based on this disclosed basic gene-editing vector.

[0144] Using the aforementioned KF80 transformation vector as the basic backbone, the SpCas9 gene was replaced by the LbCpf1 gene (SEQ ID NO 11), and the gRNA was replaced by the artificially synthesized tRNA-crRNA-OsWx-T1-tRNA-crRNA-OsWx-T2 (SEQ ID NO 13) fragment. The 5' and 3' ends of this fragment are homologous to ZmU6 pro and AtU6-26 term, respectively, and are used for Gibson cloning. Finally, the gene editing vector KF315 based on the LbCpf1 effector nuclease was formed, which consists of three gene expression units: ZmU6 pro:tRNA-crRNA-OsWx-T1-tRNA-crRNA-OsWx-T2:AtU6-26 term, ZmUBI1 pro:LbCpf1:PsE9 term, and 35S pro:HYG:35S term, which can edit the OsWx-T1 and OsWx-T2 sites in rice. Subsequently, the LbCpf1 gene in KF315 was replaced with the Cas-kf8 and Cas-kf9 (SEQ ID NO 2, 5) genes, respectively, to form gene editing vectors KF316 and KF317 based on the Cas-kf8 and Cas-kf9 effector nucleases. These vectors can also edit the OsWx-T1 and OsWx-T2 sites in rice. The composition and sequence of the KF316 vector are shown in Figure 3 and SEQ ID NO 14, respectively. Except for the effector nuclease, the composition and sequence of KF315 and KF317 are completely identical to those of KF316, and therefore will not be listed further. The KF315, KF316, and KF317 vectors were transformed into Agrobacterium EHA105 and stored at -80℃ for later use.

[0145] IV. Novel effector nuclease-mediated plant gene editing

[0146] Gene editing transformation in rice: Gene editing vectors based on novel effector nucleases must be transformed into plant cells to verify their function. Using Agrobacterium-mediated transformation, healthy mature seeds of the rice variety Kendao 32 were used as explants to induce callus tissue. Gene editing vectors KF316 and KF317 were transformed into cells, respectively. Resistance transgenic events were screened, and the DNA sequences of the target genes OsWx-T1 and OsWx-T2 sites were examined for specific changes to evaluate the gene editing effects of each effector nuclease in plants. The formulations of various basic culture media used in the genetic transformation experiments are listed in Table 1, and the specific operating procedures are as follows:

[0147] Induction of rice callus: For each transformation experiment, select approximately 200 mature, plump, healthy, and clean rice seeds, remove the husks, and place them in a 100 ml sterile glass bottle. Wash the seeds three times with 75% ethanol, add an equal volume of sterile water and 10% sodium hypochlorite (NaClO), and two drops of Tween 20. Sterilize by shaking for 20 minutes, and wash with sterile water 3-5 times until the foam is removed. Blot the seeds dry with sterile filter paper, place them flat on the induction medium, and incubate in the dark at 33℃ for 3-4 weeks until callus tissue grows from the embryo.

[0148] Subculture of rice callus: Select vigorous healthy callus tissue, place it on a new induction medium and subculture for 1-2 weeks to proliferate and maintain the vitality of the callus tissue.

[0149] Preparation of Agrobacterium: Remove the Agrobacterium glycerol storage tube carrying the gene-editing vector from the -80℃ freezer. Streak a small amount of Agrobacterium onto a solid YEP medium plate containing 50 mg / L kanamycin and incubate in the dark at 28℃ for 2-3 days. Pick a single colony and inoculate it into 5 mL of liquid YEP medium, incubating overnight in the dark at 28℃. Extract plasmid DNA from a small amount of bacterial culture for vector-specific PCR to confirm that the strain carries the correct gene-editing vector. Simultaneously, plate a portion of the bacterial culture onto a YEP medium plate for observation. The Agrobacterium used for infection must be free from contamination. The entire plate should be smooth and free of granular or other colored fungi or molds to ensure the purity of the obtained Agrobacterium. Strict aseptic technique must be maintained to ensure no contamination in subsequent work. After confirming that the Agrobacterium plate is free of contamination, the bacterial cells are directly scraped into a 100 ml Erlenmeyer flask containing 25 ml of suspension infection medium. The flask is then incubated for 2-3 hours on a rotating shaker at 100-120 rpm at a constant temperature of 25°C. The OD600 value is measured, and the Agrobacterium OD600 value is adjusted to 0.1-0.2 with suspension infection medium before being used for callus infection.

[0150] Infection of callus tissue: Select small granular, healthy callus tissue into a 250 ml sterile glass bottle, add Agrobacterium and shake to infect for 3-5 minutes, then use sterile filter paper to absorb the bacterial solution on the surface of the callus tissue, place it on sterile filter paper that has been placed on the surface of the co-culture medium (the filter paper prevents the growth of Agrobacterium), and co-culture at 28°C in the dark for 3 days.

[0151] Screening and identification of resistant callus: Infected rice callus was transferred to a sterile 250 mL glass bottle and washed several times with sterile water until the water was clear. Finally, it was soaked and washed for 0.5 hours on a shaker at 100 rpm with carbenicillin at a final concentration of approximately 250 mg / L. The callus was then transferred to a blank culture dish, and the surface moisture was blotted dry with sterile filter paper. After air-drying in a sterile laminar flow hood for 2.5 hours, the callus was aliquoted and placed on screening medium. The medium was then incubated in the dark at 33°C for 3-4 weeks. The screening medium contained 400 mg / L carbenicillin to inhibit Agrobacterium growth and 50 mg / L hygromycin to screen for transformed cells. Non-transformed cells stopped growing and gradually died on the screening medium. After 3-4 weeks of culture, successfully transformed cells grew resistant callus. After the resistant callus continued to grow, a suitable amount was sampled, genomic DNA was extracted, and PCR identification was performed using primers specific to the transformation vector. Each PCR-positive callus was considered an independent transformation event.

[0152] Table 1. Composition and preparation method of basic culture medium used in the Agrobacterium tumefaciens genetic transformation experiment of japonica rice.

[0153] 2. Analysis of gene editing effects:

[0154] Resistance transformation events need to be detected using molecular biology methods to confirm the success of genetic transformation. A commonly used, simple, and reliable detection method is PCR amplification. Rice resistance callus tissue was sampled, and genomic DNA was extracted using a rapid DNA extraction method based on SDS (Sodium dodecyl sulfate). First, vector-specific PCR detection of the HygR gene was performed to identify the transformation event. The PCR reaction system used was the Quick Taq HSDyeMix (DTM-101) kit (TOYOBO Life Science). A 20 μL PCR reaction system contained 10 μL 2x Quick Taq HSDyeMix, 1.0 μL 10 pmol / μL primer Hyg-F1 (SEQ ID NO 15) and 1.0 μL 10 pmol / μL primer Hyg-R1 (SEQ ID NO 16), 6.0 μL sterile water, and finally 2.0 μL 50 ng / μL of the sample's genomic DNA. The PCR reaction conditions were as follows: denaturation at 95°C for 5 minutes, followed by denaturation at 95°C for 30 seconds, annealing at 60°C for 1 minute, extension at 72°C for 1 minute, for a total of 30 cycles, and a final extension at 72°C for 7 minutes, followed by holding at 4°C. The PCR amplification products were separated by 1% agarose gel electrophoresis. Samples that successfully amplified a 402bp long fragment specific to the HygR gene (SEQ ID NO 17) were identified as transgenic positive.

[0155] Using a similar PCR amplification method and specific primers OsWx-F5 and OsWx-R1 (SEQ ID NO 18, 19) for OsWx(LOC_Os06g04200), a 916 bp target fragment (SEQ ID NO 20) near the OsWx-T1 and OsWx-T2 gene editing sites can be amplified from HygR-positive samples. After purification, the fragment is directly sequenced using the same upstream or downstream primers and compared with the same gene fragment sequence from the wild-type Kendao 32. If the DNA fragment sequence of the transgenic event is no different from the wild-type sequence, gene editing has not occurred; if a peak appears near the expected effector nuclease cleavage site, it indicates that the sequenced PCR fragment is a mixed template containing both edited and unedited DNA fragments, indicating heterozygous gene editing; if there is a definite sequence difference from the wild-type, it indicates homozygous gene editing. Gene editing experiments have demonstrated for the first time that both Cas-kf8 and Cas-kf9 effector nucleases have gene editing functions, but their overall gene editing efficiency differs significantly. The editing efficiency at different sites of OsWx-T1 and OsWx-T2 also differs. The gene editing efficiency results obtained by direct sequencing analysis of the target gene PCR fragment are shown in Table 2.

[0156] The TOPO PCR fragment representing the target gene of the editing event was cloned into the pCR2.1 vector (Thermo Fisher Scientific). Four clones from each sample were sequenced to accurately analyze the specific editing changes in the DNA sequence at the OsWx-T1 and OsWx-T2 sites. Different clones from the same callus editing event often exhibit multiple editing sequence changes. Some clones edited only at the OsWx-T1 or OsWx-T2 sites, while others edited at both sites simultaneously. Therefore, the number of editing types at each site does not correspond to the total number of editing events. Cloning and sequencing further confirmed that two novel effector nucleases, Cas-kf8 and Cas-kf9, both possess significant gene editing activity, but there are also significant differences between them. The editing efficiency of Cas-kf9 is approximately three times that of Cas-kf8. Cas-kf9 can edit both the OsWx-T1 and OsWx-T2 sites and exhibits multiple dual-site editing capabilities. The Cas-kf8 has relatively low editing efficiency. Although it could not edit OsWx-T1, it could edit OsWx-T2 and had a variety of editing types (Table 2).

[0157] Table 2 Gene editing efficiency of effector nucleases

[0158] The above target gene sequence analysis results revealed that the gene editing mediated by Cas-kf8 and Cas-kf9 effector nucleases were mostly small fragment deletions, all occurring downstream of the expected target gene site, consistent with other known Cas12a effector nucleases such as LbCpf1 and FnCpf1. Detailed sequencing analysis results of each effector nuclease editing event are summarized in Table 3. In the table, R234-xx events represent callus obtained by transforming the KF316 vector containing the Cas-kf8 effector nuclease, and R243-xx events represent callus obtained by transforming the KF317 vector containing the Cas-kf9 effector nuclease. For ease of description, the first T in the TTTV sequence of the 5' PAM at each target site is marked as position 1. The editing sequences of some Cas-kf8 at the OsWx-T2 site and some Cas-kf9 at the OsWx-T1 and OsWx-T2 sites are shown in Figures 4 and 5, respectively.

[0159] Table 3. Sequence changes at the editing sites in OsWx gene editing events.

[0160] This invention yielded two novel Cas12a-type effector nuclease proteins, Cas-kf8 and Cas-kf9. Through codon optimization and plant gene editing transformation experiments, these proteins demonstrated the ability to recognize and cleave target sites with TTTV sequences at the 5' end, resulting in novel effector nucleases suitable for plant gene editing. Furthermore, this invention provides a highly efficient plant gene editing vector for targeted genome modification in plants.

[0161] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Based on all the teachings that have been published, various modifications and changes can be made to the details. Any other changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention should be considered equivalent substitutions and are included within the protection scope of the present invention.

[0162] DNA sequence description in this invention

[0163] SEQ ID NO 1: Cas-kf8 protein 1368aa

[0164] SEQ ID NO 2:Cas-kf8 coding DNA 4104bp

[0165] SEQ ID NO 3:Cas-kf8 crRNA repeat DNA 36bp

[0166] SEQ ID NO 4:Cas-kf9 protein 1287aa

[0167] SEQ ID NO 5:Cas-kf9 coding DNA 3864bp

[0168] SEQ ID NO 6:Cas-kf9 crRNA repeat DNA 36bp

[0169] SEQ ID NO 7:synthesized OsWx-T1 fragment 29bp

[0170] SEQ ID NO 8:synthesized PAM-less OsWx-T1 fragment 25bp

[0171] SEQ ID NO 9:synthesized OsWx-T2 fragment 29bp

[0172] SEQ ID NO 10:synthesized PAM-less OsWx-T2 fragment 25bp

[0173] SEQ ID NO 11:Lachnospiraceae bacterium Cas12a(LbCpf1)CDS 3843bp

[0174] SEQ ID NO 12:Lachnospiraceae bacterium Cas12a(LbCpf1)crRNA repeat 36 bp

[0175] SEQ ID NO 13:synthesized tRNA-crRNA-OsWx-T1-tRNA-crRNA-OsWx-T2 DNA fragment 286bp

[0176] SEQ ID NO 14:KF316 DNA 10461 bp

[0177] SEQ ID NO 15:Hyg-F1 oligo 22 bp

[0178] SEQ ID NO 16:Hyg-R1 oligo 22 bp

[0179] SEQ ID NO 17:HygR gene PCR fragment 402 bp

[0180] SEQ ID NO 18:OsWx-F5 oligo 22 bp

[0181] SEQ ID NO 19:OsWx-R1 oligo 22 bp

[0182] SEQ ID NO 20:Oryza sativa Kendao32 Waxy gene OsWx(LOC_Os06g04200)PCR fragment 916bp

Claims

1. A Cas protein, characterized in that, The Cas protein is Cas-kf8 or Cas-kf9, the amino acid sequence of Cas-kf8 is shown in SEQ ID No 1, and the amino acid sequence of Cas-kf9 is shown in SEQ ID No 4.

2. A fusion protein comprising the Cas protein of claim 1 and other modified portions.

3. An isolated polynucleotide, characterized in that, The polynucleotide is a polynucleotide sequence encoding the Cas protein of claim 1, or a polynucleotide sequence encoding the fusion protein of claim 2.

4. The polynucleotide according to claim 3, characterized in that, The sequence is as shown in SEQ ID No 2 or 5.

5. A carrier, characterized in that, The vector comprises the polynucleotide of claim 3 or 4 and a regulatory element operatively linked thereto.

6. A CRISPR-Cas system, characterized in that, The system includes the Cas protein of claim 1 and at least one gRNA, the gRNA comprising a repeating sequence for binding the Cas protein and a spacer sequence for targeting a target sequence.

7. The CRISPR-Cas system according to claim 6, characterized in that, The repeating sequence is as shown in SEQ ID No 3 or 6.

8. A carrier system, characterized in that, The vector system includes one or more vectors, which include: a gRNA expression unit and the Cas protein expression unit of claim 1, wherein the gRNA includes a crRNA sequence capable of targeting a target sequence, and the gRNA expression unit and the Cas protein expression unit are located on the same or different vectors of the system.

9. The carrier system according to claim 8, characterized in that, The gRNA expression unit is ZmU6 pro:gRNA:AtU6-26 term.

10. The carrier system according to claim 9, characterized in that, The ZmU6 pro and gRNA have a BsaI restriction site, which allows one or more gRNA sequences to be inserted to form a gene editing vector system that targets any site of the target sequence.

11. The carrier system according to claim 10, characterized in that, The Cas protein expression unit is ZmUBI1 pro:Cas:PsE9 term.

12. The carrier system according to claim 8, characterized in that, The vector system also includes an expression unit for resistance selection genes.

13. The carrier system according to claim 12, characterized in that, The resistance selection gene expression unit is 35Spro:HYG:35S term.

14. A composition, characterized in that, The composition comprises: (i) a protein component selected from the Cas protein of claim 1 or the fusion protein of claim 2; and (ii) a nucleic acid component selected from gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA of the gRNA, or a nucleic acid encoding the precursor RNA of the gRNA; wherein the gRNA comprises a repeating sequence that binds to the Cas protein and a spacer sequence for targeting a target sequence.

15. An engineered host cell, characterized in that, The host cell comprises the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3 or 4, or the vector of claim 5, or the CRISPR-Cas system of claim 6 or 7, or the vector system of any one of claims 8-13, or the composition of claim 14.

16. The use of the Cas protein of claim 1, or the fusion protein of claim 2, or the polynucleotide of claim 3 or 4, or the vector of claim 5, or the CRISPR-Cas system of claim 6 or 7, or the vector system of any one of claims 8-13, or the composition of claim 14, or the host cell of claim 15, in any of the following: (1) Applications in gene editing or gene targeting for the purpose of diagnosis and treatment of non-human diseases; (2) Applications in the preparation of gene editing or gene targeting reagents or kits; (3) Applications of gene cutting for purposes other than disease diagnosis and treatment; (4) Applications in the preparation of gene cutting reagents or kits; (5) Application in non-specific cleavage of side branch nucleic acids; (6) Application in non-specific degradation of side-branched nucleic acids; (7) Application in plant gene editing breeding.

17. A basic editing medium, characterized in that, Using pCambia3301 as the vector backbone, three expression units were set starting from the right boundary of T-DNA: ZmU6 pro:gRNA:AtU6-26 term, ZmUBI1 pro:LbCpf1:PsE9 term, and 35S pro:HYG:35S term. The Cas-kf8 basic editing vector or the Cas-kf9 basic editing vector was obtained by replacing LbCpf1 with the Cas protein described in claim 1.

18. A target sequence editing vector, characterized in that, A recognition sequence is designed for different sites of the target sequence to be edited, and a gRNA sequence is formed with an appropriate crRNA repeat sequence. One or more gRNA sequences are cross-linked using tRNA technology. After linearizing the basic editing vector described in claim 17, one or more gRNA sequences are inserted using DNA ligase or Gibson cloning method to form a target sequence editing vector.

19. A method for genetic transformation of plants, characterized in that, The target sequence editing vector of claim 18 is introduced into recipient plant cells, and then screened to obtain heritable, non-transgenic, and stably inherited progeny plants. The preferred process includes the following steps: (1) Genetic transformation of recipient plant cells to obtain regenerated plants; (2) Detection of whether gene editing of regenerated plants was successful; (3) Screening to obtain offspring plants with homozygous editing of the target sequence trait, non-transgenic, and stably inherited.

20. A method for editing a target sequence or targeting a target sequence, characterized in that, The method is not for disease diagnosis and treatment purposes, and the object of the method does not include humans. The method includes contacting the target sequence with the CRISPR-Cas system of claim 6 or 7, or the vector system of claim 8, or the composition of claim 14, or the host cell of claim 15.

21. A method for cutting a target sequence, characterized in that, The method is not for disease diagnosis and treatment purposes, and the object of the method does not include humans. The method includes contacting the target sequence with the CRISPR-Cas system of claim 6 or 7, or the vector system of claim 8, or the composition of claim 14, or the host cell of claim 15.

22. A kit for gene editing or gene targeting, characterized in that, The kit comprises the CRISPR-Cas system of claim 6 or 7, or the vector system of any one of claims 8-13, or the composition of claim 14, or the host cell of claim 15.

23. A kit for detecting a target sequence in a sample, characterized in that, The kit contains: (a) the Cas protein of claim 1, or the nucleic acid encoding the Cas protein; and / or (b) a gRNA, or a nucleic acid encoding the gRNA, or a precursor RNA containing the gRNA, or a nucleic acid encoding the precursor RNA; the gRNA includes a crRNA sequence capable of targeting a target sequence; and / or (c) A single-stranded nucleic acid detector, wherein the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.

24. A method for detecting a target sequence in a sample, characterized in that, The method is for non-disease diagnosis and treatment purposes. The method includes contacting a sample with the Cas protein, gRNA, and single-stranded nucleic acid detector as described in claim 23, detecting a detectable signal generated by the Cas protein cleaving the single-stranded nucleic acid detector, thereby detecting a target sequence; the single-stranded nucleic acid detector is a single-stranded nucleic acid that does not hybridize with the gRNA.

Citation Information

Patent Citations

  • Plant guide template in-situ synthesis gene editing method and application

    CN111748578A

  • NOVEL Cas ENZYMES AND SYSTEMS AND USES

    CN116334037A

  • Cas13 protein and application thereof

    CN117230043A

  • Adenosine deaminase variants and uses thereof

    CN117561074A

  • CRISPR-Cas13 system and application thereof

    CN118159650A