VARIANTS OF CPF1 (Cas12a) HAVING IMPROVED ACTIVITY

By modifying the specific amino acid sequence of the Cas12a protein and binding to the heterologous domain, the problem of insufficient frequency and accuracy of site-directed nuclease editing in the prior art is solved, and efficient SDN-1 editing in the plant genome is achieved, giving plants the required characteristics.

CN120500536APending Publication Date: 2025-08-15SYNGENTA CROP PROTECITON AG +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202380090794.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-01-27
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing site-directed nuclease editing technology has limitations in improving the frequency and accuracy of nucleic acid editing, especially in the plant genome, which is difficult to achieve efficient SDN-1 editing.

Method used

By modifying the specific amino acid sequence of the Cas12a protein, artificially induced mutations, such as the substitution of cysteine ​​to serine, enhance its nuclease activity, and combine heterologous domains and nucleic acid and DNA constructs, it is applied to plant cells for editing.

Benefits of technology

The frequency and accuracy of site-directed nuclease editing is improved, and more efficient SDN-1 editing is achieved in the plant genome, which can introduce specific genomic changes to impart the required characteristics to the plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005487135320000151
    Figure BDA0005487135320000151
  • Figure BDA0005487135320000161
    Figure BDA0005487135320000161
  • Figure BDA0005487135320000162
    Figure BDA0005487135320000162
Patent Text Reader

Abstract

Provided herein are variant Cas12a proteins comprising at least one artificial evoked mutation. Also provided are fusion proteins comprising these variant Cas12a proteins and one or more heterologous domains. Related nucleic acids, DNA constructs, vectors, cells, and methods of editing nucleic acids using these variant Cas12a proteins and / or fusion proteins are also provided. Using the provided proteins can increase the frequency of desired nucleic acid editing (e.g., SDN-1 editing in the plant genome).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods for increasing site-directed nuclease editing. Reference to a sequence listing submitted as an XML file

[0002] Attached to this application is a sequence listing named 82447-SL.xml, which was created on January 19, 2023 and is approximately 149 kilobytes in size. This sequence listing is incorporated herein by reference in its entirety. Background Art

[0003] Site-directed nucleases (SDNs) (such as zinc finger nucleases, transcription activator-like effector nucleases, CRISPR-associated nucleases) are becoming more and more popular in the gene editing space. These SDNs act as endonucleases and typically produce double-strand breaks (DSBs) in specific DNA sequences, thereby activating the intrinsic repair mechanisms (such as homologous recombination) of the cell. During the repair process, it is possible to achieve site-directed modification of the specific DNA sequence. CRISPR (clustered regularly interspaced short palindromic repeats) / Cas (CRISPR-associated) systems have evolved into adaptive immune systems in bacteria and archaea to defend against viral attacks. In recent years, the CRISPR / Cas system has attracted special attention as a tool for genome editing. The CRISPR / Cas system that produces site-specific double-strand breaks (DSBs) can be used, for example, to edit the DNA in eukaryotic cells by producing deletions, insertions, and / or changes in the nucleotide sequence. Summary of the Invention

[0004] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help limit the scope of the claimed subject matter.

[0005] In one aspect, there is provided a Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO:1 and an artificially induced mutation at position C965. In some embodiments, the artificially induced mutation is a substitution of cysteine to serine. In some embodiments, the Cas12a protein is further included in an artificially induced mutation at position D156. In some embodiments, the artificially induced mutation at position D156 is a substitution of aspartic acid to arginine. In some embodiments, the sequence of the Cas12a protein comprises any one of SEQ ID NO:5-11.

[0006] On the other hand, there is provided a Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO:2 and an artificially induced mutation at position C70, C1116 and / or C1190. In some embodiments, the artificially induced mutation is a substitution of cysteine to serine. In some embodiments, the Cas12a protein is further included in an artificially induced mutation at position E184. In some embodiments, the artificially induced mutation at position E184 is a substitution of glutamic acid to arginine.

[0007] On the other hand, there is provided Cas12a protein, which comprises a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO:3 and an artificially induced mutation at position C334, C379 and / or C674. In some embodiments, the artificially induced mutation is a substitution of cysteine to serine. In some embodiments, the Cas12a protein is further included in an artificially induced mutation at position E174. In some embodiments, the artificially induced mutation at position E174 is a substitution of glutamic acid to arginine.

[0008] On the other hand, there is provided a Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO:4 and artificially induced mutations at positions C270, C583, C1068, C1099 and / or C1149. In some embodiments, the artificially induced mutation is a substitution of cysteine to serine. In some embodiments, the Cas12a protein is further included in an artificially induced mutation at position D172. In some embodiments, the artificially induced mutation at position D172 is a substitution of aspartic acid to arginine. In some embodiments, the sequence of the Cas12a protein comprises any one of SEQ ID NO:12-19.

[0009] In some embodiments of any of the above Cas12a proteins, the Cas12a protein is a catalytically inactive Cas12a (dCas12a) protein of the nickase Cas12a (nCas12a) protein.

[0010] In some embodiments of any of the above Cas12a proteins, the Cas12a protein further comprises a nuclear localization signal.

[0011] In another aspect, a fusion protein is provided, comprising any one of the above-described Cas12a proteins and a heterologous domain.

[0012] In some embodiments, the heterologous domain is a deaminase domain, a transcription factor domain, a nuclease domain, a reverse transcriptase domain, a transposase domain, an integrase domain, a uracil DNA glycosylase inhibitor domain, a recombinase domain, a nickase domain, a methyltransferase domain, a methylase domain, an acetylase domain, an acetyltransferase domain, a transcription activator domain, or a transcription repressor domain.

[0013] In some embodiments of the fusion protein, the Cas12a protein is linked to the heterologous domain via a linker sequence.

[0014] On the other hand, a nucleic acid is provided, which encodes any one of the above-mentioned Cas12a proteins or any one of the fusion proteins. In some embodiments, the nucleic acid sequence is any one of SEQ ID NO:20-34.

[0015] In another aspect, a DNA construct is provided, comprising a promoter operably linked to a nucleic acid encoding any one of the above-described Cas12a proteins or any one of the fusion proteins.

[0016] In another aspect, a vector is provided, comprising the above-described nucleic acid or DNA construct.

[0017] In another aspect, a cell is provided comprising the nucleic acid, DNA construct or vector. In certain embodiments, the cell is a plant cell. In certain embodiments, the cell is a corn plant cell, a wheat plant cell, a rice plant cell, a soybean plant cell, a sunflower plant cell or a tomato plant cell.

[0018] In another aspect, a method for editing a nucleic acid is provided, the method comprising contacting the nucleic acid with: (i) any one of the above-described Cas12a proteins or any one of the above-described fusion proteins, and (ii) a guide RNA having a region complementary to a selected portion of the nucleic acid, thereby editing the nucleic acid. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] This application includes the following figures. These figures are intended to illustrate certain embodiments and / or features of the compositions and methods and to supplement any one or more of the descriptions of the compositions and methods. These figures do not limit the scope of the compositions and methods unless the written description clearly indicates otherwise.

[0020] Figure 1It is shown that the cysteine residues in LbCas12a may potentially form intermolecular or intramolecular interactions. Left: PyMOL surface model of the LbCas12a-crRNA-DNA ternary complex (PDB entry 5XUS). The highlighted area indicated by the arrow is the thiol group of C965 and C1090 that may be exposed on the surface. Right: Four cysteine residues (C10, C805, C912, C965) dispersed in the linear amino acid sequence form a cluster inside the 3D structure of LbCas12a.

[0021] Figure 2 Two cysteine residues in FnCas12a selected for substitution according to aspects of the present disclosure are shown. The PyMOL stick model of C1190 and C1116 shows that the thiol groups (black) are close to each other in the FnCas12a 3D structure (PDB entry 5NFV) and can potentially form an intramolecular disulfide bond between the two. DETAILED DESCRIPTION

[0022] The following description lists various aspects and embodiments of the compositions and methods of the present invention. The specific embodiments are not intended to limit the scope of the compositions and methods. Rather, the embodiments merely provide non-limiting examples of various compositions and methods that are at least within the scope of the disclosed compositions and methods. The description should be read from the perspective of one of ordinary skill in the art; therefore, it does not necessarily include information that would be familiar to a skilled artisan. I. Terminology

[0023] Unless otherwise defined below, all technical and scientific terms used herein are intended to have the same meaning as commonly understood by those of ordinary skill in the art. References to the techniques employed herein are intended to refer to techniques commonly understood in the art, including modifications of those techniques and / or equivalent technical alternatives that are clear to those of ordinary skill in the art. Although it is believed that the following terms may be well understood by those of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the subject matter disclosed herein.

[0024] As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an enzyme" optionally includes combinations of two or more such molecules, etc.

[0025] As used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0026] As used herein, the term "about" refers to the usual error range for the corresponding value that is readily known to those skilled in the art, for example, ±20%, ±10% or ±5% within the intended meaning of the recited value.

[0027] As used herein, the term "comprising" or "comprises" is open ended. When used in conjunction with a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or amino acid sequence) that includes the subject sequence as a portion or as its entire sequence.

[0028] As used herein, the transition phrase "consisting essentially of" means that the scope of a claim is interpreted to encompass the specified materials or steps recited in the claim, as well as those materials or steps that do not materially affect one or more of the basic and novel characteristics of the claimed subject matter. Therefore, when used in the claims of the present disclosure, the term "consisting essentially of" is not intended to be interpreted as equivalent to "comprising."

[0029] The term "plurality" refers to more than one entity. Thus, a "majority of individuals" refers to at least two individuals. In some embodiments, the term "majority" refers to more than half of a whole. For example, in some embodiments, a "majority of a group" refers to more than half of the members of that group.

[0030] As used herein, term " plant " refers to any plant, particularly seed plant in any developmental stage. As used herein, term " plant cell " is the structure and physiological unit of plant, comprises protoplast and cell wall.Plant cell can be in the form of isolated single cell or cultured cell, or as a part of higher organization unit (such as, plant tissue, plant organ or whole plant).Plant cell can derive from angiosperms or gymnosperms or be a part of them.Plant cell can be monocotyledonous plant cell (for example, corn cell, rice cell, sorghum cell, sugarcane cell, barley cell, wheat cell, oat cell, turf grass cell or ornamental grass cell) or dicotyledonous plant cell (for example, tobacco cell, pepper cell, eggplant cell, sunflower cell, cruciferous plant cell, flax cell, potato cell, cotton cell, soybean cell, sugar beet cell or oilseed rape cell). As used herein, the term "plant cell culture" means a culture of plant units (such as, for example, protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at different developmental stages). As used herein, the term "plant tissue" refers to a group of plant cells organized into structural and functional units. Includes any plant tissue in a plant or in culture. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue cultures, and any plant cell groups organized into structural and / or functional units. The use of this term in conjunction with (or when it does not exist) any specific type of plant tissue listed above or the use of other methods encompassed by this definition is not intended to exclude any other type of plant tissue. As used herein, the term "plant part" refers to a part of a plant, including single cells and cell tissues (such as complete plant cells in a plant), cell clumps, and tissue cultures that can regenerate plants. Examples of plant parts include, but are not limited to, single cells and tissues from: pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral parts, fruits, stems, buds, cuttings, and seeds; as well as pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral parts, fruits, stems, buds, cuttings, scions, rhizomes, seeds, protoplasts, callus, and the like.

[0031] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, these terms encompass amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds.

[0032] The terms "nucleic acid" and "polynucleotide" are used interchangeably and, as used herein, refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in single-stranded or double-stranded form, and polymers thereof, and refer to both the sense and antisense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers thereof. In higher plants, DNA is the genetic material, while RNA is involved in the transfer of information contained within DNA into proteins. A "genome" is the entirety of the genetic material contained in each cell of an organism. It should be understood that when RNA is described, its corresponding cDNA is also described, with uridine represented as thymidine. In specific embodiments, a nucleotide refers to a ribonucleotide, a deoxynucleotide, or a modified form of any type of nucleotide, and combinations thereof. In addition, the polynucleotides disclosed herein may include either or both naturally occurring nucleotides and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide linkages. Nucleic acid molecules may be chemically or biochemically modified, or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those skilled in the art. Such modification includes, for example, tags, methylation, substitution of one or more naturally occurring nucleotides, internucleotide modifications such as uncharged linkages (for example, methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (for example, phosphorothioates, phosphorodithioates, etc.), side groups (for example, polypeptides), intercalators (for example, acridine, psoralens, etc.), chelating agents, alkylating agents, and modified linkages (for example, α anomeric nucleic acids, etc.). The above terms are also intended to include any topological conformation, including single-stranded, double-stranded, partially duplexed, triplexed, hairpin-shaped, circular, and padlock-shaped conformations. Unless otherwise indicated, reference to a nucleic acid sequence encompasses its complement. Therefore, reference to a nucleic acid molecule with a specific sequence is understood to encompass its complementary chain with its complementary sequence. When nucleotide sequences specifically hybridize in solution, these nucleotide sequences are "complementary" (for example, according to Watson-Crick base pairing principles). The term also includes codon-optimized nucleic acids encoding identical polypeptide sequences. It is also understood that nucleic acids can be unpurified, purified, or attached to, for example, a synthetic material such as a bead or column matrix.

[0033] In the context of nucleic acid sequences, the term "corresponding to" means that when the nucleic acid sequences of certain sequences are aligned with each other, the nucleic acids "corresponding to" certain enumerated positions in the present invention are those aligned with these positions in the reference sequence, but are not necessarily located in these precise numerical positions relative to the specific nucleic acid sequence of the present invention. Optimal alignment of sequences for comparison can be performed by computerized implementations of known algorithms or by visual inspection. Easily available sequence comparison and multiple sequence alignment algorithms are the Basic Local Alignment Search Tool (BLAST) and ClustalW / ClustalW2 / Clustal Omega programs available on the Internet (e.g., the website of EMBL-EBI), respectively. Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG software package available from Accelrys Corporation (San Diego, California, USA). See also Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; Ausubel et al., 1988; and Sambrook and Russell, 2001.

[0034] Unless otherwise specified, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as sequences explicitly specified. In particular, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is replaced by mixed bases and / or deoxyinosine residues. See Batzer et al., Nucleic Acids Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).

[0035] As used in the context of polynucleotide or polypeptide sequences described herein, the terms "identity" or "substantial identity" refer to sequences that have at least 60% sequence identity to a reference sequence. Alternatively, the percent identity can be any integer from 60% to 100%. Exemplary embodiments include: using the programs described herein, preferably BLAST using standard parameters as described below, such as at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% compared to a reference sequence. One skilled in the art will recognize that these values can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame positioning, etc.

[0036] For sequence comparison, typically, a kind of sequence serves as the reference sequence that is compared with the test sequence.When using a sequence comparison algorithm, test sequence and reference sequence are input into a computer, and subsequence coordinates are specified if necessary, and sequence algorithm program parameters are specified. Default program parameters can be used, or alternative parameters can be specified. Then, the sequence comparison algorithm will calculate the sequence identity percentage of the test sequence relative to the reference sequence based on the program parameters.

[0037] As used herein, "comparison window" includes reference to a segment having any one of the number of consecutive positions selected from the group consisting of from 20 to 600, typically about 50 to about 200, more typically about 100 to about 150, wherein a sequence can be compared to a reference sequence having the same number of consecutive positions after the two sequences are optimally aligned. Methods of sequence alignment for comparison are well known in the art. Optimal alignment of sequences for comparison can be achieved by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981); the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970); the search by similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85:2444 (1988); computer implementations of these algorithms (e.g., BLAST); or by manual alignment and visual inspection.

[0038] Suitable algorithms for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (1990) J. Mol. Biol. 215:403-410 and Altschul et al. (1977) Nucleic Acids Res. 25:3389-3402, respectively. Software for performing BLAST analysis is publicly available through the website of the National Center for Biotechnology Information (NCBI). The algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits serve as seeds for initial searches to find longer HSPs containing them. These word hits are then extended in both directions along each sequence until the cumulative alignment score can be increased. For nucleotide sequences, the cumulative score is calculated using the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatched residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. The extension of the word hit in each direction terminates when: the cumulative alignment score drops by the amount X from its maximum achieved value; the cumulative score goes to zero or lower due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses a word length (W) of 28, an expectation (E) of 10, M=1, N=-2, and a two-strand comparison as defaults. For amino acid sequences, the BLASTP program uses a word length (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix as defaults. See Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1989).

[0039] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences. See, for example, Karlin and Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability that a match would occur by chance between two nucleotide or amino acid sequences. For example, if the smallest sum probability in a comparison of a test nucleic acid to a reference nucleic acid is less than about 0.01, more preferably less than about 10, then the match is considered to be a match between two nucleotide or amino acid sequences. -5 and most preferably less than about 10 -20 , the nucleic acid is considered similar to the reference sequence.

[0040] "Recombination" is the exchange of DNA strands to produce new nucleotide sequence arrangements. The term can refer to the homologous recombination process that occurs in the repair of double-stranded DNA breaks, in which polynucleotides are used as templates to repair homologous polynucleotides. The term can also refer to the information exchange between two homologous chromosomes during meiosis. The frequency of double recombination is the product of the frequencies of single recombinants. For example, the frequency of recombinants found in a 10 cM region is 10%, and the frequency of double recombinants is found to be 10% x 10% = 1% (1 centimorgan is defined as 1% of recombinant generations in a test hybrid).

[0041] A "gene" is a defined region within a genome that, in addition to the aforementioned coding nucleic acid sequence, also contains other major regulatory nucleic acid sequences responsible for controlling the expression (i.e., transcription and translation) of the coding portion. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein, including regulatory sequences. A gene may or may not be used to produce a functional protein. In some embodiments, a gene refers only to the coding region. The term "natural gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene that contains: 1) a DNA sequence including regulatory and coding sequences that are not found together in nature, or 2) a sequence encoding a portion of a protein that is not naturally adjacent, or 3) a portion of a promoter that is not naturally adjacent. Thus, a chimeric gene can contain regulatory and coding sequences obtained from different sources, or regulatory and coding sequences obtained from the same source but arranged in a manner different from that found in nature. A gene can be "isolated," which means a nucleic acid molecule that is substantially or essentially free from components normally associated with the nucleic acid molecule in its natural state. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acid molecules.

[0042] "Gene of interest" or "nucleotide sequence of interest" refers to any gene that, when transferred to a plant, confers a desired characteristic on the plant, such as antibiotic resistance, virus resistance, insect resistance, disease resistance, or resistance to other harmful organisms, herbicide tolerance, improved nutritional value, improved performance of an industrial process, or altered reproductive capacity. A "gene of interest" may also be a gene transferred to a plant for the production of a commercially valuable enzyme or metabolite in the plant.

[0043] "Isolated" nucleic acid molecules or nucleotide sequence or "isolated" polypeptide are nucleic acid molecules, nucleotide sequence or polypeptides that exist by artificially breaking away from their natural environment and / or have different, modified, regulated and / or changed functions and therefore are not natural products.Isolated nucleic acid molecules or isolated polypeptides can exist in purified form or can be present in non-natural environments (such as, for example, recombinant host cells).Therefore, for example, with respect to polynucleotides, the meaning of the term isolated is to separate polynucleotides from their naturally occurring chromosomes and / or cells therein. If a kind of polynucleotides is separated from its naturally occurring chromosomes and / or cells therein, and then inserted into genetic background, chromosomes, chromosomal position and / or cells therein that are not naturally present in, then the polynucleotides are also isolated.Recombinant nucleic acid molecules and nucleotide sequence of the present invention can be considered as "isolated" as defined above.

[0044] Thus, an "isolated nucleic acid molecule" or "isolated nucleotide sequence" is a nucleic acid molecule or nucleotide sequence that is not immediately adjacent to a nucleotide sequence (either a sequence at the 5' end or a sequence at the 3' end) in the naturally occurring genome of the organism from which it is derived. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences that are immediately adjacent to the coding sequence. Thus, the term includes, for example, recombinant nucleic acids that are incorporated into a vector, into a self-replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or that exist as a separate molecule (e.g., a cDNA or genomic DNA fragment generated by PCR or restriction endonuclease treatment) independent of other sequences. It also includes recombinant nucleic acids that are part of a hybrid nucleic acid molecule that encodes another polypeptide or peptide sequence. An "isolated nucleic acid molecule" or "isolated nucleotide sequence" may also include a nucleotide sequence that is derived from and inserted into the same natural original cell type, but is present in a non-natural state, for example, in a different copy number, and / or under the control of regulatory sequences that are different from those found in the natural state of the nucleic acid molecule.

[0045] The term "isolated" may further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide or fragment that is substantially free of cellular material, viral material and / or culture medium (e.g., when produced by recombinant DNA techniques), or chemical precursors or other chemicals (e.g., when chemically synthesized). Additionally, an "isolated fragment" is a fragment of a nucleic acid molecule, nucleotide sequence or polypeptide that does not naturally occur as a fragment and would not so exist in nature. "Isolated" does not necessarily mean that the preparation is industrially pure (homogeneous), but that it is sufficiently pure to provide the polypeptide or nucleic acid in a form that can be used for its intended purpose.

[0046] "Homology-dependent repair" or "homology-directed repair" or "HDR" refers to a mechanism for repairing ssDNA and double-stranded DNA (dsDNA) damage in cells. This repair mechanism can be utilized by cells in the presence of an HDR template with a sequence that is significantly identical to the site of damage. The term "perfect HDR" refers to a situation in which the genomic homologous junctions in the replaced allele undergo complete HDR, and "imperfect HDR" refers to a situation in which the genomic homologous junctions in the replaced allele undergo partial or incomplete HDR. A donor DNA molecule with homology to the cleavage target DNA sequence is used as a template for repairing the cleavage target DNA sequence, thereby generating a transfer of genetic information from the donor polynucleotide to the target DNA. In this way, new nucleic acid material can be inserted / copied into the site. In some cases, the target DNA is contacted with a donor molecule (e.g., a donor DNA molecule). In some cases, the donor DNA molecule is introduced into the cell. In some cases, at least one segment of the donor DNA molecule is integrated into the genome of the cell.

[0047] "Microhomology-mediated end joining" or "MMEJ" or "alternative non-homologous end joining" (Alt-NHEJ) refers to a form of repair of double-strand breaks in DNA. This repair mechanism utilizes microhomology sequences to align the broken strands. "Non-homologous end joining" or "NHEJ" refers to a form of repair of double-strand breaks in DNA. Double-strand breaks are repaired by direct ligation of the broken ends to each other. Generally, no new nucleic acid material is inserted into the site, but some nucleic acid material may be lost or added, resulting in small deletions or insertions.

[0048] As used herein, "heterologous" refers to a nucleic acid molecule, nucleotide sequence, polypeptide or amino acid sequence that is not naturally associated with the host cell into which it is introduced, but is derived from another species or from the same species or organism, but has been modified from its original form or the form primarily expressed in the cell, including non-naturally occurring multiple copies of a naturally occurring nucleic acid sequence. Thus, an amino acid sequence derived from an organism or species different from the organism or species to which the cell into which the amino acid sequence is introduced is heterologous with respect to that cell or the progeny of the cell. Additionally, heterologous sequences include sequences that are derived from and inserted into the same natural original cell type, but exist in a non-natural state, for example, exist in different copy numbers, and / or are under the control of regulatory sequences different from those found in the native state of the polypeptide. Sequences can also be heterologous to other sequences related thereto, for example, in nucleic acid constructs, such as expression vectors. As a non-limiting example, a promoter may be present in a nucleic acid construct in combination with one or more regulatory elements and / or coding sequences that are not naturally found in association with that particular promoter, i.e., they are heterologous to the promoter. II. Introduction

[0049] In some aspects, variant Cas12a proteins are provided herein, which have increased site-directed nuclease (SDN) genome editing activity. Site-directed nuclease technology greatly improves the speed and accuracy of genome editing in a variety of organisms (including plants). Generally speaking, the desired result in SDN-mediated genome editing is 1) targeting SDN to crack DNA at a specific genomic site in the host (e.g., plant cells), and 2) using the host's natural repair mechanism to introduce specific genome changes at the cracking site. Changes can include small deletions, substitutions, or additions of some nucleotides. Such targeted editing can produce new and desired features (e.g., enhanced nutrient uptake, reduced allergen production) and / or reduce undesirable features (e.g., herbicide sensitivity). SDN applications are generally divided into three categories: SDN-1, SDN-2, and SDN-3. SDN-1 produces double-strand breaks in the genome without adding exogenous DNA. When such a break is repaired by the host (e.g., via NHEJ), mutations or deletions can be introduced. If these mutations or deletions are in a gene, the gene can be silenced or knocked out. SDN-2 uses template DNA to introduce a predicted modification at the target cleavage site (e.g., via HDR), but does not result in the insertion of recombinant DNA. SDN-3 also uses template DNA to introduce a recombinant or exogenous DNA template (e.g., a transgene) at the target cleavage site.

[0050] Cas12a is a CRISPR-associated (Cas) SDN that functions in the CRISPR (clustered regularly interspaced short palindromic repeats) / Cas system. In bacteria, this system can provide adaptive immunity against foreign DNA (Barrangou, R. et al., “CRISPR provides acquired resistance against viruses in prokaryotes,” Science (2007) 315:1709-1712; Makarova, KS et al., “Evolution and classification of the CRISPR-Cas systems,” Nat Rev Microbiol (2011) 9:467-477; Garneau, JE et al., “The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA,” Nature (2010) 468:67-71; Sapranauskas, R. et al., “The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli,” Nucleic Acids Res (2011) 39:9275-9282). CRISPR / Cas systems (e.g., modified and / or unmodified) can be used as genome engineering tools in a wide variety of organisms, including different mammals, animals, plants, microorganisms, and yeast. The CRISPR / Cas system can include a guide nucleic acid, such as a guide RNA (gRNA), complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid editing. The RNA-guided Cas protein (e.g., a Cas nuclease such as the Cas9 nuclease) can specifically bind to a target polynucleotide (e.g., DNA) in a sequence-dependent manner.Cas proteins can cleave DNA if they have nuclease activity (Gasiunas, G. et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109: E2579-E286; Jinek, M. et al. “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337: 816-821; Sternberg, SH et al. “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9 [CRISPR

[0014] "DNA interrogation with the RNA-guided endonuclease Cas9," Nature (2014) 507:62; Deltcheva, E. et al., "CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III," Nature (201 1) 471:602-607). DNA cleavage (e.g., double-strand breaks) can generate DNA break repair, thereby allowing the introduction of one or more genetic modifications (e.g., nucleic acid edits).

[0051] Cysteine residues are highly reactive residues that are modified post-translationally. Undesirable disulfide bonds and / or modified formation may affect the correct folding and / or positioning and / or enzymatic activity of the protein. There are 8-9 cysteine residues in most Cas12a straight homologues; By contrast, there are only 2 cysteine residues in the Cas9 (SpCas9) from Streptococcus pyogenes. Most of the cysteine residues in Cas12a straight homologues are not conserved. Therefore, those cysteine residues exposed on the surface are more likely to participate in intermolecular disulfide bond formation and / or post-translational modification. Conserved cysteine residues LbCas12a protein, FnCas12a protein, AsCas12a protein and Mb2Cas12a are shown in Table 1-Table 4. Table 1. Cysteine residues in LbCas12a and their aligned residues in pairwise alignments of four orthologs. Table 2. Cysteine residues in FnCas12a and their aligned residues in pairwise alignments of four orthologs. Table 3. Cysteine residues in AsCas12a and their aligned residues in pairwise alignments of four orthologs. Table 4. Cysteine residues in Mb2Cas12a and their aligned residues in pairwise alignments of four orthologs.

[0052] This disclosure is based in part on the discovery of the inventors that mutations of cysteine residues exposed on the surface of Cas12a can improve the bioavailability of Cas12a proteins. Without being bound by any particular theory, such mutations may avoid the above-mentioned undesirable modifications. Provided herein are variant Cas12a proteins comprising at least one artificially induced mutation. Also provided are fusion proteins comprising these variant Cas12a proteins and one or more heterologous domains. Also provided are methods for editing nucleic acids using related nucleic acids, DNA constructs, vectors, cells, and these variant Cas12a proteins and / or fusion proteins. In some embodiments, as demonstrated in the examples herein, the method provided increases the frequency of the desired nucleic acid editing. In some embodiments, the editor is SDN-1 editing. In some embodiments, the frequency of the desired nucleic acid editing is observed to increase at genomic sites that are difficult to edit. III. Variant Cas12a Proteins and Fusion Proteins

[0053] In one aspect, provided herein are variant Cas12a proteins comprising at least one artificially induced mutation with enhanced function (i.e., when compared to unmodified Cas12a proteins). Also provided are fusion proteins comprising the variant Cas12a proteins and at least one heterologous domain. In certain embodiments, the enhanced function of Cas12a is the increased SDN-1 genome editing activity. In certain embodiments, variant Cas12a proteins include the replacement of one or more surface-exposed cysteine residues. In certain embodiments, variant Cas12a proteins include the replacement of cysteine to serine at one or more surface-exposed cysteine residues. In certain embodiments, provided herein are variant Cas12a proteins further include the replacement of aspartic acid residues and / or glutamic acid residues to arginine residues.

[0054] Cas12a (also known as Cpf1) is a Class II, Type V CRISPR / Cas. The variant Cas12a proteins provided herein can be modified forms of Cas12a from any of a variety of bacterial species, including but not limited to Lachnospiraceae bacterium, Acidaminococcus sp., Moraxella bovoculi, Thiomicrospira sp., Moraxella lacunata, Methanomethylophilus alvus, Btyrivibrio sp., or Bacteroidetesoral sp. Unmodified Cas12a protein sequences include Lachnospiraceae Cas12a (LbCas12a; SEQ ID NO: 1), Francisella novicida U112 Cas12a (FnCas12a; SEQ ID NO: 2), Acidaminococcus species Cas12a (AsCas12a; SEQ ID NO: 3), and Moraxella bovis strain 57922 Cas12a (Mb2Cas12a; SEQ ID NO: 4).

[0055] In certain embodiments, variant Cas12a protein is a modified form of LbCas12a. In certain embodiments, Cas12a protein includes and SEQ ID NO:1 amino acid sequence has at least 60% homogeneity (for example, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence and at least one artificially induced mutation. In certain embodiments, artificially induced mutation is the replacement of surface-exposed cysteine residues. Surface-exposed cysteine residues can be identified using methods known in the art (for example, by methods described in Examples herein). In certain embodiments, one or more surface-exposed cysteine residues are replaced by another residue (e.g., serine residue). In certain embodiments, artificially induced mutation is at position C965 (i.e., at SEQ ID NO:1 position 965 cysteine residue). In certain embodiments, artificially induced mutation is the replacement of cysteine residues. In certain embodiments, artificially induced mutation is the replacement of cysteine residues to serine. In certain embodiments, Cas12a protein is further included in artificially induced mutation at position D156 (i.e., at SEQ ID NO:1 position 156 aspartic acid residue), as described in, for example, WO 2018195545 and WO 2017184768 (incorporated herein in their entirety by reference). In certain embodiments, artificially induced mutation is the replacement of aspartic acid residues. In certain embodiments, artificially induced mutation is the replacement of aspartic acid to arginine. In certain embodiments, the sequence of Cas12a protein includes SEQ ID NO:5-11 any one.

[0056] In some embodiments, variant Cas12a protein is a modified form of FnCas12a. In some embodiments, Cas12a protein includes and SEQ ID NO:2 amino acid sequence has at least 60% homogeneity (for example, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence and at least one artificially induced mutation. In some embodiments, artificially induced mutation is the replacement of surface-exposed cysteine residues. In some embodiments, one or more surface-exposed cysteine residues are replaced by another residue (for example, serine residue). In certain embodiments, artificially induced mutations are at positions C70, C1116 and / or C1190. In certain embodiments, artificially induced mutations are replacements of cysteine residues. In certain embodiments, artificially induced mutations are replacements of cysteine to serine. In certain embodiments, Cas12a protein is further included in artificially induced mutations at position E184 (i.e., at SEQ ID NO:2 position 184 glutamic acid residues), as described in, for example, WO 2018195545 (incorporated herein in its entirety by reference). In certain embodiments, artificially induced mutations are replacements of glutamic acid residues. In certain embodiments, artificially induced mutations are replacements of glutamic acid to arginine.

[0057] In some embodiments, variant Cas12a protein is a modified form of AsCas12a. In some embodiments, Cas12a protein includes and SEQ ID NO:3 amino acid sequence has at least 60% homogeneity (for example, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence and at least one artificially induced mutation. In some embodiments, artificially induced mutation is the replacement of surface-exposed cysteine residues. In some embodiments, one or more surface-exposed cysteine residues are replaced by another residue (for example, serine residue). In certain embodiments, artificially induced mutations are at positions C334, C379 and / or C674. In certain embodiments, artificially induced mutations are replacements of cysteine residues. In certain embodiments, artificially induced mutations are replacements of cysteine to serine. In certain embodiments, Cas12a protein is further included in artificially induced mutations at position E174, as described in, for example, WO 2018195545 (incorporated herein by reference in its entirety). In certain embodiments, artificially induced mutations are replacements of glutamic acid residues. In certain embodiments, artificially induced mutations are replacements of glutamic acid to arginine.

[0058] In some embodiments, variant Cas12a protein is a modified form of Mb2Cas12a. In some embodiments, Cas12a protein includes and SEQ ID NO:4 amino acid sequence has at least 60% homogeneity (for example, at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence and at least one artificially induced mutation. In some embodiments, artificially induced mutation is the replacement of surface-exposed cysteine residues. In some embodiments, one or more surface-exposed cysteine residues are replaced by another residue (for example, serine residue). In some embodiments, the artificially induced mutation is at position C270, C583, C1068, C1099, and / or C1149. In some embodiments, the artificially induced mutation is the replacement of a cysteine residue. In some embodiments, the artificially induced mutation is the replacement of cysteine to serine. In some embodiments, the Cas12a protein is further included in an artificially induced mutation at position D172. In some embodiments, the artificially induced mutation is the replacement of an aspartic acid residue. In some embodiments, the artificially induced mutation is the replacement of aspartic acid to arginine. In some embodiments, the sequence of the Cas12a protein includes any one of SEQ ID NO:12-19.

[0059] Cas protein (for example, Cas12a protein) can include one or more domains.Non-limiting examples of domains include guiding nucleic acid recognition and / or binding domains, nuclease domains (for example, DNA enzyme or RNA enzyme domains, RuvC, HNH), DNA binding domains, RNA binding domains, helicase domains, protein-protein interaction domains and dimerization domains.Guiding nucleic acid recognition and / or binding domains can interact with guiding nucleic acids. Nuclease domains can include catalytic activity for nucleic acid cracking. Nuclease domains can lack catalytic activity to prevent nucleic acid cracking. Cas protein can be a chimeric Cas protein fused with other proteins or polypeptides. Cas protein can be a chimera of various Cas proteins, for example, comprising domains from different Cas proteins.

[0060] As used herein, Cas protein (for example, Cas12a protein) can be an active variant, inactive variant or fragment of a wild-type or modified Cas protein.Relative to the wild-type form of Cas protein, Cas protein can include amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras or any combination thereof.Cas protein can be a polypeptide having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity or sequence similarity with a wild-type exemplary Cas protein.Cas protein can be a polypeptide having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% sequence identity and / or sequence similarity with a wild-type exemplary Cas protein. The variant or fragment can comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or portion thereof. The variant or fragment can be targeted to a nucleic acid locus complexed with a guide nucleic acid while lacking nucleic acid cleavage activity.

[0061] In some embodiments, the modified Cas protein has a reduced function relative to the unmodified form. In some embodiments, the modified Cas protein is a functional defective unmodified form. For example, the nuclease-deficient Cas protein retains the ability to bind DNA, but lacks or has reduced nucleic acid cleavage activity. Cas nucleases (e.g., retaining wild-type nuclease activity, having reduced nuclease activity and / or lacking nuclease activity) can act in the CRISPR / Cas system to regulate the level and / or activity of the target gene or protein (e.g., reduce, increase or eliminate). The Cas protein can bind to the target polynucleotide and prevent transcription by physical barriers or editing nucleic acid sequences to produce non-functional gene products. In some embodiments, the modified Cas protein has no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5% or no more than 1% of the function (e.g., nuclease activity) of the wild-type Cas protein (e.g., Cas12a). In some embodiments, the modified Cas protein does not have the substantial function of the wild-type Cas protein. When the Cas protein is a modified form that does not have substantial nucleic acid cleavage activity, it can be referred to as enzymatically inactive and / or "dead" (abbreviated as "d"). Dead Cas proteins (e.g., dCas, dCas12a) can bind to target polynucleotides but cannot cleave target polynucleotides. In some embodiments, the Cas12a protein provided herein is a dCas12a protein.

[0062] In some embodiments, the modified Cas protein can be a modified Cas "base editor". Base editing enables a target DNA base to be directly and irreversibly converted to another base in a programmable manner without the need for DNA cleavage or donor DNA molecules. For example, Komor et al. (2016, Nature [Nature], 533: 420-424) teach Cas9-cytidine deaminase fusions, in which Cas9 has also been engineered to be inactive and does not induce double-stranded DNA breaks. In addition, Gaudelli et al. (2017, Nature [Nature], doi: 10.1038 / nature24644) teach Cas9 fused to tRNA adenosine deaminase with impaired catalytic activity, which can mediate the conversion of A / T to G / C in the target DNA sequence. In some embodiments, the Cas12a protein provided herein is a modified Cas12a base editor.

[0063] The Cas protein can be modified to optimize the regulation of gene expression. The Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity and / or enzymatic activity. The Cas protein can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted or inactivated, or the Cas protein can be truncated to remove domains that are unnecessary for protein function or to optimize (e.g., enhance or reduce) the activity of the Cas protein in order to regulate gene expression.

[0064] One or more nuclease domains (for example, RuvC, HNH) of Cas protein can be made to lack or mutate so that they are no longer functional or comprise reduced nuclease activity.For example, in the Cas protein (for example, Cas12a) comprising at least two nuclease domains, if a nuclease domain is made to lack or mutate, the Cas protein produced that is called as nickase can produce single-strand breaks at CRISPR RNA (crRNA) recognition sequence in double-stranded DNA, but does not produce double-strand breaks. This nickase can crack complementary strands or non-complementary strands, but can not crack both at the same time. In certain embodiments, the specificity of targeting double-strand breaks is improved by making nickase target to the opposite chains at two near loci. If the single chain at nickase cracking two loci, double-strand breaks are formed and can be repaired as described herein. If all nuclease domains (for example, RuvC nuclease domain in Cas12a protein) of Cas protein are missing or mutated, the Cas protein produced can have the ability of two chains of the cracking double-stranded DNA reduced or do not have this ability. In certain embodiments, Cas12a protein provided herein is Cas12a nickase protein.

[0065] Also provided herein is a fusion protein comprising any of the above-mentioned proteins and a heterologous domain. As used throughout, a "fusion protein" is a protein comprising two different polypeptide sequences (i.e., the Cas12a protein sequence and the heterologous polypeptide sequence as described above) that are connected (join) or connected (link) to form a single polypeptide. In certain embodiments, the two amino acid sequences are encoded by separate nucleic acid sequences, and these separate nucleic acid sequences have been connected so that they are transcribed and translated to produce a single polypeptide. The Cas12a protein and the heterologous domain can be connected in any order and orientation relative to each other. For example, the C' end of the Cas12a protein can be connected to the N' end or the C' end of the heterologous domain. The Cas12a protein and the heterologous domain can also be separated by one or more other fusion protein domains, as described below.

[0066] Exemplary heterologous domains include deaminase domains, transcription factor domains, nuclease domains, reverse transcriptase domains, transposase domains, integrase domains, uracil DNA glycosylase inhibitor domains, recombinase domains, nickase domains, methyltransferase domains, methylase domains, acetylase domains, acetyltransferase domains, transcription activator domains, and transcription repressor domains. See, for example, WO 2021 / 061507, which is incorporated herein by reference in its entirety.

[0067] In some embodiments, the fusion protein provided herein comprises one or more linkers. As used herein, a linker (also referred to as a spacer) is a flexible molecule or flexible molecule segment that connects (join) or connects two parts (e.g., domains) of a fusion protein or variant Cas12a protein as provided herein. In some embodiments, the linker is a polypeptide. The protein having a domain connected by a polypeptide linker is referred to as a fusion protein. In some embodiments, the linker is a non-peptide linker. The protein having a domain connected by a polypeptide linker is referred to as a modified protein. It should be understood that in the case of discussing fusion proteins throughout this disclosure, modified proteins are generally also envisioned (where feasible).

[0068] The linker can increase the range of orientations that can be adopted by the domain of the fusion protein or variant protein. The linker can be optimized to produce the desired effect in the fusion protein or variant protein. Aspects of linker design and consideration are described in, for example, Chen, X. et al., Adv Drug Deliv Rev. [Advanced Drug Delivery Review] 2013 Oct 15; 65(10): 1357-1369; and Klein, JS et al. 2014 Protein Eng. Des. Sel. [Protein Engineering Design and Selection] 27(10): 325-330. In some embodiments, the protein provided herein comprises a peptide linker. In some embodiments, the protein provided herein comprises a non-peptide linker. In some embodiments, the protein provided herein comprises a peptide linker and a non-peptide linker. The protein provided herein may also comprise a plurality of linkers, including at least one peptide linker, at least one non-peptide linker, or at least one peptide linker and at least one non-peptide linker.

[0069] The linker can be short or long, flexible or rigid. See, for example, WO 2021 / 061507 (incorporated herein by reference in its entirety) and WO 2020 / 168102 (incorporated herein by reference in its entirety) and US 2021 / 0017506 (incorporated herein by reference in its entirety).

[0070] In certain embodiments, the length of the joint can affect one or more functions of the fusion protein.Select the joint to realize the desired length within the ability of those skilled in the art.In certain embodiments, the length of the peptide linker can be, for example, 5 to 100 or more amino acids (for example, 4aa, 5aa, 8aa, 10aa, 15aa, 18aa, 20aa, 25aa, 30aa, 35aa, 40aa, 45aa, 50aa, 55aa, 60aa, 65aa, 70aa, 75aa, 80aa, 85aa, 90aa, 95aa or 100aa).In certain embodiments, the length of the joint is about 30 amino acids.In certain embodiments, the length of the joint is about 8 amino acids.

[0071] In some embodiments, the joint sequence can have a plurality of secondary structures, such as a spiral, a beta strand, a curl / bend, and a rotation. In some cases, the joint sequence can have an extended conformation and work as an independent domain that does not interact with adjacent protein domains. The joint sequence can be flexible or rigid. Flexible joints provide a certain degree of movement or interaction between the polypeptide domains and are generally rich in little or polar amino acids such as Gly and Ser (for example, at least 90%, at least 95%, at least 98%, at least 99% or all of the amino acid residues in the joint). Rigid joints can be used for maintaining a fixed distance between the domains and help maintain their independent function. Joint attachment can be carried out by amide linkage (for example, peptide bond) or other functionalities as discussed further below.

[0072] In some embodiments, the peptide linkers described herein comprise one or more repeats (e.g., 2 repeats, 3 repeats, 4 repeats, 5 repeats, 6 repeats or more) of GSSSS (SEQ ID NO: 43), and / or one or more repeats of GGGGS (SEQ ID NO: 44), and / or one or more repeats of GSSGSS (SEQ ID NO: 45), and / or one or more repeats of SGGS (SEQ ID NO: 77). In some embodiments, the linker comprises an amino acid sequence having at least 90% sequence identity to (GSSSS)6 (SEQ ID NO: 46) or (SGGS)2 (SEQ ID NO: 78). Additional exemplary peptide linkers include, but are not limited to, peptide linkers comprising: SGSETPGTSESATPE (SEQ ID NO: 47), SGSETPGTSESATPES (SEQ ID NO: 48), (GGGGS)3 (SEQ ID NO: 49), (GGGGS)5 (SEQ ID NO: 50), (GGGGS) 10(SEQ ID NO:51), GGGGGGGG (SEQ ID NO:52), GSAGSAAGSGEF (SEQ ID NO:53), A(EAAAK)3A(SEQ IDNO:54) or A(EAAAK) 10 A (SEQ ID NO: 55). Additional non-limiting exemplary linkers that can be used include those disclosed in PCT / US2020 / 051383; Chen et al., Adv. Drug. Deliv. Rev. [Advanced Drug Delivery Review] 65(10): 1357-1369 (2014) and Rosemalen et al., Biochemistry [Biochemistry] 2017, 56, 50, 6565-6574, the entire contents of both documents are incorporated herein by reference.

[0073] In some embodiments, the non-peptide linker can include any of a plurality of known chemical linkers. Exemplary chemical linkers can include one or more units of β-alanine, 4-aminobutyric acid (GABA), (2-aminoethoxy) acetic acid (AEA), 5-aminohexanoic acid (Ahx), PEG polymers, and trioxatridecane-succinamic acid (Ttds). In some embodiments, the non-peptide linker comprises one or more units of polyethylene glycol (PEG), which is commonly used as a linker for conjugating polypeptide domains due to its water solubility, lack of toxicity, low immunogenicity, and well-defined chain length. See, for example, Ramirez-Paz, J. et al., PLoS One [Public Library of Science Comprehensive] 13 (7): e0197643 (2018). The number of PEG linkage units can be selected based on the desired linker length.

[0074] The modified protein comprising a non-peptide linker can be produced in a variety of ways. For example, Cas12a protein and a heterologous domain can be produced separately (e.g., in vitro or by expression in a host cell and purified therefrom) and chemically connected in vitro. In certain embodiments, Cas12a protein, a heterologous domain and a linker can each be produced separately and chemically connected in vitro. Various chemical linkers can be used to cross-link two amino acid residues.

[0075] Such embodiment is also contemplated herein, wherein Cas12a albumen and heterologous domain (for example, being introduced into cell alone or being applied to target nucleic acid alone) as described above are used alone and the two are approached to form complex without using joint as described above.The various methods of forming complex between two or more polypeptides are known in the art and include but are not limited to using protein-protein interaction strategy (for example, SunTag, coiled coil etc.), using RNA-aptamers and related binding proteins (for example, MS2, N22 etc.) and label: trap strategy.For example, the site-directed nuclease of present disclosure can include MS2RNA aptamers, which will promote the interaction with the non-specific end processing enzyme comprising MS2 coat protein.

[0076] In some embodiments, the fusion proteins provided herein comprise an amino acid sequence that is at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identical to any one of SEQ ID NOs: 1-4. In some embodiments, the fusion proteins provided herein comprise an amino acid sequence set forth in any one of SEQ ID NOs: 5-19.

[0077] Any protein and fusion protein described herein can further include a targeting sequence that mediates protein localization (or retention) to a subcellular location, such as a plasma membrane or a given organelle membrane, nucleus, cytosol, mitochondria, endoplasmic reticulum (ER), Golgi apparatus, chloroplast, apoplast, peroxisome or other organelle. For example, a targeting sequence can direct a protein (e.g., a nuclease) to the nucleus using a nuclear localization signal (NLS); direct a protein to the outside of the nucleus of the cell using a nuclear export signal (NES), such as to the cytoplasm; direct a protein to the mitochondria using a mitochondrial targeting signal; direct a protein to the endoplasmic reticulum (ER) using an ER retention signal; direct a protein to the peroxisome using a peroxisome targeting signal; direct a protein to the plasma membrane using a membrane localization signal; or a combination thereof. In some embodiments, a protein includes a nuclear localization signal.Non-limiting examples of NLSs include NLS sequences derived from the following: the NLS of the SV40 virus large T antigen, having the amino acid sequence PKKKRKV (SEQ ID NO: 56); an NLS from a nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 57)); a c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 58) or RQRRNELKRSP (SEQ ID NO: 59); an hRNPA1 M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 60); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 61) from the IBB domain of importin-α; the sequences VSRKRPRP (SEQ ID NO: 62) and PPKKARED (SEQ ID NO: 63) of myoma T protein. sequence of the mouse Mx1 protein (SEQ ID NO: 69); sequence of human poly (ADP-ribose) polymerase KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 70); sequence of the steroid hormone receptor (human) glucocorticoid RKCLQAGMNLEARKTKK (SEQ ID NO: 71); and sequence of the Agrobacterium VirD2 protein KRPRDRHDGELGGRKRAR (SEQ ID NO: 72).

[0078] Any of the proteins and fusion proteins described herein may further comprise a detectable portion, such as a fluorescent protein or a fragment thereof. Examples of fluorescent proteins include, but are not limited to, yellow fluorescent protein (YFP, such as Venus), green fluorescent protein (GFP), and red fluorescent protein (RFP) and derivatives of these proteins, such as mutant derivatives. See, for example, Chudakov et al. "Fluorescent Proteins and Their Applications in Imaging Living Cells and Tissues [fluorescent proteins and their applications in imaging living cells and tissues]," Physiological Reviews [Physiological Reviews] 90 (3): 1103-1163 (2010); and Specht et al., "A Critical and Comparative Review of Fluorescent Tools for Live-Cell Imaging [critical comparative review of fluorescent tools for live cell imaging]," Annual Review of Physiology [Physiological Year Review] 79: 93-117 (2017)).

[0079] For example, any of the proteins and fusion proteins described herein can further comprise an affinity tag (e.g., a polyhistidine tag (e.g., (His)6 (SEQ ID NO: 73)), an HA tag (e.g., YPYDVPDYA (SEQ ID NO: 74)), an albumin binding protein, alkaline phosphatase, an AU1 epitope, an AU5 epitope, a biotin-carboxyl carrier protein (BCCP), a FLAG epitope (e.g., DYKDDDDK (SEQ ID NO: 75), or a MYC epitope (e.g., EQKLISEEDL (SEQ ID NO: 76)). See Kimple et al., "Overview of Affinity Tags for Protein Purification," Curr. Protoc. Protein Sci. 73: Unit 9.9 (2013).

[0080] Also provided herein are variants of the polypeptides (e.g., proteins and fusion proteins) disclosed herein. Unless otherwise expressly indicated, polypeptide variants retain their corresponding biological activity. For example, variants of Cas12a polypeptides retain the biological function of the full-length native sequence site-directed Cas12a protein. In another example, variants of heterologous domains retain the biological function of the full-length native sequence heterologous domains.

[0081] The modification of any polypeptide or protein provided herein is carried out by known methods. By way of example, modification is carried out in the following manner: the nucleotides in the nucleic acid encoding the polypeptide are subjected to site-specific mutagenesis, thereby producing a DNA encoding the modification, and then the DNA is expressed in recombinant cell culture to produce the encoded polypeptide. The technology for carrying out substitution mutations at predetermined sites in DNA with known sequences is well known. For example, M13 primer mutagenesis and PCR-based mutagenesis methods can be used to produce one or more substitution mutations. Any nucleotide sequence provided herein can be codon optimized to change, for example, to maximize expression in a host cell or organism.

[0082] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D-stereoisomers of naturally occurring amino acids, non-natural amino acids, and chemically modified amino acids. Non-natural amino acids (i.e., those that are not naturally found in proteins) are also known in the art, as described, for example, in Zhang et al., "Protein engineering with unnatural amino acids," Curr. Opin. Struct. Biol. 23(4):581-587 (2013); Xie et al., "Adding amino acids to the genetic repertoire," 9(6):548-54 (2005); and all references cited therein. Beta and gamma amino acids are known in the art and are also contemplated herein as non-natural amino acids.

[0083] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, the side chain can be modified to include a signaling moiety, such as a fluorophore or a radiolabel. The side chain can also be modified to include a new functional group, such as a thiol, a carboxylic acid, or an amino group. Post-translationally modified amino acids are also included in the definition of chemically modified amino acids.

[0084] Conservative amino acid substitutions are also contemplated. By way of example, conservative amino acid substitutions can be made in one or more amino acid residues, for example, in one or more lysine residues of any of the polypeptides provided herein. Those skilled in the art will appreciate that a conservative substitution is the replacement of one amino acid residue with another amino acid residue that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservative substitutions for one another: 1) Alanine (A), glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), tyrosine (Y), tryptophan (W); 7) Serine (S), Threonine (T); and 8) Cysteine (C), methionine (M).

[0085] By way of example, when referring to arginine to serine, conservative substitutions of serine are also contemplated (e.g., threonine). Non-conservative substitutions are also contemplated, such as replacing lysine with asparagine. IV. Recombinant Nucleic Acids, Constructs, Vectors, and Host Cells

[0086] Also provided herein is a recombinant nucleic acid encoding any variant Cas12a protein or fusion protein described herein. For example, a recombinant nucleic acid encoding a polypeptide having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% or 100%) identity to any one of SEQ ID NO:20-34 is also provided. A recombinant nucleic acid having at least 70% identity to any one of SEQ ID NO:20-34 is also provided.

[0087] Also provided are DNA constructs comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein as described herein or a domain thereof. A nucleic acid is "operably linked" when it is placed in a functional relationship with another nucleic acid sequence. A variety of promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of the start of transcription and involved in the recognition and binding of RNA polymerase and other proteins to initiate transcription.

[0088] As used herein, the term "promoter" refers to a nucleotide sequence, usually upstream (5') of its coding sequence, which controls the expression of the coding sequence by providing the recognition of the RNA polymerase and other factors required for appropriate transcription." Promoter regulatory sequence "is composed of proximal and more distal upstream elements. The promoter regulatory sequence affects the transcription, RNA processing or stability or translation of the relevant coding sequence. Regulatory sequences include enhancers, promoters, non-translated leader sequences, introns and polyadenylation signal sequences. They include natural sequences and synthetic sequences and sequences that may be combinations of synthetic sequences and natural sequences. "Enhancer" is a DNA sequence that can stimulate promoter activity and can be an intrinsic element of the promoter or an inserted heterologous element to enhance the level or tissue specificity of the promoter. It can operate in two orientations (normal or flipped) and can even play a role when moved to the upstream or downstream of the promoter. The meaning of the term "promoter" includes "promoter regulatory sequence".

[0089] The choice of promoter to be included depends on several factors, including but not limited to efficiency, selectability, inducibility, desired expression level, and cell or tissue preferential expression. It is routine for those skilled in the art to regulate the expression of a sequence by appropriately selecting and positioning promoters and other regulatory regions relative to the sequence.

[0090] It has been shown that certain promoters can direct RNA synthesis at a higher rate than others. These are referred to as "strong promoters." It has been shown that certain other promoters direct RNA synthesis at higher levels only in specific types of cells or tissues, and are often referred to as "tissue-specific promoters," or if a promoter preferentially directs RNA synthesis in certain tissues (where RNA synthesis may occur at a reduced level in other tissues), it is referred to as a "tissue-preferred promoter." Because promoters are used to control the expression pattern of one (or more) chimeric genes introduced into plants, there is ongoing interest in isolating novel promoters that can control the expression of one (or more) chimeric genes at certain levels in specific tissue types or at specific plant developmental stages.

[0091] Some promoters can instruct RNA synthesis at relatively similar levels in all tissues of a plant. These are called "constitutive promoters" or "non-tissue-dependent" promoters. Constitutive promoters can be divided into strong, medium, and weak categories based on the effectiveness of their instructing RNA synthesis. Because it is necessary to express one (or more) chimeric genes simultaneously in the different tissues of a plant in many cases to obtain the desired function of one (or more) genes, constitutive promoters are particularly useful in this regard. Although many constitutive promoters have been found and characterized from plants and plant viruses, there is still a continued interest in separating more novel constitutive promoters (synthetic or natural), which can control the expression of one (or more) chimeric genes at different levels and control the expression of multiple genes in the same transgenic plant to carry out gene stacking.

[0092] The most commonly used promoters include the nopaline synthase (NOS) promoter (Ebert et al., Proc. Natl. Acad. Sci. USA 84:5745-5749 (1987)); the octopine synthase (OCS) promoter; cauliflower mosaic virus promoters such as the cauliflower mosaic virus (CaMV) 19S promoter (Lawton et al., Plant Mol. Biol. 9:315-324 (1987)); the light-inducible promoter from the small subunit of ribulose bisphosphate carboxylase / oxygenase (Pellegrineschi et al., Biochem. Soc. Trans. [Proceedings of the Biochemical Society] 23(2):247-250 (1995)); Adh promoter (Walker et al., Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences of the United States of America] 84:6624-66280 (1987)); sucrose synthase promoter (Yang et al., Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences of the United States of America] 87:414-44148 (1990)); R gene complex promoter (Chandler et al., Plant Cell [Plant Cell] 1:1175-1183 (1989)); chlorophyll a / b binding protein gene promoter; etc.

[0093] In addition, it is contemplated that promoters that combine elements from more than one promoter may be useful. For example, U.S. Patent No. 5,491,288 discloses combining a cauliflower mosaic virus promoter with a histone promoter. Thus, elements from the promoters disclosed herein can be combined with elements from other promoters. Promoters useful for plant transgenic expression include those that are inducible, viral, synthetic, constitutive (Odell Nature 313:810-812 (1985)), temporally regulated, spatially regulated, tissue-specific, and spatiotemporally regulated. Using the regulatory elements described herein, many agronomic genes can be expressed in transformed plants. More particularly, plants can be genetically engineered to express various phenotypes of agronomic interest.

[0094] In some embodiments of the DNA constructs provided herein, the promoter can be a eukaryotic or prokaryotic promoter. In certain embodiments, the promoter is an inducible promoter, a natural inducible promoter (e.g., drought-inducible Rab17), a synthetic inducible promoter (e.g., auxin-inducible DR5, estradiol-inducible XVE / pLex, dexamethasone-inducible GVG / Gal4), a constitutive promoter (e.g., ZmUbq1, OsAct1, OsTub3, EF, EF1α), an egg cell-specific promoter (e.g., EC1, EC2, EC3, EC4, EC5), a pollen-specific promoter, an apical meristem-specific promoter, or a promoter with enriched expression in a zygote. In certain embodiments, the promoter is a flower mosaic promoter (e.g., ZmBde1, OsAP1). In certain embodiments, the promoter is an ubiquitin 4 promoter (e.g., sugarcane ubiquitin 4 promoter), an actin promoter, a tubulin promoter, a MADS box promoter, or a plant virus promoter. Suitable promoters are disclosed, for example, in U.S. Patent No. 10,519,456 (incorporated herein by reference in its entirety) and PCT / US2022 / 020690 (incorporated herein by reference).

[0095] The recombinant nucleic acid provided herein can be included in an expression cassette for expressing in a target host cell or organism. The cassette will include 5' and 3' regulatory sequences operably connected to the recombinant nucleic acid provided herein for allowing the expression of the fusion protein. The cassette can contain at least one other gene or genetic element to be co-transformed into a cell or organism in addition. In the case of including other genes or elements, these components are operably connected. Alternatively, one or more other genes or elements can be provided on a plurality of expression cassettes. This expression cassette is provided with a plurality of restriction sites and / or recombination sites so that the insertion of the polynucleotide is under the transcriptional regulation of the regulatory region. The expression cassette can contain a selective marker gene in addition. The expression cassette will include in the reverse direction of 5' to 3' transcription: transcription and translation initiation region (i.e., promoter), polynucleotides of the present invention, and transcription and translation termination region (i.e., termination region) that work in the target cell or organism. The promoters of the present invention are capable of directing or driving the expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, regardless of whether the RNA is then translated to produce a protein) in a host cell. The regulatory regions (i.e., the promoter, transcriptional regulatory regions, and translational termination regions) may be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence is a sequence that is derived from a foreign species or, if derived from the same species, has been substantially modified from its native form in composition and / or genomic locus by deliberate human intervention.

[0096] Additional regulatory signals include, but are not limited to, transcription initiation start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, initiation codons, termination signals, etc. See Sambrook et al., (1992) Molecular Cloning: A Laboratory Manual, Maniatis et al., eds. (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, New York; Davis et al., eds., (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, New York, and references cited therein.

[0097] The expression cassette can also comprise a selective marker gene for selecting transformed cells. For example, marker genes include genes that confer antibiotic resistance, such as those that confer hygromycin resistance, ampicillin resistance, gentamicin resistance, neomycin resistance. Other selective markers are known, and any one can be used.

[0098] In preparing the expression cassette, the various DNA fragments may be manipulated to provide a DNA sequence in the proper orientation and, if appropriate, the proper reading frame. To this end, adapters or linkers may be employed to connect the DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, resubstitution (e.g., transitions and transversions) may be involved.

[0099] In preparing the expression cassette, the various DNA fragments can be manipulated to provide a DNA sequence in the proper orientation and, if appropriate, the proper reading frame. To this end, adapters or linkers can be employed to connect the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, resubstitution (e.g., transitions and transversions) can be used.

[0100] Further provided is a vector, which comprises the recombinant nucleic acid or DNA construct set forth herein. It is envisioned that the vector has the necessary functional elements for guiding and regulating the transcription of the inserted nucleic acid. These functional elements include, but are not limited to, a promoter, a region upstream or downstream of the promoter (such as an enhancer and terminator that can regulate the transcriptional activity of the promoter), an origin of replication, an appropriate restriction site for promoting the cloning of the inserted sequence adjacent to the promoter, an antibiotic resistance gene or other marker that can be used to select cells containing the vector or a vector containing the inserted sequence, an RNA splicing junction, a transcription termination region, or any other region that can be used to promote the expression of an inserted gene or a hybrid gene. Generally speaking, see, Sambrook et al. Molecular Cloning: A Laboratory Manual [Molecular Cloning: Laboratory Manual], 4th edition, Cold Spring Harbor Laboratory Press [Cold Spring Harbor Laboratory Press], Cold Spring Harbor [Cold Spring Harbor], 2012. The vector can be, for example, a plasmid. In some embodiments of the DNA constructs and vectors provided herein, the constructs and vectors comprise a nopaline synthase gene terminator sequence (e.g., an Agrobacterium tumefaciens nopaline synthase gene terminator sequence).

[0101] There are many E. coli expression vectors known to those of ordinary skill in the art, and these vectors can be used for the expression of nucleic acid. Other microbial hosts suitable for use include bacillus (such as bacillus subtilis (Bacillus subtilis)) and other enterobacteriaceae (such as Salmonella (Salmonella), Serratia (Senatia)) and various Pseudomonas species. In these prokaryotic hosts, expression vectors can also be prepared, and these expression vectors will typically contain expression control sequences (for example, replication origin) compatible with the host cell. In addition, there will be any number of multiple well-known promoters, such as lactose promoter system, tryptophan (Trp) promoter system, beta-lactamase promoter system or promoter system from bacteriophage lambda. In addition, yeast expression can be used. Nucleic acids encoding polypeptides of the present invention are provided herein, and wherein nucleic acids can be expressed by yeast. More particularly, nucleic acids can be expressed by pichia pastoris (Pichia pastoris) or yeast saccharomyces cerevisiae (S.cerevisiae).

[0102] Mammalian cells also allow protein expression in an environment that is conducive to important post-translational modifications, such as folding and cysteine pairing, the interpolation of complex carbohydrate structures and the secretion of activated proteins. Carriers that can be used for expressing activated proteins in mammalian cells are known in the art, and can contain genes that confer hygromycin resistance, geneticin or G418 resistance, or other genes or phenotypes suitable for use as selective markers, or methotrexate resistance for gene amplification. Developed in the art are multiple suitable host cell lines capable of secreting complete human proteins, and include Chinese hamster ovary celI, HeLa cell, HEK-293 cell, HEK-293T cell, U2OS cell or any other former or transformed cell line. Other suitable host cell lines include COS-7 cell, myeloma cell line, Jurkat cell etc. The expression vectors used for these cells can include expression control sequences, such as replication origin, promoter, enhancer and necessary information processing site, such as ribosome binding site, RNA splicing site, polyadenylation site and transcription terminator sequence. Preferred expression control sequences are promoters derived from immunoglobulin genes, SV40, adenovirus, bovine papilloma virus, and the like.

[0103] The expression vectors described herein can also include nucleic acids as described herein under the control of an inducible promoter (such as a tetracycline inducible promoter or a glucocorticoid inducible promoter). The nucleic acids of the present invention can also be under the control of a tissue-specific promoter to promote expression of the nucleic acids in specific cells, tissues, or organs. Any regulatable promoter is also contemplated, such as a metallothionein promoter, a heat shock promoter, and other regulatable promoters, many examples of which are well known in the art. In addition, the Cre-loxP inducible system and the Flp recombinase inducible promoter system can also be used, both of which are known in the art.

[0104] Insect cells also allow for the expression of polypeptides. Recombinant proteins produced in insect cells using baculovirus vectors undergo post-translational modifications similar to those of wild-type mammalian proteins.

[0105] Also provided herein are host cells comprising the recombinant nucleic acids, DNA constructs and / or vectors described herein and methods for preparing such cells. In certain embodiments, the cell is a plant cell. In certain embodiments, the plant cell is a corn plant cell, a wheat plant cell, a rice plant cell, a soybean plant cell, a sunflower plant cell, or a tomato plant cell.

[0106] Provided are host cells comprising nucleic acids or vectors as described herein. The host cell can be an in vitro, ex vivo, or in vivo host cell. The host cells as provided herein are capable of expressing fusion proteins. Also provided are cell populations of any of the host cells described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprise a recombinant nucleic acid encoding a fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprise a DNA construct encoding a protein and / or fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprise a vector comprising a recombinant nucleic acid or DNA construct encoding a protein and / or fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprise a plurality of any of the host cells described herein. In some embodiments, a plurality of cells of any of the cell populations described herein express a protein and / or fusion protein as described herein.

[0107] In certain embodiments, the cell stabilization or transient expression protein and / or fusion protein provided.The stable expression of protein and / or fusion protein in cell refers to any one of nucleic acid as herein described, DNA construct or the vector being integrated into the cell genome, thereby allowing cell expression of protein and / or fusion protein.Transient expression refers to any one of nucleic acid, DNA construct and / or the vector being directly introduced into the cell and / or fusion protein expression (that is, the gene of coded protein and / or fusion protein is not integrated into the genome of cell).

[0108] In some embodiments, provided cells express proteins and / or fusion proteins constitutively or inducibly. Constitutive expression refers to the sustained, continuous expression of a gene (i.e., a protein), while inducible expression refers to the expression of a gene (protein) in response to a stimulus. Inducible expression is typically regulated via an inducible promoter (described above).

[0109] Also provided are cell cultures comprising one or more host cells described herein. Methods for culturing and producing a variety of cells are available in the art, including cells of bacterial (e.g., E. coli and other bacterial strains), animal (particularly mammalian) and archaeal origin. Ausubel, ed. (1995) Current Protocols in Molecular Biology, John Wiley & Sons; and Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, 3rd ed., Wiley-Liss, New York, and references cited therein; Doyle and Griffiths (1997) Mammalian Cell Culture: Essential Techniques, John Wiley and Sons, NY; Humason (1979) Animal Tissue Techniques, 4th ed., WH Freeman and Company; and Ricciardelli et al., (1989) Invitro Cell Culture: Essential Techniques. Dev.Biol. [In vitro Cell and Developmental Biology] 25:1016–1024.

[0110] The host cell can be a prokaryotic cell, including, for example, a bacterial cell. Alternatively, the cell can be a eukaryotic cell, such as a mammalian cell. In certain embodiments, the cell can be a HEK-293T cell, a HEK-293 cell, a Chinese hamster ovary (CHO) cell, a U2OS cell, or any other primary or transformed cell. In certain embodiments, the cell can be a COS-7 cell, a HELA cell, an avian cell, a myeloma cell, a Pichia pastoris cell, an insect cell, or a plant cell. A variety of other suitable host cell lines have been developed, and include myeloma cell lines, fibroblast cell lines, and various tumor cell lines such as melanoma cell lines. The vector containing the nucleic acid segment of interest can be transferred or introduced into the host cell by well-known methods, and these methods vary according to the type of cell host.

[0111] As used herein, the phrase "introducing" in the context of introducing nucleic acid into a cell (e.g., a prokaryotic cell, a bacterial cell, a eukaryotic cell, a plant cell) refers to translocating a nucleic acid sequence from extracellular to intracellular. In some cases, introducing is to translocating a nucleic acid from extracellular to the nucleus of a cell. In the case of introducing more than one nucleic acid molecule, these nucleic acid molecules can be assembled into a part of a single polynucleotide or nucleic acid construct, or assembled into a separate polynucleotide or nucleic acid construct, and can be located on the same or different nucleic acid constructs. Therefore, such polynucleotides can be introduced into a cell (e.g., a plant cell) in a single transformation event, in a separate transformation event, or, for example, as a part of a breeding scheme. It is envisioned that various methods for introducing nucleic acid into a cell include, but are not limited to, electroporation, nanoparticle delivery, gene gun transformation, viral delivery, contact with nanowires or nanotubes, receptor-mediated internalization, translocation via cell-penetrating peptides, liposome-mediated translocation, DEAE dextran, lipofectamine, calcium phosphate, or any method now known or identified in the future for introducing nucleic acid into a prokaryotic or eukaryotic host. Targeted nuclease systems (e.g., RNA-guided nucleases, transcription activator-like effector nucleases (TALENs), zinc finger nucleases (ZFNs), or large-scale TALs (MTs) can also be used to introduce nucleic acids (e.g., nucleic acids encoding proteins and / or fusion proteins described herein) into host cells. See Li et al. Signal Transduction and Targeted Therapy 5, Article No. 1 (2020).

[0112] The conversion of cell can be stable or transient.Therefore, transgenic cell of the present invention, vegetable cell, plant and / or plant part can be by stably transformed or transient transformation." conversion " can refer to nucleic acid molecule being transferred in the genome of host cell, produces genetically stable heredity.In certain embodiments, be introduced into plant, plant part and / or vegetable cell and be via bacteria-mediated conversion, particle bombardment conversion, calcium phosphate-mediated conversion, cyclodextrin-mediated conversion, electroporation, liposome-mediated conversion, nanoparticle-mediated conversion, polymer-mediated conversion, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, ultrasonic treatment, infiltration, polyethylene glycol-mediated conversion, protoplast transformation or make nucleic acid be introduced into any other electrical, chemical, physical and / or biological mechanism or its any combination in plant, plant part and / or its cell and carry out.

[0113] The procedures for transforming plants are well known and conventional in the art and are generally described in the literature. Non-limiting examples of methods for plant transformation include transformation via bacteria-mediated nucleic acid delivery (e.g., via bacteria from the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical) and / or biological mechanism that allows the introduction of nucleic acids into plant cells, including any combination thereof. General guides to various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants," in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, eds. (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7:849-858 (2002)).

[0114] Agrobacterium-mediated transformation is a common method for transforming plants due to its high transformation efficiency and due to its wide applicability with many different species. Agrobacterium-mediated transformation typically involves transferring the binary vector carrying the foreign DNA of interest to an appropriate Agrobacterium strain, which may depend on the complement of the vir genes carried by the host Agrobacterium strain on a co-existing Ti plasmid or chromosomally (Uknes et al., 1993, Plant Cell [plant cells] 5: 159-169). The Escherichia coli carrying the recombinant binary vector can be used, and the auxiliary Escherichia coli strain (the auxiliary Escherichia coli strain carries a plasmid capable of moving the recombinant binary vector to the target Agrobacterium strain) is passed through a three-parent mating procedure to achieve the transfer of the recombinant binary vector to Agrobacterium. Alternatively, the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation ( and Willmitzer 1988, Nucleic Acids Res 16:9877).

[0115] Plant transformation by recombinant Agrobacterium typically involves co-cultivation of Agrobacterium with explants from the plant and follows methods well known in the art. Transformed tissue is typically regenerated on selective media carrying an antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders.

[0116] Another method for transforming plants, plant parts and plant cells involves advancing inert or biologically active particles on plant tissues and cells. See, for example, U.S. Patent Nos. 4,945,050; 5,036,006 and 5,100,792. Generally speaking, this method involves advancing inert or biologically active particles at plant cells under conditions effective to penetrate the outer surface of the cell and provide for incorporation into its interior. When utilizing inert particles, the vector can be introduced into the cell by coating the particles with a vector containing the target nucleic acid. Alternatively, one or more cells can be surrounded by the vector so that the vector is carried into the cell by stimulation of the particle. Biologically active particles (e.g., dried yeast cells, dried bacteria or bacteriophages, each containing one or more nucleic acids intended to be introduced) can also be advanced into plant tissue. As used herein, the phrase "biolistic transformation" refers to a method of introducing RNA or DNA directly into cells (e.g., plant cells) in which the RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cells (e.g., plant cells) using high-speed pressure to allow the RNA or DNA to penetrate the cells (e.g., penetrate the plant cell wall).

[0117] The CRISPR / Cas system can also be used to edit the genome of a host cell or organism. As described in detail above, the "CRISPR / Cas" system refers to a class of bacterial systems that are widely used to defend against foreign nucleic acids. Any of the CRISPR / Cas system components described herein can be used to introduce proteins, fusion proteins, recombinant nucleic acids or systems into the genome of a host cell or organism. Methods for genome editing mediated by the CRISPR / Cas system are known in the art. It should be understood that the use of the CRISPR / Cas system for introducing proteins, fusion proteins, recombinant nucleic acids or systems described herein into the genome of a host cell or organism is different from the specific methods and systems provided herein.

[0118] Any protein and / or fusion protein described herein can be purified or isolated from a host cell or host cell colony. For example, a recombinant nucleic acid encoding any protein and / or fusion protein described herein can be introduced into a host cell under conditions that allow expression of the protein and / or fusion protein. In certain embodiments, the recombinant nucleic acid is codon-optimized for expression. After expression in the host cell, purification methods known in the art can be used to separate or purify the protein and / or fusion protein. V. System

[0119] On the other hand, provided herein are systems that can be used to edit one or more nucleic acids. These systems include one or more of the above-mentioned Cas12a proteins and / or fusion proteins (or recombinant nucleic acids, constructs, vectors or host cells). In certain embodiments, these systems further include one or more other elements that can be used to edit one or more nucleic acids. For example, a system comprising a fusion protein containing a Cas nuclease can further include one or more guide nucleic acids described in detail below. The system provided herein can be used to carry out the method described in Chapter VI of this disclosure.

[0120] In some cases, the systems and methods described herein include at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein include multiple guide nucleic acids. In certain embodiments, the polynucleotide can be deoxyribonucleic acid (DNA). In some cases, the DNA sequence can be single-stranded or double-stranded. In certain embodiments, the at least one guide nucleic acid polynucleotide can be ribonucleic acid (guide RNA).

[0121] In certain embodiments, Cas12a protein can be compounded with at least one guide RNA polynucleotide.The at least one guide RNA polynucleotide can include a nucleic acid targeting region, which includes a sequence complementary to the nucleic acid sequence on a targeting polynucleotide (such as a targeted genomic locus or gene) to give the sequence specificity of the nuclease targeting. In certain embodiments, the at least one guide RNA polynucleotide can include two independent nucleic acid molecules (which can be referred to as dual-guide nucleic acids) or a single nucleic acid molecule (which can be referred to as a single guide nucleic acid (e.g., single guide RNA or sgRNA).

[0122] The Cas protein binding segment of the guidance nucleic acid can include two nucleotide segments (for example, crRNA and tracrRNA) that are complementary to each other. Two nucleotide segments (for example, crRNA and tracrRNA) that are complementary to each other can be covalently linked by intervening nucleotides (for example, joints in the case of single guidance nucleic acid). Two nucleotide segments (for example, crRNA and tracrRNA) that are complementary to each other can hybridize to form double-stranded RNA duplex or hairpin of Cas protein binding segment, thus producing stem-loop structure. CrRNA and tracrRNA can be covalently linked via 3' end of crRNA and 5' end of tracrRNA. Alternatively, tracrRNA and crRNA can be covalently linked via 5' end of tracrRNA and 3' end of crRNA. CrRNA can include the nucleic acid targeting segment (for example, spacer district) of guidance nucleic acid and can form half of the nucleotide segment of the double-stranded duplex of the Cas protein binding segment of guidance nucleic acid. CrRNA can also provide the single-stranded nucleic acid targeting segment (for example, spacer district) hybridized with target nucleic acid recognition sequence (for example, prototype spacer). Whether the nuclease requires only a crRNA molecule or both a crRNA molecule and a tracrRNA molecule (whether covalently linked or not) depends on the CRISPR-associated nuclease used. The Cas12 protein typically does not require tracrRNA.

[0123] In certain embodiments, the length of the nucleic acid targeting region of the guide nucleic acid can be between 18 and 72 nucleotides. The length of the nucleic acid targeting region (e.g., spacer region) of the guide nucleic acid can be from about 12 nucleotides to about 100 nucleotides. For example, the length of the nucleic acid targeting region (e.g., spacer region) of the guide nucleic acid can be from about 12 nucleotides (nt) to about 80nt, from about 12nt to about 50nt, from about 12nt to about 40nt, from about 12nt to about 30nt, from about 12nt to about 25nt, from about 12nt to about 20nt, from about 12nt to about 19nt, from about 12nt to about 18nt, from about 12nt to about 17nt, from about 12nt to about 16nt or from about 12nt to about 15nt. Alternatively, the length of the DNA targeting segment can be from about 18 nt to about 20 nt, from about 18 nt to about 25 nt, from about 18 nt to about 30 nt, from about 18 nt to about 35 nt, from about 18 nt to about 40 nt, from about 18 nt to about 45 nt, from about 18 nt to about 50 nt, from about 18 nt to about 60 nt, from about 18 nt to about 70 nt, from about 18 nt to about 80 nt, from about 18 nt to about 90 nt, from about 18 nt to about 100 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, from about 20 nt to about 60 nt, from about 20 nt to about 70 nt, from about 20 nt to about 80 nt, from about 20 nt to about 90 nt, or from about 20 nt to about 100 nt. The nucleic acid targeting region can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides in length. The nucleic acid targeting region (e.g., a spacer sequence) can be at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides in length.

[0124] In some embodiments, the length of the nucleic acid targeting region (e.g., a spacer) of a guide nucleic acid is 20 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 19 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 18 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 17 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 16 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 21 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 22 nucleotides.

[0125] The length of the nucleotide sequence of the guide nucleic acid that is complementary to the nucleotide sequence of the target nucleic acid (target sequence) can be, for example, at least about 12 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt. The length of the nucleotide sequence of the guide nucleic acid that is complementary to the nucleotide sequence of the target nucleic acid (target sequence) can be from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 45 nt, from about 12 nt to about 40 nt, from about 12 nt to about 35 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 19 nt to about 20 ... From about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt.

[0126] The protospacer sequence of the targeting polynucleotide can be identified by identifying the protospacer adjacent motif (PAM) within the region of interest and selecting a region of desired size upstream or downstream of the PAM as the protospacer. The corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.

[0127] Spacer sequences can be identified using a computer program (e.g., machine readable code). The computer program can use variables such as predicted melting temperature, secondary structure formation and predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, GC%, genomic occurrence frequency, methylation status, SNP presence, etc.

[0128] The percent complementarity between a nucleic acid targeting sequence (e.g., a spacer sequence of at least one guide polynucleotide as disclosed herein) and a target nucleic acid (e.g., a protospacer sequence of one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percent complementarity between a nucleic acid targeting sequence and a target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 consecutive nucleotides.

[0129] The length of the Cas protein binding segment of the guide nucleic acid can be from about 10 nucleotides to about 100 nucleotides, for example, from about 10 nucleotides (nt) to about 20nt, from about 20nt to about 30nt, from about 30nt to about 40nt, from about 40nt to about 50nt, from about 50nt to about 60nt, from about 60nt to about 70nt, from about 70nt to about 80nt, from about 80nt to about 90nt or from about 90nt to about 100nt. For example, the length of the Cas protein binding segment of the guide nucleic acid can be from about 15 nucleotides (nt) to about 80nt, from about 15nt to about 50nt, from about 15nt to about 40nt, from about 15nt to about 30nt or from about 15nt to about 25nt.

[0130] The length of the dsRNA duplex of the Cas protein binding segment of the guide nucleic acid can be from about 6 base pairs (bp) to about 50bp. For example, the length of the dsRNA duplex of the protein binding segment can be from about 6bp to about 40bp, from about 6bp to about 30bp, from about 6bp to about 25bp, from about 6bp to about 20bp, from about 6bp to about 15bp, from about 8bp to about 40bp, from about 8bp to about 30bp, from about 8bp to about 25bp, from about 8bp to about 20bp or from about 8bp to about 15bp. For example, the length of the dsRNA duplex of the Cas protein binding segment can be from about 8bp to about 10bp, from about 10bp to about 15bp, from about 15bp to about 18bp, from about 18bp to about 20bp, from about 20bp to about 25bp, from about 25bp to about 30bp, from about 30bp to about 35bp, from about 35bp to about 40bp or from about 40bp to about 50bp.

[0131] In some embodiments, the length of the dsRNA duplex of the Cas protein binding segment can be 36 base pairs. The complementarity percentage between the nucleotide sequences of the dsRNA duplexes that hybridize to form the protein binding segment can be at least about 60%. For example, the complementarity percentage between the nucleotide sequences of the dsRNA duplexes that hybridize to form the protein binding segment can be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99%. In some cases, the complementarity percentage between the nucleotide sequences of the dsRNA duplexes that hybridize to form the protein binding segment is 100%.

[0132] The guide nucleic acid of the system of the present disclosure can include modifications or sequences that provide additional desired characteristics (e.g., stability of modification or regulation; subcellular targeting; tracking with fluorescent labels; binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (7-methylguanylate cap (m7G)); a 3' polyadenylation tail (3' poly (A) tail); a riboswitch sequence (e.g., to allow regulated stability and / or regulated accessibility of proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (hairpin); a modification or sequence that targets RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a portion that facilitates fluorescence detection, a sequence that allows fluorescence detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).

[0133] The guide nucleic acid may comprise one or more modifications (e.g., base modifications, backbone modifications) to provide a nucleic acid with new or enhanced characteristics (e.g., improved stability). The guide nucleic acid may comprise a nucleic acid affinity tag. The nucleoside may be a base-sugar combination. The base portion of the nucleotide may be a heterocyclic base. The two most common types of such heterocyclic bases are purines and pyrimidines. The nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides comprising pentofuranosyl sugars, the phosphate group may be linked to the 2', 3', or 5' hydroxyl portion of the sugar. When forming the guide nucleic acid, the phosphate group may covalently link adjacent nucleosides to form a linear polymer compound. Furthermore, the corresponding ends of this linear polymer compound may be further linked to form a cyclic compound; however, a linear compound may be suitable. In addition, the linear compound may have internal nucleotide base complementarity and may therefore be folded in a manner that facilitates the production of a fully or partially double-stranded compound. In addition, within the guide nucleic acid, the phosphate group may generally be involved in forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid can be a 3' to 5' phosphodiester linkage.

[0134] The guide nucleic acid can comprise a modified backbone and / or modified internucleoside linkage. The modified backbone can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.

[0135] Suitable modified guide nucleic acid backbones containing phosphorus atoms can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkylphosphonates (such as 3'-alkylenephosphonates, 5'-alkylenephosphonates, chiral phosphonates), phosphinates, phosphoramidates (including 3'-aminophosphoramidates and aminoalkylphosphoramidates), phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphonate triesters, selenophosphoroates, and boranophosphates with normal 3'-5' linkages, 2'-5' linkage analogs, and those with reverse polarity in which one or more internucleotide linkages are 3' to 3', 5' to 5', or 2' to 2' linkages. Suitable guide nucleic acids with reverse polarity can comprise a single 3' to 3' linkage at the 3'-most internucleotide linkage (such as a single reverse nucleotide residue in which a nucleobase is deleted or has a hydroxyl group in its place). Various salts (eg, potassium chloride or sodium chloride), mixed salts, and free acid forms are also included.

[0136] The guide nucleic acid can comprise one or more phosphorothioate and / or heteroatom internucleoside linkages, particularly -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (methylene(methylimino) or MMI backbone), -CH2-ON(CH3)-CH2-, -CH2-N(CH3)-N(CH3)-CH2-, and -ON(CH3)-CH2-CH2- (where the native phosphodiester internucleoside linkage is represented as -OP(=O)(OH)-O-CH2-).

[0137] Instruct nucleic acid can comprise morpholino backbone structure.For example, nucleic acid can comprise 6 yuan of morpholino ring to replace ribose ring.In some of these embodiments, diaminophosphorothioate or other non-phosphodiester internucleoside linkage replaces phosphodiester linkage.

[0138] The guiding nucleic acid can comprise such polynucleotide backbones, which are formed by short-chain alkyl or cycloalkyl nucleoside linkages, mixed heteroatoms and alkyl or cycloalkyl nucleoside linkages or one or more short-chain heteroatoms or heterocyclic nucleoside linkages. These can include those having the following: morpholino linkages (partially formed by the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide and sulfone backbones; formyl acetyl and thioformyl acetyl backbones; methylene formyl acetyl and thioformyl acetyl backbones; ribose acetyl backbones; backbones containing olefins; sulfamate backbones; methylene imino and methylene hydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones with mixed N, O, S and CH2 components.

[0139] Instruct nucleic acid can include nucleic acid mimics.Term " mimics " can be intended to include such polynucleotide, wherein only furanose ring or furanose ring and internucleotide linkage are replaced by non-furanose group, and only the replacement of furanose ring can also be referred to as sugar substitute.Heterocyclic base part or the heterocyclic base part of modification can be retained for hybridization with appropriate target nucleic acid.A kind of such nucleic acid can be peptide nucleic acid (PNA).In PNA, the sugar backbone of polynucleotide can be replaced by containing amide backbone, particularly aminoethylglycine backbone.Nucleotide can be retained and directly or indirectly bonded to the nitrogen-nitrogen atom of the amide part of backbone.The backbone in PNA compound can comprise the aminoethylglycine unit of two or more connections that give PNA containing amide backbone.Heterocyclic base part can be directly or indirectly combined with the nitrogen-nitrogen atom of the amide part of backbone.

[0140] Guidance nucleic acid can include connected morpholino units (morpholino nucleic acids), and these morpholino units have heterocyclic bases attached to the morpholino ring. Linking groups can connect the morpholino monomer units in morpholino nucleic acids. Oligomeric compounds based on nonionic morpholinos can have less desirable interactions with cellular proteins. Polynucleotides based on morpholinos can be nonionic mimics of guidance nucleic acids. Various compounds within the morpholino species can be connected using different linking groups. Another type of polynucleotide mimics can be referred to as cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced by a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used for the synthesis of oligomeric compounds using phosphoramidite chemistry. CeNA monomers can be incorporated into nucleic acid chains to increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements, and these complexes have stability similar to natural complexes. Another modification can include locked nucleic acids (LNA), in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2. LNA and LNA analogs can exhibit very high duplex thermal stability (Tm = +3°C to +10°C) with complementary nucleic acids, stability to 3'-exonucleolytic degradation, and good solubility properties.

[0141] The guide nucleic acid may comprise one or more substituted sugar moieties. Suitable polynucleotides may comprise a sugar substituent group selected from the group consisting of: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl, wherein alkyl, alkenyl, and alkynyl groups may be substituted or unsubstituted C1 to C 10 Alkyl or C2 to C 10Particularly suitable is O((CH2) n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, wherein n and m are from 1 to about 10. The sugar substituent group may be selected from: C1 to C 10 Lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocycloalkyl, heterocycloalkylaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleavage group, reporter group, intercalator, group for improving the pharmacokinetic properties of guide nucleic acid, or group for improving the pharmacodynamic properties of guide nucleic acid and other substituents with similar properties. Suitable modifications may include 2'-methoxyethoxy (2'-O-CH2 CHOCH3, also known as 2'-O-(2-methoxyethyl) or 2'-MOE, i.e., alkoxyalkoxy group). Additional suitable modifications may include 2'-dimethylaminooxyethoxy (O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE), 2'-dimethylaminoethoxyethoxy (also known as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), or 2'-O-CH2-O-CH2-N(CH3)2.

[0142] Other suitable sugar substituent groups can include methoxy (-O-CH3), aminopropoxy (-OCH2CH2NH2), allyl (-CH2-CH=CH2), -O-allyl (-O--CH2—CH=CH2) and fluorine (F). The 2'-sugar substituent group can be in the arabinose (upper) position or the ribose (lower) position. A suitable 2'-arabinose modification is 2'-F. Similar modifications can also be made at other positions on the oligomeric compound, particularly the 3' position of the sugar on the 3' terminal nucleoside or in the 2'-5' linked nucleotide and the 5' position of the 5' terminal nucleotide. The oligomeric compound can also have a sugar mimetic, such as a cyclobutyl moiety in place of the pentofuranosyl sugar.

[0143] The guide nucleic acid can also comprise nucleobase (or "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases can include purine bases such as adenine (A) and guanine (G), and pyrimidine bases such as thymine (T), cytosine (C), and uracil (U). Modified nucleobases can include other synthetic and natural nucleobases such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halogeno, 8-amino, 8-sulfhydryl, 8-sulfanyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halogeno, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine as well as 3-deazaguanine and 3-deazaadenine. Modified nucleobases may include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).

[0144] Heterocyclic base moieties can include those in which purine or pyrimidine bases are replaced by other heterocycles such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine and 2-pyridone. Core bases can be used to increase the binding affinity of polynucleotide compounds. These core bases can include 5-substituted pyrimidines, 6-azapyrimidines and N-2, N-6 and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil and 5-propynylcytosine. 5-methylcytosine substitution can increase nucleic acid duplex stability by 0.6°C-1.2°C, and can be suitable base substitutions (e.g., when combined with 2'-O-methoxyethyl sugar modifications).

[0145] Modification of guide nucleic acids can include chemically linking one or more moieties or conjugates to the guide nucleic acid that can enhance the activity, cellular distribution, or cellular uptake of the guide nucleic acid. These moieties or conjugates can include conjugate groups covalently bound to functional groups such as primary or secondary hydroxyl groups. Conjugate groups can include, but are not limited to, intercalators, reporters, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacodynamic properties of oligomers, and groups that can enhance the pharmacodynamic properties of oligomers. Conjugate groups can include, but are not limited to, cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that enhance pharmacodynamic properties include groups that improve uptake, enhance degradation resistance, and / or strengthen sequence-specific hybridization with target nucleic acids. Groups that can enhance pharmacokinetic properties include groups that improve the uptake, distribution, metabolism, or excretion of nucleic acids. The conjugate moiety can include, but is not limited to, a lipid moiety such as a cholesterol moiety, cholic acid, a thioether (e.g., hexyl-S-tritylthiol), a thiocholesterol, an aliphatic chain (e.g., dodecandiol or undecyl residues), a phospholipid (e.g., di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glyceryl-3-H-phosphonate), a polyamine or a polyethylene glycol chain or adamantaneacetic acid, a palmityl moiety or octadecylamine or a hexylamino-carbonyl-hydroxycholesterol moiety.

[0146] In some embodiments, at least one guide RNA polynucleotide of the system or method provided herein can be combined with at least a portion of a genome (e.g., a plant genome) or a gene (e.g., a plant gene). In some cases, the at least one guide RNA polynucleotide can be complexed with Cas12a protein to guide a portion of a protein-targeted target nucleic acid (e.g., a site in a genome or gene).

[0147] In some embodiments, the systems described herein include at least one guide RNA polynucleotide capable of forming a complex with the Cas12a protein or fusion protein of the system. In some embodiments, the systems described herein include at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides capable of forming a complex with the site-directed nuclease portion of the fusion protein of the system.

[0148] In some embodiments, the guide nucleic acid comprises a nucleotide sequence that is at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to any one of SEQ ID NOs: 27-34 shown in Table 5. Table 5. Exemplary gRNA sequences

[0149] Also provided herein are kits comprising components of the systems described herein. In some embodiments, the kits comprise one or more of the fusion proteins and / or polynucleotides described herein. VI. Methods

[0150] On the other hand, provided herein is a method for editing one or more nucleic acids using Cas12a proteins, fusion proteins and / or systems as described herein. In certain embodiments, these methods include contacting nucleic acid (i.e., nucleic acid to be edited) with at least one Cas12a protein and / or fusion protein as described herein. In certain embodiments, these methods further include contacting nucleic acid with a guide RNA (e.g., as described in Section V above) having a region complementary to the selected portion of the nucleic acid. In certain embodiments, nucleic acid is contacted with Cas12a protein and / or fusion protein and guide RNA to edit nucleic acid. Nucleic acid (i.e., nucleic acid to be edited) can be any suitable nucleic acid. In certain embodiments, nucleic acid is a part of a chromosome. In certain embodiments, nucleic acid is a part of a genome (e.g., a plant genome).

[0151] As described herein and shown in the following examples, the method provided herein can increase the frequency of one or more desired nucleic acid editing results (e.g., SDN-1 editing). In certain embodiments, the SDN-1 editing efficiency can be measured in the following manner: the number of plants with insertion or deletion ("indel (indel)") is divided by the total number of transgenic plants. In certain embodiments, relative to using unmodified (i.e. wild type) Cas12a protein, the SDN-1 editing efficiency is increased using Cas12a protein or fusion protein provided herein. In certain embodiments, the homozygous editing (i.e., the same indel is present in two alleles of the target nucleic acid) and biallelic editing (i.e., different indels are present in each allele of the target nucleic acid) of the indel event can be further analyzed. In certain embodiments, the rate of homozygous editing / biallelic editing can be measured in the following manner: the number of plants with homozygous editing / biallelic editing is divided by the total number of plants with indel. In certain embodiments, the rate of homozygous editing / biallelic editing is increased using Cas12a protein or fusion protein provided herein.

[0152] The method herein includes providing Cas12a protein and / or fusion protein and nucleic acid to be edited, and may also include providing at least one guide RNA. Any suitable technology can be used to provide these various components. For example, providing Cas12a protein or fusion protein may include introducing Cas12a protein or fusion protein into a cell or introducing a recombinant nucleic acid, construct or vector encoding Cas12a protein or fusion protein into a cell. Similarly, gRNA can be provided by introducing gRNA itself or a nucleic acid sequence encoding gRNA. In certain embodiments, Cas12a protein and / or fusion protein and gRNA can be encoded by the same DNA construct or vector. Examples Example 1. C965S and D156R synergize to improve SDN1 editing efficiency at difficult target sites in maize

[0153] By analyzing the crystal structure of LbCas12a (PDB entry 5XUS), two surface-exposed cysteine residues (Cys965 and Cys1090) and another cysteine residue near the N-terminus (Cys10) of the native LbCas12a protein (SEQ ID NO: 1) were selected for site-directed mutagenesis ( Figure 1 A total of five LbCas12a variants were generated: the first variant mutated the Cys965 residue to a serine residue and was named LbCas12a-C965S; the second variant mutated both the Cys10 and Cys965 residues to serine residues and was named LbCas12a-C10S-C965S; the third variant mutated both the Cys965 and Cys1090 residues to serine residues and was named LbCas12a-C965S-C1090S; the fourth variant mutated the Cys965 residue to a serine residue and the Asp156 residue to an arginine residue and was named LbCas12a-D156R-C965S; the fifth variant mutated only the Asp156 residue to an arginine residue and was named LbCas12a-D156R. The coding sequence of LbCas12a was optimized based on maize preferred codon usage; By introducing mutations in overlapping PCR primers, the codon triplet TGC of the selected cysteine residues was mutated to TCC encoding serine. For all five variants plus wild-type LbCas12a as a control, SV40 NLS (SEQ ID NO: 56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS) 6 (SEQ ID NO: 46) peptide linker, while two SV40 NLSs separated by 8-amino acid (SGGS) 2 (SEQ ID NO: 78) peptide linkers were also fused to the C-terminus via a flexible 30-amino acid (GSSSS) 6 (SEQ ID NO: 46) peptide linker.

[0154] For each of the five variants plus a wild-type control, a binary vector was constructed to express a variant (or control) in a stable transgenic corn plant to evaluate the SDN1 production performance of the variant. In each construct, the coding sequence of a variant fused to NLS was operably connected to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in corn cells. In all constructs, the same gRNA array driven by the rice (Oryza sativa) U6 promoter was designed to express gRNA targeting the corn gene starch branching enzyme IIb (ZmSBEIIb). The gRNA is based on the mature crRNA scaffold of LbCas12a. Transgenic corn plants were generated in the following manner: callus derived from immature corn embryos was infected with an Agrobacterium tumefaciens strain containing one of the above-mentioned binary vectors, followed by a tissue culture procedure.

[0155] Leaf sheaths of regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan qPCR assay. Sequences spanning the target site were PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target site. As summarized in Tables 6 and 7, SDN1 efficiency at the SBEIIb target site and the Wx1 target site were compared. The SBEIIb target site is difficult to edit with Cas12a. Compared to the wild type, C965S alone slightly improved the overall SDN1 efficiency of SBEIIb but did not improve the rate of homozygous editing / biallelic editing. Neither C10 nor C1090S showed a positive effect on C965S. In contrast, when paired with D156R, C965S increased overall SDN1 efficiency by 4-fold compared to wild type, with more than half of these edits occurring as homozygous or biallelic edits. In contrast, D156R alone increased SDN1 efficiency at the ZmSBEIIb target site by 3-fold, with approximately half of these edits occurring as homozygous or biallelic edits. Regarding Wx1, SDN1 editing efficiency was similar to wild type (only slightly improved). Table 6. SDN1 editing efficiency of LbCas12a variants in maize. Table 7. SDN1 editing efficiency of LbCas12a variants in maize. Example 2. C965S and D156R act synergistically to improve SDN1 editing efficiency in soybean

[0156] To evaluate the efficacy of C965S in improving the SDN1 induction efficiency of LbCas12a in soybean, the SDN1 production performance of two LbCas12a variants, LbCas12a-D156R and LbCas12a-D156R-C965S, was compared. Both variants were identical to those tested in corn as described in Example 1, except that the coding sequence was optimized based on the preferred codon usage of Arabidopsis thaliana.

[0157] For each variant, two binary vectors were constructed to test the SDN1 efficiency at different target loci. In each construct, the coding sequence of a variant fused to NLS is operably linked to a promoter (e.g., Arabidopsis elongation factor 1α (EF1α) promoter) and a terminator (e.g., Agrobacterium tumefaciens nopaline synthase gene terminator) for strong constitutive expression in soybean cells. In all constructs, gRNA or gRNA arrays driven by the soybean ubiquitin 1 promoter are designed to express gRNA targeting soybean genomic sites, such as FAD2 (SEQ ID NO: 38 provides LbCas12a gRNA targeting soybean FAD2-1A gene). gRNA is based on the mature crRNA scaffold of LbCas12a and is processed by the self-cleaving ribozyme on the flank. Transgenic soybean plants are generated by infecting mature soybean seeds with an Agrobacterium tumefaciens strain containing one of the above-mentioned binary vectors, followed by a tissue culture procedure.

[0158] Leaves of regenerated plantlets will be sampled for DNA extraction, and transgenic plants will be identified by TaqMan assay.Sequences spanning each of the three target sites will be PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target site. Example 3. Generation and identification of FnCas12a cysteine substitution variants with enhanced plant SDN1 editing efficiency.

[0159] There are a total of nine cysteine residues (Cys70, Cys473, Cys568, Cys717, Cys882, Cys1086, Cys1116, Cys1190, and Cys1196) in the FnCas12a primary sequence (SEQ ID NO: 2). The crystal structure of FnCas12a (PDB entries 5NFV and 6I1K) indicates that four cysteine residues (Cys70, Cys473, Cys1116, and Cys1190) are most likely surface exposed and therefore may be prone to undesirable interactions and / or modifications. The surface topology around Cys473 suggests that interacting proteins or modifying enzymes are difficult to access, while our PyMOL analysis suggests that Cys1116 and Cys1190 may form intramolecular disulfide bonds ( Figure 2). Therefore, Cys70, Cys1116 and Cys1190 were selected for substitution.

[0160] All cysteine-substituted variants were generated based on the FnCas12a-E184R variant. Three variants carrying a single Cys to Ser substitution (FnCas12a-E184R-C70S, FnCas12a-E184R-C1116S, FnCas12a-E184R-C1190S) and three variants carrying a double Cys to Ser substitution (FnCas12a-E184R-C70S-C1116S, FnCas12a-E184R-C70S-C1190S, FnCas12a-E184R-C1116S-C1190S). The coding sequence of FnCas12a was optimized based on maize preferred codon usage; by introducing mutations in overlapping PCR primers, the codon triplet TGC of the selected cysteine residues was mutated: for C1116 site mutation to TCC encoding serine, for C70 and C1190 mutation to AGC. For all six variants plus FnCas12a-E184R as a control, SV40 NLS (SEQ ID NO: 56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS) 6 (SEQ ID NO: 46) peptide linker, while two SV40 NLS separated by an 8-amino acid (SGGS) 2 (SEQ ID NO: 78) peptide linker were also fused to the C-terminus via a flexible 30-amino acid (GSSSS) 6 (SEQ ID NO: 46).

[0161] For each of the six variants plus the FnCas12a-E184R control, a binary vector was constructed to express a variant (or control) in a stable transgenic corn plant to evaluate the SDN1 production performance of the variant. In each construct, the coding sequence of a variant fused to NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in corn cells. In all constructs, the same gRNA array driven by the rice U6 promoter was designed to express three gRNAs targeting three different corn genes: Waxy1 (ZmWx1), Glossy2 (ZmGL2) and starch branching enzyme IIb (ZmSBEIIb). gRNA is based on the mature crRNA scaffold of FnCas12a. Transgenic corn plants will be generated by infecting callus derived from immature corn embryos with an Agrobacterium tumefaciens strain containing one of the above-mentioned binary vectors, followed by a tissue culture procedure.

[0162] Leaf sheaths of regenerated plantlets will be sampled for DNA extraction, and transgenic plants will be identified by TaqMan assay. Sequences spanning each of the three target sites will be PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target site. The overall SDN1 editing efficiency and the rates of homozygous / biallelic mutants for each variant will be compared with those of the FnCas12a-E184R control to assess the efficacy of the cysteine substitution. Example 4. Generation and identification of AsCas12a cysteine substitution variants with enhanced plant SDN1 editing efficiency.

[0163] There are a total of 8 cysteine residues (Cys65, Cys205, Cys334, Cys379, Cys608, Cys674, Cys1025 and Cys1248) in the AsCas12a primary sequence (SEQ ID NO: 3). The crystal structure of AsCas12a (PDB entry 5KK5) shows that three cysteine residues: Cys334, Cys379 and Cys674 are most likely surface-exposed, and are therefore prone to undesirable interactions and / or modifications. These three residues are selected for substitution.

[0164] All cysteine-substituted variants are generated based on the AsCas12a-E174R variant. Three variants carrying single Cys to Ser substitutions (AsCas12a-E174R-C334S, AsCas12a-E174R-C379S, AsCas12a-E174R-C674S) and three variants carrying dual Cys to Ser substitutions (AsCas12a-E174R-C334S-C379S, AsCas12a-E174R-C334S-C674S, AsCas12a-E174R-C379S-C674S) were generated. The coding sequence of AsCas12a is optimized based on the preferred codon usage of corn; By introducing mutations in overlapping PCR primers, the codon triplet TGC of the selected cysteine residue is mutated to TCC encoding serine. For all six variants plus AsCas12a-E174R as a control, the SV40 NLS (SEQ ID NO: 56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO: 46) peptide linker, while two SV40 NLSs separated by an 8-amino acid (SGGS)2 (SEQ ID NO: 78) peptide linker were also fused to the C-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO: 46) peptide linker.

[0165] For each of the six variants plus an AsCas12a-E174R control, a binary vector was constructed to express a variant (or control) in a stable transgenic corn plant to evaluate the SDN1 production performance of the variant. In each construct, the coding sequence of a variant fused to NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in corn cells. In all constructs, the same gRNA array driven by the rice U6 promoter was designed to express three gRNAs targeting three different corn genes: Waxy1 (ZmWx1), Glossy2 (ZmGL2) and starch branching enzyme IIb (ZmSBEIIb). gRNA is based on the mature crRNA scaffold of AsCas12a. Transgenic corn plants will be generated by infecting callus derived from immature corn embryos with an Agrobacterium tumefaciens strain containing one of the above-mentioned binary vectors, followed by a tissue culture procedure.

[0166] Leaf sheaths of regenerated plantlets will be sampled for DNA extraction, and transgenic plants will be identified by TaqMan assay. Sequences spanning each of the three target sites will be PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target site. The overall SDN1 editing efficiency and the rates of homozygous / biallelic mutants for each variant will be compared with those of the AsCas12a-E174R control to assess the efficacy of the cysteine substitution. Example 5. Generation and identification of Mb2Cas12a cysteine substitution variants with enhanced plant SDN1 editing efficiency.

[0167] Because there is no published crystal structure of Mb2Cas12a (from M. bovis strain 57922) to date, the crystal structure of MbCas12a from M. bovis strain 22581 (PDB entry 6IV6), the closest ortholog with 94.7% amino acid identity to Mb2Cas12a, was used as a reference structure to estimate the positions of cysteine residues in Mb2Cas12a. There are a total of 8 cysteine residues (Cys270, Cys307, Cys583, Cys662, Cys1068, Cys1099, Cys1149, and Cys1162) in the primary sequence of Mb2Cas12a from strain 57922 (SEQ ID NO: 4), which correspond to Cys283, Cys320, Cys593, Cys672, Cys1078, Cys1109, Cys1159, and Tyr1172 in Mb2Cas12a from strain 22581, respectively. This estimate suggests that Cys270, Cys307, Cys583, Cys1068, Cys1099, Cys1149, and Cys1162 may be exposed on the surface of Mb2Cas12a. Since Cys1162 of Mb2Cas12a is aligned with Tyr1172 in MbCas12a, Tyr1172 was mutated in 6IV6, and the structure was reshaped using PyMOL. The resulting structural model suggests that Cys1162 may also be surface exposed in Mb2Cas12a. However, the surface topology indicates that Cys1162 is difficult to be approached by interacting proteins or modifying enzymes. Therefore, Cys270, Cys583, Cys1068, Cys1099, and Cys1149 were selected for site-directed mutagenesis.

[0168] All cysteine substituted variants were generated based on the Mb2Cas12a-D172R variant, which was a control for the new variants. Five variants carrying single Cys to Ser substitutions (Mb2Cas12a-D172R-C270S, Mb2Cas12a-D172R-C583S, Mb2Cas12a-D172R-C1068S, Mb2Cas12a-D172R-C1099S, Mb2Cas12a-D172R-C1149S), one variant carrying a quintuple Cys to Ser substitution (Mb2Cas12a-D172R-C270S-C583S-C1068S-C1099S-C1149S), and one variant carrying a quintuple Cys to Ala substitution (Mb2Cas12a-D172R-C270A-C583A-C1068A-C1099A-C1149A) were generated. The coding sequence of Mb2Cas12a is optimized based on corn preferred codon usage; By introducing mutations in overlapping PCR primers of five single mutation variants, the codon triplet TGC of the selected cysteine residue is mutated to TCC encoding serine. For variants with five mutations, Mb2Cas12a is synthesized by introducing TCC encoding serine instead of TGC or GCC encoding alanine. For all seven variants plus Mb2Cas12a-D172R as a control, SV40 NLS (SEQ ID NO: 56) is fused to N-terminus via a flexible 30-amino acid (GSSSS) 6 (SEQ ID NO: 46) peptide linker, while two SV40 NLS separated by 8-amino acid (SGGS) 2 (SEQ ID NO: 78) peptide linkers are also fused to C-terminus via a flexible 30-amino acid (GSSSS) 6 (SEQ ID NO: 46) peptide linker.

[0169] For each of the six variants plus the Mb2Cas12a-D172R control, a binary vector was constructed to express one variant (or control) in a stable transgenic corn plant to evaluate the SDN1 production performance of the variant. In each construct, the coding sequence of a variant fused to NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in corn cells. In all constructs, the same gRNA array driven by the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator was designed to express four gRNAs targeting four different corn genes: Waxy1 (ZmWx1), benzoxazinone synthesis 9 (ZmBx9), Glossy2 (ZmGL2) and ZmBINa. The gRNA is based on the mature crRNA scaffold of LbCas12a and is processed by the self-cleaving ribozyme on the flank. Transgenic maize plants were generated by infecting callus tissue derived from immature maize embryos with an Agrobacterium tumefaciens strain containing one of the above binary vectors, followed by tissue culture procedures.

[0170] Leaf sheaths of regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan assay. Sequences spanning each of the three target sites were PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target site. As summarized in Table 8, all variants with a single Cys to Ser mutation increased the rate of homozygous mutants / biallelic mutants compared to the Mb2Cas12a-D172R control. The efficacy of stacking five cysteine mutations will be similarly determined. Table 8. SDN1 editing efficiency of Mb2Cas12a variants in maize. Reference sequence list SEQ ID NO: 1-Lachnospiraceae bacterial Cas12a protein (LbCas12a) SEQ ID NO: 2—Franciscoides novicida U112 Cas12a protein (FnCas12a) SEQ ID NO: 3 - Aminoacillus sp. Cas12a protein (AsCas12a) SEQ ID NO:4—Moraxella bovis strain 57922 Cas12a protein (Mb2Cas12a) SEQ ID NO:5 – Amino acid sequence of LbCas12a+linker: SEQ ID NO:8—Amino acid sequence of LbCas12a+C10S+C965S: SEQ ID NO:9—Amino acid sequence of LbCas12a+C965S+C1090S: SEQ ID NO: 10—Amino acid sequence of LbCas12a+linker+D156R: SEQ ID NO: 13 - Amino acid sequence of Mb2Cas12a + linker + D172R + C270S: SEQ ID NO: 15 - Amino acid sequence of Mb2Cas12a + linker + D172R + C1068S: SEQ ID NO: 16 - Amino acid sequence of Mb2Cas12a + linker + D172R + C1099S: SEQ ID NO: 17 - Amino acid sequence of Mb2Cas12a + linker + D172R + C1149S: SEQ ID NO: 18 - Amino acid sequence of Mb2Cas12a + linker + D172R + C270S + C583S + C1068S + C1099S + C1149S: SEQ ID NO: 19 - Amino acid sequence of Mb2Cas12a + linker + D172R + C270A + C583A + C1068A + C1099A + C1149A: SEQ ID NO: 20—Nucleic acid sequence encoding LbCas12a+linker, maize codon optimized: SEQ ID NO: 21 - Nucleic acid sequence encoding LbCas12a + linker + D156R, maize codon optimized: SEQ ID NO: 22—Nucleic acid sequence encoding LbCas12a+linker+D156R+C965S, maize codon optimized: SEQ ID NO: 23 - Nucleic acid sequence encoding LbCas12a + linker + C10S + C965S, maize codon optimized: SEQ ID NO: 24—Nucleic acid sequence encoding LbCas12a+linker+C965S+C1090S, maize codon optimized: SEQ ID NO:25—Nucleic acid sequence encoding LbCas12a+linker+D156R, Arabidopsis thaliana codon-optimized: SEQ ID NO: 26—Nucleic acid sequence encoding LbCas12a+linker+D156R+C965S, Arabidopsis thaliana codon-optimized: SEQ ID NO:27—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R, maize codon optimized: SEQ ID NO:28—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C270S, maize codon optimized: SEQ ID NO:29—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C583S, maize codon optimized: SEQ ID NO:30—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C1068S, maize codon optimized: SEQ ID NO:31 - Nucleic acid sequence encoding Mb2Cas12a + linker + D172R + C1099S, maize codon optimized: SEQ ID NO:32—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C1149S, maize codon optimized: SEQ ID NO:33—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C270S+C583S+C1068S+C1099S+C1149S, maize codon optimized: SEQ ID NO:34—Nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C270A+C583A+C1068A+C1099A+C1149A, maize codon optimized:

[0171] All patents, patent publications, patent applications, journal articles, books, technical references, etc. discussed in this disclosure are incorporated herein by reference in their entirety for all purposes.

[0172] It should be understood that the figures and descriptions of the present disclosure have been simplified to illustrate elements relevant to a clear understanding of the present disclosure. It should be understood that the figures are presented for illustrative purposes and not as structural diagrams. Omitted details and modifications or alternative embodiments are within the knowledge of those of ordinary skill in the art.

[0173] It is understood that in certain aspects of the present disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component to provide an element or structure or to perform a given function or functions. Unless such a replacement would be inoperable to implement certain embodiments of the present disclosure, such a replacement is considered to be within the scope of the present disclosure.

[0174] The examples presented herein are intended to illustrate the potential and specific implementation of the present disclosure. It will be appreciated that these examples are primarily intended to illustrate the purpose of the present disclosure to those skilled in the art. Without departing from the spirit of the present disclosure, these figures or the operations described herein may be changed. For example, in some cases, method steps or operations may be performed or executed in different orders, or operations may be added, deleted, or modified.

[0175] Where a range of values is provided, it is understood that each intervening value between the upper and lower limits of that range (to the nearest decimal place of the lower limit unless the context clearly dictates otherwise) is also specifically disclosed. Any smaller ranges between any stated value or non-stated intervening value in a stated range and any other stated value or intervening value in that stated range are encompassed. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range in which either, neither, or both of the limits are included in the smaller range is also encompassed within the present technology, subject to any specifically excluded limits in the stated range. Where a stated range includes one or both of the limits, ranges excluding one or both of those included limits are also included.

[0176] In the foregoing description, many specific details have been set forth to provide a more thorough understanding of the present invention. However, it will be clear to those skilled in the art that the invention described in this disclosure can be practiced without one or more of these specific details. In other cases, features and procedures well known to those skilled in the art are not described to avoid obscuring the present invention. The embodiments of the present disclosure have been described for illustrative and not restrictive purposes. Although the present invention is primarily described with reference to specific embodiments, other embodiments that will become clear to those skilled in the art upon reading this disclosure are also contemplated, and such embodiments are intended to be included within the inventive method. Therefore, the present disclosure is not limited to the embodiments described above or depicted in the accompanying drawings, and various embodiments and modifications may be made without departing from the scope of the following claims.

Claims

1. A Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO: 1 and an artificially induced mutation at position C965.

2. Cas12a albumen as claimed in claim 1, wherein the artificial induction mutation is the replacement of cysteine to serine.

3. Cas12a albumen as claimed in claim 1 or 2, is further included in the artificial induction mutation at position D156.

4. Cas12a albumen as claimed in claim 3, wherein the artificial induction mutation at position D156 is the replacement of aspartic acid to arginine.

5. Cas12a protein as described in any one in claim 1 to 4, wherein the sequence includes any one of SEQ ID NO:5-11.

6. A Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO: 2 and artificially induced mutations at positions C70, C1116 and / or C1190.

7. Cas12a albumen as claimed in claim 6, wherein the artificial induction mutation is the replacement of cysteine to serine.

8. Cas12a albumen as claimed in claim 6 or 7, it is further included in the artificial induction mutation at position E184.

9. Cas12a albumen as claimed in claim 8, wherein the artificial induction mutation at position E184 is the replacement of glutamic acid to arginine.

10. A Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO: 3 and artificially induced mutations at positions C334, C379 and / or C674.

11. Cas12a albumen as claimed in claim 10, wherein the artificial induction mutation is the replacement of cysteine to serine.

12. Cas12a albumen as claimed in claim 10 or 11, it is further included in the artificial induction mutation at position E174.

13. Cas12a albumen as claimed in claim 12, wherein the artificial induction mutation at position E174 is the replacement of glutamic acid to arginine.

14. A Cas12a protein comprising a sequence having at least 80% identity to the amino acid sequence of SEQ ID NO: 4 and artificially induced mutations at positions C270, C583, C1068, C1099 and / or C1149.

15. Cas12a albumen as claimed in claim 14, wherein the artificial induction mutation is the replacement of cysteine to serine.

16. Cas12a proteins as described in claim 14 or 15, are further included in the artificial induction mutation at position D172.

17. Cas12a proteins as claimed in claim 16, wherein the artificial mutation induced at position D172 is the replacement of aspartic acid to arginine.

18. The Cas12a protein of any one of claims 14 to 17, wherein the sequence comprises any one of SEQ ID NO: 12-19.

19. Cas12a protein as described in any one of claims 1 to 18, wherein the Cas12a protein is a catalytically inactivated Cas12a (dCas12a) protein of nickase Cas12a (nCas12a) protein.

20. Cas12a albumen as described in any one in claim 1 to 19, further comprising a nuclear localization signal.

21. a fusion protein comprising the Cas12a protein and a heterologous domain as described in any one of claims 1 to 20.

22. The fusion protein of claim 21, wherein the heterologous domain is a deaminase domain, a transcription factor domain, a nuclease domain, a reverse transcriptase domain, a transposase domain, an integrase domain, a uracil DNA glycosylase inhibitor domain, a recombinase domain, a nickase domain, a methyltransferase domain, a methylase domain, an acetylase domain, an acetyltransferase domain, a transcription activator domain, or a transcription repressor domain.

23. fusion proteins as described in claim 21 or 22, wherein the Cas12a albumen is connected with the heterologous domain by a linker sequence.

24. a nucleic acid encoding the Cas12a protein as described in any one of claims 1 to 20 or the fusion protein as described in any one of claims 21 to 23.

25. The nucleic acid of claim 24, wherein the nucleic acid sequence is any one of SEQ ID NOs: 20-34.

26. A DNA construct comprising a promoter operably linked to the nucleic acid of claim 24 or 25.

27. A vector comprising the nucleic acid of claim 24 or 25 or the DNA construct of claim 26.

28. A cell comprising the nucleic acid of claim 24, the DNA construct of claim 26 or the vector of claim 27.

29. The cell of claim 28, wherein the cell is a plant cell.

30. The cell of claim 29, wherein the cell is a corn plant cell, a wheat plant cell, a rice plant cell, a soybean plant cell, a sunflower plant cell, or a tomato plant cell.

31. A method for editing nucleic acid, the method comprising: The nucleic acid is contacted with (i) a Cas12a protein as described in any one of claims 1 to 20 or a fusion protein as described in any one of claims 21 to 23, and (ii) a guide RNA having a region complementary to a selected portion of the nucleic acid, thereby editing the nucleic acid.

Citation Information

Patent Citations

  • Simultaneous gene editing and haploid induction

    US10519456B2

  • Optimized protein linkers and methods of use

    US20210017506A1

  • Method for transporting substances into living cells and tissues and apparatus therefor

    US4945050A

  • Method for transporting substances into living cells and tissues and apparatus therefor

    US5036006A

  • Method for transporting substances into living cells and tissues

    US5100792A