Variants of CPF1 (CAS12A) with improved activity

JP2026503682A5Pending Publication Date: 2026-02-05SYNGENTA CROP PROTECITON AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025543308
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-01-27
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing site-specific nucleases, such as CRISPR/Cas systems, face challenges in achieving efficient and precise genome editing due to issues with cysteine residues that can form undesired disulfide bonds and affect protein function, leading to suboptimal editing outcomes.

Method used

Mutating surface-exposed cysteine residues in Cas12a proteins to serine or arginine, and incorporating heterologous domains, enhances the editing efficiency and specificity of Cas12a proteins, particularly in plant cells.

Benefits of technology

The modified Cas12a proteins demonstrate increased frequency and precision in site-specific nucleic acid editing, especially at challenging genomic sites, facilitating improved genetic modifications in plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000082_0000
    Figure 00000082_0000
  • Figure 00000082_0001
    Figure 00000082_0001
  • Figure 00000082_0002
    Figure 00000082_0002
Patent Text Reader

Abstract

Provided herein are variant Cas12a proteins containing at least one human-induced mutation. Also provided are fusion proteins containing the variant Cas12a proteins and one or more heterologous domains. Also provided are related nucleic acids, DNA constructs, vectors, cells, and methods for editing nucleic acids using the variant Cas12a proteins and / or fusion proteins. Use of the provided proteins can increase the frequency of desired nucleic acid editing (e.g., SDN-1 editing in plant genomes).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to methods for increasing site-specific nuclease editing.

[0002] Reference to sequence listing submitted as an XML file This application is accompanied by a Sequence Listing entitled 82447-SL.xml, created on January 19, 2023, which is approximately 149 kilobytes in size. This Sequence Listing is incorporated herein by reference in its entirety. [Background technology]

[0003] Site-specific nucleases (SDNs) (e.g., zinc finger nucleases, transcription activator-like effector nucleases, and CRISPR-associated nucleases) are becoming increasingly popular in the gene editing space. These SDNs act as endonucleases, generally creating double-strand breaks (DSBs) at specific DNA sequences, thereby activating the cell's intrinsic repair mechanisms (e.g., homologous recombination). The repair process can achieve site-specific modifications to the specific DNA sequence. The CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-associated) system evolved in bacteria and archaea as an adaptive immune system to defend against viral attacks. In recent years, the CRISPR / Cas system has attracted particular attention as a genome editing tool. The CRISPR / Cas system, which generates site-specific double-strand breaks (DSBs), can be used to edit the DNA of eukaryotic cells, for example, by making deletions, insertions, and / or changes in nucleotide sequences. Summary of the Invention [Means for solving the problem]

[0004] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in limiting the scope of the claimed subject matter.

[0005] In one aspect, a Cas12a protein is provided that comprises an amino acid sequence at least 80% identical to SEQ ID NO: 1 and a human-induced mutation at position C965. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position D156. In some embodiments, the human-induced mutation at position D156 is an aspartic acid to arginine substitution. In some embodiments, the sequence of the Cas12a protein comprises any one of SEQ ID NOs: 5-11.

[0006] In another aspect, a Cas12a protein is provided that comprises an amino acid sequence at least 80% identical to SEQ ID NO:2 and a human-induced mutation at positions C70, C1116, and / or C1190. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position E184. In some embodiments, the human-induced mutation at position E184 is a glutamic acid to arginine substitution.

[0007] In another aspect, a Cas12a protein is provided that comprises an amino acid sequence at least 80% identical to SEQ ID NO: 3 and a human-induced mutation at positions C334, C379, and / or C674. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position E174. In some embodiments, the human-induced mutation at position E174 is a glutamic acid to arginine substitution.

[0008] In another aspect, provided is a Cas12a protein comprising a sequence at least 80% identical to the amino acid sequence of SEQ ID NO: 4 and a human-induced mutation at positions C270, C583, C1068, C1099, and / or C1149. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position D172. In some embodiments, the human-induced mutation at position D172 is an aspartic acid to arginine substitution. In some embodiments, the sequence of the Cas12a protein comprises any one of SEQ ID NOs: 12-19.

[0009] In some embodiments of any of the above Cas12a proteins, the Cas12a protein is a catalytically inactivated Cas12a (dCas12a) protein of a nickase Cas12a (nCas12a) protein.

[0010] In some embodiments of any of the above Cas12a proteins, the Cas12a protein further comprises a nuclear localization signal.

[0011] In another aspect, a fusion protein is provided comprising any of the above Cas12a proteins and a heterologous domain.

[0012] In some embodiments, the heterologous domain is a deaminase domain, a transcription factor domain, a nuclease domain, a reverse transcriptase domain, a transposase domain, an integrase domain, a uracil DNA glycosylase inhibitor domain, a recombinase domain, a nickase domain, a methyltransferase domain, a methylase domain, an acetylase domain, an acetyltransferase domain, a transcriptional activator domain, or a transcriptional repressor domain.

[0013] In some embodiments of the fusion protein, the Cas12a protein is linked to the heterologous domain by a linker sequence.

[0014] In another aspect, a nucleic acid encoding any of the above-described Cas12a proteins or fusion proteins is provided. In some embodiments, the nucleic acid sequence is any one of SEQ ID NOs: 20-34.

[0015] In another aspect, a DNA construct is provided that includes a promoter operably linked to a nucleic acid encoding any of the above-described Cas12a proteins or fusion proteins.

[0016] In another aspect, there is provided a vector comprising the above nucleic acid or DNA construct.

[0017] In another aspect, a cell is provided comprising the nucleic acid, DNA construct, or vector described above. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a corn plant cell, a wheat plant cell, a rice plant cell, a soybean plant cell, a sunflower plant cell, or a tomato plant cell.

[0018] In another aspect, a method of editing a nucleic acid is provided, comprising contacting the nucleic acid with (i) any one of the above-described Cas12a proteins or any one of the above-described fusion proteins, and (ii) a guide RNA having a region complementary to a selected portion of the nucleic acid, thereby effecting editing of the nucleic acid.

[0019] This application includes the following figures. These figures are intended to illustrate certain embodiments and / or features of the present compositions and methods and to supplement any one or more descriptions of the present compositions and methods. These figures are not intended to limit the scope of the present compositions and methods, except where expressly indicated to the contrary by the specification. [Brief explanation of the drawings]

[0020] [Figure 1]Cysteine ​​residues in LbCas12a potentially form intermolecular or intramolecular interactions. Left: PyMOL surface model of the LbCas12a-crRNA-DNA ternary complex (PDB entry 5XUS). Highlighted areas indicated by arrows are the potentially surface-exposed thiol groups of C965 and C1090. Right: Four cysteine ​​residues (C10, C805, C912, and C965) scattered within the linear amino acid sequence form a cluster within the 3D structure of LbCas12a. [Figure 2] Figure 1 shows two of the cysteine ​​residues in FnCas12a selected for substitution according to embodiments of the present disclosure. The PyMOL stick model of C1190 and C1116 suggests that the thiol groups (black) are close to each other in the FnCas12a 3D structure (PDB entry 5NFV), potentially forming an intramolecular disulfide bond between them. DETAILED DESCRIPTION OF THE INVENTION

[0021] The following description describes various aspects and embodiments of the present compositions and methods. The specific embodiments are not intended to define the scope of the present compositions and methods. Rather, the embodiments merely provide non-limiting examples of various compositions and methods that fall at least within the scope of the disclosed compositions and methods. The description should be read from the perspective of a person skilled in the art, and therefore may not necessarily include information that is known to a person skilled in the art.

[0022] I. Terminology All technical and scientific terms used herein are intended to have the same meaning as commonly understood by those skilled in the art, unless otherwise defined below. References to technology used herein are intended to refer to technology as commonly understood in the art, including variations of those technologies and / or equivalent technology alternatives that would be apparent to those skilled in the art. While the following terms are believed to be well understood by those skilled in the art, definitions are provided below to facilitate description of the subject matter of the present disclosure.

[0023] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an enzyme" includes, optionally, a combination of two or more such molecules, and the like.

[0024] As used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.

[0025] The term "about," as used herein, refers to the normal range of error for the respective value, readily known to one of ordinary skill in the art, e.g., ±20%, ±10%, or ±5% within the intended meaning of the recited value.

[0026] As used herein, the terms "comprising" or "comprise" are open-ended. When used in the context of a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or amino acid sequence) that includes the subject sequence as a portion or in its entirety.

[0027] As used herein, the transitional phrase "consisting essentially of" means that the claim should be construed to include within its scope the specified materials or steps recited in the claim and materials or steps that do not materially affect the basic novel characteristic or characteristics of the claimed subject matter. Thus, it is intended that the term "consisting essentially of," when used in the claims of this disclosure, should not be construed as the equivalent of "comprising."

[0028] The term "plurality" refers to more than one entity. Thus, "plurality of individuals" refers to at least two individuals. In some embodiments, the term "plurality" refers to more than half of a total. For example, in some embodiments, a "plurality of a population" refers to more than half of the members of the population.

[0029] The term "plant," as used herein, refers to any plant at any stage of development, particularly a seed plant. The term "plant cell," as used herein, refers to the structural and physiological unit of a plant, including a protoplast and a cell wall. A plant cell may be in the form of an isolated single cell or a cultured cell, or as part of a more highly organized unit, such as a plant tissue, a plant organ, or a whole plant. A plant cell may be derived from or part of an angiosperm or a gymnosperm. The plant cell may be a monocotyledonous plant cell (e.g., a corn cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turfgrass cell, or an ornamental herb cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, a eggplant cell, a sunflower cell, a cruciferous plant cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugarbeet cell, or a rapeseed cell. The term "plant cell culture," as used herein, refers to a culture of plant units at various developmental stages, such as, for example, protoplasts, cell culture cells, cells of plant tissue, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos. The term "plant tissue," as used herein, refers to a group of plant cells organized into a structural and functional unit. It includes any tissue of a plant, whether in planta or in culture. This term includes, but is not limited to, whole plants, plant organs, plant Plant tissues include seeds, tissue cultures, and any group of plant cells organized into a structural and / or functional unit. When used in conjunction with or without any specific type of plant tissue, as listed above or otherwise encompassed by this definition, this term is not intended to exclude any other type of plant tissue. The term "plant part," as used herein, refers to a part of a plant, including single cells and cell tissues, such as intact plant cells in a plant, cell clumps and tissue cultures that can regenerate plants. Examples of plant parts include, but are not limited to, pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, inflorescences, fruits, stems, shoots, cuttings, and seeds; and single cells and tissues from pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, inflorescences, fruits, stems, shoots, cuttings, scions, rootstocks, seeds, protoplasts, calluses, etc.

[0030] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, these terms encompass amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds.

[0031] The terms "nucleic acid" and "polynucleotide" are used interchangeably and, as used herein, refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in either single- or double-stranded form, as well as both sense and antisense strands of RNA, cDNA, genomic DNA, and mitochondrial DNA, as well as synthetic forms and mixed polymers of the above. In higher plants, DNA is the genetic material, while RNA is responsible for transferring the information contained within DNA to proteins. A "genome" is the entire body of genetic material contained in each cell of an organism. When RNA is described, it is understood that its corresponding cDNA is also described, and in cDNA, uridine is represented as thymidine. In certain embodiments, nucleotide refers to ribonucleotides, deoxynucleotides, or modified forms of either type of nucleotide, and combinations thereof. In addition, polynucleotides disclosed herein can include either or both naturally occurring and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide linkages. As will be readily understood by those skilled in the art, nucleic acid molecules may be chemically or biochemically modified or may contain non-natural or derivatized nucleotide bases. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications, such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, etc.), chelators, alkylating agents, and modified linkages (e.g., alpha-anomeric nucleic acids, etc.). The above terms are also intended to include any conformation of any topology, including single-stranded, double-stranded, partially duplexed, triplexed, hairpinned, circular, and padlock conformations. A reference to a nucleic acid sequence includes its complement unless otherwise specified.Thus, a reference to a nucleic acid molecule having a particular sequence should be understood to encompass its complementary strand, with its complementary sequence. Nucleotide sequences are "complementary" when they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules). The term also includes codon-optimized nucleic acids that encode the same polypeptide sequence. It is also understood that nucleic acids can be crude, purified, or attached to synthetic materials, such as beads or column matrices.

[0032] The term "corresponding to" in the context of nucleic acid sequences means that when the nucleic acid sequences of a given sequence are aligned with each other, those nucleic acids "corresponding" to a recited position in the present invention are aligned with those positions in the reference sequence, but not at their exact numerical positions relative to a particular nucleic acid sequence of the present invention. Optimal alignment of sequences for comparison can be performed by computerized implementation of known algorithms or by visual inspection. Ready-made sequence comparison and multiple sequence alignment algorithms are the Basic Local Alignment Search Tool (BLAST) and ClustalW / ClustalW2 / Clustal Omega programs, respectively, available on the Internet (e.g., the EMBL-EBI website). Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG package available from Accelrys, Inc., San Diego, Calif., United States of America. See also Smith & Waterman, 1981; Needleman & Wunsch, 1970; Pearson & Lipman, 1988; Ausubel et al., 1988; and Sambrook & Russell, 2001.

[0033] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences along with the explicitly indicated sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. See Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).

[0034] The terms "identity" or "substantial identity," when used in the context of polynucleotide or polypeptide sequences described herein, refer to a sequence having at least 60% sequence identity with a reference sequence. Alternatively, the percent identity may be any integer between 60% and 100%. Exemplary embodiments include at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% when compared to a reference sequence using the programs described herein; preferably, BLAST with standard parameters, as described below. Those skilled in the art will recognize that these values ​​can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences, taking into account codon degeneracy, amino acid similarity, reading frame alignment, and the like.

[0035] For sequence comparison, typically one sequence serves as a reference sequence, and test sequences are compared with it.When using sequence comparison algorithm, test sequences and reference sequences are input into computer, subsequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Default program parameters can be used, or alternative parameters can be designated.The sequence comparison algorithm then calculates the percent sequence identity of the test sequence to the reference sequence based on the program parameters.

[0036] As used herein, the term "comparison window" refers to any segment of a number of consecutive positions selected from the group consisting of 20 to 600, usually about 50 to about 200, more usually about 100 to about 150, where a sequence can be compared to a reference sequence of the same number of consecutive positions after the two sequences are optimally aligned. Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be achieved by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), the similarity search method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85:2444 (1988), computer implementations of these algorithms (e.g., BLAST), or by manual alignment and visual inspection.

[0037] Suitable algorithms for determining percent sequence identity and percent sequence similarity are the BLAST and BLAST 2.0 algorithms described in Altschul et al. (1990) J. Mol. Biol. 215:403-410 and Altschul et al. (1977) Nucleic Acids Res. 25:3389-3402, respectively. Software for performing BLAST analyses is publicly available through the website of the National Center for Biotechnology Information (NCBI). This algorithm involves first identifying high-scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence that, when aligned with words of the same length in a database sequence, match or meet some positive threshold score T. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as possible to increase the cumulative alignment score. For nucleotide sequences, the cumulative score is calculated using the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is ​​used to calculate the cumulative score. Extension of the word hits in each direction is stopped when the cumulative alignment score falls by an amount X from its maximum achieved value; the cumulative score falls below 0 due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as default a word size (W) of 28, an expectation (E) of 10, M=1, N=-2, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as default a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix.See Henikoff & Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989).

[0038] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences. See, e.g., Karlin & Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability that a match between two nucleotide or amino acid sequences would occur by chance. For example, a smallest sum probability of less than about 0.01, more preferably less than about 10, in a comparison of a test nucleic acid with a reference nucleic acid is preferred. -5 less than, most preferably about 10 -20 A nucleic acid is considered similar to a reference sequence if it is less than

[0039] "Recombination" is the exchange of DNA strands to create a new nucleotide sequence configuration. The term can also refer to the homologous recombination process that occurs in the repair of double-stranded DNA breaks, in which a polynucleotide is used as a template to repair a homologous polynucleotide. The term can also refer to the exchange of information between two homologous chromosomes during meiosis. The frequency of this double recombination is the product of the frequency of single recombinants. For example, recombinants in a 10 cM area can be found at a frequency of 10%, and double recombinants will be found at a frequency of 10% x 10% = 1% (1 centimorgan is defined as 1% recombinant progeny in a testcross).

[0040] A "gene" is a defined region located within a genome that contains, in addition to the aforementioned coding nucleic acid sequence, other primarily regulatory nucleic acid sequences involved in the expression of the coding portion, i.e., the control of transcription and translation. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein that contains regulatory sequences. A gene may or may not be usable to produce a functional protein. In some embodiments, a gene refers to only the coding region. The term "native gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene that 1) contains regulatory and coding sequences that are not found together in nature, or 2) contains sequences encoding portions of a protein that are not naturally contiguous, or 3) contains portions of a promoter that are not naturally contiguous. Thus, a chimeric gene may contain regulatory and coding sequences that are derived from different sources, or may contain regulatory and coding sequences that are derived from the same source but that are arranged in a manner that differs from that found in nature. A gene may also be "isolated," which refers to a nucleic acid molecule that is substantially or essentially free from components that are normally found associated with the nucleic acid molecule in nature. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in chemically synthesizing the nucleic acid molecule.

[0041] A "gene of interest" or "nucleotide sequence of interest" refers to any gene that, when transferred into a plant, confers a desired characteristic on the plant, such as antibiotic resistance, viral resistance, insect resistance, disease resistance, or resistance to other pests, herbicide resistance, improved nutritional value, improved performance in an industrial process, or altered reproductive ability. A "gene of interest" can also include a commercially valuable enzyme or metabolite that is transferred to the plant so that it is produced therein.

[0042] An "isolated" nucleic acid molecule or nucleotide sequence, or an "isolated" polypeptide, is a nucleic acid molecule, nucleotide sequence, or polypeptide that exists apart from its natural environment by the hand of man and / or has a different, modified, regulated, and / or altered function compared to its function in its natural environment, and is therefore not a product of nature. An isolated nucleic acid molecule or isolated polypeptide can exist in purified form or can exist in a non-native environment (e.g., a recombinant host cell). Thus, for example, with respect to a polynucleotide, the term isolated means that it is separated from the chromosome and / or cell in which it occurs in nature. A polynucleotide is also isolated if it is separated from the chromosome and / or cell in which it occurs in nature and then inserted into a genetic context, chromosome, chromosomal location, and / or cell that is not naturally occurring. Recombinant nucleic acid molecules and nucleotide sequences of the present invention can be considered "isolated" as defined above.

[0043] Thus, an "isolated nucleic acid molecule" or "isolated nucleotide sequence" is a nucleic acid molecule or nucleotide sequence that is not immediately adjacent to the nucleotide sequences to which it is immediately adjacent (one at the 5' end and one at the 3' end) in the naturally occurring genome of the organism from which it originates. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences immediately adjacent to the coding sequence. Thus, the term includes recombinant nucleic acids that are incorporated into, for example, a vector, an autonomously replicating plasmid, or virus, or the genomic DNA of a prokaryote or eukaryote, or that exist as a separate molecule independent of other sequences (e.g., a cDNA or genomic DNA fragment produced by PCR or restriction endonuclease treatment). It also includes recombinant nucleic acids that are part of a hybrid nucleic acid molecule that encodes an additional polypeptide or peptide sequence. An "isolated nucleic acid molecule" or "isolated nucleotide sequence" can also include nucleotide sequences that are derived from and inserted into the same natural cell type of origin, but that exist in a non-natural state, e.g., in different copy numbers and / or under the control of regulatory sequences that differ from those found in the nucleic acid molecule's natural state.

[0044] The term "isolated" can further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide, or fragment (e.g., when produced by recombinant DNA technology), or chemical precursor or other chemical (e.g., when chemically synthesized), that is substantially free of cellular material, viral material, and / or culture medium. Furthermore, an "isolated fragment" is a fragment of a nucleic acid molecule, nucleotide sequence, or polypeptide that is not naturally occurring as a fragment and, as such, would not be found in the natural state. "Isolated" does not necessarily mean that the preparation is technically pure (homogeneous), but is sufficiently pure to provide the polypeptide or nucleic acid in a form that can be used for its intended purpose.

[0045] "Homology-dependent repair" or "homologous recombination repair" or "HDR" refers to a mechanism for repairing ssDNA and double-stranded DNA (dsDNA) damage in cells. This repair mechanism can be used by cells when an HDR template with a sequence highly homologous to the damaged site is present. The term "complete HDR" refers to a situation in which the genomic homology junction in the replaced allele has undergone complete HDR, while "incomplete HDR" refers to a situation in which the genomic homology junction in the replaced allele has undergone partial or incomplete HDR. A donor DNA molecule with homology to the cut target DNA sequence is used as a template for repair of the cut target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. Thus, new nucleic acid material can be inserted / copied into the site. In some cases, the target DNA is contacted with a donor molecule, e.g., a donor DNA molecule. In some cases, the donor DNA molecule is introduced into the cell. In some cases, at least a segment of the donor DNA molecule is integrated into the cell's genome.

[0046] "Microhomology-mediated end joining" or "MMEJ" or "alternative non-homologous end joining" (Alt-NHEJ) refers to a form of double-strand break repair in DNA. This repair mechanism utilizes microhomology sequences to align the broken strands. "Non-homologous end joining" or "NHEJ" refers to a form of double-strand break repair in DNA. The double-strand break is repaired by directly ligating the broken ends to each other. Generally, there is no insertion of new nucleic acid material at the site, although some nucleic acid material may be lost or added, resulting in small deletions or small insertions.

[0047] As used herein, "heterologous" refers to a nucleic acid molecule, nucleotide sequence, polypeptide, or amino acid sequence that is not naturally associated with the host cell into which it is introduced, whether from another species or from the same species or organism but modified from its original form or the form primarily expressed in the cell (including non-naturally occurring multiple copies of a naturally occurring nucleic acid sequence). Thus, an amino acid sequence derived from an organism or species different from that of the cell into which it is introduced is heterologous with respect to that cell and its progeny. In addition, heterologous sequences include sequences derived from and inserted into the same native cell type of origin, but present in a non-native state (e.g., present in a different copy number) and / or under the control of regulatory sequences different from those found in the polypeptide's native state. A sequence can also be heterologous to other sequences with which it may be associated, for example, in a nucleic acid construct such as an expression vector. As a non-limiting example, a promoter can be present in a nucleic acid construct in combination with one or more regulatory elements and / or coding sequences that are not naturally occurring in association with that particular promoter, i.e., heterologous to the promoter.

[0048] II. Introduction In some embodiments, variant Cas12a proteins with increased site-specific nuclease (SDN) genome editing activity are provided herein. Site-specific nuclease technology has dramatically improved the speed and precision with which genome editing can be performed in various organisms, including plants. Generally, the desired outcome of SDN-mediated genome editing is 1) targeting the SDN to cleave DNA at a specific genomic site in a host (e.g., a plant cell) and 2) using the host's natural repair mechanisms to introduce specific genomic changes at the cleavage site. Changes can include small deletions, substitutions, or additions of several nucleotides. Such targeted editing can result in new and desirable traits (e.g., enhanced nutrient uptake, reduced allergen production) and / or reduced undesirable traits (e.g., herbicide sensitivity). SDN applications generally fall into three categories: SDN-1, SDN-2, and SDN-3. SDN-1 generates double-strand breaks in the genome without the addition of exogenous DNA. When such breaks are repaired by the host (e.g., via NHEJ), mutations or deletions can be introduced. If these mutations or deletions are in a gene, the gene can be silenced or knocked out. SDN-2 uses template DNA to introduce predicted repair (e.g., via HDR) at the target break site, but does not result in the insertion of recombinant DNA. SDN-3 also uses template DNA to introduce recombinant or exogenous DNA templates (e.g., transgenes) at the target break site.

[0049] Cas12a is a CRISPR-associated (Cas) SDN that functions in the CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas system. In bacteria, this system can provide adaptive immunity to foreign DNA (Barrangou, R., et al., "CRISPR provides acquired resistance against viruses in prokaryotes," Science (2007) 315:1709-1712; Makarova, K.S., et al., "Evolution and classification of the CRISPR-Cas systems," Nat Rev Microbiol (2011) 9:467-477; Garneau, J.E., et al., "The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA," Nature (2010) 468:67-71; Sapranauskas, R., et al., "The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli," Nucleic Acids Res (2011) 39:9275-9282). CRISPR / Cas systems (e.g., modified and / or unmodified) can be utilized as genome engineering tools in a wide variety of organisms, including various mammals, animals, plants, microorganisms, and yeast. CRISPR / Cas systems can include a guide nucleic acid, such as a guide RNA (gRNA), complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid editing. RNA-guided Cas proteins (e.g., Cas nucleases, such as Cas9 nuclease) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner.Cas proteins can cleave DNA when they possess nuclease activity (Gasiunas, G., et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109:E2579-E286; Jinek, M., et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337:816-821; Sternberg, SH, et al., “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9,” Nature (2014) 507:62; Deltcheva, E., et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature (2011) 471:602-607). DNA breaks (e.g., double-strand breaks) can result in DNA break repair, allowing for the introduction of one or more genetic modifications (e.g., nucleic acid editing).

[0050] Cysteine ​​residues are highly reactive residues that undergo post-translational modification. The formation of undesired disulfide bonds and / or modifications can affect the proper folding, localization, and / or enzymatic activity of proteins. Most Cas12a orthologs contain eight to nine cysteine ​​residues; in comparison, Cas9 from Streptococcus pyogenes (SpCas9) contains only two cysteine ​​residues. Most cysteine ​​residues in Cas12a orthologs are not conserved. Therefore, surface-exposed ones are more likely to be involved in intermolecular disulfide bond formation and / or post-translational modification. The conserved cysteine ​​residues in LbCas12a, FnCas12a, AsCas12a, and Mb2Cas12a are listed in Tables 1–4.

[0051] [Table 1]

[0052] [Table 2]

[0053] [Table 3]

[0054] [Table 4]

[0055] The present disclosure is based, in part, on the inventors' discovery that surface-exposed cysteine ​​residues in Cas12a can be mutated to improve the bioavailability of Cas12a proteins. Without being bound by any particular theory, such mutations likely avoid the undesirable modifications described above. Provided herein are variant Cas12a proteins containing at least one human-induced mutation. Also provided are fusion proteins comprising variant Cas12a proteins and one or more heterologous domains. Also provided are related nucleic acids, DNA constructs, vectors, cells, and methods of editing nucleic acids using variant Cas12a proteins and / or fusion proteins. In some embodiments, as demonstrated in the Examples herein, the provided methods result in an increased frequency of desired nucleic acid editing. In some embodiments, the editing is SDN-1 editing. In some embodiments, the increased frequency of desired nucleic acid editing is observed at difficult-to-edit genomic sites.

[0056] III. Variant Cas12a Proteins and Fusion Proteins In one aspect, provided herein is a variant Cas12a protein comprising at least one human-induced mutation that has enhanced function (i.e., compared to an unmodified Cas12a protein). Also provided is a fusion protein comprising the variant Cas12a protein and at least one heterologous domain. In some embodiments, the enhanced function of Cas12a is increased SDN-1 genome editing activity. In some embodiments, the variant Cas12a protein comprises a substitution of one or more surface-exposed cysteine ​​residues. In some embodiments, the variant Cas12a protein comprises a cysteine ​​to serine substitution at one or more surface-exposed cysteine ​​residues. In some embodiments, the variant Cas12a protein provided herein further comprises a substitution of an aspartic acid residue and / or a glutamic acid residue with an arginine residue.

[0057] Cas12a (also known as Cpf1) is a class II, type V CRISPR / Cas. The variant Cas12a proteins provided herein can be modified forms of Cas12a from any of a number of bacterial species, including, but not limited to, Lachnospiraceae bacteria, Acidaminococcus sp., Moraxella bovoculi, Thiomicrospira sp., Moraxella lacunata, Methanomethylophilus alvus, or Bacteroidetesoral sp. Unmodified Cas12a protein sequences include Lachnospiraceae bacterium Cas12a (LbCas12a; SEQ ID NO: 1), Francisella novicida U112 Cas12a (FnCas12a; SEQ ID NO: 2), Acidaminococcus sp. Cas12a (AsCas12a; SEQ ID NO: 3), and Moraxella bovoculi strain 57922 Cas12a (Mb2Cas12a; SEQ ID NO: 4).

[0058] In some embodiments, the variant Cas12a protein is a modified form of LbCas12a. In some embodiments, the variant Cas12a protein comprises an amino acid sequence at least 60% identical (e.g., at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to SEQ ID NO: 1 and at least one human-induced mutation. In some embodiments, the human-induced mutation is a substitution of a surface-exposed cysteine ​​residue. Surface-exposed cysteine ​​residues can be identified using methods known in the art, for example, by the methods described in the Examples herein. In some embodiments, one or more surface-exposed cysteine ​​residues are substituted with another residue (e.g., a serine residue). In some embodiments, the human-induced mutation is at position C965 (i.e., the cysteine ​​residue at position 965 of SEQ ID NO: 1). In some embodiments, the human-induced mutation is a substitution of a cysteine ​​residue. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position D156 (i.e., the aspartic acid residue at position 156 of SEQ ID NO: 1), e.g., as described in WO2018195545 and WO2017184768, which are incorporated by reference in their entireties. In some embodiments, the human-induced mutation is a substitution of an aspartic acid residue. In some embodiments, the human-induced mutation is an aspartic acid to arginine substitution. In some embodiments, the sequence of the Cas12a protein comprises any one of SEQ ID NOs: 5-11.

[0059] In some embodiments, the variant Cas12a protein is a modified form of FnCas12a. In some embodiments, the variant Cas12a protein comprises a sequence at least 60% identical (e.g., at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to the amino acid sequence of SEQ ID NO:2 and at least one human-induced mutation. In some embodiments, the human-induced mutation is a substitution of a surface-exposed cysteine ​​residue. In some embodiments, one or more surface-exposed cysteine ​​residues are substituted with another residue (e.g., a serine residue). In some embodiments, the human-induced mutation is at position C70, C1116, and / or C1190. In some embodiments, the human-induced mutation is a substitution of a cysteine ​​residue. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position E184 (i.e., a glutamic acid residue at position 184 of SEQ ID NO: 2), e.g., as described in WO2018195545, the entire contents of which are incorporated herein by reference. In some embodiments, the human-induced mutation is a substitution of a glutamic acid residue. In some embodiments, the human-induced mutation is a glutamic acid to arginine substitution.

[0060] In some embodiments, the variant Cas12a protein is a modified form of AsCas12a. In some embodiments, the variant Cas12a protein comprises a sequence at least 60% identical (e.g., at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to the amino acid sequence of SEQ ID NO: 3 and at least one human-induced mutation. In some embodiments, the human-induced mutation is a substitution of a surface-exposed cysteine ​​residue. In some embodiments, one or more surface-exposed cysteine ​​residues are substituted with another residue (e.g., a serine residue). In some embodiments, the human-induced mutation is at position C334, C379, and / or C674. In some embodiments, the human-induced mutation is a substitution of a cysteine ​​residue. In some embodiments, the human-induced mutation is a cysteine ​​to serine substitution. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position E174, e.g., as described in WO2018195545, the entire contents of which are incorporated by reference herein. In some embodiments, the human-induced mutation is a substitution of a glutamic acid residue. In some embodiments, the human-induced mutation is a glutamic acid to arginine substitution.

[0061] In some embodiments, the variant Cas12a protein is a modified form of Mb2Cas12a. In some embodiments, the variant Cas12a protein comprises a sequence at least 60% identical (e.g., at least 65%, at least 70%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99%) to the amino acid sequence of SEQ ID NO:4 and at least one human-induced mutation. In some embodiments, the human-induced mutation is a substitution of a surface-exposed cysteine ​​residue. In some embodiments, one or more surface-exposed cysteine ​​residues are substituted with another residue (e.g., a serine residue). In some embodiments, the human-induced mutation is at position C270, C583, C1068, C1099, and / or C1149. In some embodiments, the human-induced mutation is a substitution of a cysteine ​​residue. In some embodiments, the human-induced mutation is a substitution of a cysteine ​​to serine. In some embodiments, the Cas12a protein further comprises a human-induced mutation at position D172. In some embodiments, the human-induced mutation is a substitution of an aspartic acid residue. In some embodiments, the human-induced mutation is a substitution of an aspartic acid to arginine. In some embodiments, the sequence of the Cas12a protein comprises any one of SEQ ID NOs: 12-19.

[0062] A Cas protein (e.g., a Cas12a protein) may comprise one or more domains. Non-limiting examples of domains include a guide nucleic acid recognition and / or binding domain, a nuclease domain (e.g., a DNase or RNase domain, RuvC, or HNH), a DNA-binding domain, an RNA-binding domain, a helicase domain, a protein-protein interaction domain, and a dimerization domain. The guide nucleic acid recognition and / or binding domain can interact with the guide nucleic acid. The nuclease domain can comprise catalytic activity for nucleic acid cleavage. The nuclease domain may lack catalytic activity to prevent nucleic acid cleavage. The Cas protein may also be a chimeric Cas protein fused with another protein or polypeptide. The Cas protein may be a chimera of various Cas proteins, for example, comprising domains from different Cas proteins.

[0063] As used herein, a Cas protein (e.g., a Cas12a protein) can be an active variant, an inactive variant, or a fragment of a wild-type or modified Cas protein. The Cas protein can contain amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof, compared to a wild-type version of the Cas protein. The Cas protein can be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to an exemplary wild-type Cas protein. A Cas protein may be a polypeptide having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas protein. A variant or fragment may contain at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or portion thereof. A variant or fragment may lack nucleic acid cleavage activity while being capable of being complexed with a guide nucleic acid and targeted to a nucleic acid locus.

[0064] In some embodiments, the modified Cas protein has reduced function compared to the unmodified form. In some embodiments, the modified Cas protein lacks the function of the unmodified form. For example, a nuclease-deficient Cas protein retains the ability to bind to DNA but lacks or reduces nucleic acid cleavage activity. Cas nucleases (e.g., retaining wild-type nuclease activity, having reduced nuclease activity, and / or lacking nuclease activity) can function in CRISPR / Cas systems to modulate (e.g., decrease, increase, or eliminate) the level and / or activity of a target gene or protein. Cas proteins can bind to target polynucleotides and prevent transcription by physical obstruction or editing the nucleic acid sequence, resulting in a non-functional gene product. In some embodiments, the modified Cas protein has 90% or less, 80% or less, 70% or less, 60% or less, 50% or less, 40% or less, 30% or less, 20% or less, 10% or less, 5% or less, or 1% or less of the function (e.g., nuclease activity) of a wild-type Cas protein (e.g., Cas12a). In some embodiments, the modified Cas protein does not have substantial function of a wild-type Cas protein. When a Cas protein is in a modified form that does not have substantial nucleic acid cleavage activity, it can be referred to as enzymatically inactive and / or "inactivated" (abbreviated "d"). An inactivated Cas protein (e.g., dCas, dCas12a) can bind to a target polynucleotide but not cleave the target polynucleotide. In some embodiments, the Cas12a protein provided herein is a dCas12a protein.

[0065] In some embodiments, the modified Cas protein can be a modified Cas "base editor." Base editing allows for the direct, irreversible conversion of one target DNA base to another in a programmable manner, without the need for DNA cleavage or a donor DNA molecule. For example, Komor et al. (2016, Nature, 533: 420-424) teach a Cas9-cytidine deaminase fusion in which Cas9 is also inactivated and engineered to not induce double-stranded DNA breaks. Additionally, Gaudelli et al. (2017, Nature, doi:10.1038 / nature24644) teach a catalytically impaired Cas9 fused to tRNA adenosine deaminase, which can mediate the conversion of a target DNA sequence from A / T to G / C. In some embodiments, the Cas12a protein provided herein is a modified Cas12a base editor.

[0066] Cas proteins can be modified to optimize regulation of gene expression. Cas proteins can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to alter any other activity or property of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for protein function or to optimize (e.g., enhance or decrease) the activity of the Cas protein to regulate gene expression.

[0067] One or more nuclease domains of a Cas protein (e.g., RuvC, HNH) can be deleted or mutated so that they are no longer functional or contain reduced nuclease activity. For example, in a Cas protein containing at least two nuclease domains (e.g., Cas12a), if one of the nuclease domains is deleted or mutated, the resulting Cas protein, known as a nickase, can generate a single-strand break, rather than a double-strand break, at the CRISPR RNA (crRNA) recognition sequence within double-stranded DNA. Such a nickase can cleave either the complementary or non-complementary strand, but not both. In some embodiments, the targeting specificity of the double-strand break is improved by targeting the nickase to opposite strands at two nearby loci. If the nickase cleaves a single strand at both loci, a double-strand break is formed and can be repaired as described herein. If all of the nuclease domains of a Cas protein (e.g., the RuvC nuclease domain in a Cas12a protein) are deleted or mutated, the resulting Cas protein may have a reduced ability or be unable to cleave both strands of double-stranded DNA. In some embodiments, the Cas12a proteins provided herein are Cas12a nickase proteins.

[0068] Also provided herein are fusion proteins comprising any of the above proteins and a heterologous domain. As used throughout, a "fusion protein" is a protein comprising two distinct polypeptide sequences, i.e., a Cas12a protein sequence as described above and a heterologous polypeptide sequence, joined or linked to form a single polypeptide. In some embodiments, the two amino acid sequences are encoded by separate nucleic acid sequences that are joined to create a single polypeptide when transcribed and translated. The Cas12a protein and heterologous domain can be linked to each other in any order and orientation. For example, the C' terminus of the Cas12a protein can be linked to the N' or C' terminus of the heterologous domain. The Cas12a protein and heterologous domain can also be separated by one or more additional fusion protein domains, as described below.

[0069] Exemplary heterologous domains include deaminase domains, transcription factor domains, nuclease domains, reverse transcriptase domains, transposase domains, integrase domains, uracil DNA glycosylase inhibitor domains, recombinase domains, nickase domains, methyltransferase domains, methylase domains, acetylase domains, acetyltransferase domains, transcriptional activator domains, and transcriptional repressor domains. See, e.g., WO 2021 / 061507, which is incorporated herein by reference in its entirety.

[0070] In some embodiments, the fusion proteins provided herein comprise one or more linkers. As used herein, a linker, also referred to as a spacer, is a flexible molecule or a flexible stretch of molecules that connects or links two portions (e.g., domains) of a fusion protein or variant Cas12a protein as provided herein. In some embodiments, the linker is a polypeptide. Proteins in which domains are connected by a polypeptide linker are referred to as fusion proteins. In some embodiments, the linker is a non-peptide linker. Proteins in which domains are connected by a polypeptide linker are referred to as modified proteins. It will be understood that where fusion proteins are discussed throughout this disclosure, modified proteins are generally also contemplated, where feasible.

[0071] Linkers can increase the range of orientations that domains in a fusion protein or variant protein can adopt. Linkers may be optimized to produce a desired effect in a fusion protein or variant protein. Aspects and considerations of linker design are described, for example, in Chen, X. et al., Adv Drug Deliv Rev. 2013 Oct 15;65(10):1357-1369, and Klein, J. Set al. 2014 Protein Eng. Des. Sel. 27(10):325-330. In some embodiments, the proteins provided herein comprise a peptide linker. In some embodiments, the proteins provided herein comprise a non-peptide linker. In some embodiments, the proteins provided herein comprise a peptide linker and a non-peptide linker. The proteins provided herein can also comprise multiple linkers, including at least one peptide linker, at least one non-peptide linker, or at least one peptide linker and at least one non-peptide linker.

[0072] Linkers can be short or long, flexible or fixed. See, for example, WO 2021 / 061507, which is incorporated by reference in its entirety, and WO 2020 / 168102, which is incorporated by reference in its entirety, and U.S. Patent Application Publication No. 2021 / 0017506, which is incorporated by reference in its entirety.

[0073] In some embodiments, the length of the linker can affect one or more functions of the fusion protein. Selection of a linker to achieve a desired length is within the ability of one of skill in the art. In some embodiments, the peptide linker can be, for example, 5 to 100 or more amino acids in length (e.g., 4 aa, 5 aa, 8 aa, 10 aa, 15 aa, 18 aa, 20 aa, 25 aa, 30 aa, 35 aa, 40 aa, 45 aa, 50 aa, 55 aa, 60 aa, 65 aa, 70 aa, 75 aa, 80 aa, 85 aa, 90 aa, 95 aa, or 100 aa). In some embodiments, the linker is about 30 amino acids in length. In some embodiments, the linker is about 8 amino acids in length.

[0074] Depending on the length, the linker sequence may have various conformations in the secondary structure, such as helical, β-strand, coil / bend, and turn. In some cases, the linker sequence may have an extended conformation and function as an independent domain that does not interact with adjacent protein domains. The linker sequence may be flexible or rigid. Flexible linkers provide some degree of movement or interaction of the polypeptide domains and are generally rich in small or polar amino acids such as Gly and Ser (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or all of the amino acid residues in the linker are either Gly or Ser). Rigid linkers can be used to maintain a constant distance between the domains and help maintain their independent function. The linker can be attached via an amide bond (e.g., a peptide bond) or other functional groups, as discussed further below.

[0075] In some embodiments, a peptide linker described herein comprises one or more repeats (e.g., two repeats, three repeats, four repeats, five repeats, six repeats, or more repeats) of GSSSS (SEQ ID NO:43) and / or one or more repeats of GGGGS (SEQ ID NO:44) and / or one or more repeats of GSSGSS (SEQ ID NO:45) and / or one or more repeats of SGGS (SEQ ID NO:77). In some embodiments, the linker comprises an amino acid sequence having at least 90% sequence identity to (GSSSS) (SEQ ID NO:46) or (SGGS) (SEQ ID NO:78). Additional exemplary peptide linkers include SGSETPGTSESATPE (SEQ ID NO:47), SGSETPGTSESATPES (SEQ ID NO:48), (GGGGS) (SEQ ID NO:49), (GGGGS) (SEQ ID NO:50), (GGGGS) 10 (SEQ ID NO: 51), GGGGGGGG (SEQ ID NO: 52), GSAGSAAGSGEF (SEQ ID NO: 53), A(EAAAK)3A (SEQ ID NO: 54), or A(EAAAK) 10A (SEQ ID NO: 55). Additional non-limiting exemplary linkers that can be used include those disclosed in PCT / US2020 / 051383, Chen et al., Adv. Drug. Deliv. Rev. 65(10): 1357-1369 (2014), and Rosemalen et al., Biochemistry 2017, 56, 50, 6565-6574 (the entire contents of both of which are incorporated herein by reference).

[0076] In some embodiments, the non-peptide linker can comprise any of a number of known chemical linkers. Exemplary chemical linkers include one or more units of beta-alanine, 4-aminobutyric acid (GABA), (2-aminoethoxy)acetic acid (AEA), 5-aminobexanoic acid (Ahx), PEG multimers, and trioxatricdeacan-succinamic acid (Ttds). In some embodiments, the non-peptide linker comprises one or more units of polyethylene glycol (PEG). PEG is commonly used as a linker in the conjugation of polypeptide domains due to its water solubility, lack of toxicity, low immunogenicity, and well-defined chain length. See, for example, Ramirez-Paz, J., et al., PLoS One 13(7):e0197643 (2018). The number of PEG linking units can be selected based on the desired length of the linker.

[0077] Modified proteins containing non-peptide linkers can be produced in a variety of ways. For example, the Cas12a protein and heterologous domain can be produced separately (e.g., in vitro or by expression in and purification from host cells) and chemically linked in vitro. In some embodiments, the Cas12a protein, heterologous domain, and linker can each be produced separately in vitro and chemically linked. A variety of chemical linkers can be used to bridge the two amino acid residues.

[0078] Also contemplated herein are embodiments in which the Cas12a protein and heterologous domain are used separately (e.g., introduced separately into a cell or applied separately to a target nucleic acid) and brought into close proximity to form a complex without the use of a linker. Various methods for forming complexes between two or more polypeptides are known in the art, including, but not limited to, using protein-protein interaction strategies (e.g., SunTag, coiled coil, etc.), RNA aptamers and related binding proteins (e.g., MS2, N22, etc.), and Tag:catcher strategies. For example, a site-specific nuclease of the present disclosure may include an MS2 RNA aptamer, which is believed to facilitate interaction with nonspecific end-processing enzymes, including the MS2 coat protein.

[0079] In some embodiments, the fusion proteins provided herein comprise an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity to any one of SEQ ID NOs: 1-4. In some embodiments, the fusion proteins provided herein comprise an amino acid sequence as set forth in any one of SEQ ID NOs: 5-19.

[0080] Any of the proteins and fusion proteins described herein may further comprise a targeting sequence that mediates localization (or retention) of the protein to a subcellular location, such as the plasma membrane or the membrane of a given organelle, such as the nucleus, cytosol, mitochondria, endoplasmic reticulum (ER), Golgi apparatus, chloroplast, apoplast, peroxisome, or other organelle. For example, the targeting sequence can utilize a nuclear localization signal (NLS) to direct the protein (e.g., a nuclease) to the nucleus; a nuclear export signal (NES) to direct it outside the nucleus of the cell, e.g., to the cytoplasm; a mitochondrial targeting signal to direct it to the mitochondria; an ER retention signal to direct it to the endoplasmic reticulum (ER); a peroxisomal targeting signal to direct it to peroxisomes; a membrane localization signal to direct it to the plasma membrane; or a combination thereof. In some embodiments, the protein comprises a nuclear localization signal.Non-limiting examples of NLSs include NLS sequences derived from: the SV40 virus large T antigen NLS having the amino acid sequence PKKKRKV (SEQ ID NO: 56); an NLS from nucleoplasmin (e.g., the nucleoplasmin bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO: 57)); the c-myc NLS having the amino acid sequence PAAKRVKLD (SEQ ID NO: 58) or RQRRNELKRSP (SEQ ID NO: 59); the hRNPA1M9 NLS having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRNQGGY (SEQ ID NO: 60); the IBB domain from importin alpha, sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO: 61); the fibroid T protein sequences VSRKRPRP (SEQ ID NO: 62) and PPKKARED (SEQ ID NO: 63); the human p53 sequence PQPKKKPL (SEQ ID NO: 64); and the mouse c-abl the sequence of IV SALIKKKKKMAP (SEQ ID NO: 65); the sequences DRLRR (SEQ ID NO: 66) and PKQKKRK (SEQ ID NO: 67) of influenza virus NS1; the sequence RKLKKKIKKL (SEQ ID NO: 68) of hepatitis virus delta antigen; the sequence REKKKFLKRR (SEQ ID NO: 69) of mouse Mx1 protein; the sequence KRKGDEVDGVDEVAKKKSKK (SEQ ID NO: 70) of human poly(ADP-ribose) polymerase; the sequence RKCLQAGMNLEARKTKK (SEQ ID NO: 71) of steroid hormone receptor (human) glucocorticoid; and the sequence KRPRDRHDGELGGRKRAR (SEQ ID NO: 72) of Agrobacterium VirD2 protein.

[0081] Any of the proteins and fusion proteins described herein may further comprise a detectable moiety, such as a fluorescent protein or a fragment thereof. Examples of fluorescent proteins include, but are not limited to, yellow fluorescent protein (YFP, e.g., Venus), green fluorescent protein (GFP), and red fluorescent protein (RFP), as well as derivatives of these proteins, such as mutant derivatives. See, for example, Chudakov et al., "Fluorescent Proteins and Their Applications in Imaging Living Cells and Tissues," Physiological Reviews 90(3):1103-1163 (2010); and Specht et al., "A Critical and Comparative Review of Fluorescent Tools for Live-Cell Imaging," Annual Review of Physiology 79:93-117 (2017)).

[0082] Any of the proteins and fusion proteins described herein can further comprise an affinity tag, such as a polyhistidine tag (e.g., (His)6 (SEQ ID NO: 73)), an HA tag (e.g., YPYDVPDYA (SEQ ID NO: 74)), an albumin binding protein, alkaline phosphatase, an AU1 epitope, an AU5 epitope, a biotin carboxy carrier protein (BCCP), a FLAG epitope (e.g., DYKDDDDK (SEQ ID NO: 75)), or a MYC epitope (e.g., EQKLISEEDL (SEQ ID NO: 76)), to name a few. See Kimple et al. "Overview of Affinity Tags for Protein Purification," Curr. Protoc. Protein Sci. 73:Unit-9.9 (2013).

[0083] Variants of the polypeptides (e.g., proteins and fusion proteins) of the present disclosure also apply herein. Polypeptide variants retain their respective biological activity unless otherwise expressly noted. For example, variants of Cas12a polypeptides retain the biological function of the full-length native sequence site-specific Cas12a protein. In another example, variants of heterologous domains retain the biological function of the full-length native sequence heterologous domain.

[0084] Modifications to any of the polypeptides or proteins provided herein can be made by known methods. For example, modifications can be made by site-specific mutagenesis of nucleotides in a nucleic acid encoding the polypeptide, thereby generating DNA encoding the modification, which can then be expressed in a recombinant cell culture to produce the encoded polypeptide. Techniques for making substitution mutations at predetermined sites in DNA having a known sequence are well known. For example, M13 primer mutagenesis and PCR-based mutagenesis methods can be used to make one or more substitution mutations. Any of the nucleic acid sequences provided herein can be codon-optimized to alter, e.g., maximize, expression in a host cell or organism.

[0085] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D stereoisomers of naturally occurring amino acids, unnatural amino acids, and chemically modified amino acids. Unnatural amino acids (i.e., those not found in proteins in nature) are also known in the art, as shown, for example, in Zhang et al. "Protein engineering with unnatural amino acids," Curr. Opin. Struct. Biol. 23(4):581-587 (2013); Xie et al. "Adding amino acids to the genetic repertoire," 9(6):548-54 (2005); and all references cited therein. B and γ amino acids are known in the art and are also contemplated as unnatural amino acids herein.

[0086] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, the side chain can be modified to include a signaling moiety, such as a fluorophore or a radiolabel. The side chain can also be modified to include a new functional group, such as a thiol, a carboxylic acid, or an amino group. Post-translationally modified amino acids are also included in the definition of chemically modified amino acids.

[0087] Conservative amino acid substitutions are also contemplated. For example, conservative amino acid substitutions can be made at one or more amino acid residues, for example, at one or more lysine residues in any of the polypeptides provided herein. Those skilled in the art will know that a conservative substitution is the replacement of an amino acid residue with another amino acid residue that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservative substitutions for each other: 1) alanine (A), glycine (G); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) Cysteine ​​(C), methionine (M).

[0088] For example, when an arginine to serine substitution is mentioned, a conservative substitution of serine (e.g., threonine) is also contemplated. Non-conservative substitutions, such as substituting lysine with asparagine, are also contemplated.

[0089] IV. Recombinant Nucleic Acids, Constructs, Vectors, and Host Cells Also provided herein are recombinant nucleic acids encoding any of the variant Cas12a proteins or fusion proteins described herein. For example, recombinant nucleic acids encoding a polypeptide having at least 70% identity (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) to any of SEQ ID NOs: 20-34. Also provided are recombinant nucleic acids having at least 70% identity to any of SEQ ID NOs: 20-34.

[0090] Also provided are DNA constructs comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein or domain thereof as described herein. A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. Multiple promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of the start of transcription that is involved in the recognition and binding of RNA polymerase and other proteins to initiate transcription.

[0091] The term "promoter," as used herein, refers to a nucleotide sequence that controls the expression of a coding sequence by providing recognition for RNA polymerase and other factors necessary for proper transcription, typically located upstream (5') of that coding sequence. A "promoter regulatory sequence" consists of proximal and more distal upstream elements. Promoter regulatory sequences affect the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. These include natural and synthetic sequences, as well as sequences that may be a combination of natural and synthetic sequences. An "enhancer" is a DNA sequence that can stimulate promoter activity and may be an intrinsic element of the promoter or a heterologous element inserted to increase the level or tissue specificity of the promoter. It is capable of operating in both orientations (normal or inverted) and can function when moved either upstream or downstream from the promoter. The term "promoter" includes "promoter regulatory sequence."

[0092] The choice of which promoter to include depends on several factors, including, but not limited to, efficiency, selectability, inducibility, desired expression level, and cell- or tissue-preferred expression. It is routine for those of skill in the art to regulate the expression of a sequence by appropriate selection and placement of promoters and other regulatory regions relative to that sequence.

[0093] Certain promoters have been shown to be capable of directing RNA synthesis at higher rates than others. These are called "strong promoters." Certain other promoters have been shown to direct RNA synthesis to higher levels only in certain types of cells or tissues; when a promoter preferentially directs RNA synthesis to a particular tissue (while RNA synthesis may occur at lower levels in other tissues), it is often referred to as a "tissue-specific promoter" or "tissue-preferred promoter." Because the expression pattern of a chimeric gene (or genes) introduced into a plant is controlled using a promoter, there is continuing interest in isolating new promoters capable of controlling the expression of a chimeric gene (or genes) to a certain level in a specific tissue type or at a specific plant developmental stage.

[0094] Certain promoters are capable of directing RNA synthesis at relatively similar levels in all tissues of a plant. These are called "constitutive promoters" or "tissue-independent" promoters. Constitutive promoters can be divided into strong, intermediate, and weak categories based on their effectiveness in directing RNA synthesis. Constitutive promoters are particularly useful in this regard, since in many cases, a chimeric gene (or genes) must be expressed simultaneously in different plant tissues to obtain the desired function of the gene (or genes). Although many constitutive promoters have been discovered and characterized from plants and plant viruses, there is still ongoing interest in isolating more novel constitutive promoters, synthetic or natural, capable of controlling the expression of a chimeric gene (or genes) at different levels and expression levels of multiple genes in the same transgenic plant for gene stacking.

[0095] Among the most commonly used promoters are the nopaline synthase (NOS) promoter (Ebert et al., Proc. Natl. Acad. Sci. USA 84:5745-5749 (1987)); the octapine synthase (OCS) promoter; caulimovirus promoters such as the cauliflower mosaic virus (CaMV) 19S promoter (Lawton et al., Plant Mol. Biol. 9:315-324 (1987)); the light-inducible promoter from the small subunit of Rubisco (Pellegrineschi et al., Biochem. Soc. Trans. 23(2):247-250 (1995)); the Adh promoter (Walker et al., Proc. Natl. Acad. Sci. USA 84:6624-66280 (1987)); the sucrose synthase promoter (Yang et al. al., Proc. Natl. Acad. Sci. USA 87:414-44148 (1990)); R gene complex promoter (Chandler et al., Plant Cell 1:1175-1183 (1989)); chlorophyll a / b binding protein gene promoter, etc.

[0096] Furthermore, it is contemplated that promoters that combine elements from two or more promoters may be useful. For example, U.S. Patent No. 5,491,288 discloses combining a cauliflower mosaic virus promoter with a histone promoter. Thus, elements from the promoters disclosed herein may be combined with elements from other promoters. Promoters useful for plant transgene expression include inducible, viral, synthetic, constitutive (Odell Nature 313:810-812 (1985)), temporally regulated, spatially regulated, tissue-specific, and spatiotemporally regulated promoters. Using the regulatory elements described herein, numerous agronomic genes can be expressed in transformed plants. More specifically, plants can be engineered to express a variety of phenotypes of agronomic interest.

[0097] In some embodiments of the DNA constructs provided herein, the promoter can be a eukaryotic or prokaryotic promoter. In some embodiments, the promoter is an inducible promoter, a naturally occurring inducible promoter (e.g., drought-inducible Rab17), a synthetic inducible promoter (e.g., auxin-inducible DR5, estradiol-inducible XVE / pLex, dexamethasone-inducible GVG / Gal4), a constitutive promoter (e.g., ZmUbq1, OsAct1, OsTub3, EF, EF1α), an egg cell-specific promoter (e.g., EC1, EC2, EC3, EC4, EC5), a pollen-specific promoter, an apical meristem-specific promoter, or a promoter with enhanced expression in the zygote. In some embodiments, the promoter is a floral mosaic promoter (e.g., ZmBde1, OsAP1). In some embodiments, the promoter is a ubiquitin 4 promoter (e.g., sugarcane ubiquitin 4 promoter), an actin promoter, a tubulin promoter, a MADS box promoter, or a plant virus promoter. Suitable promoters are disclosed, for example, in U.S. Pat. No. 10,519,456, the entire contents of which are incorporated herein by reference, and PCT / US2022 / 020690, the entire contents of which are incorporated herein by reference.

[0098] The recombinant nucleic acids provided herein can be included in an expression cassette for expression in a host cell or organism of interest. The cassette will include 5' and 3' regulatory sequences operably linked to the recombinant nucleic acids provided herein, allowing for expression of the fusion protein. The cassette may additionally contain at least one additional gene or genetic element to be cotransformed into the cell or organism. When additional genes or elements are included, the components are operably linked. Alternatively, the additional genes or elements can be provided on multiple expression cassettes. Such expression cassettes comprise multiple restriction and / or recombination sites for insertion of polynucleotides under the transcriptional control of the regulatory regions. The expression cassette may additionally contain a selectable marker gene. The expression cassette will include, in the 5' to 3' transcriptional direction, a transcriptional and translational initiation region (i.e., promoter), a polynucleotide of the invention, and a transcriptional and translational termination region (i.e., termination region) functional in the cell or organism of interest. A promoter of the present invention is capable of directing or driving expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA, such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, whether or not that RNA is then translated to produce a protein) in a host cell. The regulatory regions (i.e., promoter, transcriptional regulatory region, and translation termination region) may be endogenous to the host cell or to each other, or heterologous. As used herein, "heterologous" with respect to a sequence is a sequence that is derived from a foreign species, or, if derived from the same species, is substantially altered in composition and / or genomic locus from its native form by deliberate human intervention.

[0099] Additional regulatory signals include, but are not limited to, a start site for transcription initiation, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.

[0100] The expression cassette may also contain a selectable marker gene for selecting transformed cells. Marker genes include genes that confer antibiotic resistance, such as those that confer hygromycin resistance, ampicillin resistance, gentamicin resistance, and neomycin resistance, to name a few. Additional selectable markers are known, and any may be used.

[0101] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences in the proper orientation and, if necessary, in the proper reading frame. To this end, adapters or linkers can be used to join DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. To this end, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, such as transitions and transversions, can be involved.

[0102] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences in the proper orientation and, if necessary, in the proper reading frame. To this end, adapters or linkers can be used to join DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. In vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transpositions and transversions, can be used for this purpose.

[0103] Also provided are vectors containing the recombinant nucleic acids or DNA constructs described herein. It is contemplated that the vectors have the necessary functional elements to direct and regulate transcription of the inserted nucleic acid. Such functional elements include, but are not limited to, a promoter, regions upstream or downstream of the promoter, such as enhancers and terminators that can regulate the transcriptional activity of the promoter, an origin of replication, appropriate restriction sites to facilitate cloning of an insert adjacent to the promoter, antibiotic resistance genes or other markers that can be useful for selecting cells containing the vector or vectors containing the insert, RNA splice junctions, transcription termination regions, or any other regions that can be useful for promoting expression of the inserted gene or hybrid gene. Generally, the functional elements described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. The vector may be, for example, a plasmid. In some embodiments of the DNA constructs and vectors provided herein, the constructs and vectors comprise a nopaline synthase gene terminator sequence (e.g., an Agrobacterium tumefaciens nopaline synthase gene terminator sequence).

[0104] Numerous E. coli expression vectors useful for expressing nucleic acids are known to those of skill in the art. Other microbial hosts suitable for use include bacilli, such as Bacillus subtilis, and other Enterobacteriaceae bacteria, such as Salmonella, Serratia, and various Pseudomonas species. Expression vectors can also be made in these prokaryotic hosts, which will typically contain expression control sequences (e.g., an origin of replication) compatible with the host cell. Additionally, several well-known promoters exist, such as the lactose promoter system, the tryptophan (Trp) promoter system, the beta-lactamase promoter system, or promoter systems from lambda phage. Additionally, yeast expression can be used. Provided herein are nucleic acids encoding the polypeptides of the invention, where the nucleic acids can be expressed by yeast cells. More specifically, the nucleic acid can be expressed by Pichia pastoris or S. cerevisiae.

[0105] Mammalian cells also allow for the expression of proteins in an environment that favors important post-translational modifications, such as folding and cysteine ​​pairing, addition of complex carbohydrate structures, and secretion of active proteins. Vectors useful for expressing active proteins in mammalian cells are known in the art and can include genes conferring hygromycin resistance, geneticin, or G418 resistance, or other genes or phenotypes suitable for use as selectable markers, or methotrexate resistance for gene amplification. A number of suitable host cell lines capable of secreting intact human proteins have been developed in the art, including CHO cells, HeLa cells, HEK-293 cells, HEK-293T cells, U2OS cells, or any other primary or transformed cell line. Other suitable host cell lines include COS-7 cells, myeloma cell lines, Jurkat cells, and the like. Expression vectors for these cells may contain expression control sequences such as a replication origin, a promoter, an enhancer, and necessary information processing sites such as a ribosome binding site, RNA splice sites, a polyadenylation site, and a transcription terminator sequence. Preferred expression control sequences are promoters derived from immunoglobulin genes, SV40, adenovirus, bovine papilloma virus, etc.

[0106] The expression vectors described herein can also include nucleic acids as described herein under the control of an inducible promoter, such as a tetracycline-inducible promoter or a glucocorticoid-inducible promoter. The nucleic acids of the present invention can also be under the control of a tissue-specific promoter that promotes expression of the nucleic acid in specific cells, tissues, or organs. Any regulatable promoter, many examples of which are known in the art, is also contemplated, such as metallothionein promoters, heat shock promoters, and other regulatable promoters. Furthermore, the Cre-loxP inducible system and the Flp recombinase inducible promoter system can also be used, both of which are known in the art.

[0107] Insect cells also allow the expression of polypeptides. Recombinant proteins produced in insect cells with baculovirus vectors undergo post-translational modifications similar to wild-type mammalian proteins.

[0108] Also provided herein are host cells comprising the recombinant nucleic acids, DNA constructs, and / or vectors described herein, as well as methods for making such cells. In some embodiments, the cell is a plant cell. In some embodiments, the plant cell is a corn plant cell, a wheat plant cell, a rice plant cell, a soybean plant cell, a sunflower plant cell, or a tomato plant cell.

[0109] Host cells comprising a nucleic acid or vector described herein are provided. The host cell may be an in vitro, ex vivo, or in vivo host cell. The host cell as provided herein is capable of expressing a fusion protein. Cell populations of any of the host cells described herein are also provided. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a recombinant nucleic acid encoding a fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a DNA construct encoding a protein and / or fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a vector comprising a recombinant nucleic acid or DNA construct encoding a protein and / or fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells comprises a plurality of any of the host cells described herein. In some embodiments, the plurality of cells of any of the cell populations described herein express a protein and / or fusion protein as described herein.

[0110] In some embodiments, the provided cells stably or transiently express the protein and / or fusion protein. Stable expression of a protein and / or fusion protein in a cell refers to the integration of any of the nucleic acids, DNA constructs, or vectors described herein into the genome of the cell, thereby enabling the cell to express the protein and / or fusion protein. Transient expression refers to the direct expression of the protein and / or fusion protein from any of the nucleic acids, DNA constructs, and / or vectors after introduction into the cell (i.e., the gene encoding the protein and / or fusion protein is not integrated into the genome of the cell).

[0111] In some embodiments, the provided cells constitutively or inducibly express proteins and / or fusion proteins. Constitutive expression refers to ongoing, continuous expression of a gene (i.e., a protein), while inducible expression refers to gene (protein) expression in response to a stimulus. Inducible expression is generally regulated by an inducible promoter, a description of which is included above.

[0112] Also provided are cell cultures comprising one or more host cells described herein. Numerous cell culture and production methods are available in the art, including cells of bacterial origin (e.g., E. coli and other bacterial strains), animal origin (especially mammalian origin), and archebacterial origin. See, for example, Sambrook, supra; Ausubel, ed. (1995) Current Protocols in Molecular Biology, John Wiley & Sons, as well as Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, 3 rdEd.,Wiley-Liss,New York and the references cited therein;Doyle and Griffiths(1997)Mammalian Cell Culture:Essential Techniques John Wiley and Sons,NY;Humason(1979)Animal Tissue Techniques,4 th See Ed. W.H. Freeman and Company; and Ricciardelli, et al., (1989) In vitro Cell Dev. Biol. 25:1016-1024.

[0113] The host cell can be a prokaryotic cell, including, for example, a bacterial cell. Alternatively, the cell can be a eukaryotic cell, such as a mammalian cell. In some embodiments, the cell can be a HEK-293T cell, a HEK-293 cell, a Chinese hamster ovary (CHO) cell, a U2OS cell, or any other primary or transformed cell. In some embodiments, the cell can be a COS-7 cell, a HELA cell, an avian cell, a myeloma cell, a Pichia cell, an insect cell, or a plant cell. A number of other suitable host cell lines have been developed, including various tumor cell lines, such as myeloma, fibroblast, and melanoma cell lines. Vectors containing the nucleic acid segment of interest can be transfected or introduced into the host cell by well-known methods that vary depending on the type of cellular host.

[0114] As used herein, the phrase "introducing," in the context of introducing a nucleic acid into a cell (e.g., a prokaryotic, bacterial, eukaryotic, or plant cell), refers to changing the location of a nucleic acid sequence from outside the cell to inside the cell. In some cases, introducing refers to changing the location of a nucleic acid from outside the cell to inside the nucleus of the cell. When two or more nucleic acid molecules are to be introduced, the nucleic acid molecules can be assembled as part of a single polynucleotide or nucleic acid construct or as separate polynucleotides or nucleic acid constructs, and can be located on the same or different nucleic acid constructs. Thus, such polynucleotides can be introduced into a cell (e.g., a plant cell) in a single transformation event, in separate transformation events, or, for example, as part of a breeding protocol. Various methods of introducing nucleic acids into cells are contemplated, including, but not limited to, electroporation, nanoparticle delivery, biolistic transformation, viral delivery, contact with nanowires or nanotubes, receptor-mediated internalization, cell-penetrating peptide-mediated transfer, liposome-mediated transfer, DEAE-dextran, lipofectamine, calcium phosphate, or any method now known or later identified for the introduction of nucleic acids into prokaryotic or eukaryotic hosts. Targeted nuclease systems (e.g., RNA-guided nucleases, transcription activator-like effector nucleases (TALENs), zinc finger nucleases (ZFNs), or megaTALs (MTs)) can also be used to introduce nucleic acids, such as nucleic acids encoding the proteins and / or fusion proteins described herein, into host cells. See Li et al., Signal Transduction and Targeted Therapy 5, Article No. 1 (2020).

[0115] Cellular transformation can be stable or transient. Thus, the transgenic cells, plant cells, plants, and / or plant parts of the present invention can be stably transformed or transiently transformed. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in stable genetic inheritance. In some embodiments, introduction into the plant, plant part, and / or plant cell is via bacterial-mediated transformation, particle bombardment transformation, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, liposome-mediated transformation, nanoparticle-mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infiltration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical, and / or biological mechanism that results in the introduction of a nucleic acid into a plant, plant part, and / or cell thereof, or any combination thereof.

[0116] Plant transformation procedures are well known and routine in the art and are described throughout this literature. Non-limiting examples of plant transformation methods include transformation by bacterial-mediated nucleic acid delivery (e.g., by bacteria from the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, particle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical), and / or biological mechanism that results in the introduction of nucleic acid into plant cells, including any combination thereof. General guides to various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants" in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, Eds. (CRC Press, Inc., Boca Raton, 1993), pages 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7:849-858 (2002)).

[0117] Agrobacterium-mediated transformation is a commonly used method for transforming plants due to its high transformation efficiency and its versatility with many different species. Agrobacterium-mediated transformation typically involves transferring a binary vector carrying the foreign DNA of interest into a suitable Agrobacterium strain, which may depend on the complement of vir genes carried by the host Agrobacterium strain either on a coexisting Ti plasmid or on the chromosome (Uknes et al. 1993, Plant Cell 5:159-169). Transfer of the recombinant binary vector into Agrobacterium can be achieved by a triparental mating procedure using Escherichia coli carrying the recombinant binary vector and a helper E. coli strain carrying a plasmid capable of mobilizing the recombinant binary vector into the target Agrobacterium strain. Alternatively, the recombinant binary vector can be transferred into Agrobacterium by nucleic acid transformation (Hoefgen and Willmitzer 1988, Nucleic Acids Res 16:9877).

[0118] Transformation of plants with recombinant Agrobacterium usually involves co-cultivation of Agrobacterium with explants from the plant, followed by methods well known in the art. The transformed tissue carries the antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders and is typically regenerated on selective media.

[0119] Another method for transforming plants, plant parts, and plant cells involves projecting inert or biologically active particles into plant tissues and cells. See, e.g., U.S. Patent Nos. 4,945,050; 5,036,006; and 5,100,792. Generally, this method involves projecting inert or biologically active particles into plant cells under conditions effective to penetrate the outer surface of the cells and cause uptake. When inert particles are used, the vector can be introduced into the cells by coating the particles with a vector containing the nucleic acid of interest. Alternatively, the vector can be surrounded by one or more cells, resulting in the particle being carried into the cells. Biologically active particles (e.g., dried yeast cells, dried bacteria, or bacteriophage, each containing one or more nucleic acids to be introduced) can also be projected into plant tissue. As used herein, the phrase "biolistic transformation" refers to a method of directly introducing RNA or DNA into a cell (e.g., a plant cell) in which the RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cell (e.g., a plant cell) using high-velocity pressure, allowing the RNA or DNA to penetrate the cell (e.g., penetrate the plant cell wall).

[0120] CRISPR / Cas systems can also be used to edit the genome of a host cell or organism. As detailed above, the "CRISPR / Cas" system refers to a broad class of bacterial systems for defense against foreign nucleic acids. Any of the CRISPR / Cas system components described herein can be used to introduce a protein, fusion protein, recombinant nucleic acid, or system into the genome of a host cell or organism. CRISPR / Cas system-mediated genome editing methods are known in the art. It will be understood that the introduction of a protein, fusion protein, recombinant nucleic acid, or system described herein into the genome of a host cell or organism using a CRISPR / Cas system differs from the specific methods and systems provided herein.

[0121] Any of the proteins and / or fusion proteins described herein can be purified or isolated from a host cell or population of host cells. For example, a recombinant nucleic acid encoding any of the proteins and / or fusion proteins described herein can be introduced into a host cell under conditions that allow expression of the protein and / or fusion protein. In some embodiments, the recombinant nucleic acid is codon-optimized for expression. After expression in the host cell, the protein and / or fusion protein can be isolated or purified using purification methods known in the art.

[0122] V.System In another aspect, provided herein are systems useful for editing one or more nucleic acids. The systems include one or more of the Cas12a proteins and / or fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above. In some embodiments, the systems further include one or more additional elements useful for editing one or more nucleic acids. For example, a system including a fusion protein comprising a Cas nuclease may further include one or more guide nucleic acids, as described in more detail below. The systems provided herein are useful for practicing the methods described in Section VI of this disclosure.

[0123] In some cases, the systems and methods described herein include at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein include multiple guide nucleic acids. In some embodiments, the polynucleotide may be deoxyribonucleic acid (DNA). In some cases, the DNA sequence may be single-stranded or double-stranded. In some embodiments, the at least one guide nucleic acid polynucleotide may be a ribonucleic acid (guide RNA).

[0124] In some embodiments, the Cas12a protein can be complexed with at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide can include a nucleic acid targeting region that includes a sequence complementary to a nucleic acid sequence on a targeted polynucleotide, such as a targeted genomic locus or gene, to confer sequence specificity for nuclease targeting. In some embodiments, the at least one guide RNA polynucleotide can include two separate nucleic acid molecules, which can be referred to as a double-guide nucleic acid or a single nucleic acid molecule, which can be referred to as a single-guide nucleic acid (e.g., single-guide RNA or sgRNA).

[0125] The Cas protein-binding segment of a guide nucleic acid may comprise two consecutive nucleotides (e.g., crRNA and tracrRNA) that are complementary to each other. The two consecutive nucleotides (e.g., crRNA and tracrRNA) that are complementary to each other may be covalently linked by an intervening nucleotide (e.g., a linker in the case of a single guide nucleic acid). The two consecutive nucleotides (e.g., crRNA and tracrRNA) that are complementary to each other may hybridize to form a double-stranded RNA duplex or hairpin of the Cas protein-binding segment, resulting in a stem-loop structure. The crRNA and tracrRNA may be covalently linked via the 3' end of the crRNA and the 5' end of the tracrRNA. Alternatively, the tracrRNA and crRNA may be covalently linked via the 5' end of the tracrRNA and the 3' end of the crRNA. The crRNA can include a nucleic acid targeting segment (e.g., a spacer region) of the guide nucleic acid and a stretch of nucleotides that can form one half of a double-stranded duplex of the Cas protein binding segment of the guide nucleic acid. The crRNA can also provide a single-stranded nucleic acid targeting segment (e.g., a spacer region) that hybridizes to a target nucleic acid recognition sequence (e.g., a protospacer). Whether the nuclease requires only a crRNA molecule or both a crRNA molecule and a tracrRNA molecule (covalently linked or not) depends on the CRISPR-associated nuclease used. The Cas12 protein typically does not require a tracrRNA.

[0126] In some embodiments, the nucleic acid targeting region of the guide nucleic acid can be 18 to 72 nucleotides in length. The nucleic acid targeting region of the guide nucleic acid (e.g., spacer region) can have a length of about 12 nucleotides to about 100 nucleotides. For example, the nucleic acid targeting region of the guide nucleic acid (e.g., spacer region) can have a length of about 12 nucleotides (nt) to about 80 nt, about 12 nt to about 50 nt, about 12 nt to about 40 nt, about 12 nt to about 30 nt, about 12 nt to about 25 nt, about 12 nt to about 20 nt, about 12 nt to about 19 nt, about 12 nt to about 18 nt, about 12 nt to about 17 nt, about 12 nt to about 16 nt, or about 12 nt to about 15 nt. Alternatively, the DNA targeting segment may have a length of about 18 nt to about 20 nt, about 18 nt to about 25 nt, about 18 nt to about 30 nt, about 18 nt to about 35 nt, about 18 nt to about 40 nt, about 18 nt to about 45 nt, about 18 nt to about 50 nt, about 18 nt to about 60 nt, about 18 nt to about 70 nt, about 18 nt to about 80 nt, about 18 nt to about 90 nt, about 18 nt to about 100 nt, about 20 nt to about 25 nt, about 20 nt to about 30 nt, about 20 nt to about 35 nt, about 20 nt to about 40 nt, about 20 nt to about 45 nt, about 20 nt to about 50 nt, about 20 nt to about 60 nt, about 20 nt to about 70 nt, about 20 nt to about 80 nt, about 20 nt to about 90 nt, or about 20 nt to about 100 nt. The length of a nucleic acid targeting region can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or more nucleotides. The length of a nucleic acid targeting region (e.g., a spacer sequence) can be at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, or more nucleotides.

[0127] In some embodiments, the nucleic acid targeting region of the guide nucleic acid (e.g., a spacer) is 20 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 19 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 18 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 17 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 16 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 21 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 22 nucleotides in length.

[0128] The nucleotide sequence of the guide nucleic acid complementary to the nucleotide sequence of the target nucleic acid (target sequence) can have a length of, for example, at least about 12 nucleotides (nt) to about 80 nt, about 12 nt to about 50 nt, about 12 nt to about 45 nt, about 12 nt to about 40 nt, about 12 nt to about 35 nt, about 12 nt to about 30 nt, about 12 nt to about 25 nt, about 12 nt to about 20 nt, about 12 nt to about 19 nt, about 19 nt to about 20 ... The length may be about 25 nt, about 19 nt to about 30 nt, about 19 nt to about 35 nt, about 19 nt to about 40 nt, about 19 nt to about 45 nt, about 19 nt to about 50 nt, about 19 nt to about 60 nt, about 20 nt to about 25 nt, about 20 nt to about 30 nt, about 20 nt to about 35 nt, about 20 nt to about 40 nt, about 20 nt to about 45 nt, about 20 nt to about 50 nt, or about 20 nt to about 60 nt.

[0129] The protospacer sequence of a targeted polynucleotide can be identified by identifying a protospacer adjacent motif (PAM) within the region of interest and selecting a region of desired size upstream or downstream of the PAM as the protospacer. The corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.

[0130] Spacer sequences can be identified using a computer program (e.g., machine-readable code) that can use variables such as predicted melting temperature, secondary structure formation, and predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, %GC, genomic frequency, methylation status, and the presence of SNPs.

[0131] The percent complementarity between a nucleic acid targeting sequence (e.g., a spacer sequence of at least one guide polynucleotide as disclosed herein) and a target nucleic acid (e.g., a protospacer sequence of one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percent complementarity between a nucleic acid targeting sequence and a target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 contiguous nucleotides.

[0132] The Cas protein-binding segment of the guide nucleic acid can have a length of about 10 nucleotides to about 100 nucleotides, e.g., about 10 nucleotides (nt) to about 20 nt, about 20 nt to about 30 nt, about 30 nt to about 40 nt, about 40 nt to about 50 nt, about 50 nt to about 60 nt, about 60 nt to about 70 nt, about 70 nt to about 80 nt, about 80 nt to about 90 nt, or about 90 nt to about 100 nt. For example, the Cas protein-binding segment of the guide nucleic acid can have a length of about 15 nucleotides (nt) to about 80 nt, about 15 nt to about 50 nt, about 15 nt to about 40 nt, about 15 nt to about 30 nt, or about 15 nt to about 25 nt.

[0133] The dsRNA duplex of the Cas protein-binding segment of the guide nucleic acid can have a length of about 6 base pairs (bp) to about 50 bp. For example, the dsRNA duplex of the protein-binding segment can have a length of about 6 bp to about 40 bp, about 6 bp to about 30 bp, about 6 bp to about 25 bp, about 6 bp to about 20 bp, about 6 bp to about 15 bp, about 8 bp to about 40 bp, about 8 bp to about 30 bp, about 8 bp to about 25 bp, about 8 bp to about 20 bp, or about 8 bp to about 15 bp. For example, the dsRNA duplex of the Cas protein binding segment can have a length of about 8 bp to about 10 bp, about 10 bp to about 15 bp, about 15 bp to about 18 bp, about 18 bp to about 20 bp, about 20 bp to about 25 bp, about 25 bp to about 30 bp, about 30 bp to about 35 bp, about 35 bp to about 40 bp, or about 40 bp to about 50 bp.

[0134] In some embodiments, the dsRNA duplex of the Cas protein-binding segment may have a length of 36 base pairs. The percent complementarity between the nucleotide sequences that hybridize to form the dsRNA duplex of the protein-binding segment may be at least about 60%. For example, the percent complementarity between the nucleotide sequences that hybridize to form the dsRNA duplex of the protein-binding segment may be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some cases, the percent complementarity between the nucleotide sequences that hybridize to form the dsRNA duplex of the protein-binding segment may be 100%.

[0135] Guide nucleic acids of the disclosed systems may include modifications or sequences that provide additional desirable characteristics (e.g., modified or controlled stability; intracellular targeting; tracking by fluorescent labels; binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (7-methylguanylate cap (m7G)); a 3' polyadenylation tail (3' poly(A) tail); a riboswitch sequence (e.g., allowing for regulated stability and / or regulated accessibility by proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (hairpin); a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows fluorescent detection, etc.); a modification or sequence that provides binding sites for proteins that act on DNA, including proteins (e.g., transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).

[0136] A guide nucleic acid may contain one or more modifications (e.g., base modifications, backbone modifications) to provide the nucleic acid with novel or enhanced characteristics (e.g., improved stability). A guide nucleic acid may contain a nucleic acid affinity tag. A nucleoside may be a base-sugar combination. The base portion of a nucleotide may be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For nucleosides that include a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming a guide nucleic acid, the phosphate group can covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound can then be further joined to form a circular compound; however, linear compounds may also be preferred. In addition, linear compounds can have internal nucleotide base complementarity and therefore can fold in such a way as to produce fully or partially double-stranded compounds. Furthermore, within a guide nucleic acid, the phosphate groups can be generally referred to as forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid can be a 3'→5' phosphodiester linkage.

[0137] The guide nucleic acid can have a modified backbone and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone.

[0138] Suitable modified guide nucleic acid backbones containing a phosphorus atom therein include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphate triesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates, e.g., 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates, including 3'-aminophosphoramidate and aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates, including those with normal 3'-5' linkages, 2'-5' linked analogs, and those with reversed polarity such that one or more internucleotide linkages are 3'→3', 5'→5', or 2'→2' linkages. Suitable guide nucleic acids with inverted polarity can include a single 3'→3' linkage at the 3'-most internucleotide linkage (such as a single inverted nucleoside residue lacking a nucleobase or having a hydroxyl group instead). Various salts (e.g., potassium chloride or sodium chloride), mixed salts, and free acid forms can also be included.

[0139] The guide nucleic acid may contain one or more phosphorothioate and / or heteroatom internucleoside linkages, specifically, -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (methylene(methylimino) or MMI backbone), -CH2-ON(CH3)-CH2-, -CH2-N(CH3)-N(CH3)-CH2- and -ON(CH3)-CH2-CH2- (where the natural phosphodiester internucleotide linkage is represented as -OP(=O)(OH)-O-CH2-).

[0140] The guide nucleic acid may include a morpholino backbone structure. For example, the nucleic acid may include a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages replace the phosphodiester linkages.

[0141] Guide nucleic acids may contain polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including those with morpholino linkages (some formed in the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and others with a mixture of N, O, S, and CH moieties.

[0142] The guide nucleic acid may include a nucleic acid mimic. The term "mimetic" may be intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring may also be referred to as a sugar surrogate. A heterocyclic base moiety or a modified heterocyclic base moiety can be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid may be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of a polynucleotide can be replaced with an amide-containing backbone, specifically an aminoethylglycine backbone. Nucleotides can be maintained directly or indirectly and linked to the aza nitrogen atoms of the amide portion of the backbone. The backbone in a PNA compound may contain two or more linked aminoethylglycine units, giving the PNA an amide-containing backbone. The heterocyclic base moiety can be directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone.

[0143] Guide nucleic acids can contain linked morpholino units (morpholino nucleic acids) with heterocyclic bases attached to morpholino rings. Linking groups can link the morpholino monomer units of morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds can have less undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of guide nucleic acids. Various compounds within the morpholino class can be joined using different linking groups. A further class of polynucleotide mimics can be termed cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used to synthesize oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with complementary nucleic acids with stability similar to that of native complexes. Further modifications include locked nucleic acids (LNAs), in which the 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom (where n is 1 or 2). LNAs and LNA analogs can exhibit extremely high thermal stability of duplexes with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility.

[0144] The guide nucleic acid may contain one or more substituted sugar moieties. Suitable polynucleotides include those containing OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-alkynyl; or O-alkyl-O-alkyl (where alkyl, alkenyl, and alkynyl are substituted or unsubstituted C1-C6). 10 Alkyl or C2-C 10 In particular, O((CH)n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m are from 1 to about 10. The sugar substituents are C1 to C 10 The group may be selected from lower alkyl, substituted lower alkyl, alkenyl, alkynyl, alkaryl, aralkyl, O-alkaryl or O-aralkyl, SH, SCH, OCN, Cl, Br, CN, CF, OCF, SOCH, SOCH, ONO, NO, N, NH, heterocycloalkyl, heterocycloalkaryl, aminoalkylamino, polyalkylamino, substituted silyl, RNA cleaving group, reporter group, intercalator, group that improves the pharmacokinetic properties of the guide nucleic acid, or group that improves the pharmacodynamic properties of the guide nucleic acid, and other substituents with similar properties. Suitable modifications may include 2'-methoxyethoxy (2'-O-CHCHOCH, also known as 2'-O-(2-methoxyethyl) or 2'-MOE, an alkoxyalkoxy group). Further suitable modifications may include 2'-dimethylaminooxyethoxy, (O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE), 2'-dimethylaminoethoxyethoxy (also known as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), or 2'-O-CH2-O-CH2-N(CH3)2.

[0145] Other suitable sugar substituents include methoxy (-O-CH), aminopropoxy (-OCHCHNH), allyl (-CH-CH=CH), -O-allyl (-O-CH-CH=CH), and fluoro (F). The 2'-sugar substituent may be in the arabino (up) or ribo (down) position. A preferred 2'-arabino modification is 2'-F. Similar modifications can be made at other positions on the oligomeric compound, particularly the 3' position of the sugar on the 3'-terminal nucleoside or in a 2'-5'-linked nucleotide and the 5' position of the 5'-terminal nucleotide. Oligomeric compounds can also have sugar mimetics, such as a cyclobutyl moiety, in place of the pentofuranosyl sugar.

[0146] A guide nucleic acid can also include nucleobase (or "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases can include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine and thymine, 5-uracil (cytosine), Other synthetic and natural nucleobases may be mentioned, such as 8-hydroxyuracil, 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenine and guanine, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-amino-adenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine. Modified nucleobases can include tricyclic pyrimidines, such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), and pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimidin-2-one).

[0147] Heterocyclic base moieties include those in which the purine or pyrimidine base is replaced with other heterocycles, such as 7-deaza-adenine, 7-deazaguanosine, 2-aminopyridine, and 2-pyridone. Nucleobases can be useful for increasing the binding affinity of polynucleotide compounds. These may include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-Methylcytosine substitutions can increase nucleic acid duplex stability by 0.6-1.2°C and may be a preferred base substitution (e.g., when combined with a 2'-O-methoxyethyl sugar modification).

[0148] Modification of the guide nucleic acid may involve chemically linking to the guide nucleic acid one or more moieties or conjugates capable of improving the activity, cellular distribution, or cellular uptake of the guide nucleic acid. These moieties or conjugates may include conjugate groups covalently attached to functional groups, such as primary or secondary hydroxyl groups. Conjugate groups may include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that improve the pharmacodynamic properties of oligomers, and groups that can improve the pharmacokinetic properties of oligomers. Conjugate groups may include, but are not limited to, cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinone, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that improve pharmacodynamic properties include groups that improve uptake, enhance degradation resistance, and / or enhance sequence-specific hybridization with the target nucleic acid. Groups that can improve the pharmacokinetic properties include groups that improve uptake, distribution, metabolism, or excretion of nucleic acids. The conjugate moiety may include, but is not limited to, a lipid moiety, such as a cholesterol moiety, cholic acid, a thioether (e.g., hexyl-S-tritylthiol), a thiocholesterol, an aliphatic chain (e.g., dodecanediol or undecyl residue), a phospholipid (e.g., di-hexadecyl-rac-glycerol or triethylammonium 1,2-di-O-hexadecyl-rac-glycero-3-H-phosphonate), a polyamine or polyethylene glycol chain, or an adamantane acetic acid, a palmityl moiety, or an octadecylamine or hexylamino-carbonyl-oxycholesterol moiety.

[0149] In some embodiments, the at least one guide RNA polynucleotide of the systems or methods provided herein can bind to at least a portion of a genome (e.g., a plant genome) or gene (e.g., a plant gene). In some cases, the at least one guide RNA polynucleotide can form a complex with a Cas12a protein and guide the protein to a portion of a target nucleic acid (e.g., a site in the genome or gene).

[0150] In some embodiments, the systems described herein comprise at least one guide RNA polynucleotide capable of forming a complex with a Cas12a protein or fusion protein of the systems. In some embodiments, the systems described herein comprise at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides capable of forming a complex with a site-specific nuclease portion of a fusion protein of the systems.

[0151] In some embodiments, the guide nucleic acid comprises a nucleotide sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity to any one of SEQ ID NOs: 27-34, as set forth in Table 5.

[0152] [Table 5]

[0153] Also provided herein are kits that include components of the systems described in this disclosure, in some embodiments, the kits include one or more of the fusion proteins and / or polynucleotides described herein.

[0154] VI. Method In another aspect, provided herein are methods for editing one or more nucleic acids using the Cas12a protein, fusion protein, and / or system described herein. In some embodiments, the method comprises contacting the nucleic acid (i.e., the nucleic acid to be edited) with at least one Cas12a protein and / or fusion protein as described herein. In some embodiments, the method further comprises contacting the nucleic acid with a guide RNA (e.g., as described in Section V above) having a region complementary to a selected portion of the nucleic acid. In some embodiments, contacting the nucleic acid with the Cas12a protein and / or fusion protein and the guide RNA results in editing of the nucleic acid. The nucleic acid (i.e., the nucleic acid to be edited) can be any suitable nucleic acid. In some embodiments, the nucleic acid is part of a chromosome. In some embodiments, the nucleic acid is part of a genome (e.g., a plant genome).

[0155] As described herein and demonstrated in the Examples below, the methods provided herein can result in an increased frequency of one or more desired nucleic acid editing outcomes (e.g., SDN-1 editing). In some embodiments, SDN-1 editing efficiency can be measured by dividing the number of plants with an insertion or deletion ("indel") by the total number of transgenic plants. In some embodiments, use of a Cas12a protein or fusion protein provided herein results in an increased SDN-1 editing efficiency compared to use of an unmodified (i.e., wild-type) Cas12a protein. In some embodiments, indel events can be further analyzed for the occurrence of homozygous editing (i.e., the same indel is present in both alleles of the target nucleic acid) and biallelic editing (i.e., a different indel is present in each allele of the target nucleic acid). In some embodiments, the rate of homozygous / biallelic editing can be measured by dividing the number of plants with homozygous / biallelic editing by the total number of plants with indels. In some embodiments, use of a Cas12a protein or fusion protein provided herein results in an increased rate of homozygous / biallelic editing.

[0156] The methods herein include providing a Cas12a protein and / or fusion protein and a nucleic acid to be edited, and may also include providing at least one guide RNA. These various components can be provided using any suitable technique. For example, providing a Cas12a protein or fusion protein can include introducing a Cas12a protein or fusion protein into a cell, or introducing a recombinant nucleic acid, construct, or vector encoding the Cas12a protein or fusion protein into a cell. Similarly, a gRNA can be provided by introducing the gRNA itself or a nucleic acid sequence encoding the gRNA. In some embodiments, the Cas12a protein and / or fusion protein and the gRNA can be encoded by the same DNA construct or vector. [Example]

[0157] Example 1. C965S acts synergistically with D156R to improve the efficiency of SDN1 editing at difficult target sites in maize By analyzing the crystal structure of LbCas12a (PDB entry 5XUS), two surface-exposed cysteine ​​residues (Cys965 and Cys1090) and another (Cys10) near the N-terminus of the native LbCas12a protein (SEQ ID NO: 1) were selected for site-directed mutagenesis (Figure 1). A total of five LbCas12a variants were generated: in the first variant, the Cys965 residue was mutated to a serine residue and designated LbCas12a-C965S; in the second variant, both the Cys10 and Cys965 residues were mutated to serine residues and designated LbCas12a-C10S-C965S; and in the third variant, the Cys965 and Cys1090 residues were mutated to serine residues and designated LbCas12a-C10S-C965S. In the fourth variant, Cys965 was mutated to a serine residue, while Asp156 was mutated to an arginine residue, resulting in LbCas12a-C965S-C1090S; in the fifth variant, only Asp156 was mutated to an arginine residue, resulting in LbCas12a-D156R-C965S. The coding sequence of LbCas12a was optimized based on preferred codon usage in maize, and the codon triplet of the selected cysteine ​​residue, TGC, was mutated to TCC, encoding serine, by introducing mutations in overlapping PCR primers. For all five variants and wild-type LbCas12a as a control, an SV40 NLS (SEQ ID NO: 56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO: 46) peptide linker, while two SV40 NLSs separated by an 8-amino acid (SGGS)2 (SEQ ID NO: 78) peptide linker were similarly fused to the C-terminus via flexible 30-amino acid (GSSSS)6 (SEQ ID NO: 46) peptide linkers.

[0158] For each of the five variants and the wild-type control, we constructed a binary vector to express one variant (or control) in stable transgenic maize plants and evaluated the SDN1 production performance of the variants. In each construct, the coding sequence of one variant fused to an NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in maize cells. In all constructs, the same gRNA array, driven by the rice (Oryza sativa) U6 promoter, was designed to express a gRNA targeting the maize gene Starch Branching Enzyme IIb (ZmSBEIIb). The gRNA was based on the mature crRNA scaffold of LbCas12a. Transgenic maize plants were generated by infecting callus derived from immature maize embryos with Agrobacterium tumefaciens strains carrying one of the binary vectors mentioned above, followed by tissue culture procedures.

[0159] Leaf sheaths of regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan qPCR assay. To determine the genotype and SDN1 efficiency at the target site, sequences spanning the target site were PCR amplified and Sanger sequenced. SDN1 efficiency at the SBEIIb and Wx1 target sites was compared, as summarized in Tables 6 and 7. The SBEIIb target site is difficult to edit with Cas12a. Compared to the wild type, C965S alone slightly improved the overall SDN1 efficiency of SBEIIb but did not improve the rate of homozygous / biallelic editing. Neither C10 nor C1090S showed any positive effect in addition to C965S. In contrast, when paired with D156R, C965S increased overall SDN1 efficiency fourfold over wild-type, with over half resulting in homozygous or biallelic editing, compared with D156R alone, which increased SDN1 efficiency at ZmSBEIIb target sites threefold, with approximately half resulting in homozygous or biallelic editing. For Wx1, SDN1 editing efficiency was similar, if only slightly improved, compared to wild-type.

[0160] [Table 6]

[0161] [Table 7]

[0162] Example 2. C965S acts synergistically with D156R to improve the efficiency of SDN1 editing in soybean To evaluate the effectiveness of C965S in improving the SDN1 induction efficiency of LbCas12a in soybean, two LbCas12a variants, LbCas12a-D156R and LbCas12a-D156R-C965S, were compared for their SDN1 production performance. The two variants are identical to those tested in maize as described in Example 1, except that the coding sequences were optimized based on Arabidopsis preferred codon usage.

[0163] For each variant, two binary vectors were constructed to test SDN1 efficiency at different target loci. In each construct, the coding sequence of one variant fused to an NLS was operably linked to a promoter, such as the Arabidopsis elongation factor 1 alpha (EF1α) promoter, and a terminator, such as the Agrobacterium tumefaciens nopaline synthase gene terminator, for strong constitutive expression in soybean cells. In all constructs, a gRNA or gRNA array driven by the soybean ubiquitin 1 promoter was designed to express one or more gRNAs, such as FAD2, targeted to a site in the soybean genome (SEQ ID NO: 38 provides the LbCas12a gRNA targeting the soybean FAD2-1A gene). One or more gRNAs were based on the mature crRNA scaffold of LbCas12a and processed by self-cleaving ribozymes on the flanks. Transgenic soybean plants are generated by infecting mature soybean seeds with Agrobacterium tumefaciens strains carrying one of the binary vectors described above, followed by tissue culture procedures.

[0164] Leaves from the regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan assay. Sequences spanning each of the three target sites were PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target sites.

[0165] Example 3. Generation and identification of FnCas12a cysteine ​​substitution variants with improved SDN1 editing efficiency in plants A total of nine cysteine ​​residues (Cys70, Cys473, Cys568, Cys717, Cys882, Cys1086, Cys1116, Cys1190, and Cys1196) are present in the FnCas12a primary sequence (SEQ ID NO: 2). The crystal structure of FnCas12a (PDB entries 5NFV and 6I1K) suggested that four cysteine ​​residues (Cys70, Cys473, Cys1116, and Cys1190) are most likely surface-exposed and therefore prone to undesired interactions and / or modifications. Although the surface topography around Cys473 suggests that it is difficult for interacting proteins or modifying enzymes to access, our PyMOL analysis suggests that Cys1116 and Cys1190 likely form intramolecular disulfide bonds (Figure 2). Therefore, Cys70, Cys1116, and Cys1190 were selected for substitution.

[0166] All cysteine ​​substitution variants were generated based on the FnCas12a-E184R variant. Three variants carrying a single Cys to Ser substitution (FnCas12a-E184R-C70S, FnCas12a-E184R-C1116S, and FnCas12a-E184R-C1190S) were generated, as well as three variants carrying a double Cys to Ser substitution (FnCas12a-E184R-C70S-C1116S, FnCas12a-E184R-C70S-C1190S, and FnCas12a-E184R-C1116S-C1190S). The coding sequence of FnCas12a was optimized based on maize codon usage, and the codon triplet of selected cysteine ​​residues, TGC, was mutated to serine-encoding TCC for C1116 and AGC for C70 and C1190 by introducing mutations into overlapping PCR primers. For all six variants and the control FnCas12a-E184R, an SV40 NLS (SEQ ID NO:56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO:46) peptide linker, while two SV40 NLSs, separated by an 8-amino acid (SGGS)2 (SEQ ID NO:78) peptide linker, were similarly fused to the C-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO:46) peptide linker.

[0167] For each of the six variants and the FnCas12a-E184R control, we constructed binary vectors to express one variant (or control) in stable transgenic maize plants and evaluated the SDN1 production performance of the variants. In each construct, the coding sequence of one variant fused to an NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in maize cells. In all constructs, the same gRNA array driven by the rice (Oryza sativa) U6 promoter was designed to express three gRNAs targeting three different maize genes: Waxy1 (ZmWx1), Glossy2 (ZmGL2), and Starch Branching Enzyme IIb (ZmSBEIIb). The gRNA was based on the mature crRNA scaffold of FnCas12a. Transgenic maize plants were generated by infecting callus derived from immature maize embryos with Agrobacterium tumefaciens strains carrying one of the binary vectors mentioned above, followed by a tissue culture procedure.

[0168] Leaf sheaths of regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan assay. Sequences spanning each of the three target sites were PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target sites. The overall SDN1 editing efficiency of each variant and the homozygous / biallelic mutants were compared to those of the FnCas12a-E184R control to assess the effectiveness of cysteine ​​substitution.

[0169] Example 4. Generation and identification of AsCas12a cysteine-substituted variants with improved SDN1 editing efficiency in plants A total of eight cysteine ​​residues (Cys65, Cys205, Cys334, Cys379, Cys608, Cys674, Cys1025, and Cys1248) are present in the AsCas12a primary sequence (SEQ ID NO: 3). The crystal structure of AsCas12a (PDB entry 5KK5) suggests that three cysteine ​​residues: Cys334, Cys379, and Cys674, are most likely surface-exposed and therefore prone to undesired interactions and / or modifications. These three residues were chosen for substitution.

[0170] All cysteine-substituted variants were generated based on the AsCas12a-E174R variant. Three variants carrying a single Cys to Ser substitution (AsCas12a-E174R-C334S, AsCas12a-E174R-C379S, and AsCas12a-E174R-C674S) were generated, as well as three variants carrying a double Cys to Ser substitution (AsCas12a-E174R-C334S-C379S, AsCas12a-E174R-C334S-C674S, and AsCas12a-E174R-C379S-C674S). The coding sequence of AsCas12a was optimized based on preferred codon usage in maize, and the codon triplet of a selected cysteine ​​residue, TGC, was mutated to TCC, encoding serine, by introducing mutations in overlapping PCR primers. For all six variants and the control AsCas12a-E174R, an SV40 NLS (SEQ ID NO:56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO:46) peptide linker, while two SV40 NLSs, separated by an 8-amino acid (SGGS)2 (SEQ ID NO:78) peptide linker, were similarly fused to the C-terminus via flexible 30-amino acid (GSSSS)6 (SEQ ID NO:46) peptide linkers.

[0171] For each of the six variants and the AsCas12a-E174R control, a binary vector was constructed to express one variant (or control) in stable transgenic maize plants, and the SDN1 production performance of the variants was evaluated. In each construct, the coding sequence of one variant fused to an NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in maize cells. In all constructs, the same gRNA array driven by the rice (Oryza sativa) U6 promoter was designed to express three gRNAs targeting three different maize genes: Waxy1 (ZmWx1), Glossy2 (ZmGL2), and Starch Branching Enzyme IIb (ZmSBEIIb). The gRNA was based on the mature crRNA scaffold of AsCas12a. Transgenic maize plants were generated by infecting callus derived from immature maize embryos with Agrobacterium tumefaciens strains carrying one of the binary vectors described above, followed by a tissue culture procedure.

[0172] Leaf sheaths of regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan assay. Sequences spanning each of the three target sites were PCR amplified and Sanger sequenced to determine the genotype and SDN1 efficiency at the target sites. The overall SDN1 editing efficiency of each variant and the homozygous / biallelic mutants were compared to those of the AsCas12a-E174R control to assess the efficacy of cysteine ​​substitution.

[0173] Example 5. Generation and identification of Mb2Cas12a cysteine ​​substitution variants with improved SDN1 editing efficiency in plants Because no crystal structure of Mb2Cas12a (or from Moraxella bovoculi strain 57922) has been published to date, we used the crystal structure of MbCas12a from M. bovoculi strain 22581 (PDB entry 6IV6), which is the closest ortholog with 94.7% amino acid identity to Mb2Cas12a, as a reference structure to predict the locations of cysteine ​​residues in Mb2Cas12a. A total of eight cysteine ​​residues (Cys270, Cys307, Cys583, Cys662, Cys1068, Cys1099, Cys1149, and Cys1162) are present in the primary sequence of Mb2Cas12a from strain 57922 (SEQ ID NO: 4), which correspond to Cys283, Cys320, Cys593, Cys672, Cys1078, Cys1109, Cys1159, and Tyr1172, respectively, in Mb2Cas12a from strain 22581. This prediction suggests that Cys270, Cys307, Cys583, Cys1068, Cys1099, Cys1149, and Cys1162 are likely surface-exposed in Mb2Cas12a. Because Cys1162 in Mb2Cas12a matches Tyr1172 in MbCas12a, Tyr1172 was mutated in 6IV6, and the structure was reconstructed using PyMOL. The resulting structural model suggests that Cys1162 is also likely surface-exposed in Mb2Cas12a. However, the surface topology suggests that Cys1162 is difficult to access for interacting proteins or modifying enzymes. Therefore, Cys270, Cys583, Cys1068, Cys1099, and Cys1149 were selected for site-directed mutagenesis.

[0174] All cysteine ​​substitution variants were generated based on the Mb2Cas12a-D172R variant, which served as a control for the new variants. Five variants carrying a single Cys to Ser substitution (Mb2Cas12a-D172R-C270S, Mb2Cas12a-D172R-C583S, Mb2Cas12a-D172R-C1068S, Mb2Cas12a-D172R-C1099S, Mb2Cas12a-D172R-C1149S) were generated, as well as five variants carrying a quintuple Cys to Ser substitution (Mb2Cas12a-D172R-C270S, Mb2Cas12a-D172R-C583S, Mb2Cas12a-D172R-C1068S, Mb2Cas12a-D172R-C1099S, Mb2Cas12a-D172R-C1149S). We generated one variant carrying a Cys to Ser substitution (Mb2Cas12a-D172R-C270S-C583S-C1068S-C1099S-C1149S) and one variant carrying a quintuple Cys to Ala substitution (Mb2Cas12a-D172R-C270A-C583A-C1068A-C1099A-C1149A). The coding sequence of Mb2Cas12a was optimized based on maize preferred codon usage, and the codon triplet for the selected cysteine ​​residue, TGC, was mutated to TCC, encoding serine, by introducing mutations into overlapping PCR primers for five single mutation variants. For the variants with quintuple mutations, Mb2Cas12a was synthesized by replacing TCC, encoding serine, with TGC or GCC, encoding alanine. For all seven variants and the control Mb2Cas12a-D172R, an SV40 NLS (SEQ ID NO: 56) was fused to the N-terminus via a flexible 30-amino acid (GSSSS)6 (SEQ ID NO: 46) peptide linker, while two SV40 NLSs separated by 8-amino acid (SGGS)2 (SEQ ID NO: 78) peptide linkers were similarly fused to the C-terminus via flexible 30-amino acid (GSSSS)6 (SEQ ID NO: 46) peptide linkers.

[0175] For each of the six variants and the Mb2Cas12a-D172R control, we constructed binary vectors to express one variant (or control) in stable transgenic maize plants and evaluated the SDN1 production performance of the variants. In each construct, the coding sequence of one variant fused to an NLS was operably linked to the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator for strong constitutive expression in maize cells. In all constructs, the same gRNA, driven by the sugarcane ubiquitin 4 gene promoter and the Agrobacterium tumefaciens nopaline synthase gene terminator, was designed to express four gRNAs targeting four different maize genes: Waxy1 (ZmWx1), Benzoxazinone synthesis 9 (ZmBx9), Glossy2 (ZmGL2), and ZmBINa. The gRNAs were based on the mature crRNA scaffold of LbCas12a and processed by self-cleaving ribozymes on the flanks. Transgenic maize plants were generated by infecting calli derived from immature maize embryos with Agrobacterium tumefaciens strains carrying one of the binary vectors described above, followed by tissue culture procedures.

[0176] Leaf sheaths of regenerated plantlets were sampled for DNA extraction, and transgenic plants were identified by TaqMan assay. To determine the genotype and SDN1 efficiency at the target site, sequences spanning each of the three target sites were PCR amplified and Sanger sequenced. As summarized in Table 8, all variants with a single Cys-to-Ser mutation increased the rate of homozygous / biallelic mutants compared to the Mb2Cas12a-D172R control. The efficacy of stacking five cysteine ​​mutations was also determined.

[0177] [Table 8]

[0178] Reference sequence list SEQ ID NO:1 - Lachnospiraceae bacterial Cas12a protein (LbCas12a) [ka] SEQ ID NO:2 - Francisella novicida U112 Cas12a protein (FnCas12a) [ka] SEQ ID NO:3 - Acidaminococcus sp. Cas12a protein (AsCas12a) [ka] SEQ ID NO:4 - Moraxella bovoculi strain 57922 Cas12a protein (Mb2Cas12a) [ka] SEQ ID NO:5 - Amino acid sequence of LbCas12a+ linker: [ka] SEQ ID NO:6 - Amino acid sequence of LbCas12a D156R: [ka] SEQ ID NO:7 - Amino acid sequence of LbCas12a+D156R+C965S: [ka] SEQ ID NO:8 - Amino acid sequence of LbCas12a+C10S+C965S: [ka] SEQ ID NO:9 - Amino acid sequence of LbCas12a+C965S+C1090S: [ka] SEQ ID NO:10 - Amino acid sequence of LbCas12a+linker+D156R: [ka] SEQ ID NO:11 - Amino acid sequence of LbCas12a+linker+D156R+C965S: [ka] SEQ ID NO:12 - Amino acid sequence of Mb2Cas12a+linker+D172R: [ka] SEQ ID NO:13 - Amino acid sequence of Mb2Cas12a+linker+D172R+C270S: [ka] SEQ ID NO:14 - Amino acid sequence of Mb2Cas12a+linker+D172R+C583S: [ka] SEQ ID NO:15 - Amino acid sequence of Mb2Cas12a+linker+D172R+C1068S: [ka] SEQ ID NO:16 - Amino acid sequence of Mb2Cas12a+linker+D172R+C1099S: [ka] SEQ ID NO:17 - Amino acid sequence of Mb2Cas12a+linker+D172R+C1149S: [ka] SEQ ID NO:18 - Amino acid sequence of Mb2Cas12a+linker+D172R+C270S+C583S+C1068S+C1099S+C1149S: [ka] SEQ ID NO:19 - Amino acid sequence of Mb2Cas12a+linker+D172R+C270A+C583A+C1068A+C1099A+C1149A: [ka] SEQ ID NO:20 - Maize codon-optimized nucleic acid sequence encoding LbCas12a+ linker: [ka] [ka] [ka] SEQ ID NO:21 - Maize codon-optimized nucleic acid sequence encoding LbCas12a+linker+D156R: [ka] [ka] SEQ ID NO:22 - Maize codon-optimized nucleic acid sequence encoding LbCas12a+linker+D156R+C965S: [ka] [ka] SEQ ID NO:23 - Maize codon-optimized nucleic acid sequence encoding LbCas12a+linker+C10S+C965S: [ka] [ka] SEQ ID NO:24 - Maize codon-optimized nucleic acid sequence encoding LbCas12a+linker+C965S+C1090S: [ka] [ka] SEQ ID NO:25 - Arabidopsis codon-optimized nucleic acid sequence encoding LbCas12a+linker+D156R: [ka] [ka] [ka] SEQ ID NO:26 - Arabidopsis codon-optimized nucleic acid sequence encoding LbCas12a + linker + D156R + C965S: [ka] [ka] SEQ ID NO:27 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R: [ka] [ka] SEQ ID NO:28 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C270S: [ka] [ka] [ka] SEQ ID NO:29 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C583S: [ka] [ka] SEQ ID NO:30 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C1068S: [ka] [ka] [ka] SEQ ID NO:31 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C1099S: [ka] [ka] SEQ ID NO:32 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C1149S: [ka] [ka] SEQ ID NO:33 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C270S+C583S+C1068S+C1099S+C1149S: [ka] [ka] SEQ ID NO:34 - Maize codon-optimized nucleic acid sequence encoding Mb2Cas12a+linker+D172R+C270A+C583A+C1068A+C1099A+C1149A: [ka] [ka]

[0179] [Table 9]

[0180] [Table 10]

[0181] [Table 11]

[0182] [Table 12]

[0183] All patents, patent publications, patent applications, journal articles, books, technical references, etc. discussed in this disclosure are hereby incorporated by reference in their entirety for all purposes.

[0184] It should be understood that the figures and descriptions of the present disclosure are simplified to illustrate elements relevant to a clear understanding of the present disclosure. It should be recognized that these figures are presented for illustrative purposes and not as structural diagrams. Omitted details and variations or alternative embodiments are within the understanding of those skilled in the art.

[0185] It can be recognized that in certain aspects of the present disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component, to provide an element or structure or to perform one or more given functions. Except to the extent that such substitution would not be functional in practicing a particular embodiment of the present disclosure, such substitution is considered to be within the scope of the present disclosure.

[0186] The examples shown herein are intended to illustrate possible specific implementations of the present disclosure. It will be understood that the examples are intended primarily to illustrate the present disclosure for those skilled in the art. There may be variations in these diagrams or in the operations described herein without departing from the spirit of the present disclosure. For example, in certain cases, method steps or operations may be performed or executed in a different order, or operations may be added, deleted, or modified.

[0187] Where a range of values ​​is provided, it is understood that each intervening value between the upper and lower limits of that range, to the smallest fraction of the lower limit, is also specifically disclosed, unless the context clearly dictates otherwise. Any narrower range between any stated value or unstated intervening value in a stated range and any other stated value or intervening value in that stated range is encompassed. The upper and lower limits of those smaller ranges may independently be included or excluded, and each range where either, neither, or both of the smaller ranges are included is also encompassed within the scope of the present technology, subject to any specifically excluded upper or lower limit in the stated range. Where a stated range includes one or both of its limits, ranges excluding either or both of those included limits are also included.

[0188] The foregoing description sets forth numerous specific details to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the invention described in this disclosure may be practiced without one or more of these specific details. In other instances, well-known features and procedures known to those skilled in the art have not been described to avoid obscuring the invention. The embodiments of the present disclosure have been described for purposes of illustration and not limitation. While the present invention has been described primarily with reference to specific embodiments, it is anticipated that other embodiments will become apparent to those skilled in the art upon reading this disclosure, and such embodiments are intended to be included within the scope of the present method. Therefore, the present disclosure is not limited to the embodiments described above or depicted in the drawings, and various embodiments and modifications can be made without departing from the scope of the following claims.

Claims

1. A Cas12a protein, A sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 1 and contains a human-induced mutation at position C965; or A sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO:2 and comprises human-induced mutations at positions C70, C1116, and / or C1190; or A sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO: 3 and comprises human-induced mutations at positions C334, C379, and / or C674; or A Cas12a protein comprising a sequence that is at least 80% identical to the amino acid sequence of SEQ ID NO:4 and human-induced mutations at positions C270, C583, C1068, C1099, and / or C1149.

2. 2. The Cas12a protein of claim 1, wherein the human-induced mutation is a cysteine ​​to serine substitution.

3. The Cas12a protein described in claim 1, further comprising a human-induced mutation at position D156 of the amino acid sequence of SEQ ID NO:

1.

4. The Cas12a protein described in Claim 3, wherein the human-induced mutation at position D156 of the amino acid sequence of SEQ ID NO: 1 is a substitution of aspartic acid with arginine.

5. The Cas12a protein of any one of claims 1, 2 or 3, wherein the sequence comprises any one of SEQ ID NOs: 5 to 11.

6. The Cas12a protein described in claim 1, further comprising a human-induced mutation at position E184 of the amino acid sequence of SEQ ID NO:

2.

7. The Cas12a protein described in Claim 6, wherein the human-induced mutation at position E184 of the amino acid sequence of SEQ ID NO: 2 is a substitution of glutamic acid with arginine.

8. The Cas12a protein described in claim 1, further comprising a human-induced mutation at position E174 of the amino acid sequence of SEQ ID NO:

3.

9. The Cas12a protein described in Claim 8, wherein the human-induced mutation at position E174 of the amino acid sequence of SEQ ID NO: 3 is a substitution of glutamic acid with arginine.

10. The Cas12a protein described in claim 1, further comprising a human-induced mutation at position D172 of the amino acid sequence of SEQ ID NO:

4.

11. The Cas12a protein described in Claim 10, wherein the human-induced mutation at position D172 of the amino acid sequence of SEQ ID NO: 4 is a substitution of aspartic acid with arginine.

12. 11. The Cas12a protein of any one of claims 1, 2 or 10, wherein the sequence comprises any one of SEQ ID NOs: 12-19.

13. The Cas12a protein of claim 1, wherein the Cas12a protein is a catalytically inactivated Cas12a (dCas12a) protein of a nickase Cas12a (nCas12a) protein.

14. The Cas12a protein of claim 1, further comprising a nuclear localization signal.

15. A fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain.

16. 16. The fusion protein of claim 15, wherein the heterologous domain is a deaminase domain, a transcription factor domain, a nuclease domain, a reverse transcriptase domain, a transposase domain, an integrase domain, a uracil DNA glycosylase inhibitor domain, a recombinase domain, a nickase domain, a methyltransferase domain, a methylase domain, an acetylase domain, an acetyltransferase domain, a transcription activator domain, or a transcription repressor domain.

17. 16. The fusion protein of claim 15, wherein the Cas12a protein is linked to the heterologous domain by a linker sequence.

18. A nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain.

19. The nucleic acid of claim 18, wherein the nucleic acid sequence is any one of SEQ ID NOs: 20 to 34.

20. 20. A DNA construct comprising a promoter operably linked to the nucleic acid of claim 18.

21. A nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain; or A vector comprising a DNA construct comprising a promoter operably linked to a nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain.

22. A nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain. A DNA construct comprising a promoter operably linked to a nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain; or A vector comprising a nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain, or a DNA construct comprising a promoter operably linked to a nucleic acid encoding the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain. including, cells.

23. 23. The cell of claim 22, wherein the cell is a plant cell.

24. 24. The cell of claim 23, wherein the cell is a corn plant cell, a wheat plant cell, a rice plant cell, a soybean plant cell, a sunflower plant cell, or a tomato plant cell.

25. 1. A method for editing nucleic acids, comprising:

10. A method comprising contacting the nucleic acid with (i) the Cas12a protein of claim 1 or a fusion protein comprising the Cas12a protein of claim 1 and a heterologous domain, and (ii) a guide RNA having a region complementary to a selected portion of the nucleic acid, thereby resulting in editing of the nucleic acid.