Mb2Cas12a variants with flexible PAM spectra
Mutant Mb2Cas12a polypeptides with specific amino acid substitutions improve PAM recognition, expanding genome editing targets in plants by enhancing PAM flexibility and enabling effective genome editing.
Patent Information
- Application Number
- JP2025543361
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-27
- Filing Date
- 2024-01-24
- Publication Date
- 2026-01-29
AI Technical Summary
Mb2Cas12a from Moraxella bovoculi AAX08 has a limited PAM site flexibility, restricting its ability to edit targets in eukaryotic genomes compared to Cas9, necessitating the development of mutant variants with broader PAM recognition.
Development of mutant Mb2Cas12a polypeptides with specific amino acid substitutions at positions such as D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, and T831, enabling recognition of alternative PAM sites.
The mutant Mb2Cas12a polypeptides exhibit enhanced PAM recognition, allowing for broader genome editing capabilities in plants, including constructs and RNP complexes for targeted genome editing.
Smart Images

Figure 2026503701000001 
Figure 2026503701000002 
Figure 2026503701000003
Abstract
Description
[Technical Field]
[0001] Cas12a mutants derived from Moraxella bovoculi AAX08 and methods for their use are described. These mutants have a broader PAM spectrum than the wild-type enzyme.
[0002] Priority claims This application claims priority under 35 U.S.C. § 119 of PCT Application No. PCT / CN2023 / 073487, filed January 27, 2023, the entire contents of which are incorporated herein by reference.
[0003] Sequence Listing This application is accompanied by a Sequence Listing entitled 82857_PCT.xml, created on January 11, 2024, which is approximately 80 kilobytes in size. This Sequence Listing is incorporated herein by reference in its entirety. [Background technology]
[0004] Although Mb2Cas12a from Moraxella bovoculi AAX08 has been demonstrated to be capable of genome editing in plants (Zhang et al., 2021), its PAM site is less flexible than that of Cas9. Therefore, fewer targets in eukaryotic genomes can be edited with Mb2Cas12a compared to Cas9. Mb2Cas12a, with its unique low-temperature performance, requires a flexible PAM. Summary of the Invention [Means for solving the problem]
[0005] Thus, there is a need for wild-type mutant Mb2Cas12a polypeptide variants that recognize alternative PAM sites. These variants may include at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. The substitutions may occur at the following positions: D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
[0006] Described herein are mutant Mb2Cas12a polypeptides comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q. In another embodiment, the mutant Mb2Cas12a polypeptide comprises a sequence selected from the group consisting of SEQ ID NOs: 2-14 and 64-66. In one embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV. In another embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG. In another embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG. In another embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
[0007] Described herein are methods for editing a plant genome, comprising contacting the plant genome with the mutant Mb2Cas12a polypeptide described above. In one embodiment, the method further comprises a guide RNA. In another embodiment, the guide RNA is encoded by a sequence comprising any of SEQ ID NOS: 16-60.
[0008] Described herein are edited plants obtained by the above methods.
[0009] Described herein are constructs or plasmids comprising polynucleotide sequences encoding the mutant Mb2Cas12a polypeptides described above. One embodiment is a non-human cell comprising the construct or plasmid.
[0010] Described herein is an RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
[0011] Brief description of the sequences in the sequence listing SEQ ID NO: 1 is the amino acid sequence of wild-type Mb2Cas12a, also referred to throughout as "wild-type" or "WT." SEQ ID NO: 2 is the amino acid sequence of the Mb2Cas12a variant designated "46(172R)," which contains D172R, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions. SEQ ID NO: 3 is the amino acid sequence of the Mb2Cas12a variant designated "46(172Y)," which contains D172Y, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions. SEQ ID NO: 4 is the amino acid sequence of an Mb2Cas12a variant designated "46" containing N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions. SEQ ID NO: 5 is the amino acid sequence of the Mb2Cas12a variant designated "51(172R)," which contains the following amino acid substitutions: D172R, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y. SEQ ID NO: 6 is the amino acid sequence of the Mb2Cas12a variant designated "51(172Y)," which contains the following amino acid substitutions: D172Y, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y. SEQ ID NO: 7 is the amino acid sequence of an Mb2Cas12a variant designated "51," which contains N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions. SEQ ID NO: 8 is the amino acid sequence of an Mb2Cas12a variant designated "18" that contains D172R, G672S, E678D, and Q725E amino acid substitutions. SEQ ID NO: 9 is the amino acid sequence of an Mb2Cas12a variant designated "21," which contains the following amino acid substitutions: D172R, N573D, A629S, L633I, Y636F, L642I, L695I, Q708K, Q725E, F753W, S758D, L782I, and L826F. SEQ ID NO: 10 is the amino acid sequence of an Mb2Cas12a variant designated "28," which contains K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions. SEQ ID NO: 11 is the amino acid sequence of the Mb2Cas12a variant designated "28(172Y)," which contains the following amino acid substitutions: D172Y, K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y. SEQ ID NO: 12 is the amino acid sequence of an Mb2Cas12a variant designated "44" containing N563R, K569V, N573R, S758E, E795C, T788F, and F824M amino acid substitutions. SEQ ID NO: 13 is the amino acid sequence of the Mb2Cas12a variant designated "44(172R)", which contains the following amino acid substitutions: D172R, N563R, K569V, N573R, S758E, E759C, T788F, and F824M. SEQ ID NO: 14 is the amino acid sequence of the Mb2Cas12a variant designated "44(172Y)", which contains the following amino acid substitutions: D172Y, N563R, K569V, N573R, S758E, E759C, T788F, and F824M. SEQ ID NO: 15 is the amino acid sequence of wild-type LbCas12a. SEQ ID NOs: 16 to 60 are the gRNA sequences used. See Table 5. SEQ ID NO: 61 is the direct repeat of Mb2Cas12a crRNA. SEQ ID NO: 62 is a synthetic ZmGL2 DNA substrate. SEQ ID NO: 63 is a synthetic ZmmiR528 DNA substrate. SEQ ID NO: 64 is the amino acid sequence of an Mb2Cas12a variant designated "27" containing K569T, E570R, D572R, N573D, A671M, Q705H, S758E, and E795C amino acid substitutions. SEQ ID NO: 65 is the amino acid sequence of an Mb2Cas12a variant designated "43," which contains N563R, K569V, N573R, S758C, D783Q, T788F, D802L, and F824Y amino acid substitutions. SEQ ID NO: 66 is the amino acid sequence of an Mb2Cas12a variant designated "45" containing N563R, K569V, N573R, K625R, S696F, S758E, E759T, and T831N amino acid substitutions.
[0012] definition All technical and scientific terms used herein are intended to have the same meaning as commonly understood by those skilled in the art, unless otherwise defined below. References to technology used herein are intended to refer to technology as commonly understood in the art, including variations of those technologies and / or equivalent technology alternatives that would be apparent to those skilled in the art. While the following terms are believed to be well understood by those skilled in the art, definitions are provided below to facilitate description of the subject matter of the present disclosure.
[0013] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an enzyme" includes, optionally, a combination of two or more such molecules, and the like.
[0014] As used herein, "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.
[0015] The term "about," as used herein, refers to the normal range of error for the respective value, readily known to one of ordinary skill in the art, e.g., ±20%, ±10%, or ±5% within the intended meaning of the recited value.
[0016] As used herein, the terms "comprising" or "comprise" are open-ended. When used in the context of a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or amino acid sequence) that includes the subject sequence as a portion or in its entirety.
[0017] As used herein, the transitional phrase "consisting essentially of" means that the claim should be construed to include the specified materials or steps recited in the claim and materials or steps that do not materially affect the basic novel characteristic or characteristics of the claimed subject matter. Thus, it is intended that the term "consisting essentially of," when used in the claims of this disclosure, should not be construed as the equivalent of "comprising."
[0018] The term "plurality" refers to more than one entity. Thus, "plurality of individuals" refers to at least two individuals. In some embodiments, the term "plurality" refers to more than half of a total. For example, in some embodiments, a "plurality of a population" refers to more than half of the members of the population.
[0019] The term "plant," as used herein, refers to any plant, particularly a seed plant, at any stage of development. The term "plant cell," as used herein, refers to the structural and physiological unit of a plant, including a protoplast and a cell wall. A plant cell may be in the form of an isolated single cell or a cultured cell, or as part of a more highly organized unit, such as a plant tissue, a plant organ, or a whole plant. A plant cell may be derived from or part of an angiosperm or a gymnosperm. The plant cell may be a monocotyledonous plant cell (e.g., a corn cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turfgrass cell, or an ornamental herb cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, a eggplant cell, a sunflower cell, a cruciferous plant cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugar beet cell, or a rapeseed cell). The term "plant cell culture" as used herein refers to a culture of plant units at various developmental stages, such as, for example, protoplasts, cell culture cells, cells of plant tissue, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos. The term "plant tissue" as used herein refers to a cell culture of a plant, including, for example, a cell of a plant, including, for example, a protoplast, a cell culture cell, a cell of a plant tissue, pollen, a pollen tube, an ovule, an embryo sac, a zygote, and an embryo, at various stages of development. The term "plant tissue" as used herein refers to a cell culture of a plant, including, for example, a plant cell ... "Plant part" refers to a group of plant cells organized into a structural and / or functional unit. It includes any tissue of a plant, whether in planta or in culture. The term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue cultures, and any group of plant cells organized into a structural and / or functional unit. When used in conjunction with or without any specific type of plant tissue, as listed above or otherwise encompassed by this definition, it is not intended to exclude any other type of plant tissue. The term "plant part," as used herein, refers to a part of a plant, including single cells and cell tissues, such as intact plant cells in a plant, cell clumps and tissue cultures that can regenerate plants.Examples of plant parts include, but are not limited to, pollen, ovule, zygote, leaf, embryo, root, root tip, anther, flower, inflorescence, fruit, stem, shoot, cutting, and seed; and single cells and tissues from pollen, ovule, egg cell, zygote, leaf, embryo, root, root tip, anther, flower, inflorescence, fruit, stem, shoot, cutting, scion, rootstock, seed, protoplast, callus, etc.
[0020] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, these terms encompass amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds.
[0021] "Nucleic acid" and "polynucleotide" are used interchangeably and, as used herein, refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) and polymers thereof in either single- or double-stranded form, as well as both sense and antisense strands of RNA, cDNA, genomic DNA, and mitochondrial DNA, as well as synthetic forms and mixed polymers of the above. In higher plants, DNA is the genetic material, while RNA is responsible for transferring the information contained within DNA to proteins. A "genome" is the entire body of genetic material contained in each cell of an organism. When RNA is described, it is understood that its corresponding cDNA is also described, and in cDNA, uridine is represented as thymidine. In certain embodiments, nucleotide refers to ribonucleotides, deoxynucleotides, or modified forms of either type of nucleotide, and combinations thereof. Additionally, polynucleotides disclosed herein can include either or both naturally occurring and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide linkages. As those skilled in the art will readily understand, nucleic acid molecules may be chemically or biochemically modified or may contain non-natural or derivatized nucleotide bases. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications, such as uncharged linkages (e.g., methylphosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), pendant moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, etc.), chelators, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). The above terms are also intended to include any conformation of any topology, including single-stranded, double-stranded, partially duplexed, triplexed, hairpinned, circular, and padlock conformations. A reference to a nucleic acid sequence includes its complement unless otherwise specified.Thus, a reference to a nucleic acid molecule having a particular sequence should be understood to encompass its complementary strand, with its complementary sequence. Nucleotide sequences are "complementary" when they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules). The term also includes codon-optimized nucleic acids that encode the same polypeptide sequence. It is also understood that nucleic acids can be crude, purified, or attached to synthetic materials, such as beads or column matrices.
[0022] Nucleic acid residues may be referred to by individual letters: "A" is adenine; "C" is cytosine; "G" is guanine; "T" is thymine; "N" is any nucleotide; and "V" is the nucleotide A, C, or G.
[0023] The term "corresponding to" in the context of nucleic acid sequences means that when the nucleic acid sequences of a given sequence are aligned with each other, those nucleic acids "corresponding" to a given recited position in the present invention align with those positions in the reference sequence, but are not at their exact numerical positions relative to a particular nucleic acid sequence of the present invention. Optimal alignment of sequences for comparison can be performed by computerized implementation of known algorithms or by visual inspection. Ready-made sequence comparison and multiple sequence alignment algorithms are the Basic Local Alignment Search Tool (BLAST) and ClustalW / ClustalW2 / Clustal Omega programs, respectively, available on the Internet (e.g., the EMBL-EBI website). Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG package available from Accelrys, Inc., San Diego, Calif., United States of America. See also Smith & Waterman, 1981; Needleman & Wunsch, 1970; Pearson & Lipman, 1988; Ausubel et al., 1988; and Sambrook & Russell, 2001.
[0024] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences along with the explicitly indicated sequence. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. See Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).
[0025] The terms "identity" or "substantial identity," when used in the context of polynucleotide or polypeptide sequences described herein, refer to a sequence having at least 60% sequence identity with a reference sequence. Alternatively, the percent identity may be any integer between 60% and 100%. Exemplary embodiments include at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% when compared to a reference sequence using the programs described herein; preferably, BLAST with standard parameters, as described below. One of skill in the art will recognize that these values can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences, taking into account codon degeneracy, amino acid similarity, reading frame alignment, and the like.
[0026] For sequence comparison, typically one sequence serves as a reference sequence, and test sequences are compared with it.When using sequence comparison algorithm, test sequences and reference sequences are input into computer, subsequence coordinates are designated as necessary, and sequence algorithm program parameters are designated.Default program parameters can be used, or alternative parameters can be designated.The sequence comparison algorithm then calculates the percent sequence identity of the test sequence to the reference sequence based on the program parameters.
[0027] As used herein, the term "comparison window" refers to any segment of a number of consecutive positions selected from the group consisting of 20 to 600, usually about 50 to about 200, more usually about 100 to about 150, where a sequence can be compared to a reference sequence of the same number of consecutive positions after the two sequences are optimally aligned. Methods for aligning sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be achieved by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), the similarity search method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85:2444 (1988), computer implementations of these algorithms (e.g., BLAST), or by manual alignment and visual inspection.
[0028] A "gene" is a defined region located within a genome that contains, in addition to the aforementioned coding nucleic acid sequence, other primarily regulatory nucleic acid sequences involved in the expression, i.e., transcription and translation, of the coding portion. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein containing regulatory sequences. A gene may or may not be usable for the production of a functional protein. In some embodiments, a gene refers to only the coding region. The term "native gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene that: 1) contains regulatory and coding sequences that are not found together in nature, or 2) encodes portions of a protein that are not naturally contiguous, or 3) contains portions of a promoter that are not naturally contiguous. Thus, a chimeric gene can contain regulatory and coding sequences from different sources, or regulatory and coding sequences from the same source but arranged in a manner different from that found in nature. A gene may be "isolated," which refers to a nucleic acid molecule that is substantially or essentially free from components normally found associated with the nucleic acid molecule in nature, including other cellular material, culture medium from recombinant production, and / or various chemicals used to chemically synthesize the nucleic acid molecule.
[0029] An "isolated" nucleic acid molecule or nucleotide sequence, or an "isolated" polypeptide, is a nucleic acid molecule, nucleotide sequence, or polypeptide that exists apart from its natural environment by the hand of man and / or has a different, modified, regulated, and / or altered function compared to its function in its natural environment, and is therefore not a product of nature. An isolated nucleic acid molecule or isolated polypeptide can exist in purified form or can exist in a non-native environment (e.g., a recombinant host cell). Thus, for example, with respect to a polynucleotide, the term isolated means that it is separated from the chromosome and / or cell in which it occurs in nature. A polynucleotide is also isolated if it is separated from the chromosome and / or cell in which it occurs in nature and then inserted into a genetic context, chromosome, chromosomal location, and / or cell that is not naturally occurring. Recombinant nucleic acid molecules and nucleotide sequences of the present invention can be considered "isolated" as defined above.
[0030] Thus, an "isolated nucleic acid molecule" or "isolated nucleotide sequence" is a nucleic acid molecule or nucleotide sequence that is not immediately adjacent to the nucleotide sequences (one at the 5' end and one at the 3' end) to which it is immediately adjacent in the naturally occurring genome of the organism from which it originates. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences immediately adjacent to the coding sequence. Thus, the term includes a recombinant nucleic acid that is incorporated into, for example, a vector, an autonomously replicating plasmid, or virus, or the genomic DNA of a prokaryote or eukaryote, or exists as a separate molecule independent of other sequences (e.g., a cDNA or genomic DNA fragment produced by PCR or restriction endonuclease treatment). It also includes a recombinant nucleic acid that is part of a hybrid nucleic acid molecule that encodes an additional polypeptide or peptide sequence. An "isolated nucleic acid molecule" or "isolated nucleotide sequence" can also include a nucleotide sequence that is derived from and inserted into the same natural cell type of origin, but exists in a non-natural state, e.g., in a different copy number and / or under the control of regulatory sequences that differ from those found in the nucleic acid molecule's natural state.
[0031] The term "isolated" can further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide, or fragment (e.g., when produced by recombinant DNA technology), or chemical precursor or other chemical (e.g., when chemically synthesized), that is substantially free of cellular material, viral material, and / or culture medium. Furthermore, an "isolated fragment" is a fragment of a nucleic acid molecule, nucleotide sequence, or polypeptide that is not naturally occurring as a fragment and, as such, would not be found in the natural state. "Isolated" does not necessarily mean that the preparation is technically pure (homogeneous), but is sufficiently pure to provide the polypeptide or nucleic acid in a form that can be used for its intended purpose.
[0032] "Homology-dependent repair" or "homologous recombination repair" or "HDR" refers to a mechanism that repairs ssDNA and double-stranded DNA (dsDNA) damage in cells. This repair mechanism can be used by cells when there is an HDR template with a sequence highly homologous to the damaged site. The term "complete HDR" refers to a situation in which the genomic homology junction in the replaced allele has undergone complete HDR, while "incomplete HDR" refers to a situation in which the genomic homology junction in the replaced allele has undergone partial or incomplete HDR. A donor DNA molecule with homology to the cut target DNA sequence is used as a template to repair the cut target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. Thus, new nucleic acid material can be inserted / copied into the site. In some cases, the target DNA is contacted with a donor molecule, e.g., a donor DNA molecule. In some cases, the donor DNA molecule is introduced into the cell. In some cases, at least a segment of the donor DNA molecule is integrated into the cell's genome.
[0033] "Microhomology-mediated end joining" or "MMEJ" or "alternative non-homologous end joining" (Alt-NHEJ) refers to a form of double-strand break repair in DNA. This repair mechanism utilizes microhomology sequences to align the broken strands. "Non-homologous end joining" or "NHEJ" refers to a form of double-strand break repair in DNA. The double-strand break is repaired by directly ligating the broken ends to each other. Generally, there is no insertion of new nucleic acid material at the site, although some nucleic acid material may be lost or added, resulting in small deletions or small insertions.
[0034] Proteins provided herein include site-specific polypeptides. Site-specific modifying polypeptides modify target DNA (e.g., by cleaving or methylating the target DNA) and / or modify polypeptides associated with the target DNA (e.g., methylation or acetylation of histone tails). In some embodiments, the site-specific modifying polypeptide interacts with a guide RNA, either a single RNA molecule or an RNA duplex of at least two RNA molecules, and is guided to a DNA sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, e.g., an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by virtue of its association with the guide RNA. In some embodiments, the site-specific polypeptide is a site-specific nuclease, which is capable of cleaving one or both strands of DNA at a designated target sequence.
[0035] The term "cleavage" or "cleaving" refers to the breaking of the covalent phosphodiester bond in the ribosyl phosphate diester backbone of a polynucleotide and encompasses both single-strand and double-strand breaks. Double-strand breaks can occur as a result of two separate single-strand cleavage events. Cleavage can result in the creation of either blunt ends or overhanging ends (also known as sticky ends). A "nuclease cleavage site" or "genomic nuclease cleavage site" is a region of nucleotides within which a site-specific nuclease cleaves (e.g., upon binding to a proximal binding site). When the polynucleotide is DNA (e.g., genomic DNA), one or both strands can be cleaved at the nuclease cleavage site. Such cleavage by nuclease enzymes triggers DNA repair mechanisms within the cell, thereby establishing an environment for homologous recombination to occur.
[0036] The site-specific nuclease can be a naturally occurring site-specific nuclease. Exemplary naturally occurring site-specific nucleases are known in the art (see, for example, Makarova et al., 2017, Cell 168:328-328.e1, and Shmakov et al., 2017, Nat Rev Microbiol 15(3):169-182 (both of which are incorporated herein by reference)). In some embodiments, the site-specific nuclease binds to a DNA targeting polynucleotide (e.g., a guide RNA), thereby being guided to a specific sequence within the target DNA, and cleaves the target DNA.
[0037] In some embodiments, a site-specific nuclease is modified from its native sequence (e.g., by mutation or one or more amino acid residues) to alter its function. For example, a site-specific nuclease may be modified to be enzymatically inactive. The term "enzymatically inactive" may refer to a site-specific nuclease that can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner but does not cleave the target polynucleotide. A polypeptide that targets an enzymatically inactive site may contain an enzymatically inactive domain (e.g., a nuclease domain). Enzymatically inactive may refer to no activity. Enzymatically inactive may refer to substantially no activity. Enzymatically inactive may refer to essentially no activity. Enzymatically inactive may refer to no more than 1% activity, no more than 2% activity, no more than 3% activity, no more than 4% activity, no more than 5% activity, no more than 6% activity, no more than 7% activity, no more than 8% activity, no more than 9% activity, or no more than 10% activity compared to the exemplary activity of the wild-type.
[0038] In some embodiments, the site-specific nuclease comprises a CRISPR-associated (Cas) protein or Cas nuclease that functions in a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas system. In bacteria, this system can provide adaptive immunity to foreign DNA (Barrangou, R., et al., "CRISPR provides acquired resistance against viruses in prokaryotes," Science (2007) 315:1709-1712; Makarova, K.S., et al., "Evolution and classification of the CRISPR-Cas systems," Nat Rev Microbiol (2011) 9:467-477; Garneau, J.E., et al., "The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA," Nature (2010) 468:67-71; Sapranauskas, R., et al., "The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli," Nucleic Acids Res (2011) 39:9275-9282). CRISPR / Cas systems (e.g., modified and / or unmodified) can be utilized as genome engineering tools in a wide variety of organisms, including various mammals, animals, plants, microorganisms, and yeast. CRISPR / Cas systems can include a guide nucleic acid, such as a guide RNA (gRNA), complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid editing. RNA-guided Cas proteins (e.g., Cas nucleases, such as Cas9 nuclease) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner.Cas proteins can cleave DNA when they possess nuclease activity (Gasiunas, G., et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109:E2579-E286; Jinek, M., et al., “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337:816-821; Sternberg, SH, et al., “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9,” Nature (2014) 507:62; Deltcheva, E., et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature (2011) 471:602-607). DNA breaks (e.g., double-strand breaks) can result in DNA break repair, allowing for the introduction of one or more genetic modifications (e.g., nucleic acid editing). DNA break repair can occur by non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), or homology-directed repair (HDR). In some embodiments, donor nucleic acids are used to facilitate HDR, as described in more detail in the "System" section below.The CRISPR-Cas system has been widely used for programmable genome editing in various organisms and model systems (Cong, L., et al., "Multiplex genome engineering using CRISPR-Cas systems," Science (2013) 339:819-823; Jiang, W., et al., "RNA-guided editing of bacterial genomes using CRISPR-Cas systems," Nat. Biotechnol. (2013) 31:233-239; Sander, JD & Joung, JK, "CRISPR-Cas systems for editing, regulating, and targeting genomes," Nature Biotechnol. (2014) 32:347-355).
[0039] In some embodiments, the site-specific nucleases described herein comprise a Cas protein complexed with a guide nucleic acid, such as a guide RNA (further described in the "Systems" section below). In some embodiments, the site-specific nuclease comprises a Cas protein complexed with a single guide nucleic acid, such as a single guide RNA (gRNA). In some embodiments, the site-specific nuclease comprises an RNA-binding protein (RBP) optionally complexed with a guide nucleic acid, such as a guide RNA (e.g., sgRNA), capable of forming a complex with the Cas protein. In some examples, the RNA-guided Cas protein recognizes a DNA target complementary to a portion of the gRNA known as the CRISPR RNA (crRNA) sequence. The target sequence is often referred to as the protospacer, and the portion of the crRNA sequence complementary to the protospacer is often referred to as the spacer. To function (e.g., to cleave DNA), many Cas nucleases also require a specific protospacer adjacent motif (PAM), a DNA sequence of approximately 2-6 base pairs immediately following the protospacer sequence.
[0040] As used herein, a Cas protein may be an active variant, an inactive variant, or a fragment of a wild-type or modified Cas protein. The Cas protein may contain amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof, compared to the wild-type version of the Cas protein. The Cas protein may be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to an exemplary wild-type Cas protein. The Cas protein may be a polypeptide with at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or similarity to an exemplary wild-type Cas protein. A variant or fragment may comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or similarity to a wild-type or modified Cas protein or portion thereof. A variant or fragment may lack nucleic acid cleavage activity while being capable of being complexed with a guide nucleic acid and targeted to a nucleic acid locus.
[0041] Cas proteins can be modified to optimize regulation of gene expression. Cas proteins can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to alter any other activity or property of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains that are not essential for protein function or to optimize (e.g., enhance or decrease) the activity of the Cas protein to regulate gene expression.
[0042] Also provided herein are variants of the polypeptides of the present disclosure.Unless otherwise explicitly stated, polypeptide variants retain their respective biological activity.For example, variants of site-specific nuclease polypeptides retain the biological function of full-length native sequence site-specific nucleases.In another example, variants of non-specific endo-processing enzymes retain the biological function of full-length native sequence non-specific endo-processing enzymes.
[0043] Modifications to any of the polypeptides or proteins provided herein can be made by known methods. For example, modifications can be made by site-specific mutagenesis of nucleotides in a nucleic acid encoding the polypeptide, thereby generating DNA encoding the modification, which can then be expressed in recombinant cell culture to produce the encoded polypeptide. Techniques for making substitution mutations at predetermined sites in DNA having a known sequence are well known. For example, M13 primer mutagenesis and PCR-based mutagenesis methods can be used to make one or more substitution mutations. Any of the nucleic acid sequences provided herein can be codon-optimized to alter, e.g., maximize, expression in a host cell or organism.
[0044] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D stereoisomers of naturally occurring amino acids, unnatural amino acids, and chemically modified amino acids. Unnatural amino acids (i.e., those not found in proteins in nature) are also known in the art, as shown, for example, in Zhang et al. "Protein engineering with unnatural amino acids," Curr. Opin. Struct. Biol. 23(4):581-587 (2013); Xie et al. "Adding amino acids to the genetic repertoire," 9(6):548-54 (2005); and all references cited therein. β- and γ-amino acids are known in the art and are also contemplated herein as unnatural amino acids.
[0045] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, the side chain can be modified to include a signaling moiety, such as a fluorophore or a radiolabel. The side chain can also be modified to include a new functional group, such as a thiol, a carboxylic acid, or an amino group. Post-translationally modified amino acids are also included in the definition of chemically modified amino acids.
[0046] Conservative amino acid substitutions are also contemplated. For example, conservative amino acid substitutions can be made at one or more amino acid residues, for example, at one or more lysine residues in any of the polypeptides provided herein. Those skilled in the art will know that a conservative substitution is the replacement of an amino acid residue with another amino acid residue that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservative substitutions for each other: 1) Alanine (A), Glycine (G); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) Cysteine (C), methionine (M).
[0047] For example, when arginine is substituted with serine, conservative substitutions of serine (e.g., threonine) are also contemplated. Non-conservative substitutions, such as substituting lysine with asparagine, are also contemplated.
[0048] Also provided are DNA constructs comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein or domain thereof as described herein. A nucleic acid is "operably linked" when it is placed into a functional relationship with another nucleic acid sequence. Many promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of the start of transcription that is involved in the recognition and binding of RNA polymerase and other proteins to initiate transcription.
[0049] The term "promoter," as used herein, refers to a nucleotide sequence that controls expression of a coding sequence by providing recognition for RNA polymerase and other factors necessary for proper transcription, typically located upstream (5') of that coding sequence. A "promoter regulatory sequence" consists of proximal and more distal upstream elements. Promoter regulatory sequences affect transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. These include natural and synthetic sequences, as well as sequences that may be a combination of natural and synthetic sequences. An "enhancer" is a DNA sequence that can stimulate promoter activity and may be an intrinsic element of a promoter or a heterologous element inserted to increase the level or tissue specificity of a promoter. It is capable of operating in both orientations (e.g., forward or reverse) and can function whether moving upstream or downstream from the promoter. The term "promoter" includes "promoter regulatory sequence."
[0050] The choice of which promoter to include depends on several factors, including, but not limited to, efficiency, selectability, inducibility, desired expression level, and cell- or tissue-preferred expression. It is routine for those of skill in the art to regulate the expression of a sequence by appropriate selection and placement of promoters and other regulatory regions relative to that sequence.
[0051] Certain promoters have been shown to be capable of directing RNA synthesis at higher rates than others. These are called "strong promoters." Certain other promoters have been shown to direct RNA synthesis to higher levels only in specific cell types or tissues; when a promoter preferentially directs RNA synthesis to a particular tissue (while RNA synthesis may occur at lower levels in other tissues), it is often referred to as a "tissue-specific promoter" or "tissue-preferred promoter." Because the expression pattern of a chimeric gene (or genes) introduced into plants is controlled using promoters, there is ongoing interest in isolating new promoters capable of controlling the expression of a chimeric gene (or genes) to a certain level in specific tissue types or at specific plant developmental stages.
[0052] Certain promoters are capable of directing RNA synthesis at relatively similar levels in all tissues of a plant. These are called "constitutive promoters" or "tissue-independent" promoters. Constitutive promoters can be divided into strong, intermediate, and weak categories based on their effectiveness in directing RNA synthesis. Constitutive promoters are particularly useful in this regard, since in many cases, a chimeric gene (or genes) must be expressed simultaneously in different plant tissues to obtain the desired function of the gene (or genes). Although many constitutive promoters have been discovered and characterized from plants and plant viruses, there is still ongoing interest in isolating more novel constitutive promoters, synthetic or natural, capable of controlling the expression of a chimeric gene (or genes) at different levels and expression levels of multiple genes in the same transgenic plant for gene stacking.
[0053] The recombinant nucleic acids provided herein can be included in an expression cassette for expression in a host cell or organism of interest. The cassette will include 5' and 3' regulatory sequences operably linked to the recombinant nucleic acids provided herein, allowing for expression of the fusion protein. The cassette may additionally contain at least one additional gene or genetic element to be cotransformed into the cell or organism. When additional genes or elements are included, the components are operably linked. Alternatively, the additional genes or elements can be provided on multiple expression cassettes. Such expression cassettes comprise multiple restriction and / or recombination sites for insertion of polynucleotides under the transcriptional control of the regulatory regions. The expression cassette may additionally contain a selectable marker gene. The expression cassette will include, in the 5' to 3' transcriptional direction, a transcriptional and translational initiation region (i.e., promoter), a polynucleotide of the invention, and a transcriptional and translational termination region (i.e., termination region) functional in the cell or organism of interest. A promoter of the invention is capable of directing or driving expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA, such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, whether or not that RNA is then translated to produce a protein) in a host cell. The regulatory regions (i.e., promoter, transcriptional regulatory region, and translation termination region) may be endogenous to the host cell or to each other, or heterologous. As used herein, "heterologous" with respect to a sequence is a sequence that is derived from a foreign species, or, if derived from the same species, is substantially altered from its native form in composition and / or genomic locus by deliberate human intervention.
[0054] Additional regulatory signals include, but are not limited to, a start site for transcription initiation, operators, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, termination signals, etc. See Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, NY, and references cited therein.
[0055] The expression cassette may also contain a selectable marker gene for selecting transformed cells. Marker genes include genes that confer antibiotic resistance, such as those that confer hygromycin resistance, ampicillin resistance, gentamicin resistance, and neomycin resistance, to name a few. Additional selectable markers are known, and any may be used.
[0056] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences in the proper orientation and, if necessary, in the proper reading frame. To this end, adapters or linkers can be used to join DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. To this end, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, such as transitions and transversions, can be involved.
[0057] Also provided are vectors containing the recombinant nucleic acids or DNA constructs described herein. It is contemplated that the vectors have the necessary functional elements to direct and regulate the transcription of the inserted nucleic acid. Such functional elements include, but are not limited to, a promoter, a region upstream or downstream of the promoter, such as an enhancer that can regulate the transcriptional activity of the promoter, an origin of replication, a restriction site suitable for facilitating the cloning of an insert adjacent to the promoter, an antibiotic resistance gene or other marker that can be useful for selecting cells containing the vector or a vector containing the insert, an RNA splice junction, a transcription termination region, or any other region that can be useful for promoting the expression of the inserted gene or hybrid gene. Generally, the functional elements described in Sambrook et al., Molecular Cloning: A Laboratory Manual, 4 th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. The vector may be, for example, a plasmid.
[0058] Cellular transformation can be stable or transient. Thus, the transgenic cells, plant cells, plants, and / or plant parts of the present invention can be stably transformed or transiently transformed. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in stable genetic inheritance. In some embodiments, introduction into the plant, plant part, and / or plant cell is via bacterial-mediated transformation, particle bombardment transformation, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, liposome-mediated transformation, nanoparticle-mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infiltration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical, and / or biological mechanism that results in the introduction of a nucleic acid into a plant, plant part, and / or cell thereof, or any combination thereof.
[0059] Plant transformation procedures are well known and routine in the art and are described throughout this literature. Non-limiting examples of plant transformation methods include transformation by bacterial-mediated nucleic acid delivery (e.g., by bacteria from the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical), and / or biological mechanism that results in the introduction of nucleic acid into plant cells, including any combination thereof. General guides to various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants" in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, Eds. (CRC Press, Inc., Boca Raton, 1993), pages 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7:849-858 (2002)).
[0060] Agrobacterium-mediated transformation is a commonly used method for transforming plants due to its high transformation efficiency and its versatility with many different species. Agrobacterium-mediated transformation typically involves transferring a binary vector carrying the foreign DNA of interest into a suitable Agrobacterium strain, which may depend on the complement of vir genes carried by the host Agrobacterium strain either on a coexisting Ti plasmid or on the chromosome (Uknes et al. 1993, Plant Cell 5:159-169). Introduction of the recombinant binary vector into Agrobacterium can be achieved by a triparental mating procedure using Escherichia coli carrying the recombinant binary vector and a helper E. coli strain carrying a plasmid capable of mobilizing the recombinant binary vector into the target Agrobacterium strain. Alternatively, the recombinant binary vector can be transferred into Agrobacterium by nucleic acid transformation (Hoefgen and Willmitzer 1988, Nucleic Acids Res 16:9877).
[0061] Transformation of plants with recombinant Agrobacterium usually involves co-cultivation of Agrobacterium with explants from the plant, followed by methods well known in the art. The transformed tissue carries the antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders and is typically regenerated on selective media.
[0062] Another method for transforming plants, plant parts, and plant cells involves projecting inert or biologically active particles into plant tissues and cells. See, e.g., U.S. Patent Nos. 4,945,050; 5,036,006; and 5,100,792. Generally, this method involves projecting inert or biologically active particles into plant cells under conditions effective to penetrate the outer surface of the cells and cause uptake. When inert particles are used, the vector can be introduced into the cells by coating the particles with a vector containing the nucleic acid of interest. Alternatively, the vector can be surrounded by one or more cells, resulting in the particle being carried into the cells. Biologically active particles (e.g., dried yeast cells, dried bacteria, or bacteriophage, each containing one or more nucleic acids to be introduced) can also be projected into plant tissue. As used herein, the phrase "biolistic transformation" refers to a method of directly introducing RNA or DNA into a cell (e.g., a plant cell) in which the RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cell (e.g., a plant cell) using high-velocity pressure, allowing the RNA or DNA to penetrate the cell (e.g., penetrate the plant cell wall).
[0063] CRISPR / Cas systems can also be used to edit the genome of a host cell or organism. As detailed above, the "CRISPR / Cas" system refers to a broad class of bacterial systems for defense against foreign nucleic acids. Any of the CRISPR / Cas system components described herein can be used to introduce a fusion protein, recombinant nucleic acid, or system into the genome of a host cell or organism. CRISPR / Cas system-mediated genome editing methods are known in the art. It will be understood that the introduction of the fusion proteins, recombinant nucleic acids, or systems described herein into the genome of a host cell or organism using a CRISPR / Cas system differs from the detailed methods and systems provided herein.
[0064] In another aspect, provided herein is a system useful for editing one or more nucleic acids. The system comprises one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above. In some embodiments, the system further comprises one or more additional elements useful for editing one or more nucleic acids. For example, the systems provided herein can further comprise a donor polynucleotide. As another example, a system comprising a fusion protein comprising a Cas nuclease can further comprise one or more guide nucleic acids and / or one or more donor polynucleotide sequences.
[0065] In some cases, the systems and methods described herein include at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein include multiple guide nucleic acids. In some embodiments, the polynucleotide may be deoxyribonucleic acid (DNA). In some cases, the DNA sequence may be single-stranded or double-stranded. In some embodiments, the at least one guide nucleic acid polynucleotide may be a ribonucleic acid (guide RNA).
[0066] In some embodiments, the nuclease can be complexed with at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide can include a nucleic acid targeting region that includes a sequence complementary to a nucleic acid sequence on a targeted polynucleotide, such as a targeted genomic locus or gene, to confer sequence specificity for nuclease targeting. In some embodiments, the guide nucleic acid is a single guide nucleic acid that includes a crRNA. In some embodiments, the guide nucleic acid is a single guide nucleic acid that includes a crRNA but lacks a tracrRNA. The crRNA can include a nucleic acid targeting segment (e.g., a spacer region) of the guide nucleic acid and a stretch of nucleotides that can form one half of a double-stranded duplex of the Cas protein-binding segment of the guide nucleic acid.
[0067] In some embodiments, the nucleic acid targeting region of the guide nucleic acid (e.g., a spacer) is 20 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 19 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 18 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 17 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 16 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 21 nucleotides in length. In some embodiments, the nucleic acid targeting region of the guide nucleic acid is 22 nucleotides in length.
[0068] The nucleotide sequence of the guide nucleic acid complementary to the nucleotide sequence of the target nucleic acid (target sequence) can have a length of, for example, at least about 12 nucleotides (nt) to about 80 nt, about 12 nt to about 50 nt, about 12 nt to about 45 nt, about 12 nt to about 40 nt, about 12 nt to about 35 nt, about 12 nt to about 30 nt, about 12 nt to about 25 nt, about 12 nt to about 20 nt, about 12 nt to about 19 nt, about 19 nt to about 20 ... The length may be about 25 nt, about 19 nt to about 30 nt, about 19 nt to about 35 nt, about 19 nt to about 40 nt, about 19 nt to about 45 nt, about 19 nt to about 50 nt, about 19 nt to about 60 nt, about 20 nt to about 25 nt, about 20 nt to about 30 nt, about 20 nt to about 35 nt, about 20 nt to about 40 nt, about 20 nt to about 45 nt, about 20 nt to about 50 nt, or about 20 nt to about 60 nt.
[0069] The protospacer sequence of a targeted polynucleotide can be identified by identifying a protospacer adjacent motif (PAM) within the region of interest and selecting a region of desired size upstream or downstream of the PAM as the protospacer. The corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.
[0070] Spacer sequences can be identified using a computer program (e.g., machine-readable code) that can use variables such as predicted melting temperature, secondary structure formation, and predicted annealing temperature, sequence identity, genomic context, chromatin exposure, %GC, genomic frequency, methylation status, and the presence of SNPs.
[0071] The percent complementarity between a nucleic acid targeting sequence (e.g., a spacer sequence of at least one guide polynucleotide as disclosed herein) and a target nucleic acid (e.g., a protospacer sequence of one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percent complementarity between a nucleic acid targeting sequence and a target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 contiguous nucleotides.
[0072] Guide nucleic acids of the disclosed systems may include modifications or sequences that provide additional desirable characteristics (e.g., modified or controlled stability; intracellular targeting; tracking by fluorescent labels; binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (7-methylguanylate cap (m7G)); a 3' polyadenylation tail (3' poly(A) tail); a riboswitch sequence (e.g., allowing for regulated stability and / or regulated accessibility by proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (hairpin); a modification or sequence that targets the RNA to a subcellular location (e.g., the nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, etc.); a modification or sequence that provides binding sites for proteins that act on DNA, including proteins (e.g., transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).
[0073] A guide nucleic acid may contain one or more modifications (e.g., base modifications, backbone modifications) to provide the nucleic acid with novel or enhanced characteristics (e.g., improved stability). A guide nucleic acid may contain a nucleic acid affinity tag. A nucleoside may be a base-sugar combination. The base portion of a nucleotide may be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For nucleosides that include a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming a guide nucleic acid, the phosphate group can covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound can then be further joined to form a circular compound; however, linear compounds may also be preferred. In addition, linear compounds can have internal nucleotide base complementarity and thus can fold in such a way as to produce fully or partially double-stranded compounds. Furthermore, within a guide nucleic acid, the phosphate groups can be generally referred to as forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid can be a 3'→5' phosphodiester linkage.
[0074] In some embodiments, the at least one guide RNA polynucleotide of the systems or methods provided herein can bind to at least a portion of a genome (e.g., a plant genome) or gene (e.g., a plant gene). In some cases, the at least one guide RNA polynucleotide is capable of forming a complex with a site-specific nuclease and directing the site-specific nuclease to target a portion of the target nucleic acid (e.g., a site in the genome or gene).
[0075] In some embodiments, the systems described herein comprise at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides capable of forming a complex with a site-specific nuclease moiety. DETAILED DESCRIPTION OF THE INVENTION
[0076] The following description describes various aspects and embodiments of the present compositions and methods. The specific embodiments are not intended to define the scope of the present compositions and methods. Rather, the embodiments merely provide non-limiting examples of various compositions and methods that fall at least within the scope of the disclosed compositions and methods. The description should be read from the perspective of a person skilled in the art, and therefore may not necessarily include information that is known to a person skilled in the art.
[0077] Provided herein are mutant Mb2Cas12a polypeptides comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q. In another embodiment, the mutant Mb2Cas12a polypeptide comprises a sequence selected from the group consisting of SEQ ID NOs: 2-14 and 64-66. In one embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV. In another embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG. In another embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG. In another embodiment, a mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
[0078] Described herein are methods for editing a plant genome, comprising contacting the plant genome with the mutant Mb2Cas12a polypeptide described above. In one embodiment, the method further comprises a guide RNA. In another embodiment, the guide RNA is encoded by a sequence comprising any of SEQ ID NOS: 16-60.
[0079] Described herein are edited plants obtained by the above methods.
[0080] Described herein are constructs or plasmids comprising polynucleotide sequences encoding the mutant Mb2Cas12a polypeptides described above. One embodiment is a non-human cell comprising the construct or plasmid.
[0081] Described herein is an RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66. [Example]
[0082] Example 1. Modifying amino acid residues can increase the PAM spectrum of Mb2Cas12a. To generate a diverse library of mutants, the nucleotide sequence of the wild-type sequence of Mb2Cas12a (encoding the amino acid sequence of SEQ ID NO: 1) was subjected to various mutagenesis methods, including targeted mutagenesis, domain variants, and molecular breeding (also known as family DNA shuffling). Mutants were initially screened using an E. coli plasmid clearance assay in a first-tier screening process, which includes two stages: a first stage and a second stage. This assay uses an E. coli host harboring a resident plasmid containing the conditional lethal gene ccdB under the control of the arabinose-inducible araC promoter. In the first stage, nine versions of the resident plasmid were required to encompass all variations of the VTV PAM (i.e., ATA, ATC, ATG, CTA, CTC, CTG, GTA, GTC, GTG) flanking the common target sequence. The corresponding crRNAs were expressed from the same plasmid. Each of the nine E. coli strains contained one of the PAM-targeting resident plasmids. A library of Mb2Cas12a variants was transformed into a pool of nine strains under the control of an IPTG-inducible T7 promoter. Negative (no Mb2Cas12a) and positive (wild-type Mb2Cas12a) controls were also transformed in parallel. The transformation mixture was plated on medium containing ampicillin, kanamycin, IPTG, and arabinose. Only cells expressing active Mb2Cas12a would be able to cleave the resident plasmid, which would then degrade it, thus eliminating ccdB expression and allowing those cells to survive. The negative control should not yield any colonies. Similarly, if wild-type Mb2Cas12a has an absolute requirement for the TTV and TTTV PAMs, it should not allow cleavage of any of the VTV mutant PAMs, and no colonies would survive. Mb2Cas12a expression plasmids encoding variants capable of utilizing the NVTV PAM were recovered and subjected to second-stage selection.
[0083] The second stage utilized a similar screening pool containing three Escherichia coli (E. coli) strains carrying resident plasmids containing the degenerate PAM TTVs (TTA, TTC, and TTG). The results of this two-stage process identified variants that reacted with the noncanonical PAM but also with the wild-type PAM. Mb2Cas12a expression plasmids were recovered from surviving colonies either individually or as pools. As a further enrichment step for Mb2Cas12a variants with altered PAM preferences in an iterative process, pooled plasmid DNA from surviving colonies was retransformed into the screening host. Plasmid DNA from individual colonies was also subjected to DNA sequencing to determine which amino acid positions were altered.
[0084] For the second-tier screen, we used a yeast system, in part because yeast cells can grow over a wider temperature range (e.g., 20–37°C). We generated yeast strains in which the Ade2 gene was interrupted by the Ura3 gene, flanking the Mb2Cas12a target site, preceded by each of 12 PAM sequences, including the NTV. The Ade2 gene was split to leave 200 bp of Ade2 sequence homology on either side of the insertion. These homology arms constitute the site for recombination. In the absence of functional Ade2, yeast colonies exhibit cell-autonomous red pigmentation. Screening hosts consisted of a pool of strains, as in the bacterial assays described above. Mb2Cas12a nuclease expression was controlled by the GAL1 galactose-inducible promoter on a plasmid that also expressed crRNA.
[0085] Individual Mb2Cas12a variants were transformed into screening pools. Upon Mb2Cas12a cleavage, single-strand annealing repair yielded an intact Ade2 gene and functional enzyme. Negative control or inactive Mb2Cas12a variants would yield 100% red colonies. Wild-type Mb2Cas12a, if it had strict TTV PAM requirements, would yield 25% white colonies. An ideal Mb2Cas12a variant with a PAM preference equivalent to NTV would yield 100% white colonies. Sectioned colonies represent Mb2Cas12a variants and / or PAM sequences that result in incomplete cleavage. Thus, the frequency of red / white / sectioned colonies provides an indication of the PAM preference and activity of each Mb2Cas12a variant. Sequence analysis of colonies with various red / white / sectioning phenotypes demonstrated the PAM requirement for each Mb2Cas12a variant.
[0086] The third tier assay quantified the PAM preference of Mb2Cas12a variants using a single crRNA. This strategy allowed for quantification of Mb2Cas12a efficiency at different PAM sites, using the same target sequence to examine each variant PAM. This approach is based on a set of yeast strains in which a variant PAM sequence containing NTV (or NNTV) was introduced immediately upstream of the Kozak sequence preceding the Ade2 coding sequence. The Mb2Cas12a nuclease cleaves 18 / 23 bases downstream of the PAM, generating an indel that places the Ade2 coding sequence out of frame. Readout and quantification of the red / white phenotype will be similar to previously described assays.
[0087] [Table 1]
[0088] Example 2. In vitro assay for Mb2Cas12a screening. To comprehensively screen the activity of Mb2Cas12a against different PAMs, variants were assayed in vitro. Target DNA containing randomly combined four nucleotides (NNNN) followed by a validated gRNA sequence was synthesized. The DNA sequence was then incubated with Mb2Cas12a variants. Carefully designed PCR and NGS sequencing determined the cleavage efficiency of each NNNN (a total of 256 combinations), indicating Mb2Cas12a recognition of that sequence.
[0089] The DNA substrates used contained, from 5' to 3', a 5' barcode sequence, four random nucleotide residues, a validated target sequence, and a 3' barcode sequence. These DNA substrates were synthesized by Azenta Inc. and are represented by SEQ ID NO: 62 or SEQ ID NO: 63 (ZmGL2 and ZmmiR528, respectively).
[0090] RNPs were assembled for each variant by incubating the following recipe at 25°C for 30 minutes:
[0091] [Table 2]
[0092] Reaction buffer and DNA substrate were added to each assembled RNP and incubated at 37°C for 1 hour:
[0093] [Table 3]
[0094] The DNA was recovered from the reaction tubes and mixed into one tube for next-generation sequencing.
[0095] Example 3. Assaying PAM spectra in prokaryotic cells. To determine the PAM recognition spectrum of engineered Mb2Cas12a variants containing wild-type and dead (inactive) Mb2Cas12a, we performed a PAM determination assay (PAMDA). PAMDA consists of two components. The first component is an ampicillin-resistant plasmid library of gRNA target sequences containing 48 PAMs (NNTVs) with the same target sequence: GTGATAAGTGGAATGGCATGTGGG. The second component is an Escherichia coli (E. coli) strain harboring a low-copy-number p15A-origin chloramphenicol-resistant plasmid overexpressing one of the engineered Mb2Cas12a variants, wild-type Mb2Cas12a, or dead Mb2Cas12a (D864A, E958A), driven by a T7 promoter, and a gRNA expression cassette that recognizes the target sequence in the PAM library. The PAM plasmid library was transformed into competent E. coli cells prepared from the overexpressing Mb2Cas12a variant and gRNA expression cassette, which were further selected on Lb agar plates containing chloramphenicol and ampicillin antibiotics. Thus, if the gRNA recognizes the target sequence on the plasmid containing the correct PAM, the plasmid will be cleaved by Cas12a and further degraded. Surviving plasmids contain a PAM sequence that cannot be recognized by the co-transformed Mb2Cas12a variant. Surviving colonies were pooled and the remaining PAM abundance analyzed using NGS. Higher PAM abundance indicates less cleavage, while lower PAM abundance indicates more cleavage. The PAM abundance of each Mb2Cas12a variant was normalized to dMb2Cas12a.
[0096] The E. coli screening test was carried out as described in Example 1.
[0097] [Table 4-1]
[0098] [Table 4-2]
[0099] Example 4. Assay of PAM spectra in eukaryotic cells. Maize mesophyll protoplasts were prepared for transfection with vectors expressing Mb2Cas12a variants and gRNAs. See, e.g., M.R. Coy, et al., Protoplast Isolation and Transfection in Maize, in Protoplast Technology Methods and Protocols, 91-104 (K. Wang & F. Zhang, eds., 2022) (detailing standard protoplast transfection protocols).
[0100] Construction of Mb2 Cas12a expression vector: Mb2 Cas12a containing specific positional mutations was maize codon-optimized and commercially synthesized (GenScript, Nanjing, China) and cloned under the sugarcane ubiquitin 4 (SoUbi4) gene promoter to generate a plant expression vector.
[0101] Construction of guide RNA expression vector: The CRISPR / Cas12a guide RNA transcript was expressed under the control of the OsU6 promoter, and the CRISPR / Cas12a guide RNA transcript targeted a native maize gene. It also contained a direct repeat of the Mb2Cas12a crRNA AATTTCTACTGTTTGTAGAT (SEQ ID NO: 61) as a scaffold.
[0102] Maize protoplast transformation of expression vectors: Plasmid DNA was purified by plasmid DNA purification using the QIAGEN Plasmid Plus Midi kit (catalog no. 12945). 20 μg of Cas12a expression plasmid and 7.5 μg of gRNA expression plasmid (2 pmol each) were mixed in 20 μL of solution and diluted to 5 × 10 by PEG method. 4Maize protoplast cells were transfected with each treatment, with three transformations and replicates.
[0103] Genomic DNA extraction after cell harvest: Cells were harvested after 48 hours, centrifuged at 100g for 5 minutes, and 250 μL of supernatant was removed. 100 μL of lysis buffer was added and placed on a shaker at room temperature ("RT") for 5 minutes. Centrifuged at 4000 rpm for 15 minutes. 100 μL of supernatant was transferred to a new round-bottom well and 5 μL of beads were added. The sample was placed on a shaker at room temperature for 5 minutes. Washed twice and air-dried for 5 minutes. DNA was eluted in 100 μL of low Tris-EDTA ("TE").
[0104] PCR amplification of target DNA fragments: PCR was performed using appropriate primers to amplify a DNA fragment containing the target site according to the following recipe and PCR conditions.
[0105] [Table 5]
[0106] [Table 6]
[0107] PCR product purification protocol: Add Agencourt AMPure XP (54 μl beads to 30 μl PCR). For maximum recovery, incubate the mixed sample at room temperature for 5 minutes. Place the reaction plate on a magnet for 2 minutes to separate the beads from the solution. Dispense 200 μl of 70% ethanol into each well of the reaction plate and incubate at room temperature for 30 seconds. Aspirate and discard the ethanol. Remove the reaction plate from the magnet and add 35 μl of warm water. Incubate for 2 minutes. Transfer the eluate to another clean plate.
[0108] T7EI detection of edits: Anneal PCR products using the following recipe and thermocycler conditions: 95°C for 5 min; 95 to 85°C, -2°C / sec; 85 to 25°C, -0.1°C / sec; hold at 4°C. 10x NEBB buffer 2: 1.5μl Purified PCR product: 11 μl Then, 2.5 ul of diluted 1 U / ul T7EI (10-fold dilution) is added to the above annealed PCR for 60 minutes at 37 degrees. 15 ul was subjected to 2% Agrose gel electrophoresis or 2 ul was subjected to Agilent 5300 fragment analyzer.
[0109] [Table 7-1]
[0110] [Table 7-2]
[0111] Example 5. List of enzymes and their mutations used, as well as gRNA sequences.
[0112] [Table 8-1]
[0113] [Table 8-2]
[0114] [Table 9-1]
[0115] [Table 9-2]
Claims
1. A mutant Mb2Cas12a polypeptide comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO:
1.
2. 2. The mutant Mb2Casl2a polypeptide of claim 1, wherein the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
3. 3. The mutant Mb2Casl2a polypeptide of claim 2, wherein the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q.
4. 4. The mutant Mb2Cas12a polypeptide of claim 3, wherein the polypeptide comprises a sequence selected from the group consisting of SEQ ID NOs: 2-14 and 64-66.
5. 5. The mutant Mc2Casl2a polypeptide of claim 4, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV.
6. 6. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG.
7. 6. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG.
8. 6. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.
9. 6. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
10. 10. A method for editing a plant genome, comprising contacting the plant genome with a mutant Mb2Cas12a polypeptide of claims 1-5.
11. 11. The method of claim 10, further comprising a guide RNA.
12. 12. The method of claim 11, wherein the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
13. An edited plant obtained by the method of claims 10 to 12.
14. A construct or plasmid comprising a polynucleotide sequence encoding a mutant Mb2Cas12a polypeptide according to claims 1 to 5.
15. A non-human cell comprising the construct or plasmid of claim 10.
16. An RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.