Mb2Cas12a variants with flexible PAM spectrum
By introducing specific amino acid substitutions in the Mb2Cas12a polypeptide sequence, mutants with a wider PAM recognition spectrum were developed, which solved the problem of insufficient target editing ability of Mb2Cas12a in eukaryotic genomes and achieved efficient editing in plant genomes.
Patent Information
- Application Number
- CN202480009033.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-27
- Filing Date
- 2024-01-24
- Publication Date
- 2025-09-19
AI Technical Summary
The PAM site of Mb2Cas12a has poor flexibility, which limits its ability to edit targets in eukaryotic genomes, especially its high performance at lower temperatures. Therefore, mutants with more flexible PAM sites are needed to expand its application range.
By introducing specific amino acid substitutions in the Mb2Cas12a polypeptide sequence, a variety of mutant Mb2Cas12a polypeptide variants were developed, expanding its PAM recognition spectrum, including novel PAM recognition sequences such as NATV, NCTV, and NGTV, and combined with guide RNA for plant genome editing.
The target editing capability of Mb2Cas12a in eukaryotic genomes has been expanded, the efficiency and flexibility of plant genome editing have been improved, and it can adapt to genome manipulation under different temperature conditions.
Smart Images

Figure BDA0005514635870000291 
Figure BDA0005514635870000301 
Figure BDA0005514635870000321
Abstract
Description
Technical Field
[0001] Described are Cas12a mutants from Moraxella bovoculi AAX08 and methods for their use. These mutants have a broader PAM spectrum than the wild-type enzyme.
[0002] Claim priority
[0003] This application claims priority under 35 U.S.C. §119 to PCT application No. PCT / CN2023 / 073487, filed on January 27, 2023 (the contents of which are incorporated herein by reference in their entirety).
[0004] Sequence Listing
[0005] Attached to this application is a sequence listing named 82857_PCT.xml, which was created on January 11, 2024 and is approximately 80 kilobytes in size. This sequence listing is incorporated herein by reference in its entirety. Background Art
[0006] Mb2Cas12a from Moraxella bovis AAX08 has been shown to have genome editing capabilities in plants (Zhang et al., 2021). However, its PAM site is less flexible than that of Cas9. Therefore, fewer targets can be edited with Mb2Cas12a in eukaryotic genomes compared to Cas9. Mb2Cas12a has the unique property of high performance at lower temperatures, necessitating a flexible PAM. Summary of the Invention
[0007] Therefore, it is necessary to recognize wild-type mutant Mb2Cas12a polypeptide variants of alternative PAM sites. The variant may include at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. The substitution may occur at the following positions: D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
[0008] Described herein are mutant Mb2Cas12a polypeptides comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F and F834Q. In another embodiment, the mutant Mb2Cas12a polypeptide comprises a sequence selected from the group consisting of SEQ ID NOs: 2-14 and 64-66. In one embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV. In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG. In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG. In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of: ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
[0009] Described herein is a method for editing a plant genome, comprising contacting the plant genome with a mutant Mb2Cas12a polypeptide as described above. In one embodiment, the method further comprises a guide RNA. In another embodiment, the guide RNA is encoded by a sequence comprising SEQ ID NO: 16-60.
[0010] Described herein are edited plants obtained by the methods outlined above.
[0011] Described herein are constructs or plasmids comprising polynucleotide sequences encoding the mutant Mb2Cas12a polypeptides listed above. One embodiment is a non-human cell comprising a construct or plasmid.
[0012] Described herein are RNP complexes comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
[0013] Brief description of the sequences in the sequence listing
[0014] SEQ ID NO: 1 is the amino acid sequence of wild-type Mb2Cas12a (also referred to as "wild-type" or "WT" throughout).
[0015] SEQ ID NO: 2 is the amino acid sequence of an Mb2Cas12a variant designated "46(172R)", which comprises D172R, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
[0016] SEQ ID NO: 3 is the amino acid sequence of an Mb2Cas12a variant designated "46(172Y)", which comprises D172Y, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
[0017] SEQ ID NO: 4 is the amino acid sequence of an Mb2Cas12a variant designated "46," which comprises N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
[0018] SEQ ID NO: 5 is the amino acid sequence of an Mb2Cas12a variant designated "51(172R)", which comprises D172R, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
[0019] SEQ ID NO: 6 is the amino acid sequence of an Mb2Cas12a variant designated "51(172Y)", which comprises D172Y, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
[0020] SEQ ID NO:7 is the amino acid sequence of an Mb2Cas12a variant designated "51," which comprises N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
[0021] SEQ ID NO: 8 is the amino acid sequence of an Mb2Cas12a variant designated "18," which comprises D172R, G672S, E678D, and Q725E amino acid substitutions.
[0022] SEQ ID NO:9 is the amino acid sequence of an Mb2Cas12a variant designated "21," comprising D172R, N573D, A629S, L633I, Y636F, L642I, L695I, Q708K, Q725E, F753W, S758D, L782I, and L826F amino acid substitutions.
[0023] SEQ ID NO: 10 is the amino acid sequence of an Mb2Cas12a variant designated "28," which comprises K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions.
[0024] SEQ ID NO: 11 is the amino acid sequence of an Mb2Cas12a variant designated "28(172Y)", which comprises D172Y, K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions.
[0025] SEQ ID NO: 12 is the amino acid sequence of an Mb2Cas12a variant designated "44," which comprises N563R, K569V, N573R, S758E, E795C, T788F, and F824M amino acid substitutions.
[0026] SEQ ID NO: 13 is the amino acid sequence of an Mb2Cas12a variant designated "44(172R)", which comprises D172R, N563R, K569V, N573R, S758E, E759C, T788F, and F824M amino acid substitutions.
[0027] SEQ ID NO: 14 is the amino acid sequence of an Mb2Cas12a variant designated "44(172Y)", which comprises D172Y, N563R, K569V, N573R, S758E, E759C, T788F, and F824M amino acid substitutions.
[0028] SEQ ID NO: 15 is the amino acid sequence of wild-type LbCas12a.
[0029] SEQ ID NO: 16-60 are the gRNA sequences used. See Table 5.
[0030] SEQ ID NO: 61 is a direct repeat of Mb2Cas12a crRNA.
[0031] SEQ ID NO:62 is a synthetic ZmGL2 DNA substrate.
[0032] SEQ ID NO:63 is the synthetic ZmmiR528 DNA substrate.
[0033] SEQ ID NO: 64 is the amino acid sequence of an Mb2Cas12a variant designated "27," comprising K569T, E570R, D572R, N573D, A671M, Q705H, S758E, and E795C amino acid substitutions.
[0034] SEQ ID NO: 65 is the amino acid sequence of an Mb2Cas12a variant designated "43," which comprises N563R, K569V, N573R, S758C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
[0035] SEQ ID NO: 66 is the amino acid sequence of an Mb2Cas12a variant designated "45," which comprises N563R, K569V, N573R, K625R, S696F, S758E, E759T, and T831N amino acid substitutions.
[0036] definition
[0037] Unless otherwise defined below, all technical and scientific terms used herein are intended to have the same meaning as commonly understood by those of ordinary skill in the art. References to the techniques employed herein are intended to refer to techniques commonly understood in the art, including modifications of those techniques and / or equivalent technical alternatives that are clear to those of ordinary skill in the art. Although it is believed that the following terms may be well understood by those of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the subject matter disclosed herein.
[0038] As used herein, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an enzyme" optionally includes combinations of two or more such molecules, etc.
[0039] As used herein, "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0040] As used herein, the term "about" refers to the usual error range for the corresponding value that is readily known to those skilled in the art, for example, ±20%, ±10% or ±5% within the intended meaning of the recited value.
[0041] As used herein, the term "comprising" or "comprises" is open ended. When used in conjunction with a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or amino acid sequence) that includes the subject sequence as a portion or as its entire sequence.
[0042] As used herein, the transition phrase "consisting essentially of" means that the scope of a claim is interpreted to encompass the specified materials or steps recited in the claim, as well as those materials or steps that do not materially affect one or more of the basic and novel characteristics of the claimed subject matter. Therefore, when used in the claims of the present disclosure, the term "consisting essentially of" is not intended to be interpreted as equivalent to "comprising."
[0043] The term "plurality" refers to more than one entity. Thus, a "majority of individuals" refers to at least two individuals. In some embodiments, the term "majority" refers to more than half of a whole. For example, in some embodiments, a "majority of a group" refers to more than half of the members of that group.
[0044] As used herein, term " plant " refers to any plant, particularly seed plant in any developmental stage. As used herein, term " plant cell " refers to the structure and physiological unit of plant, including protoplast and cell wall. Plant cell can be in the form of isolated single cell or cultured cell, or as a part of higher organization unit (such as, plant tissue, plant organ or whole plant). Plant cell can derive from angiosperms or gymnosperms or be a part of them. Plant cell can be monocotyledonous plant cell (for example, maize cell, rice cell, sorghum cell, sugarcane cell, barley cell, wheat cell, oat cell, lawn grass cell or ornamental grass cell) or dicotyledonous plant cell (for example, tobacco cell, pepper cell, eggplant cell, sunflower cell, cruciferous plant cell, flax cell, potato cell, cotton cell, soybean cell, sugar beet cell or rape cell). As used herein, the term "plant cell culture" means a culture of plant units (such as, for example, protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at different developmental stages). As used herein, the term "plant tissue" refers to a group of plant cells organized into structural and functional units. Includes any plant tissue in a plant or in culture. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue cultures, and any plant cell groups organized into structural and / or functional units. The term is not intended to exclude any other type of plant tissue by combining or applying alone with any specific type of plant tissue as listed above or otherwise encompassed by this definition. As used herein, the term "plant part" refers to a part of a plant, including single cells and cell tissues (such as complete plant cells in a plant), cell clumps, and tissue cultures that can regenerate plants. Examples of plant parts include, but are not limited to, single cells and tissues from: pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral parts, fruits, stems, buds, cuttings, and seeds; as well as pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral parts, fruits, stems, buds, cuttings, scions, rhizomes, seeds, protoplasts, callus, and the like.
[0045] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, these terms encompass amino acid chains of any length, including full-length proteins, in which the amino acid residues are linked by covalent peptide bonds.
[0046] The terms "nucleic acid" and "polynucleotide" are used interchangeably and, as used herein, refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in single-stranded or double-stranded form, and polymers thereof, and refer to both the sense and antisense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers thereof. In higher plants, DNA is the genetic material, while RNA is involved in the transfer of information contained within DNA into proteins. A "genome" is the entirety of the genetic material contained in each cell of an organism. It should be understood that when RNA is described, its corresponding cDNA is also described, with uridine represented as thymidine. In specific embodiments, a nucleotide refers to a ribonucleotide, a deoxynucleotide, or a modified form of any type of nucleotide, and combinations thereof. In addition, the polynucleotides disclosed herein may include either or both naturally occurring nucleotides and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide linkages. Nucleic acid molecules may be chemically or biochemically modified, or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those skilled in the art. Such modification includes, for example, tags, methylations, substitutions of one or more naturally occurring nucleotides, internucleotide modifications such as uncharged linkages (for example, methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (for example, phosphorothioates, phosphorodithioates, etc.), side groups (for example, polypeptides), intercalators (for example, acridine, psoralens, etc.), chelating agents, alkylating agents, and modified linkages (for example, α-anomeric nucleic acids, etc.). The above terms are also intended to include any topological conformation, including single-stranded, double-stranded, partially duplexed, triplexed, hairpin-shaped, circular, and padlock-shaped conformations. Unless otherwise indicated, reference to a nucleic acid sequence encompasses its complementary sequence. Therefore, reference to a nucleic acid molecule with a specific sequence is understood to encompass its complementary chain with its complementary sequence. When nucleotide sequences specifically hybridize in solution, these nucleotide sequences are "complementary" (for example, according to Watson-Crick base pairing principles). The term also includes codon-optimized nucleic acids encoding identical polypeptide sequences. It is also understood that nucleic acids can be unpurified, purified, or attached to, for example, a synthetic material such as a bead or column matrix.
[0047] Nucleic acid residues can be represented by individual letters: "A" is adenine; "C" is cytosine; "G" is guanine; "T" is thymine; "N" is any nucleotide; and "V" is the nucleotide A, C, or G.
[0048] In the context of nucleic acid sequences, the term "corresponding to" means that when the nucleic acid sequences of certain sequences are aligned with each other, the nucleic acids "corresponding to" certain enumerated positions in the present invention are those aligned with these positions in the reference sequence, but are not necessarily located in these precise numerical positions relative to the specific nucleic acid sequence of the present invention. Optimal alignment of sequences for comparison can be performed by computerized implementations of known algorithms or by visual inspection. Easily available sequence comparison and multiple sequence alignment algorithms are the Basic Local Alignment Search Tool (BLAST) and ClustalW / ClustalW2 / Clustal Omega programs available on the Internet (e.g., the website of EMBL-EBI), respectively. Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG software package available from Accelrys Corporation (San Diego, California, USA). See also Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; Ausubel et al., 1988; and Sambrook and Russell, 2001.
[0049] Unless otherwise specified, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences, as well as sequences explicitly specified. In particular, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is replaced by mixed bases and / or deoxyinosine residues. See Batzer et al., Nucleic Acids Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).
[0050] As used in the context of polynucleotide or polypeptide sequences described herein, the terms "identity" or "substantial identity" refer to sequences that have at least 60% sequence identity to a reference sequence. Alternatively, the percent identity can be any integer from 60% to 100%. Exemplary embodiments include at least 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% compared to a reference sequence using the programs described herein, preferably BLAST using standard parameters as described below. One skilled in the art will recognize that these values can be appropriately adjusted to determine the corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame positioning, etc.
[0051] For sequence comparison, typically, a kind of sequence serves as the reference sequence that is compared with the test sequence.When using a sequence comparison algorithm, test sequence and reference sequence are input into a computer, and subsequence coordinates are specified if necessary, and sequence algorithm program parameters are specified. Default program parameters can be used, or alternative parameters can be specified. Then, the sequence comparison algorithm will calculate the sequence identity percentage of the test sequence relative to the reference sequence based on the program parameters.
[0052] As used herein, "comparison window" includes reference to a segment having any one of the number of consecutive positions selected from the group consisting of from 20 to 600, typically about 50 to about 200, more typically about 100 to about 150, wherein a sequence can be compared to a reference sequence having the same number of consecutive positions after the two sequences are optimally aligned. Methods of sequence alignment for comparison are well known in the art. Optimal alignment of sequences for comparison can be achieved by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981); the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970); the search by similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. 85:2444 (1988); computer implementations of these algorithms (e.g., BLAST); or by manual alignment and visual inspection.
[0053] A "gene" is a defined region within a genome that, in addition to the aforementioned coding nucleic acid sequence, also contains other major regulatory nucleic acid sequences responsible for controlling the expression (i.e., transcription and translation) of the coding portion. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein, including regulatory sequences. A gene may or may not be used to produce a functional protein. In some embodiments, a gene refers only to the coding region. The term "natural gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene that contains: 1) a DNA sequence including regulatory and coding sequences that are not found together in nature, or 2) a sequence encoding a portion of a protein that is not naturally adjacent, or 3) a portion of a promoter that is not naturally adjacent. Thus, a chimeric gene can contain regulatory and coding sequences obtained from different sources, or regulatory and coding sequences obtained from the same source but arranged in a manner different from that found in nature. A gene can be "isolated," which means a nucleic acid molecule that is substantially or essentially free from components normally associated with the nucleic acid molecule in its natural state. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acid molecules.
[0054] "Isolated" nucleic acid molecules or nucleotide sequence or "isolated" polypeptide are nucleic acid molecules, nucleotide sequence or polypeptides that exist by artificially breaking away from their natural environment and / or have different, modified, regulated and / or changed functions and therefore are not natural products.Isolated nucleic acid molecules or isolated polypeptides can exist in purified form or can be present in non-natural environments (such as, for example, recombinant host cells).Therefore, for example, with respect to polynucleotides, the meaning of the term isolated is to separate polynucleotides from their naturally occurring chromosomes and / or cells therein. If a kind of polynucleotides is separated from its naturally occurring chromosomes and / or cells therein, and then inserted into genetic background, chromosomes, chromosomal position and / or cells therein that are not naturally present in, then the polynucleotides are also isolated.Recombinant nucleic acid molecules and nucleotide sequence of the present invention can be considered as "isolated" as defined above.
[0055] Thus, an "isolated nucleic acid molecule" or "isolated nucleotide sequence" is a nucleic acid molecule or nucleotide sequence that is not immediately adjacent to the nucleotide sequences immediately adjacent to it (the sequences at the 5' end and the sequences at the 3' end) in the naturally occurring genome of the organism from which it is derived. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences that are immediately adjacent to the coding sequence. Thus, the term includes, for example, recombinant nucleic acids that are incorporated into a vector, into a self-replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or that exist as a separate molecule (e.g., a cDNA or genomic DNA fragment generated by PCR or restriction endonuclease treatment) independent of other sequences. It also includes recombinant nucleic acids that are part of a hybrid nucleic acid molecule that encodes another polypeptide or peptide sequence. An "isolated nucleic acid molecule" or "isolated nucleotide sequence" may also include a nucleotide sequence that is derived from and inserted into the same natural original cell type, but is present in a non-natural state, for example, in a different copy number, and / or under the control of regulatory sequences that are different from those found in the natural state of the nucleic acid molecule.
[0056] The term "isolated" may further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide or fragment that is substantially free of cellular material, viral material and / or culture medium (e.g., when produced by recombinant DNA techniques), or chemical precursors or other chemicals (e.g., when chemically synthesized). Additionally, an "isolated fragment" is a fragment of a nucleic acid molecule, nucleotide sequence or polypeptide that does not naturally occur as a fragment and would not so exist in nature. "Isolated" does not necessarily mean that the preparation is industrially pure (homogeneous), but that it is sufficiently pure to provide the polypeptide or nucleic acid in a form that can be used for its intended purpose.
[0057] "Homology-dependent repair" or "homology-directed repair" or "HDR" refers to a mechanism for repairing ssDNA and double-stranded DNA (dsDNA) damage in cells. This repair mechanism can be utilized by cells in the presence of an HDR template with a sequence that has significant homology to the site of damage. The term "perfect HDR" refers to a situation in which the genomic homology junctions in the replaced allele undergo complete HDR, and "imperfect HDR" refers to a situation in which the genomic homology junctions in the replaced allele undergo partial or incomplete HDR. A donor DNA molecule with homology to the cleavage target DNA sequence is used as a template for repairing the cleavage target DNA sequence, thereby generating a transfer of genetic information from the donor polynucleotide to the target DNA. In this way, new nucleic acid material can be inserted / copied into the site. In some cases, the target DNA is contacted with a donor molecule (e.g., a donor DNA molecule). In some cases, the donor DNA molecule is introduced into the cell. In some cases, at least one segment of the donor DNA molecule is integrated into the genome of the cell.
[0058] "Microhomology-mediated end joining" or "MMEJ" or "alternative non-homologous end joining" (Alt-NHEJ) refers to a form of repair of double-strand breaks in DNA. This repair mechanism utilizes microhomology sequences to align the broken strands. "Non-homologous end joining" or "NHEJ" refers to a form of repair of double-strand breaks in DNA. Double-strand breaks are repaired by direct ligation of the broken ends to each other. Generally, no new nucleic acid material is inserted into the site, but some nucleic acid material may be lost or added, resulting in small deletions or insertions.
[0059] The proteins provided herein include site-directed polypeptides. A "site-directed modifying polypeptide" modifies a target DNA (e.g., via cleavage or methylation of the target DNA) and / or a polypeptide associated with the target DNA (e.g., methylation or acetylation of a histone tail). In some embodiments, due to the association of the site-directed modifying polypeptide with the guide RNA, the site-directed modifying polypeptide interacts with the guide RNA (the guide RNA is a single RNA molecule or an RNA duplex of at least two RNA molecules) and is guided to a DNA sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, such as an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.). In some embodiments, the site-directed polypeptide is a site-directed nuclease that is capable of cleaving one or both chains of DNA at a specified target sequence.
[0060] The term "cleavage" or "cleavage" refers to the fracture of the covalent phosphodiester linkage in the ribosylphosphodiester backbone of a polynucleotide, and encompasses both single-strand breaks and double-strand breaks. As a result of two different single-strand cleavage events, double-strand cleavage can occur. Cracks can result in blunt ends or staggered ends (also referred to as sticky ends). A "nuclease cleavage site" or "genomic nuclease cleavage site" is a region of nucleotides where a site-directed nuclease cracks (e.g., when bound to a proximal binding site). When a polynucleotide is DNA (e.g., genomic DNA), one or both chains can be cracked at the nuclease cleavage site. This cracking performed by nucleases triggers a DNA repair mechanism in the cell, which establishes an environment where homologous recombination occurs.
[0061] The site-directed nuclease can be a naturally occurring site-directed nuclease. Exemplary naturally occurring site-directed nucleases are known in the art (see, for example, Makarova et al., 2017, Cell 168:328-328.e1 and Shmakov et al., 2017, Nat Rev Microbiol 15(3):169-182, both of which are incorporated herein by reference). In some embodiments, the site-directed nuclease binds to a DNA-targeting polynucleotide (e.g., a guide RNA) and is thereby directed to a specific sequence within the target DNA and cleaves the target DNA.
[0062] In some embodiments, the site-directed nuclease is modified relative to its native sequence (e.g., via mutation of one or more amino acid residues) to alter its function. For example, the site-directed nuclease can be modified to be enzymatically inactive. The term "enzymatically inactive" can mean that the site-directed nuclease can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner, but cannot cleave the target polynucleotide. An enzymatically inactive site-directed polypeptide can comprise an enzymatically inactive domain (e.g., a nuclease domain). Enzymatically inactive can mean inactive. Enzymatically inactive can mean substantially inactive. Enzymatically inactive can mean essentially inactive. Enzymatically inactive can mean an activity that is no more than 1%, no more than 2%, no more than 3%, no more than 4%, no more than 5%, no more than 6%, no more than 7%, no more than 8%, no more than 9%, or no more than 10% active compared to an exemplary wild-type activity.
[0063] In some embodiments, the site-directed nuclease comprises a CRISPR-associated (Cas) protein or a Cas nuclease that functions in a CRISPR (clustered regularly interspaced short palindromic repeats) / Cas system. In bacteria, this system can provide adaptive immunity against foreign DNA (Barrangou, R. et al., “CRISPR provides acquired resistance against viruses in prokaryotes,” Science (2007) 315:1709-1712; Makarova, KS et al., “Evolution and classification of the CRISPR-Cas systems,” Nat Rev Microbiol (2011) 9:467-477; Garneau, JE et al., “The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA,” Nature (2010) 468:67-71; Sapranauskas, R. et al., “The Streptococcus thermophilus CRISPR / Cas system provides immunity in Escherichia coli,” Nucleic Acids Res (2011) 39:9275-9282). CRISPR / Cas systems (e.g., modified and / or unmodified) can be used as genome engineering tools in a wide variety of organisms, including different mammals, animals, plants, microorganisms, and yeast. The CRISPR / Cas system can include a guide nucleic acid, such as a guide RNA (gRNA), complexed with a Cas protein for targeted regulation of gene expression and / or activity or nucleic acid editing. The RNA-guided Cas protein (e.g., a Cas nuclease such as the Cas9 nuclease) can specifically bind to a target polynucleotide (e.g., DNA) in a sequence-dependent manner.Cas proteins can cleave DNA if they have nuclease activity (Gasiunas, G. et al., “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109: E2579-E286; Jinek, M. et al. “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337: 816-821; Sternberg, SH et al. “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9 [DNA interrogation by the CRISPR RNA-guided endonuclease Cas9],” Nature (2014) 507:62; Deltcheva, E. et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature (2011) 471:602-607). DNA cleavage (e.g., double-strand break) can generate DNA break repair, thereby allowing the introduction of one or more genetic modifications (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), or homology-directed repair (HDR). In some embodiments, donor nucleic acids are used to promote HDR, as described in detail below in the “Systems” section.CRISPR-Cas systems have been widely used for programmable genome editing in a variety of organisms and model systems (Cong, L. et al. “Multiplex genome engineering using CRISPR Cas systems,” Science (2013) 339:819-823; Jiang, W. et al. “RNA-guided editing of bacterial genomes using CRISPR-Cas systems,” Nat. Biotechnol. (2013) 31:233-239; Sander, JD and Joung, JK, “CRISPR-Cas systems for editing, regulating and targeting genomes,” Nature Biotechnol. (2014) 32:347-355).
[0064] In some embodiments, the site-directed nuclease described herein includes a Cas protein (further described below in the "system" section) that forms a complex with a guide nucleic acid (such as a guide RNA). In some embodiments, the site-directed nuclease includes a Cas protein that forms a complex with a single guide nucleic acid (such as a single guide RNA (sgRNA)). In some embodiments, the site-directed nuclease includes an RNA binding protein (RBP) that is optionally compounded with a guide nucleic acid (such as a guide RNA (e.g., sgRNA)), which is capable of forming a complex with the Cas protein. In some cases, the RNA-guided Cas protein recognizes a DNA target that is complementary to a portion of a gRNA (referred to as a CRISPRRNA (crRNA) sequence). The target sequence is often referred to as a protospacer, and the crRNA sequence portion that is complementary to the protospacer is often referred to as a spacer. In order to work (e.g., to crack DNA), many Cas nucleases also require a specific protospacer adjacent motif (PAM) (a kind of approximately 2 to 6 base pair DNA sequence) immediately following the protospacer sequence.
[0065] As used herein, the Cas protein can be an active variant, an inactive variant or a fragment of a wild-type or modified Cas protein. Relative to the wild-type form of the Cas protein, the Cas protein can include amino acid changes, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras or any combination thereof. The Cas protein can be a polypeptide having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity or sequence similarity with a wild-type exemplary Cas protein. The Cas protein can be a polypeptide having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% sequence identity and / or sequence similarity with a wild-type exemplary Cas protein. The variant or fragment can comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or portion thereof. The variant or fragment can be targeted to a nucleic acid locus complexed with a guide nucleic acid while lacking nucleic acid cleavage activity.
[0066] The Cas protein can be modified to optimize the regulation of gene expression. The Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity and / or enzymatic activity. The Cas protein can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted or inactivated, or the Cas protein can be truncated to remove domains that are unnecessary for protein function or to optimize (e.g., enhance or reduce) the activity of the Cas protein in order to regulate gene expression.
[0067] Also provided herein are variants of the polypeptides disclosed herein. Unless otherwise expressly indicated, polypeptide variants retain their corresponding biological activity. For example, variants of site-directed nuclease polypeptides retain the biological function of the full-length native sequence site-directed nuclease. In another example, variants of non-specific end processing enzymes retain the biological function of the full-length native sequence non-specific end processing enzyme.
[0068] The modification of any polypeptide or protein provided herein is carried out by known methods. By way of example, modification is carried out in the following manner: the nucleotides in the nucleic acid encoding the polypeptide are subjected to site-specific mutagenesis, thereby producing a DNA encoding the modification, and then the DNA is expressed in recombinant cell culture to produce the encoded polypeptide. The technology for carrying out substitution mutations at predetermined sites in DNA with known sequences is well known. For example, M13 primer mutagenesis and PCR-based mutagenesis methods can be used to produce one or more substitution mutations. Any nucleotide sequence provided herein can be codon optimized to change, for example, to maximize expression in a host cell or organism.
[0069] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D-stereoisomers of naturally occurring amino acids, non-natural amino acids, and chemically modified amino acids. Non-natural amino acids (i.e., those that are not naturally found in proteins) are also known in the art, as described, for example, in Zhang et al., "Protein engineering with unnatural amino acids," Curr. Opin. Struct. Biol. 23(4):581-587 (2013); Xie et al., "Adding amino acids to the genetic repertoire," 9(6):548-54 (2005); and all references cited therein. Beta and gamma amino acids are known in the art and are also contemplated herein as non-natural amino acids.
[0070] As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, the side chain can be modified to include a signaling moiety, such as a fluorophore or a radiolabel. The side chain can also be modified to include a new functional group, such as a thiol, a carboxylic acid, or an amino group. Post-translationally modified amino acids are also included in the definition of chemically modified amino acids.
[0071] Conservative amino acid substitutions are also contemplated. By way of example, conservative amino acid substitutions can be made in one or more amino acid residues, for example, in one or more lysine residues of any of the polypeptides provided herein. Those skilled in the art will appreciate that a conservative substitution is the replacement of one amino acid residue with another amino acid residue that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservative substitutions for one another:
[0072] 1) Alanine (A), glycine (G);
[0073] 2) Aspartic acid (D), glutamic acid (E);
[0074] 3) Asparagine (N), glutamine (Q);
[0075] 4) Arginine (R), Lysine (K);
[0076] 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);
[0077] 6) Phenylalanine (F), tyrosine (Y), tryptophan (W);
[0078] 7) Serine (S), Threonine (T); and
[0079] 8) Cysteine (C), methionine (M).
[0080] By way of example, when referring to arginine to serine, conservative substitutions of serine are also contemplated (e.g., threonine). Non-conservative substitutions are also contemplated, such as replacing lysine with asparagine.
[0081] Also provided are DNA constructs comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein as described herein or a domain thereof. A nucleic acid is "operably linked" when it is placed in a functional relationship with another nucleic acid sequence. A variety of promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of the start of transcription and involved in the recognition and binding of RNA polymerase and other proteins to initiate transcription.
[0082] As used herein, the term "promoter" refers to a nucleotide sequence, usually upstream (5') of its coding sequence, which controls the expression of the coding sequence by providing the identification of the RNA polymerase and other factors required for appropriate transcription." Promoter regulatory sequence "is composed of proximal and more distal upstream elements. The promoter regulatory sequence affects the transcription, RNA processing or stability or translation of the relevant coding sequence. Regulatory sequences include enhancers, promoters, non-translated leader sequences, introns and polyadenylation signal sequences. They include natural sequences and synthetic sequences and sequences that may be combinations of synthetic sequences and natural sequences. "Enhancer" is a DNA sequence that can stimulate promoter activity and can be an intrinsic element of the promoter or an inserted heterologous element to enhance the level or tissue specificity of the promoter. It can operate in two orientations (forward or reverse) and can even play a role when moved to the upstream or downstream of the promoter. The meaning of the term "promoter" includes "promoter regulatory sequence".
[0083] The choice of promoter to be included depends on several factors, including but not limited to efficiency, selectability, inducibility, desired expression level, and cell or tissue preferential expression. It is routine for those skilled in the art to regulate the expression of a sequence by appropriately selecting and positioning promoters and other regulatory regions relative to the sequence.
[0084] It has been shown that certain promoters can direct RNA synthesis at a higher rate than others. These are referred to as "strong promoters." It has been shown that certain other promoters direct RNA synthesis at higher levels only in specific types of cells or tissues, and are often referred to as "tissue-specific promoters," or if a promoter preferentially directs RNA synthesis in certain tissues (where RNA synthesis may occur at a reduced level in other tissues), it is referred to as a "tissue-preferred promoter." Because promoters are used to control the expression pattern of one (or more) chimeric genes introduced into plants, there is ongoing interest in isolating novel promoters that can control the expression of one (or more) chimeric genes at certain levels in specific tissue types or at specific plant developmental stages.
[0085] Some promoters can instruct RNA synthesis at relatively similar levels in all tissues of a plant. These are called "constitutive promoters" or "non-tissue-dependent" promoters. Constitutive promoters can be divided into strong, medium, and weak categories based on the effectiveness of their instructing RNA synthesis. Because it is necessary to express one (or more) chimeric genes simultaneously in the different tissues of a plant in many cases to obtain the desired function of one (or more) genes, constitutive promoters are particularly useful in this regard. Although many constitutive promoters have been found and characterized from plants and plant viruses, there is still a continued interest in separating more novel constitutive promoters (synthetic or natural), which can control the expression of one (or more) chimeric genes at different levels and control the expression of multiple genes in the same transgenic plant to carry out gene stacking.
[0086] The recombinant nucleic acid provided herein can be included in an expression cassette for expressing in a target host cell or organism. The cassette will include 5' and 3' regulatory sequences operably connected to the recombinant nucleic acid provided herein for allowing the expression of the fusion protein. The cassette can contain at least one other gene or genetic element to be co-transformed into a cell or organism in addition. In the case of including other genes or elements, these components are operably connected. Alternatively, one or more other genes or elements can be provided on a plurality of expression cassettes. This expression cassette is provided with a plurality of restriction sites and / or recombination sites so that the insertion of the polynucleotide is under the transcriptional regulation of the regulatory region. The expression cassette can contain a selective marker gene in addition. The expression cassette will include in the reverse direction of 5' to 3' transcription: transcription and translation initiation region (i.e., promoter), polynucleotides of the present invention, and transcription and translation termination region (i.e., termination region) that work in the target cell or organism. The promoters of the present invention are capable of directing or driving the expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, regardless of whether the RNA is then translated to produce a protein) in a host cell. The regulatory regions (i.e., the promoter, transcriptional regulatory regions, and translational termination regions) may be endogenous or heterologous to the host cell or to each other. As used herein, "heterologous" with respect to a sequence is a sequence that is derived from a foreign species or, if derived from the same species, has been substantially modified from its native form in composition and / or genomic locus by deliberate human intervention.
[0087] Additional regulatory signals include, but are not limited to, transcription initiation start sites, operators, activators, enhancers, other regulatory elements, ribosome binding sites, initiation codons, termination signals, etc. See Sambrook et al., (1992) Molecular Cloning: A Laboratory Manual, Maniatis et al., eds. (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, New York; Davis et al., eds., (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, New York, and references cited therein.
[0088] The expression cassette can also comprise a selective marker gene for selecting transformed cells. For example, marker genes include genes that confer antibiotic resistance, such as those that confer hygromycin resistance, ampicillin resistance, gentamicin resistance, neomycin resistance. Other selective markers are known, and any one can be used.
[0089] In preparing the expression cassette, the various DNA fragments may be manipulated to provide a DNA sequence in the proper orientation and, if appropriate, the proper reading frame. To this end, adapters or linkers may be employed to connect the DNA fragments, or other manipulations may be involved to provide convenient restriction sites, remove excess DNA, remove restriction sites, etc. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, resubstitution (e.g., transitions and transversions) may be involved.
[0090] Further provided is a vector, which comprises the recombinant nucleic acid or DNA construct set forth herein. It is envisioned that the vector has the necessary functional elements for guiding and regulating the transcription of the inserted nucleic acid. These functional elements include, but are not limited to, a promoter, a region upstream or downstream of the promoter (such as an enhancer that can regulate the transcriptional activity of the promoter), an origin of replication, an appropriate restriction site for promoting the cloning of the inserted sequence adjacent to the promoter, an antibiotic resistance gene or other marker that can be used to select cells containing the vector or a vector containing the inserted sequence, an RNA splicing junction, a transcription termination region, or any other region that can be used to promote the expression of an inserted gene or a hybrid gene. Generally referring to, Sambrook et al. Molecular Cloning: A Laboratory Manual [Molecular Cloning: Laboratory Manual], 4th edition, Cold Spring Harbor Laboratory Press [Cold Spring Harbor Laboratory Press], Cold Spring Harbor [Cold Spring Harbor], 2012. The vector can be, for example, a plasmid.
[0091] The conversion of cell can be stable or transient.Therefore, transgenic cell of the present invention, vegetable cell, plant and / or plant part can be by stably transformed or transient transformation." conversion " can refer to nucleic acid molecule being transferred in the genome of host cell, produces genetically stable heredity.In certain embodiments, be introduced into plant, plant part and / or vegetable cell and be via bacteria-mediated conversion, particle bombardment conversion, calcium phosphate-mediated conversion, cyclodextrin-mediated conversion, electroporation, liposome-mediated conversion, nanoparticle-mediated conversion, polymer-mediated conversion, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, ultrasonic treatment, infiltration, polyethylene glycol-mediated conversion, protoplast transformation or make nucleic acid be introduced into any other electrical, chemical, physical and / or biological mechanism or its any combination in plant, plant part and / or its cell and carry out.
[0092] The procedures for transforming plants are well known and conventional in the art and are generally described in the literature. Non-limiting examples of methods for plant transformation include transformation via bacteria-mediated nucleic acid delivery (e.g., via bacteria from the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical) and / or biological mechanism that allows the introduction of nucleic acids into plant cells, including any combination thereof. General guides to various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants," in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, eds. (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7:849-858 (2002)).
[0093] Agrobacterium-mediated transformation is a common method for transforming plants due to its high transformation efficiency and due to its wide applicability with many different species. Agrobacterium-mediated transformation typically involves transferring the binary vector carrying the foreign DNA of interest to an appropriate Agrobacterium strain, which may depend on the complement of the vir genes carried by the host Agrobacterium strain on a co-existing Ti plasmid or chromosomally (Uknes et al., 1993, Plant Cell [plant cells] 5: 159-169). Escherichia coli carrying a recombinant binary vector can be used, and an auxiliary Escherichia coli strain (the auxiliary Escherichia coli strain carries a plasmid capable of moving the recombinant binary vector into the target Agrobacterium strain) is passed through a three-parent mating procedure to achieve the transfer of the recombinant binary vector to Agrobacterium. Alternatively, the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation ( and Willmitzer 1988, Nucleic Acids Res 16:9877).
[0094] Plant transformation by recombinant Agrobacterium typically involves co-cultivation of Agrobacterium with explants from the plant and follows methods well known in the art. Transformed tissue is typically regenerated on selective media carrying an antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders.
[0095] Another method for transforming plants, plant parts and plant cells involves advancing inert or biologically active particles on plant tissues and cells. See, for example, U.S. Patent Nos. 4,945,050; 5,036,006 and 5,100,792. Generally speaking, this method involves advancing inert or biologically active particles at plant cells under conditions effective to penetrate the outer surface of the cell and provide for incorporation into its interior. When utilizing inert particles, the vector can be introduced into the cell by coating the particles with a vector containing the target nucleic acid. Alternatively, one or more cells can be surrounded by the vector so that the vector is carried into the cell by stimulation of the particle. Biologically active particles (e.g., dried yeast cells, dried bacteria or bacteriophages, each containing one or more nucleic acids intended to be introduced) can also be advanced into plant tissue. As used herein, the phrase "biolistic transformation" refers to a method of introducing RNA or DNA directly into cells (e.g., plant cells) in which the RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cells (e.g., plant cells) using high-speed pressure to allow the RNA or DNA to penetrate the cells (e.g., penetrate the plant cell wall).
[0096] The CRISPR / Cas system can also be used to edit the genome of a host cell or organism. As described in detail above, the "CRISPR / Cas" system refers to a class of bacterial systems that are widely used to defend against foreign nucleic acids. Any of the CRISPR / Cas system components described herein can be used to introduce fusion proteins, recombinant nucleic acids or systems into the genome of a host cell or organism. Methods for genome editing mediated by the CRISPR / Cas system are known in the art. It should be understood that the use of the CRISPR / Cas system for introducing fusion proteins, recombinant nucleic acids or systems described herein into the genome of a host cell or organism is different from the specific methods and systems provided herein.
[0097] On the other hand, provided herein are systems that can be used to edit one or more nucleic acids. These systems include one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors or host cells) described above. In certain embodiments, these systems further include one or more other elements that can be used to edit one or more nucleic acids. For example, provided herein are systems that can further include donor polynucleotides. As another example, systems including fusion proteins containing Cas nucleases can further include one or more guide nucleic acids and / or one or more donor polynucleotide sequences.
[0098] In some cases, the systems and methods described herein include at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein include multiple guide nucleic acids. In certain embodiments, the polynucleotide can be deoxyribonucleic acid (DNA). In some cases, the DNA sequence can be single-stranded or double-stranded. In certain embodiments, the at least one guide nucleic acid polynucleotide can be ribonucleic acid (guide RNA).
[0099] In certain embodiments, nuclease can be compounded with at least one guide RNA polynucleotide.The at least one guide RNA polynucleotide can include a nucleic acid targeting region, and the nucleic acid targeting region includes a sequence complementary to the nucleic acid sequence on the targeting polynucleotide (such as a targeting genomic locus or a gene) to give the sequence specificity of the nuclease targeting. In certain embodiments, the guide nucleic acid is a single guide nucleic acid including crRNA. In certain embodiments, the guide nucleic acid is a single guide nucleic acid including crRNA but lacking tracrRNA. The crRNA can include a nucleic acid targeting section (e.g., a spacer region) of the guide nucleic acid and a nucleotide segment of half the double-stranded duplex of the Cas protein binding section of the guide nucleic acid that can form a segment.
[0100] In some embodiments, the length of the nucleic acid targeting region (e.g., a spacer) of a guide nucleic acid is 20 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 19 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 18 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 17 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 16 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 21 nucleotides. In some embodiments, the length of the nucleic acid targeting region of a guide nucleic acid is 22 nucleotides.
[0101] The length of the nucleotide sequence of the guide nucleic acid that is complementary to the nucleotide sequence of the target nucleic acid (target sequence) can be, for example, at least about 12 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt. The length of the nucleotide sequence of the guide nucleic acid that is complementary to the nucleotide sequence of the target nucleic acid (target sequence) can be from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 45 nt, from about 12 nt to about 40 nt, from about 12 nt to about 35 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 19 nt to about 20 ... From about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt.
[0102] The protospacer sequence of the targeting polynucleotide can be identified by identifying the protospacer adjacent motif (PAM) in the region of interest and selecting a region of desired size upstream or downstream of the PAM as the protospacer. The corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.
[0103] Spacer sequences can be identified using a computer program (e.g., machine readable code). The computer program can use variables such as predicted melting temperature, secondary structure formation and predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, GC%, genomic occurrence frequency, methylation status, SNP presence, etc.
[0104] The percent complementarity between a nucleic acid targeting sequence (e.g., a spacer sequence of at least one guide polynucleotide as disclosed herein) and a target nucleic acid (e.g., a protospacer sequence of one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percent complementarity between a nucleic acid targeting sequence and a target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 consecutive nucleotides.
[0105] The guide nucleic acid of the system of the present disclosure can include modifications or sequences that provide additional desired characteristics (e.g., stability of modification or regulation; subcellular targeting; tracking with fluorescent labels; binding sites for proteins or protein complexes, etc.). Examples of such modifications include, for example, a 5' cap (7-methylguanylate cap (m7G)); a 3' polyadenylation tail (3' poly (A) tail); a riboswitch sequence (e.g., to allow regulated stability and / or regulated accessibility of proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (hairpin); a modification or sequence that targets RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, etc.); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a portion that facilitates fluorescence detection, a sequence that allows fluorescence detection, etc.); a modification or sequence that provides a binding site for a protein (e.g., a protein that acts on DNA, including transcriptional activators, transcriptional repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).
[0106] The guide nucleic acid may comprise one or more modifications (e.g., base modifications, backbone modifications) to provide a nucleic acid with new or enhanced characteristics (e.g., improved stability). The guide nucleic acid may comprise a nucleic acid affinity tag. The nucleoside may be a base-sugar combination. The base portion of the nucleotide may be a heterocyclic base. The two most common types of such heterocyclic bases are purines and pyrimidines. The nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides comprising pentofuranosyl sugars, the phosphate group may be linked to the 2', 3', or 5' hydroxyl portion of the sugar. When forming the guide nucleic acid, the phosphate group may covalently link adjacent nucleosides to form a linear polymer compound. Furthermore, the corresponding ends of this linear polymer compound may be further linked to form a cyclic compound; however, a linear compound may be suitable. In addition, the linear compound may have internal nucleotide base complementarity and may therefore be folded in a manner that facilitates the production of a fully or partially double-stranded compound. In addition, within the guide nucleic acid, the phosphate group may generally be involved in forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid can be a 3' to 5' phosphodiester linkage.
[0107] In some embodiments, the at least one guide RNA polynucleotide of the systems or methods provided herein can bind to at least a portion of a genome (e.g., a plant genome) or a gene (e.g., a plant gene). In some cases, the at least one guide RNA polynucleotide is capable of complexing with a site-directed nuclease to direct the site-directed nuclease to target a portion of a target nucleic acid (e.g., a site in a genome or gene).
[0108] In some embodiments, the systems described herein include at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides capable of forming a complex with a site-directed nuclease. DETAILED DESCRIPTION
[0109] The following description and examples illustrate various aspects and embodiments of the compositions and methods of the present invention. The specific examples are not intended to limit the scope of the compositions and methods. Rather, the examples merely provide non-limiting examples of various compositions and methods that are at least within the scope of the disclosed compositions and methods. The descriptions should be read from the perspective of one of ordinary skill in the art; therefore, they do not necessarily include information that would be familiar to a skilled artisan.
[0110] Described herein are mutant Mb2Cas12a polypeptides comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F and F834Q. In another embodiment, the mutant Mb2Cas12a polypeptide comprises a sequence selected from the group consisting of SEQ ID NOs: 2-14 and 64-66. In one embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV. In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG. In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG. In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.In another embodiment, the mutant Mc2Cas12a polypeptide has a PAM recognition sequence selected from the group consisting of: ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
[0111] Described herein is a method for editing a plant genome, comprising contacting the plant genome with a mutant Mb2Cas12a polypeptide as described above. In one embodiment, the method further comprises a guide RNA. In another embodiment, the guide RNA is encoded by a sequence comprising SEQ ID NO: 16-60.
[0112] Described herein are edited plants obtained by the methods outlined above.
[0113] Described herein are constructs or plasmids comprising polynucleotide sequences encoding the mutant Mb2Cas12a polypeptides listed above. One embodiment is a non-human cell comprising a construct or plasmid.
[0114] Described herein are RNP complexes comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
[0115] Examples Example 1. Changing amino acid residues can increase the PAM spectrum of Mb2Cas12a.
[0116] The nucleotide sequence of the wild-type sequence of Mb2Cas12a (encoding SEQ ID NO:1 amino acid sequence) is subjected to various mutagenesis methods, including directed mutagenesis, domain variants and molecular breeding (also referred to as family DNA shuffling), to generate a library of different mutants. In the first layer of screening (including the first and second phases), mutants are first screened using an E. coli plasmid removal assay. The assay uses an E. coli host with a resident plasmid containing the conditional lethal gene ccdB under the control of an arabinose-inducible araC promoter. In the first phase, nine versions of the resident plasmid are needed to cover all variants (i.e., ATA, ATC, ATG, CTA, CTC, CTG, GTA, GTC, GTG) of the VTV PAM flanking the common target sequence. The corresponding crRNA is expressed by the same plasmid. Each of the nine E. coli strains will contain one of the PAM target resident plasmids. The library of the Mb2Cas12a variant under the control of the IPTG-inducible T7 promoter is transformed into a pool of nine strains. In parallel, negative (no Mb2Cas12a) and positive (wild-type Mb2Cas12a) controls were also transformed. The transformation mixture was plated on a culture medium containing ampicillin, kanamycin, IPTG and arabinose. Only cells expressing active Mb2Cas12a can cut the resident plasmid, which is then degraded, thereby eliminating ccdB expression and allowing those cells to survive. The negative control should not produce colonies. Similarly, if wild-type Mb2Cas12a has an absolute requirement for TTV and TTTV PAM, then VTV variant PAM should not be allowed to cut, and no colonies will survive. The Mb2Cas12a expression plasmid encoding a variant that can utilize NVTV PAM is recovered and enters the second stage of selection.
[0117] The second phase utilized a similar screening pool in which three E. coli strains carried resident plasmids containing the degenerate PAM TTV (TTA, TTC, TTG). The results of this two-step process identified variants that reacted with non-classical PAMs but also with wild-type PAMs. The Mb2Cas12a expression plasmids were recovered from surviving colonies as individuals or pools. The pooled plasmid DNA from the surviving colonies was re-transformed into the screening host as a further enrichment step for Mb2Cas12a variants with altered PAM preferences in an iterative process. Plasmid DNA from individual colonies was also subjected to DNA sequencing to determine which amino acid positions had changed.
[0118] In the second-tier screening, a yeast system was used, in part because yeast cells can grow over a wider temperature range (e.g., 20°C-37°C). A yeast strain was generated in which the Ade2 gene was interrupted by the Ura3 gene adjacent to the Mb2Cas12a target site, which was preceded by each of the 12 PAM sequences containing NTV. The Ade2 gene was split in such a way as to leave 200bp of Ade2 sequence homology on either side of the insertion. These homology arms constitute the recombination site. In the absence of functional Ade2, yeast colonies exhibited cell-autonomous red pigmentation. The screening host consisted of a pool of strains as in the above-mentioned bacterial assay. Mb2Cas12a nuclease expression was controlled by the GAL1 galactose-inducible promoter on a plasmid that also expressed crRNA.
[0119] A single Mb2Cas12a variant is transformed into a screening pool. After Mb2Cas12a cutting, single-strand annealing repair produces a complete Ade2 gene and a functional enzyme. The negative control or inactive Mb2Cas12a variant will produce 100% red colonies. Wild-type Mb2Cas12a (if it has strict TTV PAM requirements) will produce 25% white colonies. The ideal Mb2Cas12a variant has a preference for PAM equal to NTV and will produce 100% white colonies. Sectored colonies represent Mb2Cas12a variants and / or PAM sequences that produce incomplete cutting. Therefore, the frequency of red / white / sectored colonies provides an indication of the PAM preference and activity of each Mb2Cas12a variant. Sequence analysis of colonies with various red / white / sectored phenotypes indicates the PAM requirements of each Mb2Cas12a variant.
[0120] The third layer of the assay uses a single crRNA to quantify the PAM preferences of Mb2Cas12a variants. This strategy allows the use of the same target sequence to interrogate each variant PAM, thereby quantifying the efficiency of Mb2Cas12a at different PAM sites. The method is based on a group of yeast strains in which variant PAM sequences containing NTV (or NNTV) are introduced upstream of the Kozak sequence just before the Ade2 coding sequence. The Mb2Cas12a nuclease cuts at 18 / 23 bases downstream of the PAM and produces insertions and deletions, which will make the Ade2 coding sequence out of frame. The red / white phenotype reading and quantification are similar to the previously described assay.
[0121] Table 1. Amino acid residue differences between variants and wild-type Mb2Cas12a.
[0122]
[0123]
[0124] Example 2. In vitro assay for Mb2Cas12a screening.
[0125] To comprehensively screen the activity of Mb2Cas12a on different PAMs, the variants were assayed in vitro. Target DNA containing a random combination of four nucleotides (NNNN) and subsequently verified gRNA sequences were synthesized. The DNA sequences were then incubated with Mb2Cas12a variants. Through carefully designed PCR and NGS sequencing, the cutting efficiency of each NNNN (256 combinations in total) was determined, indicating that Mb2Cas12a recognized the sequence.
[0126] The DNA substrates used contained a 5' barcode sequence, four random nucleotide residues, a verified target sequence, and a 3' barcode sequence in the 5' to 3' direction. These DNA substrates were synthesized by Azenta Inc. and are represented by SEQ ID NO: 62 or SEQ ID NO: 63 (ZmGL2 and ZmmiR528, respectively).
[0127] Assemble the RNPs of each variant by incubating the following formulation at 25 °C for 30 min:
[0128] 10x NEB Buffer 3 2ul Cas12a variants 2 pmol sgRNA 1 pmol RNase inhibitors (NEB) 1 μl Nuclease-free water 8μl
[0129] Add reaction buffer and DNA substrate to each assembled RNP and incubate at 37 °C for one hour:
[0130] 10x NEB Buffer 3 0.5ul Assembled RNPs (number 1-number 15) 20ul DNA substrates (No. 1-No. 15) 4.5ul (about 20ng)
[0131] DNA was recovered from the reaction tubes and pooled into one tube for next-generation sequencing.
[0132] Example 3. Determination of PAM profiles in prokaryotic cells.
[0133] In order to determine the PAM recognition spectrum of engineered Mb2Cas12a variants and wild-type and inactive Mb2Cas12a, a PAM determination assay (PAMDA) was performed. PAMDA consists of two components. The first component is an ampicillin-resistant plasmid library of gRNA target sequences, which contains 48 PAMs (NNTV) with identical target sequences: GTGATAAGTGGAATGGCATGTGGG. The second component is an Escherichia coli strain carrying a chloramphenicol-resistant plasmid derived from a low copy number p15A. The plasmid is driven by a T7 promoter to overexpress one of engineered Mb2Cas12a variants, wild-type Mb2Cas12a or inactive dMb2Cas12 (D864A, E958A), and a gRNA expression cassette that recognizes the target sequence in the PAM library. The PAM plasmid library is transformed into competent Escherichia coli cells prepared by overexpressing Mb2Cas12a variants and gRNA expression cassettes and further selected on Lb agar plates containing chloramphenicol and ampicillin antibiotics. Therefore, if gRNA recognizes the targeting sequence on the plasmid containing the correct PAM, the plasmid will be cut by Cas12a and further degraded. The surviving plasmid contains a PAM sequence that the co-transformed Mb2Cas12a variant cannot recognize. The surviving colonies are merged and the remaining PAM abundance is analyzed using NGS. The more abundant the PAM, the less cutting, and the less abundant the PAM, the more cutting. The PAM abundance of each Mb2Cas12a variant is normalized to dMb2Cas12a.
[0134] The E. coli screening assay was performed as described in Example 1.
[0135] Table 2. Mb2Cas12a variants with expanded PAM spectra in E. coli assays.
[0136]
[0137]
[0138] ++ indicates good editing activity at that PAM; + indicates some editing activity was observed at that PAM; - indicates no editing was observed at that PAM.
[0139] Example 4. Determination of PAM profiles in eukaryotic cells.
[0140] Maize mesophyll protoplasts were prepared for transfection with vectors expressing Mb2Cas12a variants and gRNAs. See, e.g., MRCoy, et al., Protoplast Isolation and Transfection in Maize, in Protoplast Technology Methods and Protocols, 91-104 (K. Wang and F. Zhang, eds., 2022) for a detailed description of standard protoplast transfection protocols.
[0141] Construction of Mb2 Cas12a expression vector: Mb2 Cas12a containing mutations at specific positions was codon-optimized for maize and commercially synthesized (GenScript, Nanjing, China) and cloned under the promoter of the sugarcane ubiquitin-4 (SoUbi4) gene to generate a plant expression vector.
[0142] Construction of guide RNA expression vector: CRISPR / Cas12a guide RNA transcript is expressed under the control of OsU6 promoter, which targets the native gene of maize. It also includes the direct repeat of Mb2Cas12a crRNA AATTTCTACTGTTTGTAGAT (SEQ ID NO: 61) as a scaffold.
[0143] Transformation of maize protoplasts with expression vectors: Plasmid DNA was purified by plasmid DNA purification using the QIAGEN Plasmid Plus Midi Kit (Cat. No. 12945). 20 μg of Cas12a expression plasmid and 7.5 μg of gRNA expression plasmid (2 pmol each) were mixed into 20 μl solution and transfected into 5 × 10 4 Maize protoplasts were grown in triplicate for each treatment.
[0144] Genomic DNA extraction after cell harvest: After 48 hours, cells were collected, centrifuged at 100g for 5 minutes, and 250μl of supernatant was removed. 100μl of lysis buffer was added and placed on a shaker for 5 minutes at room temperature ("RT"). Centrifuged at 4000rpm for 15 minutes. 100μl of supernatant was transferred to a new round-bottom well and 5μl of beads were added. The sample was placed on a shaker for 5 minutes at RT. Washed twice and air-dried for 5 minutes. DNA was eluted in 100μl of low Tris-EDTA ("TE").
[0145] PCR amplification of target site DNA fragments:
[0146] PCR to amplify a DNA fragment containing the target site using appropriate primers was performed according to the following recipe and PCR conditions.
[0147] gDNA 10 μl Forward primer 0.3 μl (10 μmol) Reverse primer 0.3 μl (10 μmol) 2X PCR mixture 15 μl <![CDATA[H2O]]> 4 μl total 30 μl
[0148] PCR recipe.
[0149]
[0150] PCR thermal cycler conditions.
[0151] PCR product purification protocol: Add Agencourt AMPure XP (54 μl beads to 30 μl PCR). Incubate the mixed sample at room temperature for 5 minutes to achieve maximum recovery. Place the reaction plate on a magnetic plate for 2 minutes to separate the beads from the solution. Dispense 200 μl of 70% ethanol into each well of the reaction plate and incubate at room temperature for 30 seconds. Aspirate and discard the ethanol. Remove the reaction plate from the magnetic plate and add 35 μl of warm water. Incubate for 2 minutes. Transfer the eluate to another clean plate.
[0152] Edited T7EI Assay: Anneal PCR products using the following recipe and thermal cycler conditions: 95°C for 5 min; 95°C to 85°C at 2°C / sec; 85°C to 25°C at 0.1°C / sec; hold at 4°C.
[0153] 10x NEB buffer 2: 1.5 μl
[0154] Purified PCR product: 11 μl
[0155] Then, 2.5ul of diluted 1U / ul T7EI (10X dilution) was added to the above annealed PCR and incubated at 37 degrees for 60 minutes.
[0156] Take 15ul for 2% agarose gel electrophoresis or transfer 2ul to Agilent 5300 fragment analyzer.
[0157] Table 3. PAM profiles of Mb2Cas12a variants in maize protoplasts.
[0158]
[0159]
[0160] Cells with numerical entries show data normalized activity compared to LbCas12a(TTTG). Empty cells reflect no observed editing. "NT" indicates not tested.
[0161] Example 5. List of enzymes used, their mutations and gRNA sequences.
[0162] Table 4. Mutant Mb2Cas12a variants.
[0163]
[0164]
[0165] Table 5. gRNA sequences.
[0166]
[0167]
Claims
1. A mutant Mb2Cas12a polypeptide comprising at least one amino acid substitution introduced into the wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO:
1.
2. The mutant Mb2Cas12a polypeptide of claim 1, wherein the at least one amino acid substitution occurs at a position selected from the group consisting of: D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
3. The mutant Mb2Cas12a polypeptide of claim 2, wherein the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q.
4. mutant Mb2Cas12a polypeptide as claimed in claim 3, wherein the polypeptide comprises a sequence selected from the group comprising SEQ ID NO:2-14 and 64-66.
5. mutant Mc2Cas12a polypeptide as claimed in claim 4, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of: NATV, NCTV, NGTV and NTTV.
6. mutant Mc2Cas12a polypeptide as claimed in claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of: AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC and TATG.
7. mutant Mc2Cas12a polypeptide as claimed in claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of: ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC and TCTG.
8. mutant Mc2Cas12a polypeptide as claimed in claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of: AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC and TGTG.
9. The mutant Mc2Cas12a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
10. A method of editing a plant genome, comprising contacting the plant genome with a mutant Mb2Cas12a polypeptide as described in claims 1-5.
11. The method of claim 10, further comprising a guide RNA.
12. The method of claim 11, wherein the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
13. An edited plant obtained by the method according to claims 10-12.
14. A construct or plasmid comprising a polynucleotide sequence encoding a mutant Mb2Cas12a polypeptide as described in claims 1-5.
15. A non-human cell comprising the construct or plasmid of claim 10.
16. An RNP complex comprising a sequence selected from the group consisting of SEQ ID NOs: 2-14 and 64-66.
Citation Information
Patent Citations
Method for transporting substances into living cells and tissues and apparatus therefor
US4945050A
Method for transporting substances into living cells and tissues and apparatus therefor
US5036006A
Method for transporting substances into living cells and tissues
US5100792A