CAS fusion proteins for site-specific integration and related methods
By binding the site-directed nuclease in the fusion protein to the recruitment domain, the problem of the lack of precision of site-directed nucleases in gene editing has been solved, and the efficiency and accuracy of genome editing have been improved. In particular, in the CRISPR/Cas system, the frequency of HDR-mediated repair and the insertion efficiency of donor polynucleotides have been increased.
Patent Information
- Application Number
- CN202380096697.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-10
- Publication Date
- 2025-11-21
AI Technical Summary
Existing site-directed nucleases lack precision in gene editing, especially in CRISPR/Cas systems where the specificity of DSB targeting can vary and the frequency of desired insertion events is low, necessitating improvements in the efficiency of targeted genome editing.
The fusion protein contains a recruitment domain that includes a site-specific nuclease and a site-specific DNA-binding domain. By binding to the recruitment sequence of the donor polynucleotide, it increases the frequency of homology-dependent repair (HDR) and promotes the tethering of the donor polynucleotide to the target site by utilizing the binding site of the recruitment domain to the donor polynucleotide template.
It improves the efficiency and precision of gene editing, increases the possibility of HDR-mediated cleavage repair, promotes the insertion of donor polynucleotides into target sites, and enhances the accuracy and efficiency of editing.
Smart Images

Figure BDA0005620135200000661 
Figure BDA0005620135200000681 
Figure BDA0005620135200000781
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to methods of increasing site-specific integration. The methods presented herein are applicable to both non-homologous end joining (NHEJ) mechanisms and homology-dependent repair (HDR) mechanisms. SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted in ASCII format via EFS-Web and is named 82448SL.xml, created on March 6, 2023, and having a size of about 62.5 kilobytes. The sequence listing is incorporated by reference herein in its entirety. BACKGROUND
[0003] Site-directed nucleases (SDNs) such as zinc finger nucleases, transcription activator-like effector nucleases, CRISPR-associated nucleases are becoming increasingly popular in the space of gene editing. These SDNs act as endonucleases and typically create double-stranded breaks (DSBs) in specific DNA sequences, thereby activating the cell’s intrinsic repair mechanisms (e.g., homologous recombination). During the repair process, site-directed modifications to the specific DNA sequence can be achieved. The CRISPR (clustered regularly interspaced palindromic repeat) / Cas (CRISPR-associated) system evolved in bacteria and archaea as an adaptive immune system to defend against viral attacks. In recent years, the CRISPR / Cas system has attracted particular attention as a tool for genome editing. CRISPR / Cas systems that generate site-specific double-stranded breaks (DSBs) can be used to edit DNA in eukaryotic cells, for example, by creating deletions, insertions, and / or changes in nucleotide sequences.
[0004] Site-directed modifications induced by SDNs often lack precision (e.g., off-target editing can occur), and they often occur at low frequencies. For example, in cases where a CRISPR / Cas system is configured to use a donor template to cause site-specific integration, the specificity of DSB targeting can vary, and the frequency of the desired insertion event can be low. As such, methods for increasing the efficiency of targeted genome editing using SDNs are needed. SUMMARY
[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the DETAILED DESCRIPTION. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0006] In one aspect, the present disclosure provides a fusion protein comprising a site-directed nuclease fused to a recruitment domain comprising a site-specific DNA binding domain.
[0007] In some embodiments, the site-directed nuclease comprises a CRISPR-associated nuclease. In some embodiments, the CRISPR-associated nuclease is selected from the group consisting of Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, Cas14, and a nickase or inactivated form thereof. In some embodiments, the CRISPR-associated nuclease is a Cas9 enzyme. In some embodiments, the CRISPR-associated nuclease is a Cas12a enzyme.
[0008] In some embodiments, the recruitment domain is a Cro repressor family protein. In some embodiments, the Cro repressor family protein comprises N15 Cro, lambda Cro, P22 Cro, 434 Cro, or any combination thereof. In some embodiments, the recruitment domain comprises an amino acid sequence that is at least 90% identical to any one of SEQ ID NOs: 1-4. In some embodiments, the recruitment domain comprises a dimerization domain.
[0009] In some embodiments, the fusion protein comprises a linker positioned between the site-directed nuclease and the recruitment domain. In some embodiments, the linker comprises any one of SEQ ID NOs: 6, 7, 15, or 16. In some embodiments, the fusion protein comprises a nuclear localization signal.
[0010] In some embodiments, the fusion protein comprises an amino acid sequence that is at least 90% identical to SEQ ID NO: 11 or 13.
[0011] In another aspect, the present disclosure provides a recombinant nucleic acid encoding a fusion protein comprising a site-directed nuclease fused to a recruitment domain comprising a site-specific DNA binding domain.
[0012] In another aspect, the present disclosure provides a DNA construct comprising a promoter operably linked to the recombinant nucleic acid. In some embodiments, the promoter comprises at least one of an inducible promoter, a constitutive promoter, an egg cell-specific promoter, a pollen-specific promoter, or a terminal meristem-specific promoter. In some embodiments, the promoter is a ubiquitin 4 promoter, an actin promoter, a tubulin promoter, a MADS-box promoter, or a plant virus promoter.
[0013] In another aspect, the present disclosure provides a vector comprising the recombinant nucleic acid or the DNA construct.
[0014] In another aspect, the disclosure provides a cell comprising a recombinant nucleic acid, DNA construct, or vector. In some embodiments, the cell is a plant cell. In some embodiments, the plant cell is a maize plant cell, a soybean plant cell, a rice plant cell, a wheat plant cell, or a sunflower plant cell.
[0015] In another aspect, the disclosure provides a method of editing a nucleic acid, the method comprising: (a) providing at least one fusion protein as described herein; (b) providing the nucleic acid, wherein the nucleic acid comprises a first binding site and a target region comprising a portion of the nucleic acid, wherein the first binding site is within or adjacent to the target region; (c) providing a donor polynucleotide comprising a donor nucleotide region and at least one recruitment sequence that specifically binds to a recruitment domain of the at least one fusion protein; and (d) contacting the nucleic acid and the donor polynucleotide with the at least one fusion protein, wherein the at least one fusion protein specifically binds to the first binding site of the nucleic acid and specifically binds to the recruitment sequence of the donor polynucleotide, thereby generating an edit to the target region of the nucleic acid.
[0016] In some embodiments, the first binding site is adjacent to the 5’ end or the 3’ end of the target region. In some embodiments, the edit to the target region of the nucleic acid is a replacement of at least a portion of the target region with at least a portion of the donor polynucleotide.
[0017] In some embodiments, the nucleic acid further comprises a second binding site, wherein the second binding site is within or adjacent to the target region, and wherein the at least one fusion protein specifically binds to the first binding site and the second binding site of the nucleic acid. In some embodiments, the second binding site is adjacent to the 5’ end or the 3’ end of the target region.
[0018] In some embodiments, the recruitment domain of the fusion protein comprises a Cro repressor family protein, and the at least one recruitment sequence comprises a Cro OR3 operon sequence. In some embodiments, the Cro OR3 operon sequence comprises an N15 OR3 operon sequence (optionally, SEQ ID NO: 18), a lambda OR3 operon sequence, a P22 OR3 operon sequence, a 434 OR3 operon sequence, or a combination thereof.
[0019] In some embodiments, the donor polynucleotide comprises at least one homology arm, wherein the at least one homology arm comprises a nucleotide sequence having complementarity to a portion of the target region of the nucleic acid. In some embodiments, the donor polynucleotide comprises at least two recruitment sequences. In some embodiments, the donor polynucleotide comprises a first recruitment sequence adjacent to the 5’ end of the donor nucleotide region and a second recruitment sequence adjacent to the 3’ end of the donor nucleotide region. In some embodiments, the at least two recruitment sequences are not within the donor nucleotide region.
[0020] In some embodiments, the site-directed nuclease of the at least one fusion protein comprises a CRISPR-associated nuclease, and the method further comprises providing at least one guide RNA, wherein the at least one guide RNA comprises a nucleotide sequence having complementarity to the first binding site and / or the second binding site of the nucleic acid. BRIEF DESCRIPTION OF DRAWINGS
[0021] The application includes the following figures. The figures are intended to illustrate certain embodiments and / or features of the compositions and methods and supplement any descriptions of any one or more of the compositions and methods. The figures do not limit the scope of the compositions and methods, except as explicitly claimed.
[0022] Figure 1 A schematic depiction of several embodiments of the donor polynucleotides described herein is shown, in accordance with aspects of the present disclosure. Depicted is a donor DNA polynucleotide (“donor DNA”) having a portion designed for insertion into a target site (“insertion / replacement sequence”), as described herein. Also shown are left and right homology arms (“LHA” and “RHA”, respectively) and a Cro OR3 recruitment sequence (“Cro OR3”). R 3 recruitment sequence (“Cro OR3”).
[0023] Figure 2 A schematic depiction of an embodiment of the methods provided herein is shown, in accordance with aspects of the present disclosure. Depicted is a nucleic acid having a target region (“genomic DNA”) and a donor DNA polynucleotide having a portion designed for insertion into the target region (“insertion / replacement sequence”), as described herein. Also shown are left and right homology arms (“LHA” and “RHA”, respectively) and a Cro OR3 recruitment sequence (“Cro OR3”). Indicated on the genomic DNA locus are potential cleavage sites that are more or less useful in the particular embodiments provided herein (“On target cut” and “Off target cut”, respectively). R 3 recruitment sequence (“Cro OR3”). Indicated on the genomic DNA locus are potential cleavage sites that are more or less useful in the particular embodiments provided herein (“On target cut” and “Off target cut”, respectively).
[0024] Figure 3 A schematic depiction of two fusion proteins tethered to a target site of a genomic DNA sequence and to a donor DNA sequence is shown, in accordance with aspects of the present disclosure.
[0025] Figure 4 A schematic depiction of a Cas9-N15 Cro fusion protein is shown, in accordance with aspects of the present disclosure.
[0026] Figure 5 A schematic depiction of a Cas12a-N15 Cro fusion protein is shown, in accordance with aspects of the present disclosure. DETAILED DESCRIPTION
[0027] The following description sets forth various aspects and embodiments of the compositions and methods of the application. The particular embodiments are not intended to limit the scope of the compositions and methods. Rather, the embodiments provide non-limiting examples of various compositions and methods included within the scope of the disclosed compositions and methods. The description will be read from the perspective of one of ordinary skill in the art; thus, it will not necessarily include information that is well known to the skilled artisan. I. Terminology
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. References to techniques employed herein are intended to refer to the techniques as commonly understood by one of ordinary skill in the art, including variations and / or substitutions of those techniques that are apparent to the skilled artisan. Although the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the disclosed subject matter.
[0029] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "an enzyme" optionally includes a combination of two or more such molecules, and the like.
[0030] As used herein, "and / or" means and / or, and includes any and all possibilities of the combination of one or more of the associated listed items.
[0031] As used herein, the term "about" means the usual error for the particular value or range given, of which one of ordinary skill in the art is readily aware, e.g., ±20%, ±10%, or ±5%, within the expected meaning of the values recited.
[0032] As used herein, the term "comprising" or "comprise" is open-ended. When used in conjunction with a subject nucleic acid (or amino acid sequence), it means a nucleic acid sequence (or amino acid sequence) that includes the subject sequence as part of, or as the entire sequence of.
[0033] As used herein, the transitional phrase "consisting essentially of means the scope of a claim is interpreted to encompass the specified materials or steps recited in the claim, and those that do not materially affect the basic and novel characteristics of the claimed subject matter. Thus, the term "consisting essentially of when used in the claims of the present disclosure is not intended to be interpreted as equivalent to "comprising."
[0034] The term "plurality" refers to more than one entity. Thus, "a plurality of individuals" refers to at least two individuals. In some embodiments, the term plurality refers to more than half of the whole. For example, in some embodiments, "a plurality in a population" refers to more than half of the members of that population.
[0035] As used herein, the term "plant" refers to any plant, particularly seed plants, at any stage of development. As used herein, the term "plant cell" is a structural and physiological unit of a plant, including a protoplast and a cell wall. A plant cell can be in the form of an isolated single cell or a cultured cell, or as part of a higher organized unit such as, for example, a plant tissue, a plant organ, or a whole plant. A plant cell can be derived from or be part of a gymnosperm or an angiosperm. A plant cell can be a monocotyledonous plant cell (e.g., a maize cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turfgrass cell, or an ornamental grass cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, a tomato cell, a sunflower cell, a Brassicaceae cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugar beet cell, or an oilseed rape cell). As used herein, the term "plant cell culture" means a culture of plant units such as, for example, protoplasts, cells of a cell culture, cells in a plant tissue, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at various stages of development. As used herein, the term "plant tissue" refers to a group of plant cells organized into a structural and functional unit. Any plant tissue in a plant or in culture is included. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue cultures, and any group of plant cells organized into a structural and / or functional unit. The use of this term in conjunction with (or in the absence of) the use of or coverage by definition of any particular type of plant tissue as listed above is not intended to exclude any other type of plant tissue. As used herein, the term "plant part" refers to a portion of a plant, including single cells and cellular tissues such as intact plant cells in a plant, cell clumps, and tissue cultures that can regenerate a plant. Examples of plant parts include, but are not limited to, single cells and tissues from pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral organ parts, fruits, stems, shoots, cuttings, and seeds; and pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral organ parts, fruits, stems, shoots, cuttings, scions, rhizomes, seeds, protoplasts, callus, and the like.
[0036] The terms "polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. As used herein, these terms encompass amino acid chains of any length including full-length proteins, wherein the amino acid residues are linked via covalent peptide bonds.
[0037] The terms "nucleic acid" and "polynucleotide" are used interchangeably and as used herein refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in either single- or double-stranded form, and polymers thereof, and refer to both the sense and antisense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers of the above. In higher plants, DNA is the genetic material, while RNA is involved in the transfer of information contained within DNA into proteins. A "genome" is the totality of genetic material contained in each cell of an organism. It will be understood that when an RNA is described, its corresponding cDNA is also described, with uridines represented as thymidines. In particular embodiments, nucleotides refer to ribonucleotides, deoxynucleotides, or modified forms of either type of nucleotide, and combinations thereof. In addition, the polynucleotides disclosed herein can include either or both of naturally-occurring nucleotides and modified nucleotides linked together by naturally-occurring and / or non-naturally-occurring nucleotide linkages. Nucleic acid molecules can be chemically or biochemically modified, or can contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more of the naturally occurring nucleotides with an analog, internucleotide modifications such as uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (e.g., phosphonates, phosphodiester, phosphorodithioates, etc.), pendent moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, tetrahydropyranyl, benzophenanthrine, etc.), chelators, alkylators and modified linkages (e.g., alpha anomeric nucleic acids, etc.). The above list of modifications is not exhaustive and is intended to include any modification which modifies a nucleic acid. The above terms are also intended to include any topological conformation, including single stranded, double stranded, partially duplexed, triplexed, hairpinned, circular and padlocked configurations. Unless otherwise indicated, a reference to a nucleic acid sequence is intended to encompass its complement. Thus, a reference to a nucleic acid molecule having a particular sequence is understood to encompass its complementary strand having its complementary sequence. Nucleotide sequences are "complementary" when they specifically hybridize to each other in solution (e.g., according to the Watson-Crick base pairing rules). The term also includes codon-optimized nucleic acids encoding the same polypeptide sequence. It will also be understood that a nucleic acid can be unpurified, purified, or attached to, for example, a synthetic material such as a bead or column matrix.
[0038] In the context of nucleic acid sequences, the term "corresponding to" means that the nucleic acids which "correspond to" certain enumerated positions in the present application are those which align with these positions in the reference sequence when the nucleic acid sequences of certain sequences are aligned with each other, but are not necessarily located in these exact numerical positions with respect to the particular nucleic acid sequence of the application. Optimal alignment of sequences for comparison can be conducted by computerized implementations of known algorithms or by visual inspection. Readily available sequence comparison and multiple sequence alignment algorithms are the Basic Local Alignment Search Tool (BLAST) and the ClustalW / ClustalW2 / ClustalOmega programs, respectively, available on the internet (e.g., at the website of EMBL-EBI). Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG software package available from Accelrys, Inc., San Diego, CA, USA. See also Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; Ausubel et al., 1988; and Sambrook and Russell, 2001.
[0039] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues. See, Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).
[0040] The term "identity" or "substantial identity" as used in the context of polynucleotide or polypeptide sequences as described herein refers to a sequence having at least 60% sequence identity to a reference sequence. Alternatively, the percent identity can be any integer from 60% to 100%. Exemplary embodiments include at least: 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% when compared to a reference sequence, as using the programs described herein, preferably using the standard parameters as described below for BLAST. Those skilled in the art will recognize that these values can be adjusted appropriately by taking into account codon degeneracy, amino acid similarity, reading frame positioning, and the like, in order to determine the corresponding identity of the proteins encoded by two nucleotide sequences.
[0041] For sequence comparison, typically, one sequence acts as the reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are input into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identity for the test sequence(s) relative to the reference sequence, based on the program parameters.
[0042] As used herein, "comparison window" includes reference to a segment of a polynucleotide sequence that is selected from any one of the contiguous positions from 20 to 600, typically from about 50 to about 200, more typically from about 100 to about 150, wherein, following the alignment, the sequences can be compared to a reference sequence of the same number of contiguous positions. Methods of alignment of sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be conducted by the local homology algorithm of Smith and Waterman Add. APL Math. 2:482 (1981); by the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970); by the search for similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. (USA) 85:2444 (1988); by computerized implementations (e.g., BLAST) of these algorithms; or by manual alignment and visual inspection.
[0043] Altschul et al. (1977) Nucleic Acids Res. 25:3389-3402. Software for performing BLAST analyses is publicly available through the web site of the National Center for Biotechnology Information (NCBI). This algorithm involves first identifying high scoring sequence pairs (HSPs) that are considered to be matches or significant alignments over some length W, and then calculating the statistical significance of those matches. One measure of the significance of a match is the number of comparisons needed to be made to find an HSP of that score. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 28, an expectation value (E) of 10, M=1, N=-2 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength of 3, and expectation value (E) of 10, and the BLOSUM62 scoring matrix. See Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915 (1989).
[0044] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences. See, e.g., Karlin and Altschul, Proc. Nat'l. Acad. Sci. USA 90:5873-5787 (1993). One measure of similarity provided by the BLAST algorithm is the smallest sum probability, P(N), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.01, more preferably less than about 10 -5 and most preferably less than about 10 -20 .
[0045] "Recombination" is the exchange of DNA strands to produce a new arrangement of nucleotide sequences. The term can refer to the process of homologous recombination that occurs in the repair of double-stranded DNA breaks, in which a polynucleotide is used as a template to repair a homologous polynucleotide. The term can also refer to the exchange of information between two homologous chromosomes during meiosis. The frequency of double recombinants is the product of the frequencies of single recombinants. For example, if the frequency of recombinants found in a 10 cM region is 10%, and the frequency of double recombinants is found to be 10% x 10% = 1% (1 centiMorgan is defined as 1% of the recombinants in a test cross).
[0046] A "gene" is a defined region located within a genome and, in addition to the aforementioned coding nucleic acid sequence, it also contains other primary regulatory nucleic acid sequences responsible for controlling the expression (that is, transcription and translation) of the coding portion. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein, including regulatory sequences. A gene can or can not be used to produce a functional protein. In some embodiments, a gene refers only to the coding region. The term "native gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene comprising 1) DNA sequences that are not found together in nature, or 2) DNA sequences derived from different sources, or 3) a DNA sequence encoding a protein having a different structure from that of a naturally occurring protein. Thus, a chimeric gene can comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences derived from one source, and coding sequences derived from another source, or regulatory and coding sequences derived from the same source, but arranged in a manner different from that found in nature. A gene can be "isolated" by substantially or essentially free from components that normally accompany or interact with it when it is found in the natural state. Such components include other cellular material, culture medium from recombinant production, and / or chemicals used in chemically synthesizing the gene.
[0047] A "gene of interest" or "nucleotide sequence of interest" is any gene that when transferred to a plant confers a desirable characteristic on the plant, such as antibiotic resistance, viral resistance, insect resistance, disease resistance, or resistance to other pests, herbicide tolerance, improved nutritional value, improved performance of an industrial process, or altered reproductive ability. A "gene of interest" can also be a gene transferred to a plant for the production of a commercially valuable enzyme or metabolite in the plant.
[0048] An "isolated" nucleic acid molecule or nucleotide sequence or "isolated" polypeptide is a nucleic acid molecule, nucleotide sequence, or polypeptide that exists apart from its natural environment and / or has different, modified, regulated, and / or altered functions when compared to its function in its natural environment and is therefore not a naturally occurring product. An isolated nucleic acid molecule or isolated polypeptide can exist in purified form or can exist in a non-natural environment such as, for example, a recombinant host cell. Thus, for example, the term isolated with respect to a polynucleotide means that the polynucleotide is separated from the chromosome and / or the cell in which it naturally occurs. A polynucleotide is also isolated if it is separated from the chromosome and / or the cell in which it naturally occurs and then inserted into a genetic background, chromosome, chromosomal location, and / or cell in which it does not naturally occur. The recombinant nucleic acid molecules and nucleotide sequences of the present application can be considered "isolated" as defined above.
[0049] Thus, an "isolated nucleic acid molecule" or "isolated nucleotide sequence" is a nucleic acid molecule or nucleotide sequence that is not immediately adjacent to the nucleotide sequences immediately flanking it in the genome of the organism from which it is derived. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences immediately flanking the coding sequence. The term therefore includes, for example, a recombinant nucleic acid incorporated into a vector, incorporated into a self-replicating plasmid or virus, or incorporated into the genomic DNA of a prokaryote or eukaryote, or which exists as a separate molecule (e.g., a cDNA or a genomic DNA fragment produced by PCR or restriction analysis) independent of other sequences. It also includes a recombinant nucleic acid which is part of a hybrid nucleic acid molecule encoding additional polypeptide or peptide sequences. An "isolated nucleic acid molecule" or "isolated nucleotide sequence" can also include a nucleotide sequence derived from the same natural original cell type and inserted into that same natural original cell type, but existing in a non-natural state, for example, in a different copy number, and / or under the control of different regulatory sequences than those found in the natural state of the nucleic acid molecule.
[0050] The term "isolated" can further refer to nucleic acid molecules, nucleotide sequences, polypeptides, peptides, or fragments that are substantially free of cellular material, viral material, and / or culture medium (e.g., when produced by recombinant DNA technology), or chemical precursors or other chemicals (e.g., when produced by chemical synthesis). Furthermore, an "isolated fragment" is a fragment of a nucleic acid molecule, nucleotide sequence, or polypeptide that does not exist as a fragment in nature and would not exist in nature in such a state. "Isolated" does not necessarily mean industrially pure (homogeneous), but it is sufficiently pure to provide the polypeptide or nucleic acid in a form that can be used for the intended purpose.
[0051] "Homology dependent repair" or "homology directed repair" or "HDR" refers to a mechanism for repairing ssDNA and double stranded DNA (dsDNA) damage in a cell. This repair mechanism can be utilized by a cell in the presence of an HDR template having sequence with substantial identity to the damaged site. The term "perfect HDR" refers to a situation in which the genomic homology junction in the replaced allele is subject to complete HDR, and "imperfect HDR" refers to a situation in which the genomic homology junction in the replaced allele is subject to partial or incomplete HDR. In some embodiments, a donor polynucleotide molecule having homology to a cleaved target DNA sequence is used as a template for repair of the cleaved target DNA sequence, such that genetic information is transferred from the donor polynucleotide to the target DNA. As such, new nucleic acid material can be inserted / copied into the site. In some cases, the target DNA is contacted with a donor molecule, e.g., a donor polynucleotide molecule. In some cases, a donor polynucleotide molecule is introduced into a cell. In some cases, at least one segment of the donor polynucleotide molecule is integrated into the genome of the cell.
[0052] "Microhomology-mediated end joining" or "MMEJ" or "alternative non-homologous end joining" (Alt-NHEJ) refers to a form of repairing double-strand breaks in DNA. This repair mechanism utilizes microhomologous sequences to align the broken strands. "Non-homologous end joining" or "NHEJ" refers to a form of repairing double-strand breaks in DNA. The double-strand break is repaired by direct ligation of the ends to each other. Typically, in the absence of a donor polynucleotide, no new nucleic acid material is inserted into the site, but some nucleic acid material can be lost or added, generating a small deletion or a small insertion. In some embodiments, a donor polynucleotide molecule can be provided (e.g., a donor polynucleotide molecule can be introduced into a cell), and a portion of the donor polynucleotide can be inserted into the genome by MMEJ or NHEJ. Some embodiments of the methods provided herein increase the likelihood of donor polynucleotide insertion by tethering of the donor polynucleotide to the target site, as described below. II. Introduction
[0053] Provided herein are fusion proteins, and related recombinant nucleic acids, systems, and methods that increase the efficiency of genome editing using SDNs and donor polynucleotide tethering methods. The present disclosure is based in part on the inventors’ discovery that 1) fusing an SDN to a recruitment domain comprising a site-specific DNA binding domain, and 2) using a donor polynucleotide homologous repair template comprising a binding site for the recruitment domain results in an increased frequency of HDR, as demonstrated by the examples herein. Without being bound by any particular theory, the recruitment domain can bind to the binding site in the donor polynucleotide template and tether the donor polynucleotide to the cleavage site (i.e., by the fusion of the recruitment domain to the SDN that forms the cleavage). Additionally, this tethering can increase the likelihood of HDR-mediated repair of the cleavage (e.g., by facilitating spatial proximity between the cleavage site and the donor polynucleotide template). III. Fusion Proteins
[0054] In one aspect, provided herein are fusion proteins comprising a site-directed nuclease linked to a recruitment domain comprising a site-specific DNA binding domain. As used throughout, a “fusion protein” is a protein comprising two different polypeptide sequences (i.e., a site-directed nuclease polypeptide sequence and a recruitment domain polypeptide sequence) that are joined or linked to form a single polypeptide. In some embodiments, the two amino acid sequences are encoded by separate nucleic acid sequences that have been joined such that they are transcribed and translated to produce a single polypeptide. The site-directed nuclease and the recruitment domain can be joined in any order and orientation relative to each other. For example, the C’ terminus of the site-directed nuclease can be linked to the N’ terminus or C’ terminus of the recruitment domain. The site-directed nuclease and the recruitment domain can be separated by one or more additional fusion protein domains, as described below. A. Site-Directed Nucleases
[0055] The fusion proteins provided herein comprise a site-directed modification polypeptide (e.g., a site-directed nuclease). A “site-directed modification polypeptide” modifies a target DNA (e.g., via cleavage or methylation of the target DNA) and / or a polypeptide associated with the target DNA (e.g., methylation or acetylation of a histone tail). In some embodiments, a site-directed modification polypeptide interacts with a guide RNA (which is a single RNA molecule or an RNA duplex of at least two RNA molecules) due to the association of the site-directed modification polypeptide with the guide RNA, and is directed to a DNA sequence (e.g., a chromosomal sequence or an extrachromosomal sequence, such as a episome sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.). In some embodiments, the site-directed polypeptide is a site-directed nuclease that is capable of cleaving one or both strands of DNA at a specified target sequence.
[0056] The term "cleavage" or "cleaving" refers to the breaking of a covalent phosphodiester linkage in the ribosyl phosphate diester backbone of a polynucleotide and encompasses both single-strand breaks and double-strand breaks. Double-strand cleavage can occur as a result of two distinct single-strand cleavage events. Cleavage can result in either a blunt end or a staggered end (also known as a sticky end). A "nuclease cleavage site" or "genomic nuclease cleavage site" is a region of nucleotides in which a site-directed nuclease cleaves (e.g., when bound to a proximal binding site). When the polynucleotide is DNA (e.g., genomic DNA), one or both strands can be cleaved at the nuclease cleavage site. This cleavage by the nuclease triggers DNA repair mechanisms within the cell that set up an environment for homologous recombination to occur.
[0057] Various site-directed nucleases can be used in the fusion proteins, systems, and methods disclosed herein. Suitable nucleases include, but are not limited to, a CRISPR-associated (Cas) protein or Cas nuclease; a zinc finger nuclease (ZFN); a transcription activator-like effector nuclease (TALEN); a meganuclease; an RNA-binding protein (RBP); a CRISPR-associated RNA-binding protein; a recombinase; a flippase; a transposase; an Argonaute (Ago) protein (e.g., a prokaryotic Argonaute (pAgo), an archaeal Argonaute (aAgo), an eukaryotic Argonaute (eAgo), and a Natronobacterium gregoryi Argonaute (NgAgo); an adenosine deaminase acting on RNA (ADAR); a CRISPR-Cas-primed RNA targeting (CIRT) system; a Pumilio / fem-3 binding factor (PUF), a homing endonuclease, or any functional fragment thereof, any derivative thereof; any variant thereof; and any fragment thereof. Exemplary site-directed nucleases suitable for use in the fusion proteins, systems, and methods disclosed herein are further described below.
[0058] In some embodiments, the site-directed nuclease is a naturally occurring site- directed nuclease. Exemplary naturally occurring site-directed nucleases are known in the art (see, e.g., Makarova et al., 2017, Cell 168:328-328.e1 and Shmakov et al., 2017, Nat Rev Microbiol 15(3): 169-182, both of which are incorporated by reference herein). In some embodiments, the site-directed nuclease binds a DNA-targeting polynucleotide (e.g., a guide RNA) and is thereby directed to a specific sequence within a target DNA and cleaves the target DNA.
[0059] In some embodiments, site-directed nucleases are modified relative to their native sequence (e.g., via mutations of one or more amino acid residues) to alter their function. For example, site-directed nucleases can be modified to be non-enzymatically active. The term "non-enzymatically active" can mean that a site-directed nuclease can bind to a nucleic acid sequence in a sequence-specific manner in a polynucleotide, but cannot cleave the target polynucleotide. Non-enzymatically active site-directed peptides may contain a non-enzymatically active domain (e.g., a nuclease domain). Non-enzymatically active can mean inactive. Non-enzymatically active can mean substantially inactive. Non-enzymatically active can mean substantially inactive. Non-enzymatically active can mean activity that is no more than 1%, no more than 2%, no more than 3%, no more than 4%, no more than 5%, no more than 6%, no more than 7%, no more than 8%, no more than 9%, or no more than 10% of the activity compared to the exemplary activity of the wild type (e.g., nucleic acid cleavage activity, wild-type Cas9 activity).
[0060] In some embodiments, a site-directed nuclease (e.g., a site-directed nuclease without enzymatic activity) is fused with one or more transcriptional repressor domains, activator domains, epigenetic domains, recombinase domains, transposase domains, flipase domains, nickase domains, cleavage domains, or any combination thereof. The activator domain may include one or more tandem activation domains located at the carboxyl terminus of the enzyme. In other cases, the actuator portion includes one or more tandem repressor domains located at the carboxyl terminus of the protein. Non-limiting exemplary activation domains include GAL4, herpes simplex activation domain VP16, VP64 (a tetramer of herpes simplex activation domain VP16), the NF-KB p65 subunit, and Epstein-Barr virus R transactivator (Rta), and are described in Chavez et al., Nat Methods, 2015, 12(4):326-328 and U.S. Patent Application Publication No. 20140068797. Non-limiting exemplary repression domains include the KRAB (Kruppel-associated box) domain of Koxl, the Mad mSIN3 interaction domain (SID), and the ERF repression factor domain (ERD), and are described in Chavez et al., Nat Methods, 2015, 12(4):326-328 and U.S. Patent Application Publication No. 20140068797. Nucleases can also be fused with heterologous peptides to provide increased or decreased stability. The fusion domain or heterologous peptide can be located at the N-terminus, C-terminus, or inside the nuclease. CRISPR / Cas nuclease
[0061] In some embodiments, site-directed nucleases include CRISPR-associated (Cas) proteins or Cas nucleases that function in the CRISPR (clustered regularly spaced short palindromic repeats) / Cas system. In bacteria, this system can provide adaptive immunity against foreign DNA (Barrangou, R. et al., “CRISPR provides acquired resistance against viruses in prokaryotes”, Science (2007) 315:1709-1712; Makarova, KS et al., “Evolution and classification of the CRISPR-Cas systems”, Nat Rev Microbiol (2011) 9:467-477; Garneau, JE et al., “The CRISPR / Cas bacterial immune system cleaves bacteriophage and plasmid DNA”, Nature (2010) 468:67-71; Sapranauskas, R. et al., “The Streptococcus thermophilus CRISPR / Cas system provides…”). Immunity in *Escherichia coli* [CRISPR / Cas system provides immunity in *Streptococcus thermophilus*], *Nucleic Acids Res* (2011) 39:9275-9282. In a wide variety of organisms, including different mammals, animals, plants, microorganisms, and yeast, the CRISPR / Cas system (e.g., modified and / or unmodified) can be used as a tool for genome engineering. CRISPR / Cas systems can include guide nucleic acids such as guide RNA (gRNA), which complex with Cas proteins for targeted regulation of gene expression and / or activity or nucleic acid editing. RNA-guided Cas proteins (e.g., Cas nucleases such as Cas9 nuclease) can specifically bind to target polynucleotides (e.g., DNA) in a sequence-dependent manner.If Cas proteins possess nuclease activity, they can cleave DNA (Gasiunas, G. et al., "Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria," Proc Natl Acad SciUSA (2012) 109: E2579-E286; Jinek, M. et al., "A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity," Science (2012) 337: 816-821; Sternberg, SH et al., "DNA interrogation by the CRISPR RNA-guided endonuclease"). Cas9 [CRISPR RNA-directed DNA interrogation by the endonuclease Cas9], “Nature [Nature] (2014) 507:62; Deltcheva, E. et al., “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III”, “Nature [Nature] (2011) 471:602-607.” DNA cleavage (e.g., double-strand breaks) can generate DNA break repair, thereby allowing the introduction of one or more gene modifications (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ), microhomologous-mediated end joining (MMEJ), or homology-directed repair (HDR). In some embodiments, donor polynucleotides are used to facilitate HDR, as detailed below in the “Systems” section.CRISPR-Cas systems have been widely used for programmable genome editing in various organisms and model systems (Cong, L. et al., “Multiplex genome engineering using CRISPR-Cas systems,” Science (2013) 339:819-823; Jiang, W. et al., “RNA-guided editing of bacterial genomes using CRISPR-Cas systems,” Nat. Biotechnol. (2013) 31:233-239; Sander, JD and Joung, JK, “CRISPR-Cas systems for editing, regulating and targeting genomes,” Nature Biotechnol. (2014) 32:347-355).
[0062] In some embodiments, the site-directed nucleases described herein include Cas proteins that form complexes with a guide nucleic acid (such as guide RNA) (described further below in the "Systems" section). In some embodiments, the site-directed nucleases include Cas proteins that form complexes with a single guide nucleic acid (such as a single guide RNA (sgRNA)). In some embodiments, the site-directed nucleases include RNA-binding proteins (RBPs) that optionally complex with a guide nucleic acid (such as guide RNA (e.g., sgRNA)) capable of forming complexes with the Cas protein. In some cases, RNA-directed Cas proteins recognize DNA targets complementary to a portion of the gRNA (called a CRISPRRNA (crRNA) sequence). The target sequence is often referred to as the prototype spacer, and the portion of the crRNA sequence complementary to the prototype spacer is often referred to as the spacer. To function (e.g., to cleave DNA), many Cas nucleases also require a specific prototype spacer adjacent motif (PAM) (typically a 2 to 6 base pair DNA sequence) immediately following the prototype spacer sequence.
[0063] Various site-directed Cas nucleases (e.g., Cas proteins from different species) can be used in the fusion proteins, systems, and methods provided herein based on the diverse enzymatic characteristics of different Cas proteins (e.g., different prototypical spacer adjacent motif (PAM) sequence preferences; increased or decreased enzyme activity; increased or decreased cytotoxicity levels; a tendency to produce one or more of NHEJ, homology-directed repair, single-strand breaks, double-strand breaks, etc.). Cas proteins from multiple species (e.g., those disclosed in Shmakov et al., 2017, or polypeptides derived from them) may require different PAM sequences in the target DNA. Therefore, for a particular Cas enzyme, the PAM sequence requirement may differ from the known 5'-N GG-3' sequence (where N is A, T, C, or G) required for Cas9 activity. Many Cas9 orthologs from a wide variety of species have been identified, and these proteins share only a few identical amino acids. All identified Cas9 orthologs have the same domain construction as the central HNH endonuclease domain and the split-type RuvC / RNase H domain. The Cas9 protein shares four key motifs with conserved structures; motifs 1, 2, and 4 are RuvC-like motifs, while motif 3 is an HNH motif. In contrast, Cas12a proteins from different species may have different PAM sequence requirements compared to the TTTV-compliant LbCas12a PAM.
[0064] Any suitable CRISPR / Cas system can be used. Multiple nomenclature systems can be used to refer to CRISPR / Cas systems. Exemplary nomenclature systems are provided in Makarova, KS et al., “An updated evolutionary classification of CRISPR-Cas systems,” Nat Rev Microbiol (2015) 13:722-736 and Shmakov, S. et al., “Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems,” Mol Cell (2015) 60:1-13. The CRISPR / Cas system can be type I, type II, type III, type IV, type V, type VI, or any other suitable CRISPR / Cas system. The CRISPR / Cas system used herein can be class 1, class 2, or any other suitable classification of CRISPR / Cas systems. Class 1 or class 2 identification can be based on genes encoding effector modules. Class 1 CRISPR / Cas systems typically possess multi-subunit crRNA-effector complexes, while Class 2 systems typically possess a single protein, such as Cas9, Cpfl (also known as Cas12a), C2c1, C2c2, C2c3, or a crRNA-effector complex. Class 1 CRISPR / Cas systems can utilize complexes of multiple Cas proteins for regulation. Class 1 CRISPR / Cas systems can include, for example, type I (e.g., I, IA, IB, IC, ID, IE, IF, IU), type III (e.g., III, IIIA, IIIB, IIIC, IIID), and type IV (e.g., IV, IVA, IVB) CRISPR / Cas types. Class 2 CRISPR / Cas systems can utilize a single large Cas protein for regulation. Class 2 CRISPR / Cas systems can include, for example, type II (e.g., II, IIA, IIB) and type V CRISPR / Cas types. CRISPR systems can be complementary to each other and / or can utilize trans-functional units to facilitate CRISPR locus targeting.
[0065] Cas proteins can originate from any suitable organism. Non-limiting examples include *Streptococcus pyogenes*, *Streptococcus thermophilus*, *Streptococcus* sp., *Staphylococcus aureus*, *Nocardiopsis dassonvillei*, *Streptomyces pristinae spiralis*, *Streptomyces viridochromogenes*, *Streptomyces viridochromogenes*, *Streptosporangium roseum*, *Streptosporangium roseum*, *AlicyclobacHlus acidocaldarius*, *Bacillus pseudomycoides*, *Bacillus sselenitireducens*, and *Exiguobacterium*. *Lactobacillus sibiricum*, *Lactobacillus delbrueckii*, *Lactobacillus salivarius*, *Microscilla marina*, *Burkholderiales bacterium*, *Polaromonas nap hthalenivorans*, *Polaromonas* sp., *Crocosphaera watsonii*, *Cyanothece* sp., *Microcystis aeruginosa*, *Pseudomonas aeruginosa*, and *Synechococcus* sp.), Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vinosum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus vannamei watsoni, Pseudoalteromonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbyasp., Microcoleus chthonoplastes, Oscillatoriasp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii) and the new culprit, Francisella novicida. In some respects, the organism is Streptococcus pyogenes (S. ).(pyogenes). In some respects, the organism is Staphylococcus aureus. In other respects, the organism is Streptococcus thermophilus.
[0066] Cas proteins can originate from a variety of bacterial species, including but not limited to Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, and Bifidobacterium. Lactobacillus bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Lactobacillus daphnegoldi, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium doolichum, and Lactobacillus coryniformis subsp.Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium, Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas The following bacteria are listed: *Plasmustris*, *Prevotella micans*, *Prevotella ruminicola*, *Flavobacterium columnnare*, *Aminomonas paucivorans*, *Rhodospirillum rubrum*, *Candidatus Puniceispirillum marinum*, *Verminephrobactereiseniae*, *Ralstonia syzygii*, *Dinoroseobactershibae*, *Azospirillum*, *Nitrobacter hamburgensis*, *Bradyrhizobium*, *Wolinella succinogenes*, and *Campylobacter jejuni subsp.*The following bacteria are listed: *Jejuni*, *Helicobacter mustelae*, *Bacillus cereus*, *Acidovorax ebreus*, *Clostridium perfringens*, *Parvibaculum lavamentivorans*, *Roseburia intestinalis*, *Neisseria meningitidis*, *Pasteurella multocida subsp. Multocida*, *Sutterella wadsworthensis*, *Proteobacterium*, *Legionella pneumophila*, *Parasutterella excrementihominis*, *Wolinella succinogenes*, and *Francisella catarrhalis*.
[0067] Non-limiting examples of Cas proteins include c2c1, C2c2, c2c3, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cash, Cas6e, Cas6f, Cas7, Cas8a, Cas8a1, Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csx12), Cas10, Cas10d, CasF, CasG, CasH, Cpfl (also known as Cas12a), Csyl, Csy2, Csy3, and Csel. (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csxl, Csx15, Csfl, Csf2, Csf3, Csf4 and Cul966 and their homologs or modified forms. In some embodiments, the site-directed nuclease of the fusion protein provided herein includes a CRISPR-associated nuclease, wherein the CRISPR-associated nuclease is Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, or Cas14. In some embodiments, the CRISPR-associated nuclease is a Cas9 enzyme. In some embodiments, the CRISPR-associated nuclease is a Cas12a enzyme. In some embodiments, the CRISPR-associated nuclease is a nicking enzyme or an inactivated form of a CRISPR-associated nuclease.
[0068] Cpf1 (LbCpf1) of Lachnospiraceae bacterium is one of many Cpf1 proteins in a large group. The terms “Cpf1” and “Cas12a” are used interchangeably throughout this disclosure. Cpf1 is a Cas protein. In some embodiments, the site-directed nuclease is a catalytically active Cas12a from Lachnospiraceae bacterium (“LbCas12a”) or Moraxella bovoculi AAX08_00205 (“Mb2Cas12a”). In some embodiments, the site-directed nuclease domain of the fusion protein is a Cas12a protein derived from any of the following: bacteria of the family Trichophyceae, species of the genus Aminococcus, species of Moraxella bubalana, species of the genus Thiomicrospira, species of Moraxella lacunata, species of Methanomethylophilus alvus, species of the genus Btyrivibrio, or species of Bacteroidetesoral.
[0069] Cas proteins may contain one or more domains. Non-limiting examples of domains include domains that guide nucleic acid recognition and / or binding, nuclease domains (e.g., DNase or RNase domains, RuvC, HNH), DNA-binding domains, RNA-binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. The domains that guide nucleic acid recognition and / or binding may interact with the guiding nucleic acid. The nuclease domains may include catalytic activity for nucleic acid cleavage. The nuclease domains may lack catalytic activity to prevent nucleic acid cleavage. Cas proteins may be chimeric Cas proteins fused with other proteins or peptides. Cas proteins may be chimeras of various Cas proteins, such as containing domains from different Cas proteins.
[0070] The Cas protein used herein may be an active variant, inactive variant, or fragment of a wild-type or modified Cas protein. Relative to the wild-type form of the Cas protein, the Cas protein may contain amino acid alterations, such as deletions, insertions, substitutions, variants, mutations, fusions, chimeras, or any combination thereof. The Cas protein may be a polypeptide having at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to the wild-type exemplary Cas protein. The Cas protein may be a polypeptide having at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to the wild-type exemplary Cas protein. Variant or fragment may contain at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to wild-type or modified Cas proteins or portions thereof. Variant or fragment may be targeted to nucleic acid loci that guide nucleic acid complexation while lacking nucleic acid cleavage activity.
[0071] In some embodiments, the modified Cas protein has reduced functionality relative to its unmodified form. In some embodiments, the modified Cas protein is a functionally defective form of the unmodified form. For example, a nuclease-deficient Cas protein retains the ability to bind DNA but lacks or has reduced nucleic acid cleavage activity. Cas nucleases (e.g., retaining wild-type nuclease activity, having reduced nuclease activity, and / or lacking nuclease activity) can function in a CRISPR / Cas system to regulate the level and / or activity (e.g., decrease, increase, or eliminate) of target genes or proteins. Cas proteins can bind target polynucleotides and prevent transcription to produce nonfunctional gene products by physical barriers or editing nucleic acid sequences. In some embodiments, the modified Cas protein has no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the function (e.g., nuclease activity) of a wild-type Cas protein (e.g., Cas9 from Streptococcus pyogenes). In some embodiments, the modified Cas protein does not have the substantial function of the wild-type Cas protein. When a Cas protein is a modified form that does not possess substantial nucleic acid cleavage activity, it can be referred to as non-enzymatic and / or "dead" (abbreviated as "d"). Dead Cas proteins (e.g., dCas, dCas9) can bind to target polynucleotides but cannot cleave them. In some respects, dead Cas proteins are dead Cas9 proteins or dead Cas12a proteins.
[0072] In some embodiments, the modified Cas protein can be a modified Cas "base editor." Base editing enables the direct and irreversible conversion of one target DNA base to another in a programmable manner, without the need for DNA cleavage or donor polynucleotide molecules. For example, Komor et al. (2016, Nature, 533:420-424) taught a Cas9-cytidine deaminase fusion in which Cas9 was engineered to be inactive and did not induce double-strand DNA breaks. Additionally, Gaudelli et al. (2017, Nature, doi:10.1038 / nature24644) taught a Cas9 fused to tRNA adenosine deaminase with impaired catalytic activity, which can mediate the conversion of A / T to G / C in a target DNA sequence. Another class of engineered Cas9 nucleases that can be used as site-directed nucleases in the fusion proteins disclosed herein are variants that can recognize a wide range of PAM sequences (including NG, GAA, and GAT) (Hu et al., 2018, Nature, doi:10.1038 / nature26155).
[0073] Cas proteins can be modified to optimize the regulation of gene expression. Cas proteins can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzyme activity. Cas proteins can also be modified to alter any other activity or property of the protein, such as stability. For example, one or more nuclease domains of a Cas protein can be modified, deleted, or inactivated, or the Cas protein can be truncated to remove domains unnecessary for protein function or to optimize (e.g., enhance or reduce) the activity of the Cas protein to regulate gene expression.
[0074] One or more nuclease domains (e.g., RuvC, HNH) of a Cas protein can be deleted or mutated, rendering them nonfunctional or containing reduced nuclease activity. For example, in a Cas protein containing at least two nuclease domains (e.g., Cas9), if one nuclease domain is deleted or mutated, the resulting Cas protein, called a nickase, can produce single-strand breaks at CRISPR RNA (crRNA) recognition sequences within double-stranded DNA, but not double-strand breaks. This nickase can cleave either the complementary or non-complementary strands, but not both simultaneously. In some embodiments, the specificity for targeting double-strand breaks is improved by targeting the nickase to opposite strands at two loci. If the nickase cleaves the single strands at both loci, a double-strand break is formed and can be repaired via HR as described herein. If all nuclease domains of a Cas protein (e.g., both the RuvC and HNH nuclease domains in the Cas9 protein; the RuvC nuclease domain in the Cpfl protein) are deleted or mutated, the resulting Cas protein may have a reduced ability to cleave both strands of double-stranded DNA or no such ability. Zinc finger nuclease
[0075] In some embodiments, a site-directed nuclease suitable for use in the fusion protein or method described herein is a “zinc finger nuclease” or “ZFN”. A ZFN is a fusion between a cleavage domain (such as the cleavage domain of Fokl) and at least one zinc finger motif (e.g., at least 2, 3, 4, or 5 zinc finger motifs) that can bind polynucleotides (such as DNA and RNA). Heterodimerization at certain sites in the polynucleotides of two separate ZFNs at specific orientations and intervals can lead to cleavage of the polynucleotides. For example, a ZFN binding DNA can induce double-strand breaks in DNA. To dimerize the two cleavage domains and cleave DNA, two separate ZFNs can bind to the opposite strand of DNA at their C-termini, spaced apart. In some cases, the linker sequence between the zinc finger domain and the cleavage domain may require the 5' edges of each binding site to be separated by approximately 5-7 base pairs. In some cases, the cleavage domain is fused to the C-terminus of each zinc finger domain. Exemplary ZFNs include, but are not limited to, those described in the following literature: Urnov et al., Nature Reviews Genetics, 2010, 11:636-646; Gaj et al., Nat Methods, 2012, 9(8):805-7; U.S. Patent Nos. 6,534,261; 6,607,882; 6,746,838; 6,794,136; 6,824,978; 6,866,997; 6,933,113; 6,979,539; 7,013,219; 7,030,215; 7,220,719; 7,241,573; 7,241,574; 7,585,849; 7,595,376; 6,903,185; 6,479,626; and U.S. Publication Nos. 2003 / 0232410 and 2009 / 0203140.
[0076] In some embodiments, a ZFN-containing nuclease can induce double-strand breaks in target polynucleotides, such as DNA. Double-strand breaks in DNA can lead to DNA break repair, thereby allowing the introduction of one or more gene modifications (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ) or homologous directed repair (HDR). In HDR, a donor polynucleotide repair template or template polynucleotide can be provided containing a homologous arm flanking the target DNA. In some embodiments, the ZFN is a zinc finger nicking enzyme that induces site-specific single-strand DNA breaks or nicks, thus generating HR. Descriptions of zinc finger nicking enzymes can be found, for example, in Ramirez et al., Nucl Acids Res, 2012, 40(12):5560-8; Kim et al., Genome Res, 2012, 22(7):1327-33. In some embodiments, the ZFN binds to polynucleotides (e.g., DNA and / or RNA) but does not cleave polynucleotides.
[0077] In some embodiments, the cleavage domain of the ZFN-containing nuclease includes a modified form of the wild-type cleavage domain. The modified form of the cleavage domain may include amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the cleavage domain. For example, the modified form of the cleavage domain may have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of the wild-type cleavage domain. The modified form of the cleavage domain may not have substantial nucleic acid cleavage activity. In some embodiments, the cleavage domain is non-enzymatically active. TAL effector nuclease
[0078] In some embodiments, the site-directed nuclease suitable for use in the fusion proteins, systems, or methods described herein is “TALEN” or “TAL effector nuclease.” TALEN refers to an engineered transcription activator-like effector nuclease, which typically contains a central domain and a cleavage domain of a DNA-binding tandem repeat sequence. TALENs can be generated by fusing the TAL effector DNA-binding domain with the DNA cleavage domain. In some cases, the DNA-binding tandem repeat sequence comprises 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13 that can recognize at least one specific DNA base pair. Transcription activator-like effector (TALE) proteins can be fused with nucleases such as wild-type or mutant Fok1 endonuclease or the catalytic domain of Fok1. Several mutations have been made in Fok1 for use in TALENs, which, for example, improve cleavage specificity or activity. Such TALENs can be engineered to bind any desired DNA sequence. TALENs can be used to generate gene modifications (e.g., nucleic acid sequence editing) by creating double-strand breaks in a target DNA sequence, subsequently undergoing NHEJ or HR. Double-strand breaks in DNA can trigger DNA break repair, allowing the introduction of one or more gene modifications (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ) or homology-directed repair (HDR). In HDR, a donor polynucleotide repair template or template polynucleotide can be provided, containing a homologous arm flanking the target DNA. In some cases, a single-stranded donor polynucleotide repair template is provided to facilitate HR. Detailed descriptions of TALEN and its use in gene editing can be found in, for example, the following: U.S. Patent Nos. 8,440,431; 8,440,432; 8,450,471; 8,586,363; and 8,697,853; Scharenberg et al., Curr Gene Ther, 2013, 13(4):291-303; Gaj et al., Nat Methods, 2012, 9(8):805-7; Beurdeley et al., Nat Commun, 2013, 4:1762; and Joung and Sander, Nat Rev MolCellBiol, 2013, 14(1):49-55.
[0079] In some embodiments, the TALEN is engineered to reduce nuclease activity. In some embodiments, the nuclease domain of the TALEN comprises a modified form of the wild-type nuclease domain. The modified form of the nuclease domain may include amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the nuclease domain. For example, the modified form of the nuclease domain may have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified form of the nuclease domain may not have substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is non-enzymatically active.
[0080] In some embodiments, a transcription activator-like effector (TALE) protein is fused with a domain capable of regulating transcription and not containing a nuclease. In some embodiments, the TALE protein is programmed to function as a transcription activator. In some embodiments, the TALE protein is programmed to function as a transcription repressor. For example, the DNA-binding domain of the TALE protein can be fused (e.g., linked) with one or more transcription activation domains or with one or more transcription repressor domains. Non-limiting examples of transcription activation domains include the herpes simplex VP16 activation domain and tetrameric repeat sequences of the VP16 activation domain (e.g., the VP64 activation domain). Non-limiting examples of transcription repressor domains include the Kruppel-associated box domain. Meganuclease
[0081] In some embodiments, the site-specific nuclease suitable for use in the fusion proteins, systems, or methods described herein is a broad-spectrum nuclease. A broad-spectrum nuclease generally refers to a rare-cutting endonuclease or a homing endonuclease that can be highly specific. A broad-spectrum nuclease can recognize DNA target sites ranging in length from at least 12 base pairs, such as 12 to 40, 12 to 50, or 12 to 60 base pairs. A broad-spectrum nuclease can be a modular DNA-binding nuclease, such as any fusion protein containing at least one catalytic domain of an endonuclease and at least one DNA-binding domain, or a protein specifying a nucleic acid target sequence. The DNA-binding domain may contain at least one motif that recognizes single-stranded or double-stranded DNA. A broad-spectrum nuclease can induce double-strand breaks. Double-strand breaks in DNA can induce DNA break repair, thereby allowing the introduction of one or more gene modifications (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ) or homologous directed repair (HDR). In HDR, a donor polynucleotide template containing a homologous arm with a side-attached target DNA site can be provided. The macronuclease can be monomeric or dimer. In some embodiments, the macronuclease is naturally occurring (found in nature) or wild-type, and in others, the macronuclease is non-natural, artificial, engineered, synthetic, rationally designed, or man-made. In some embodiments, the macronucleases disclosed herein include I-CreI macronuclease, I-CeuI macronuclease, I-Msol macronuclease, I-SceI macronuclease, variants thereof, derivatives thereof, and fragments thereof. A wide range of useful nucleases and their applications in gene editing can be described in, for example, the following literature: Silva et al., Curr Gene Ther, 2011, 11(1):11-27; Zaslavoskiy et al., BMC Bioinformatics, 2014, 15:191; Takeuchi et al., Proc Natl Acad Sci USA, 2014, 111(11):4061-4066 and U.S. Patent Nos. 7,842,489; 7,897,372; 8,021,867; 8,163,514; 8,133,697; 8,021,867; 8,119,361; 8,119,381; 8,124,36; and 8,129,134.
[0082] In some embodiments, the nuclease domain of the macronase includes a modified form of the wild-type nuclease domain. The modified form of the nuclease domain may include amino acid alterations (e.g., deletions, insertions, or substitutions) that reduce the nucleic acid cleavage activity of the nuclease domain. For example, the modified form of the nuclease domain may have no more than 90%, 80%, 70%, 60%, 50%, 40%, 30%, 20%, 10%, 5%, or 1% of the nucleic acid cleavage activity of the wild-type nuclease domain. The modified form of the nuclease domain may not have substantial nucleic acid cleavage activity. In some embodiments, the nuclease domain is non-enzymatically active. In some embodiments, the macronase can bind DNA but cannot cleave DNA. B. Recruitment Structure Domain
[0083] The fusion protein described in this article contains a recruitment domain that includes a site-specific DNA-binding domain. The recruitment domain of the fusion protein described in this article can contain any polypeptide with a unique recognition motif.
[0084] In some embodiments, the recruitment domain is a Cro repressor family protein. Cro repressor family proteins are phage transcription factors that function as homodimers. See, for example, MS Dubrava et al., N15 Cro and λCro: Orthologous DNA-binding domains with completely different but equally effective homodimer interfaces, Protein Science 17:803–812 (2008). In some embodiments, the recruitment domain comprises N15 Cro protein, λCro protein, P22 Cro protein (see, for example, ARPoteete et al., Bacteriophage P22 Cro Protein: Sequence, Purification, and Properties, Biochemistry 25:251–256 (1986)), 434Cro protein (see, for example, C. Wolberger et al., Structure of a phage 434Cro / DNA complex, Nature 335:789–795 (1988)), or combinations thereof. Further information on Cro family proteins from N15, λ, P22, and 434 phages can be found in BMHall et al., Extreme divergence between one-to-one orthologs: the structure of N15 Cro bound to operator DNA and its relationship to the λCro complex, Nucleic Acids Research, 47(13):7118–7129 (2019).
[0085] In some embodiments, the recruitment domain of the fusion protein provided herein comprises all or most of the polypeptide sequence of a naturally occurring DNA-binding protein. In some embodiments, the recruitment domain comprises only the DNA-binding domain of a naturally occurring DNA-binding protein. In some embodiments, the recruitment domain comprises one or more modifications relative to the protein from which it is derived (e.g., as described in the “Variations” section below). In some embodiments, the recruitment domain comprises a synthetic DNA-binding polypeptide sequence.
[0086] In some embodiments of the fusion proteins provided herein, the recruitment domain functions as an oligomer (i.e., binds to a specific DNA sequence). In some embodiments, the recruitment domain functions as a homooligomer. In some embodiments, the recruitment domain functions as a homodimer, homotrimer, homotetramer, or a higher-order homooligomer. In such embodiments, the fusion proteins provided herein may comprise more than one monomer (e.g., two, three, four, or more monomers) of the recruitment domain. In some embodiments, the monomers are separated by a linker (e.g., any linker described herein). For example, Cro repressor family proteins typically function as homodimers. In some embodiments, the fusion proteins provided herein comprise two or more monomers of a Cro repressor family protein recruitment domain sequence. In some embodiments, two or more monomers are linked by a flexible linker, thereby allowing the monomers to interact and form a homooligomer. Exemplary fusion proteins comprising two monomers of a Cro repressor family protein are described in the examples herein and depicted in [the following text is missing from the original extract]. Figure 4 and Figure 5 middle.
[0087] In some embodiments, the recruitment domain functions as a heterooligomer. In such embodiments, the fusion protein provided herein may comprise at least one monomer of two or more recruitment domain proteins.
[0088] In some embodiments, the recruitment domain comprises an amino acid sequence that is at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identical to any one of SEQ ID NO:1-4. C. Other fusion protein domains
[0089] In some embodiments, the fusion proteins provided herein comprise one or more linkers. As used herein, a linker (also referred to as a spacer) is a flexible molecule or flexible molecular segment that joins or connects two parts (e.g., domains) of a fusion protein or modified protein as provided herein. In some embodiments, the linker is a polypeptide. A protein having domains linked by a polypeptide linker is called a fusion protein. In some embodiments, the linker is a non-peptide linker. A protein having domains linked by a polypeptide linker is called a modified protein. It should be understood that, in the context of fusion proteins discussed throughout this disclosure, modified proteins are generally also contemplated (where feasible).
[0090] Linkers can increase the range of orientations that can be adopted by the domains of fusion proteins or modified proteins. Linkers can be optimized to produce the desired effect in fusion proteins or modified proteins. Aspects of linker design and consideration are described, for example, in Chen, X. et al., Adv Drug Deliv Rev. 2013 Oct 15; 65(10):1357-1369; and Klein, JS et al. 2014 Protein Eng. Des. Sel. 27(10):325-330. In some embodiments, the proteins provided herein contain peptide linkers. In some embodiments, the proteins provided herein contain non-peptide linkers. In some embodiments, the proteins provided herein contain both peptide and non-peptide linkers. The proteins provided herein may also contain multiple linkers, including at least one peptide linker, at least one non-peptide linker, or at least one peptide linker and at least one non-peptide linker.
[0091] The joints can be short or long, flexible or rigid. See, for example, PCT / US 2020 / 051383 (incorporated herein by reference in its entirety), WO 2020 / 168102 (incorporated herein by reference in its entirety), and US2021 / 0017506 (incorporated herein by reference in its entirety).
[0092] In some embodiments, the length of the linker can affect one or more functions of the fusion protein. Selecting the linker to achieve the desired length is within the capabilities of a person skilled in the art. In some embodiments, the length of the peptide linker can be, for example, 5 to 100 or more amino acids (e.g., 5 aa, 10 aa, 15 aa, 20 aa, 25 aa, 30 aa, 35 aa, 40 aa, 45 aa, 50 aa, 55 aa, 60 aa, 65 aa, 70 aa, 75 aa, 80 aa, 85 aa, 90 aa, 95 aa, or 100 aa).
[0093] Depending on their length, linker sequences can have various conformations of secondary structure, such as helices, β-chains, coils / bends, and rotations. In some cases, linker sequences can have an extended conformation and function as independent domains that do not interact with adjacent protein domains. Linker sequences can be flexible or rigid. Flexible linkers provide a degree of mobility or interaction between polypeptide domains and are typically rich in small or polar amino acids such as Gly and Ser (e.g., at least 90%, at least 95%, at least 98%, at least 99%, or all of the amino acid residues in the linker are Gly or Ser). Rigid linkers can be used to maintain a fixed distance between domains and help maintain their independent function. Linker attachment can be performed via amide bonding (e.g., peptide bonding) or other functionalities discussed further below.
[0094] In some embodiments, the peptide linker described herein comprises an amino acid sequence having at least 90% sequence identity with any one of SEQ ID NO: 6, 7, 15, or 16. In some embodiments, the linker comprises an XTEN linker sequence. See, for example, X. Li, et al., Base editing with a Cpf1–cytidine deaminase fusion, Nature Biotechnology 36:324–327 (2018); Y. Zong, et al., Precise base editing in rice, wheat and maize with a Cas9-cytidine deaminase fusion, Nature Biotechnology 35:438–440 (2017); and V. Schellenberger, et al., A recombinant polypeptide extends the in vivo half-life of peptides and proteins in a tunable manner, Nature Biotechnology 27:1189–1190 (2009). In some embodiments, the connector includes one or more repetitions of SGGS (SEQ ID NO:22), GGGS (SEQ ID NO:23), GGGGS (SEQ ID NO:24) (e.g., 2 repetitions, 3 repetitions, 4 repetitions, 5 repetitions, 6 repetitions or more) and / or one or more repetitions of GSSGSS (SEQ ID NO:25).Other exemplary peptide linkers include, but are not limited to, peptide linkers comprising: SGSETPGTSESATPE (“XTEN-03”, SEQ ID NO:27), SGSETPGTSESATPES (“XTEN-01”, SEQ ID NO:28), SGSETPGTSESATPELK (“XTEN-02”, SEQ ID NO:29), (GGGGS)3 (SEQ ID NO:30), (GGGGS)5 (SEQ ID NO:31), (GGGGS)10 (SEQ ID NO:32), GGGGGGGG (SEQ ID NO:33), GSAGSAAGSGEF (SEQ ID NO:34), A(EAAAK)3A (SEQ ID NO:35), or A(EAAAK)10A (SEQ ID NO:36). Other non-limiting exemplary connectors that may be used include those disclosed in the following literature: PCT / US 2020 / 051383; Chen et al., Adv. Drug. Deliv. Rev. [Advanced Drug Delivery Review] 65(10):1357-1369 (2014); and Rosemaren et al., Biochemistry [Biochemistry] 2017, 56, 50, 6565-6574, the entire contents of which are incorporated herein by reference.
[0095] In some embodiments, the non-peptide linker may comprise any plurality of known chemical linkers. Exemplary chemical linkers may comprise one or more units of β-alanine, 4-aminobutyric acid (GABA), (2-aminoethoxy)acetic acid (AEA), 5-aminohexanoic acid (Ahx), PEG polymers, and trioxane-tetrazine-succinic acid (Ttds). In some embodiments, the non-peptide linker comprises one or more units of polyethylene glycol (PEG), which is commonly used as a linker for conjugating peptide domains due to its water solubility, lack of toxicity, low immunogenicity, and well-defined chain length. See, for example, Ramirez-Paz, J. et al., PLoSOne [PLOS ONE] 13(7):e0197643 (2018). The number of PEG-linking units may be selected based on the desired linker length.
[0096] Proteins containing non-peptide linker modifications can be produced in a variety of ways. For example, the site-directed nuclease and recruitment domain can be produced separately (e.g., in vitro or by expression and purification from host cells) and chemically linked in vitro. In some embodiments, the site-directed nuclease / recruitment domain and linker can each be produced separately and chemically linked in vitro. Various chemical linkers can be used to crosslink two amino acid residues.
[0097] This document also envisions embodiments in which the site-directed nuclease and recruitment domain described above are used alone (e.g., introduced into the cell alone or applied to the target nucleic acid alone) and brought into contact to form a complex without using the adapter described above. Various methods for forming complexes between two or more peptides are known in the art and include, but are not limited to, the use of protein-protein interaction strategies (e.g., SunTag, coil-coil, etc.), the use of RNA-aptamers and associated binding proteins (e.g., MS2, N22, etc.), and tag:capture strategies. For example, the site-directed nuclease disclosed herein may contain an MS2 RNA aptamer that will facilitate interaction with a recruitment domain containing an MS2 shell protein.
[0098] In some embodiments, the fusion protein provided herein includes a targeting sequence that mediates protein localization (or retention) to a subcellular location, such as the plasma membrane or a given organelle membrane, nucleus, cytosol, mitochondria, endoplasmic reticulum (ER), Golgi apparatus, chloroplasts, apoplasts, peroxisomes, or other organelles. For example, the targeting sequence may use a nuclear localization signal (NLS) to direct the protein (e.g., a nuclease) to the nucleus; a nuclear export signal (NES) to direct the protein to the extranuclear region of the cell, such as to the cytoplasm; a mitochondrial targeting signal to direct the protein to the mitochondria; an ER retention signal to direct the protein to the endoplasmic reticulum (ER); a peroxisome targeting signal to direct the protein to the peroxisome; a membrane localization signal to direct the protein to the plasma membrane; or a combination thereof. In some embodiments, the fusion protein includes a nuclear localization signal.Non-limiting examples of NLS include NLS sequences derived from: NLS of the SV40 viral large T antigen, having the amino acid sequence PKKKRKV (SEQ ID NO:37); NLS from nucleoplasmic proteins (e.g., nucleoplasmic protein bipartite NLS having the sequence KRPAATKKAGQAKKKK (SEQ ID NO:38)); c-myc NLS, having the amino acid sequence PAAKRVKLD (SEQ ID NO:39) or RQRRNELKRSP (SEQ ID NO:40); hRNPA1 M9 NLS, having the sequence NQSSNFGPMKGGNFGGRSSGPYGGGGQYFAKPRN QGGY (SEQ ID NO:41); the sequence RMRIZFKNKGKDTAELRRRRVEVSVELRKAKKDEQILKRRNV (SEQ ID NO:42) from the IBB domain of the input protein-α; and the sequences VSRKRPRP (SEQ ID NO:43) and PPKKARED (SEQ ID NO:48) from the fibroid T protein. Sequence of human p53: PQPKKKPL (SEQ ID NO:45); Sequence of mouse c-abl IV: SALIKKKKKMAP (SEQ ID NO:46); Sequence of influenza virus NS1: DRLRR (SEQ ID NO:47) and PKQKKRK (SEQ ID NO:48); Sequence of hepatitis virus delta antigen: RKLKKKIKKL (SEQ ID NO:49); Sequence of mouse Mx1 protein: REKKKFLKRR (SEQ ID NO:50); Sequence of human poly(ADP-ribose) polymerase: KRKGDEVDGVDEVAKKKSKK (SEQ ID NO:51); Sequence of steroid hormone receptor (human) glucocorticoid: RKCLQAGMNLEARKTKK (SEQ ID NO:52); and Sequence of Agrobacterium VirD2 protein: KRPRDRHDGELGGRKRAR (SEQ ID NO:53).
[0099] In some embodiments, the fusion protein provided herein comprises an amino acid sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with SEQ ID NO:11 or 13.
[0100] Any polypeptide and fusion protein described herein may further include a detectable portion, such as a fluorescent protein or a fragment thereof. Examples of fluorescent proteins include, but are not limited to, yellow fluorescent protein (YFP, e.g., Venus), green fluorescent protein (GFP), and red fluorescent protein (RFP), as well as derivatives of these proteins, such as mutant derivatives. See, for example, Chudakov et al., “Fluorescent Proteins and Their Applications in Imaging Living Cells and Tissues,” Physiological Reviews 90(3):1103-1163 (2010); and Spetht et al., “A Critical and Comparative Review of Fluorescent Tools for Live-Cell Imaging,” Annual Review of Physiology 79:93-117 (2017).
[0101] For example, any polypeptide described herein may further comprise an affinity tag (e.g., a multihistidine tag (e.g., (His)6 (SEQ ID NO:54)), an HA tag (e.g., YPYDVPDYA (SEQ ID NO:55)), an albumin-binding protein, an alkaline phosphatase, an AU1 epitope, an AU5 epitope, a biotin-carboxyl carrier protein (BCCP), a FLAG epitope (e.g., DYKDDDDK (SEQ ID NO:56)), or a MYC epitope (e.g., EQKLISEEDL (SEQ ID NO:57)). See Kimple et al., “Overview of Affinity Tags for Protein Purification,” Curr. Protoc. Protein Sci. 73: Unit 9.9 (2013). D. Variant
[0102] This document also provides variants of the peptides disclosed herein. Unless otherwise expressly indicated, peptide variants retain their corresponding biological activities. For example, variants of site-directed nuclease peptides retain the biological functions of the full-length natural sequence site-directed nuclease. In another instance, variants of recruitment domains retain the biological functions of the full-length natural sequence recruitment domain.
[0103] Modification of any of the polypeptides or proteins provided herein is performed using known methods. By way of example, modification is performed by site-specific mutagenesis of nucleotides in a nucleic acid encoding the polypeptide, thereby producing DNA encoding the modification, and then expressing the DNA in a recombinant cell culture to produce the polypeptide. Techniques for substitution mutations at predetermined sites in DNA having a known sequence are well known. For example, one or more substitution mutations can be generated using M13 primer mutagenesis and PCR-based mutagenesis methods. Any of the nucleic acid sequences provided herein can be codon-optimized to alter, for example, to maximize expression in a host cell or organism.
[0104] The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D-stereoisomers of naturally occurring amino acids, non-natural amino acids, and chemically modified amino acids. Non-natural amino acids (i.e., those not naturally found in proteins) are also known in the art, as illustrated in, for example, the following references: Zhang et al., “Protein engineering with unnatural amino acids,” Curr. Opin. Struct. Biol. 23(4):581-587 (2013); Xie et al., “Adding amino acids to the genetic repertoire,” 9(6):548-54 (2005); and all references cited therein. β and γ amino acids are known in the art and are also conceived as non-natural amino acids in this paper.
[0105] As used herein, a chemically modified amino acid is one whose side chain has been chemically modified. For example, the side chain can be modified to include a signal transduction motif, such as a fluorophore or radiolabel. It can also be modified to include new functional groups, such as thiols, carboxylic acids, or amino groups. Post-translational modified amino acids are also included in the definition of chemically modified amino acids.
[0106] Conservative amino acid substitutions are also envisioned. By way of example, conserved amino acid substitutions can be carried out in one or more amino acid residues, for example, in one or more lysine residues of any of the polypeptides provided herein. Those skilled in the art will appreciate that a conserved substitution is the replacement of one amino acid residue with another amino acid residue that is biologically and / or chemically similar. The following eight groups each contain amino acids that are conservedly substituted for each other: 1) Alanine (A), glycine (G); 2) Aspartic acid (D), glutamic acid (E); 3) Asparagine (N), glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W); 7) Serine (S), threonine (T); and 8) Cysteine (C), Methionine (M).
[0107] By way of example, when referring to arginine to serine, conserved substitutions of serine (e.g., threonine) are also envisioned. Non-conserved substitutions are also envisioned, such as replacing lysine with asparagine. IV. Recombinant nucleic acids, constructs, vectors, and host cells
[0108] This document also provides recombinant nucleic acids encoding any of the polypeptides described herein. For example, recombinant nucleic acids encoding polypeptides having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with SEQ ID NO:12 or 14 are also provided. Recombinant nucleic acids having at least 70% identity with SEQ ID NO:12 or 14 are also provided.
[0109] DNA constructs are also provided that contain a promoter operatively linked to a recombinant nucleic acid encoding a fusion protein or its domain as described herein. The nucleic acid is “operatively linked” when it is positioned to have a functional relationship with another nucleic acid sequence. A variety of promoters can be used in the constructs described herein. A promoter is a region or sequence located upstream and / or downstream of transcription initiation and involved in the recognition and binding of RNA polymerases and other proteins to initiate transcription.
[0110] As used herein, the term "promoter" refers to a nucleotide sequence, typically upstream (5') of its coding sequence, that controls the expression of the coding sequence by providing recognition of RNA polymerases and other factors required for proper transcription. "Promoter regulatory sequences" consist of proximal and more distal upstream elements. Promoter regulatory sequences influence transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. These include native and synthetic sequences, as well as sequences that may be combinations of synthetic and native sequences. An "enhancer" is a DNA sequence that can stimulate promoter activity and can be an intrinsic element of the promoter or an inserted heterologous element to enhance the promoter's level or tissue specificity. It is capable of operating in both orientations (normal or inverted) and can function even when moved upstream or downstream of the promoter. The term "promoter" includes the meaning of "promoter regulatory sequences."
[0111] The selection of promoters to be included depends on several factors, including but not limited to efficiency, selectivity, inducibility, desired expression level, and cell or tissue preference. It is common practice for those skilled in the art to regulate the expression of a sequence by appropriately selecting and locating promoters and other regulatory regions relative to the sequence.
[0112] It has been shown that some promoters can direct RNA synthesis at a higher rate than others. These are called “strong promoters.” Other promoters have been shown to direct RNA synthesis at higher levels only in specific cell or tissue types and are often called “tissue-specific promoters,” or “tissue-preferred promoters” if a promoter preferentially directs RNA synthesis in some tissues (where RNA synthesis may occur at reduced levels in others). Because promoters are used to control the expression patterns of one (or more) chimeric genes introduced into plants, there is ongoing interest in isolating novel promoters capable of controlling the expression of one (or more) chimeric genes at certain levels in specific tissue types or at specific plant developmental stages.
[0113] Certain promoters are capable of directing RNA synthesis at relatively similar levels in all plant tissues. These are called “constitutive promoters” or “tissue-independent” promoters. Constitutive promoters can be categorized into strong, moderate, and weak classes based on their effectiveness in directing RNA synthesis. Constitutive promoters are particularly useful in this regard because it is often necessary to simultaneously express one (or more) chimeric genes in different plant tissues to achieve the desired function of one (or more) genes. Although many constitutive promoters have been identified and characterized from plants and plant viruses, there remains a continuing interest in isolating more novel constitutive promoters (synthetic or natural) capable of controlling the expression of one (or more) chimeric genes at different levels and controlling the expression of multiple genes in the same transgenic plant for gene stacking.
[0114] The most commonly used promoters include the cauliflower mosaic virus (NOS) promoter (Ebert et al., Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences of the United States of America] 84:5745-5749 (1987)); the octopus cauliflower mosaic virus (OCS) promoter; promoters of the cauliflower mosaic virus group such as the cauliflower mosaic virus (CaMV) 19S promoter (Lawton et al., Plant Mol. Biol. [Plant Molecular Biology] 9:315-324 (1987)); and photoinducible promoters derived from the small subunit of ribulose diphosphate carboxylase / oxygenase (Pellegrineschi et al., Biochem. Soc. Trans.). [Biochemical Society Report] 23(2):247-250 (1995)); Adh promoter (Walker et al., Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences of the United States of America] 84:6624-66280 (1987)); Sucrose synthase promoter (Yang et al., Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences of the United States of America] 87:414-44148 (1990)); R gene complex promoter (Chandler et al., Plant Cell [Plant Cell] 1:1175-1183 (1989)); Chlorophyll a / b binding protein gene promoter; etc.
[0115] Furthermore, it is envisioned that promoters combining elements from more than one promoter may be useful. For example, U.S. Patent No. 5,491,288 discloses the combination of a cauliflower mosaic virus promoter with a histone promoter. Therefore, elements from the promoters disclosed herein can be combined with elements from other promoters. Promoters that can be used for transgenic plant expression include inducible, viral, synthetic, constitutive (Odell Nature 313:810-812 (1985)), temporally regulated, spatially regulated, tissue-specific, and spatiotemporally regulated promoters. Using the regulatory elements described herein, many agronomic genes can be expressed in transformed plants. More specifically, plants can be genetically engineered to express a variety of phenotypes for agronomic purposes.
[0116] In some embodiments of the DNA constructs provided herein, the promoter may be a eukaryotic or prokaryotic promoter. In some embodiments, the promoter is an inducible promoter, a naturally inducible promoter (e.g., drought-inducible Rab17), a synthetic inducible promoter (e.g., auxin-inducible DR5, estradiol-inducible XVE / pLex, dexamethasone-inducible GVG / Gal4), a constitutive promoter (e.g., ZmUbq1, OsAct1, OsTub3, EF), an oocyte-specific promoter (e.g., EC1, EC2, EC3, EC4, EC5), a pollen-specific promoter, an apical meristem-specific promoter, or a promoter with enriched expression in the conjugate. In some embodiments, the promoter is a floral mosaic promoter (e.g., ZmBde1, OsAP1). In some embodiments, the promoter is a ubiquitin 4 promoter, an actin promoter, a tubulin promoter, a MADS box promoter, or a plant virus promoter. Suitable promoters are disclosed, for example, in U.S. Patent No. 10,519,456 (the entire contents of which are incorporated herein by reference) and PCT / US2022 / 020690 (incorporated herein by reference).
[0117] The recombinant nucleic acid provided herein may be included in an expression cassette for expression in a target host cell or organism. The cassette will include 5' and 3' regulatory sequences operatively linked to the recombinant nucleic acid provided herein that allows expression of the fusion protein. The cassette may additionally contain at least one additional gene or genetic element to be co-transformed into a cell or organism. Where additional genes or elements are included, these components are operatively linked. Alternatively, one or more additional genes or elements may be provided on multiple expression cassettes. Such expression cassettes are provided with multiple restriction sites and / or recombination sites to allow the insertion of polynucleotides under transcriptional regulation of the regulatory region. The expression cassette may additionally contain a selective marker gene. The expression cassette will include, in reverse 5' to 3' of transcription: a transcription and translation initiation region (i.e., promoter) that functions in the target cell or organism, the polynucleotide of the present invention, and a transcription and translation termination region (i.e., termination region). The promoters of this invention are capable of directing or driving the expression of coding sequences (i.e., nucleic acid sequences transcribed into RNA such as mRNA, rRNA, tRNA, snRNA, ncRNA, lncRNA, sense RNA, or antisense RNA, regardless of whether the RNA is then translated to produce a protein) in host cells. For the host cell or for each other, the regulatory regions (i.e., promoters, transcriptional regulatory regions, and translation termination regions) can be endogenous or heterologous. As used herein, "heterogeneous" in relation to a sequence is a sequence that originates from a foreign species, or, if from the same species, has undergone substantial modifications in its natural form in terms of composition and / or genomic loci through intentional human intervention.
[0118] Other regulatory signals include, but are not limited to, transcription initiation sites, operons, activators, enhancers, other regulatory elements, ribosome binding sites, start codons, and termination signals. See Sambrook et al., (1992) *Molecular Cloning: A Laboratory Manual*, edited by Maniatis et al. (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, New York; Davis et al., edited by Davis et al., (1980) *Advanced Bacterial Genetics* (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, New York, and the references cited therein.
[0119] Expression cassettes may also contain selective marker genes for selecting transformed cells. For example, marker genes include those conferring antibiotic resistance, such as those conferring resistance to hygromycin, ampicillin, gentamicin, and neomycin. Other selective markers are known, and any of them may be used.
[0120] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences that are appropriately oriented and, where appropriate, form the appropriate reading frame. For this purpose, adaptors or linkers can be used to ligate the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, or eliminate restriction sites. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, and re-substitution (e.g., conversion and transversion) can be employed.
[0121] In preparing expression cassettes, various DNA fragments can be manipulated to provide DNA sequences that are appropriately oriented and, where appropriate, form the appropriate reading frame. For this purpose, adaptors or linkers can be used to ligate the DNA fragments, or other manipulations can be involved to provide convenient restriction sites, remove excess DNA, or eliminate restriction sites. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, and substitution (e.g., conversion and transversion) can be used.
[0122] Furthermore, vectors containing the recombinant nucleic acid or DNA constructs described herein are provided. The vectors are envisioned to possess essential functional elements that guide and regulate the transcription of the inserted nucleic acid. These functional elements include, but are not limited to, promoters, regions upstream or downstream of the promoter (such as enhancers that can regulate the transcriptional activity of the promoter), origins of replication, appropriate restriction sites for promoting the cloning of the inserted sequence adjacent to the promoter, antibiotic resistance genes or other markers that can be used to select cells containing the vector or vectors containing the inserted sequence, RNA splicing junctions, transcription termination regions, or any other regions that can be used to promote the expression of the inserted gene or heterozygous gene. See generally Sambrook et al., *Molecular Cloning: A Laboratory Manual*, 4th edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. Vectors may be, for example, plasmids.
[0123] Numerous *E. coli* expression vectors are known to those skilled in the art and can be used for nucleic acid expression. Other suitable microbial hosts include bacilli (such as *Bacillus subtilis*) and other Enterobacteriaceae (such as *Salmonella* and *Senatia*) and various *Pseudomonas* species. Expression vectors can also be prepared in these prokaryotic hosts, and these vectors will typically contain expression control sequences (e.g., origin of replication) compatible with the host cell. Additionally, any number of well-known promoters will be available, such as the lactose promoter system, the tryptophan (Trp) promoter system, the β-lactamase promoter system, or promoter systems derived from bacteriophage λ. Yeast expression can also be used. This document provides nucleic acids encoding the polypeptides of the present invention, wherein the nucleic acids can be expressed by yeast. More specifically, the nucleic acids can be expressed by *Pichia pastoris* or *Saccharomyces cerevisiae*.
[0124] Mammalian cells also allow proteins to be expressed in environments conducive to important post-translational modifications such as folding and cysteine pairing, the addition of complex carbohydrate structures, and the secretion of active proteins. Vectors for expressing active proteins in mammalian cells are known in the art and may contain genes conferring resistance to hygromycin, genimycin, or G418, or other genes or phenotypes suitable for use as selectivity markers, or methotrexate resistance for gene amplification. A variety of suitable host cell lines capable of secreting fully human proteins have been developed in the art, including CHO cells, HeLa cells, HEK-293 cells, HEK-293T cells, U2OS cells, or any other primary or transformed cell lines. Other suitable host cell lines include COS-7 cells, myeloma cell lines, Jurkat cells, etc. Expression vectors for these cells may include expression control sequences such as origin of replication, promoters, enhancers, and essential information processing sites such as ribosome binding sites, RNA splicing sites, polyadenylation sites, and transcription terminator sequences. The preferred expression control sequence is a promoter derived from immunoglobulin genes, SV40, adenovirus, bovine papillomavirus, etc.
[0125] The expression vectors described herein may also include nucleic acids as described herein, controlled by inducible promoters (such as tetracycline-inducible promoters or glucocorticoid-inducible promoters). The nucleic acids of the present invention may also be controlled by tissue-specific promoters to promote nucleic acid expression in specific cells, tissues, or organs. Any regulated promoters, such as metallothionein promoters, heat shock promoters, and other regulated promoters, many examples of which are well known in the art, are also contemplated. Furthermore, the Cre-loxP inducible system and the Flp recombinase inducible promoter system, both of which are known in the art, may also be used.
[0126] Insect cells also allow for peptide expression. Recombinant proteins generated in insect cells using baculovirus vectors undergo post-translational modifications similar to those of wild-type mammalian proteins.
[0127] This document also provides host cells comprising the recombinant nucleic acids, DNA constructs, and / or vectors described herein, as well as methods for preparing such cells. In some embodiments, the cells are plant cells. In some embodiments, the plant cells are corn plant cells, soybean plant cells, rice plant cells, wheat plant cells, or sunflower plant cells.
[0128] Host cells comprising the nucleic acids or vectors described herein are provided. The host cells may be in vitro, ex vivo, or in vivo. Host cells as provided herein are capable of expressing fusion proteins. Cell populations of any of the host cells described herein are also provided. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells contain a recombinant nucleic acid encoding a fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells contain a DNA construct encoding a fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells contains a vector containing a recombinant nucleic acid or DNA construct encoding a fusion protein as described herein. In some embodiments, the cell population comprises a plurality of cells, wherein the plurality of cells includes a plurality of any of the host cells described herein. In some embodiments, the plurality of cells in any of the cell populations described herein express a fusion protein as described herein.
[0129] In some embodiments, the provided cells exhibit stable or transient expression of the fusion protein. Stable expression of the fusion protein in cells refers to the integration of any of the nucleic acids, DNA constructs, or vectors described herein into the cell genome, thereby allowing the cell to express the fusion protein. Transient expression refers to the expression of the fusion protein directly from any of the nucleic acids, DNA constructs, and / or vectors introduced into the cell (i.e., the gene encoding the fusion protein is not integrated into the cell genome).
[0130] In some embodiments, the provided cells constitutively or inducibly express the fusion protein. Constitutive expression refers to the sustained, continuous expression of a gene (i.e., a protein), while inducible expression refers to gene (protein) expression in response to a stimulus. Inducible expression is typically regulated via an inducible promoter (described above).
[0131] Cell cultures comprising one or more of the host cells described herein are also provided. Methods for culturing and producing numerous cells are available in the art, including cells derived from bacteria (e.g., Escherichia coli and other bacterial strains), animals (especially mammals), and archaea. See, for example, Sambrook, ibid.; Ausubel, ed. (1995) Current Protocols in Molecular Biology, John Wiley & Sons; and Freshney (1994) Culture of Animal Cells, a Manual of Basic Technique, 3rd ed., Wiley-Liss, New York, and the references cited therein; Doyle and Griffiths (1997) Mammalian Cell Culture: Essential Techniques, John Wiley and Sons, New York; Humason (1979) Animal Tissue Techniques, 4th ed., WH Freeman and Company; and Ricciardelli, et al. (1989) Invitro Cell Dev.Biol. [In Vitro Cell and Developmental Biology] 25:1016-1024.
[0132] The host cell can be a prokaryotic cell, including, for example, bacterial cells. Alternatively, the cell can be a eukaryotic cell, such as a plant cell, yeast cell, insect cell, avian cell, or mammalian cell. The plant cell can be a field or greenhouse crop plant cell, including but not limited to large-scale crop plants, fruits and vegetables, perennial trees, and ornamental plants. In some embodiments, the plant cell is the plant cell of sugarcane, squash, corn, wheat, rice, cassava, soybean, hay, potato, cotton, tomato, alfalfa, and green algae. In some embodiments, the plant cell is the plant cell of any vegetable (e.g., cabbage, turnip, carrot, parsley, beetroot, lettuce, legume, broad bean, pea, potato, eggplant, tomato, cucumber, squash, squash, onion, garlic, leek, pepper, spinach, yam, sweet potato, and cassava). In some embodiments, the plant cell is the plant cell of corn, soybean, sunflower, tomato, rice, or wheat. In some embodiments, the mammalian cells may be HEK-293T cells, HEK-293 cells, Chinese hamster ovary (CHO) cells, U2OS cells, COS-7 cells, HELA cells, or any other primary or transformed cells. A variety of other suitable host cell lines have been developed, including myeloma cell lines, fibroblast cell lines, and various tumor cell lines such as melanoma cell lines. Vectors containing the target nucleic acid segment can be transferred or introduced into host cells using well-known methods that vary depending on the type of cell host.
[0133] As used herein, the phrase “introduction” in the context of introducing nucleic acids into cells (e.g., prokaryotic cells, bacterial cells, eukaryotic cells, plant cells) refers to the translocation of a nucleic acid sequence from the extracellular space to the intracellular space. In some cases, introduction is the translocation of nucleic acids from the extracellular space to the cell nucleus. In cases where more than one nucleic acid molecule is to be introduced, these nucleic acid molecules may be assembled as part of a single polynucleotide or nucleic acid construct, or as separate polynucleotides or nucleic acid constructs, and may be located on the same or different nucleic acid constructs. Thus, such polynucleotides may be introduced into cells (e.g., plant cells) in a single transformation event, in a separate transformation event, or, for example, as part of a breeding program. Various methods for introducing nucleic acids into cells are contemplated, including but not limited to electroporation, nanoparticle delivery, gene gun transformation, viral delivery, contact with nanowires or nanotubes, receptor-mediated internalization, translocation via cell-penetrating peptides, liposome-mediated translocation, DEAE dextran, lipofectamine, calcium phosphate, or any method now known or to be identified in the future for introducing nucleic acids into prokaryotic or eukaryotic cell hosts. Targeted nuclease systems (e.g., RNA-directed nucleases, transcription activator-like effector nucleases (TALENs), zinc finger nucleases (ZFNs), or large-scale TAL (MT)) can also be used to introduce nucleic acids (e.g., nucleic acids encoding the fusion proteins described herein) into host cells. See Li et al., Signal Transduction and Targeted Therapy, 5, Article 1 (2020).
[0134] Cellular transformation can be stable or transient. Therefore, the transgenic cells, plant cells, plants, and / or plant parts of the present invention can be stably or transiently transformed. "Transformation" can refer to the transfer of nucleic acid molecules into the genome of a host cell, resulting in genetically stable inheritance. In some embodiments, the introduction into plants, plant parts, and / or plant cells is carried out via bacterial-mediated transformation, particle bombardment transformation, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, liposome-mediated transformation, nanoparticle-mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infiltration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical, and / or biological mechanism or any combination thereof that introduces nucleic acids into plants, plant parts, and / or their cells.
[0135] Procedures for transforming plants are well-known and conventional in the art and are commonly described in the literature. Non-limiting examples of methods for plant transformation include transformation via: bacterial-mediated nucleic acid delivery (e.g., via bacteria from the genus Agrobacterium), virus-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome-mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrical, chemical, physical (mechanical), and / or biological mechanisms, including any combination thereof, that introduce nucleic acids into plant cells. General guidelines for plant transformation methods known in the art include Miki et al. (“Procedures for Introducing Foreign DNA into Plants” in Methods in Plant Molecular Biology and Biotechnology, edited by Glick, BR and Thompson, JE (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Letters 7:849-858 (2002)).
[0136] Agrobacterium-mediated transformation is a common method for transforming plants due to its high transformation efficiency and its wide applicability to many different species. Agrobacterium-mediated transformation typically involves the transfer of a binary vector carrying the target foreign DNA into a suitable Agrobacterium strain, which may depend on the complement of the vir gene carried by the host Agrobacterium strain on a co-existing Ti plasmid or chromosomally (Uknes et al., 1993, Plant Cell 5:159-169). The transfer of the recombinant binary vector into Agrobacterium can be achieved using Escherichia coli carrying the recombinant binary vector, an auxiliary E. coli strain (carrying a plasmid capable of moving the recombinant binary vector into the target Agrobacterium strain), via a three-parental mating procedure. Alternatively, the recombinant binary vector can be transferred into Agrobacterium via nucleic acid transformation. And Willmitzer 1988, Nucleic Acids Res [Nucleic Acid Research] 16:9877).
[0137] Plant transformation via recombinant Agrobacterium typically involves co-culturing Agrobacterium with explants from plants, following methods well-known in the art. Typically, the transformed tissues are regenerated on selective media carrying antibiotic or herbicide resistance markers located between the boundaries of the binary plasmid T-DNA.
[0138] Another method for transforming plants, plant parts, and plant cells involves advancing inert or biologically active particles onto plant tissues and cells. See, for example, U.S. Patent Nos. 4,945,050; 5,036,006, and 5,100,792. Typically, this method involves advancing inert or biologically active particles onto plant cells under conditions that are effective in penetrating the outer surface of the cells and providing incorporation within them. When using inert particles, the particles can be introduced into the cells by coating the particles with a carrier containing the target nucleic acid. Alternatively, one or more cells can be surrounded by a carrier such that the carrier is carried into the cells by excitation of the particles. Biologically active particles (e.g., dried yeast cells, dried bacteria, or bacteriophages, each containing one or more nucleic acids to be introduced) can also be advanced into plant tissues. As used herein, the phrase “gene gun conversion” refers to a method of directly introducing RNA or DNA into cells (e.g., plant cells), wherein the RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cells (e.g., plant cells) using high-speed pressure to allow the RNA or DNA to penetrate the cells (e.g., penetrate the plant cell wall).
[0139] The CRISPR / Cas system can also be used to edit the genome of a host cell or organism. As detailed above, the “CRISPR / Cas” system refers to a class of bacterial systems widely used to defend against foreign nucleic acids. Any of the CRISPR / Cas system components described herein can be used to introduce fusion proteins, recombinant nucleic acids, or systems into the genome of a host cell or organism. Methods for CRISPR / Cas system-mediated genome editing are known in the art. It should be understood that the use of the CRISPR / Cas system for introducing the fusion proteins, recombinant nucleic acids, or systems described herein into the genome of a host cell or organism differs from the specific methods and systems provided herein.
[0140] Any of the fusion proteins described herein can be purified or isolated from host cells or populations of host cells. For example, a recombinant nucleic acid encoding any of the fusion proteins described herein can be introduced into a host cell under conditions that allow for the expression of the fusion protein. In some embodiments, the recombinant nucleic acid is optimized for expression codons. After expression in host cells, the fusion protein can be isolated or purified using purification methods known in the art. V. System
[0141] On the other hand, this document provides systems that can be used to edit one or more nucleic acids. These systems include one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above. In some embodiments, these systems further include one or more additional elements that can be used to edit one or more nucleic acids. For example, the systems provided herein may further include donor polynucleotides. As another example, systems comprising fusion proteins containing Cas nucleases may further include one or more guide nucleic acids and / or one or more donor polynucleotide sequences. Donor polynucleotides and guide nucleic acids are described in detail below. The systems provided herein can be used to perform the methods described in Section VI of this disclosure. A. Donor polynucleotides
[0142] The systems and methods disclosed herein may include donor polynucleotides. A “donor polynucleotide,” “donor molecule,” or “donor template” is a polymer or oligomer of nucleotides intended for insertion at a target polynucleotide (typically a target genomic site). The donor sequence may be one or more target transgenes, expression cassettes, or nucleotide sequences. The donor molecule may be a single-stranded, partially double-stranded, or double-stranded donor DNA molecule. The donor polynucleotide may be a natural or modified polynucleotide, an RNA-DNA chimera, or a DNA fragment, a single-stranded, or at least partially double-stranded, or fully double-stranded DNA molecule, or PGR-amplified ssDNA, or at least partially dsDNA fragment. In some embodiments, the donor DNA molecule is part of a circularized DNA molecule. In some cases, fully double-stranded donor DNA may provide increased stability because dsDNA fragments are generally more resistant to nuclease degradation than ssDNA.
[0143] In some embodiments, the donor polynucleotide comprises at least one recruitment sequence that binds to a recruitment domain of at least one fusion protein provided herein. In some embodiments, the donor polynucleotide comprises two, three, four, five, six, seven, eight, or more recruitment sequences. In some embodiments, two or more recruitment sequences are identical sequences. In some embodiments, two or more recruitment sequences are different sequences. In some embodiments, the recruitment sequence comprises at least 10 (e.g., at least 12, at least 14, at least 16, at least 18, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, or more) consecutive nucleotides that are at least 70% identical to a recognition motif that specifically binds to the recruitment domain. In some embodiments, the recruitment domain comprises a naturally occurring sequence that specifically binds to a DNA-binding protein. In some embodiments, the recruitment domain comprises one or more modifications relative to the naturally occurring sequence. In some embodiments, the recruitment domain comprises a synthetic sequence that specifically binds to the recruitment domain described herein.
[0144] In some embodiments, the recruitment sequence comprises a Cro repressor family protein operon sequence (“Cro O”). R 3” sequence). In some embodiments, the recruitment sequence contains N15 O R 3. Operator sequence, λO R 3. Operator sequence, P22 O R 3 operon sequences, 434O R 3. Operator sequence or combination thereof. In some embodiments, the recruited sequence comprises a nucleotide sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with any of SEQ ID NO:5-10.
[0145] The donor molecule may contain at least 10 consecutive nucleotides (often called homologous arms) in which the nucleic acid molecule is at least 70% identical to the genomic nucleotide sequence, such that these consecutive nucleotides are sufficient to homologously recombine the donor polynucleotide molecule into the cell’s genome at the target genomic DNA sequence after cleavage, for example, by site-directed nuclease. In some embodiments, the donor polynucleotide molecule may comprise at least about 10, 20, 30, 50, 70, 80, 100, 150, 200, 250, 300, 250, 400, 450, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, 7500, 10000, 15,000, or 20,000 nucleotides, including any value within this range not explicitly stated herein, wherein the donor polynucleotide molecule is at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identical to the genomic nucleic acid sequence. In some embodiments, the donor molecule includes at least one homologous arm. In some embodiments, the donor molecule includes two homologous arms (which may be referred to as the left homologous arm and the right homologous arm).
[0146] In some embodiments, the donor polynucleotide molecule may be substantially complementary to a genomic nucleic acid sequence. In some embodiments, the donor polynucleotide molecule comprises a heterologous nucleic acid sequence. In some embodiments, the donor polynucleotide molecule comprises at least one expression cassette. In some embodiments, the donor polynucleotide molecule may comprise a transgene containing at least one expression cassette. In some embodiments, the donor polynucleotide molecule comprises an allelic modification of a gene that is natural to the target genome. The allelic modification may include at least one nucleotide insertion, at least one nucleotide deletion, and / or at least one nucleotide substitution. In some embodiments, the allelic modification may include a small insertion or deletion.
[0147] The donor polynucleotide can be any suitable nucleic acid. In some embodiments, the donor polynucleotide is part of a donor template. In some embodiments, the donor template is part of a plasmid or linear nucleic acid. In some embodiments, the donor polynucleotide is part of a chromosome.
[0148] The various sequences of donor polynucleotides described in this article can be organized in any suitable configuration. Figure 1 Exemplary embodiments of donor polynucleotides are shown, which contain CroO R 3. A recruitment sequence and left and right homologous arms located flanking the insert sequence (designed for integration into the target site (e.g., according to the methods described herein)). In some embodiments, the donor polynucleotide comprises two homologous arms located flanking the insert sequence. In embodiments comprising one or more recruitment sequences, the recruitment sequence may be upstream and / or downstream of the insert sequence. In some embodiments, the recruitment sequence is between the homologous arms and the insert sequence. In some embodiments, the recruitment sequence is outside the homologous arms. In some embodiments, the donor polynucleotide comprises more than one recruitment sequence containing tandem identical sequences.
[0149] In some embodiments, the donor polynucleotide comprises a nucleotide sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with SEQ ID NO:20 or 21. B. Guiding nucleic acid testing
[0150] In some cases, the systems and methods described herein include at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein include multiple guide nucleic acids. In some embodiments, the polynucleotide may be deoxyribonucleic acid (DNA). In some embodiments, the DNA sequence may be single-stranded or double-stranded. In some embodiments, at least one guide nucleic acid polynucleotide may be ribonucleic acid (guide RNA).
[0151] In some embodiments, a nuclease may be complexed with at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide may contain a nucleic acid targeting region containing a sequence complementary to a nucleic acid sequence on a targeting polynucleotide (such as a targeting genomic locus or gene) to confer sequence specificity for the nuclease targeting. In some embodiments, the at least one guide RNA polynucleotide may comprise two separate nucleic acid molecules (which may be referred to as dual guide RNAs) or a single nucleic acid molecule (which may be referred to as a single guide RNA (e.g., a single guide RNA or sgRNA)). In some embodiments, the guide RNA is a single guide RNA comprising a fused CRISPR RNA (crRNA) and a trans-activating crRNA (tracrRNA). In some embodiments, the guide RNA is a single guide RNA comprising crRNA. In some embodiments, the guide RNA is a single guide RNA comprising crRNA but lacking tracrRNA. In some embodiments, the guide RNA is a dual guide RNA comprising a non-fused crRNA and tracrRNA. Exemplary dual guide RNAs may comprise crRNA-like molecules and tracrRNA-like molecules. Exemplary single guide RNAs may comprise crRNA-like molecules. Exemplary single guide RNAs may comprise fused crRNA-like molecules and tracrRNA-like molecules.
[0152] crRNA may contain a nucleic acid targeting segment (e.g., a spacer region) that guides the nucleic acid and a nucleotide segment that can form half of a double-stranded double helix that guides the nucleic acid Cas protein binding segment.
[0153] tracrRNA can contain the other half of a double-stranded nucleotide segment that forms the Cas protein-binding domain of gRNA. The nucleotide segment of crRNA can be complementary to and hybridize with the nucleotide segment of tracrRNA to form a double-stranded nucleotide segment that guides the Cas protein-binding domain of nucleic acid.
[0154] crRNA and tracrRNA can hybridize to form a guide nucleic acid. crRNA can also provide a single-stranded nucleic acid targeting segment (e.g., a spacer region) that hybridizes with a target nucleic acid recognition sequence (e.g., a prototype spacer). The sequence of a crRNA or tracrRNA molecule that includes a spacer region can be designed to be species-specific for the guide nucleic acid to be used.
[0155] Whether a nuclease requires only crRNA molecules or both crRNA and tracrRNA molecules (whether covalently linked or not) depends on the CRISPR-associated nuclease used.
[0156] In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid can be between 18 and 72 nucleotides. The length of the nucleic acid targeting region guiding the nucleic acid (e.g., a spacer region) can be from about 12 nucleotides to about 100 nucleotides. For example, the length of the nucleic acid targeting region guiding the nucleic acid (e.g., a spacer region) can be from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 40 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 12 nt to about 18 nt, from about 12 nt to about 17 nt, from about 12 nt to about 16 nt, or from about 12 nt to about 15 nt. Alternatively, the length of the DNA targeting segment can be from about 18 nt to about 20 nt, from about 18 nt to about 25 nt, from about 18 nt to about 30 nt, from about 18 nt to about 35 nt, from about 18 nt to about 40 nt, from about 18 nt to about 45 nt, from about 18 nt to about 50 nt, from about 18 nt to about 60 nt, from about 18 nt to about 70 nt, from about 18 nt to about 80 nt, from about 18 nt to about 90 nt, from about 18 nt to about 100 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, from about 20 nt to about 60 nt, from about 20 nt to about 70 nt, from about 20 nt to about 80 nt, from about 20 nt to about 90 nt, or from about 20 nt to about 100 nt. The length of the nucleic acid target region can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The length of the nucleic acid target region (e.g., the spacer sequence) can be at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides.
[0157] In some embodiments, the length of the nucleic acid targeting region (e.g., a spacer) guiding the nucleic acid is 20 nucleotides. In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid is 19 nucleotides. In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid is 18 nucleotides. In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid is 17 nucleotides. In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid is 16 nucleotides. In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid is 21 nucleotides. In some embodiments, the length of the nucleic acid targeting region guiding the nucleic acid is 22 nucleotides.
[0158] The length of the guide nucleic acid complementary to the nucleotide sequence of the target nucleic acid (target sequence) can be, for example, at least about 12 nucleotides (nt), at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt. The length of the guide nucleic acid complementary to the nucleotide sequence of the target nucleic acid (target sequence) can be from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 45 nt, from about 12 nt to about 40 nt, from about 12 nt to about 35 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 19 nt to about 20 nt, or from about 19 nt... From about 25nt, from about 19nt to about 30nt, from about 19nt to about 35nt, from about 19nt to about 40nt, from about 19nt to about 45nt, from about 19nt to about 50nt, from about 19nt to about 60nt, from about 20nt to about 25nt, from about 20nt to about 30nt, from about 20nt to about 35nt, from about 20nt to about 40nt, from about 20nt to about 45nt, from about 20nt to about 50nt, or from about 20nt to about 60nt.
[0159] Prototype spacer sequences targeting polynucleotides can be identified by identifying the prototype spacer adjacent motif (PAM) within the target region and selecting a region of desired size upstream or downstream of the PAM as the prototype spacer. The corresponding spacer sequence can be designed by determining the complementary sequence of the prototype spacer region.
[0160] Spacer subsequences can be identified using computer programs (e.g., machine-readable code). These programs can use variables such as predicted melting temperature, secondary structure formation and predicted annealing temperature, sequence identity, genomic background, chromatin accessibility, GC percentage, genomic frequency, methylation status, SNP presence, etc.
[0161] The percentage of complementarity between the nucleic acid targeting sequence (e.g., at least one spacer sequence of a guiding polynucleotide as disclosed herein) and the target nucleic acid (e.g., a prototype spacer sequence of one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percentage of complementarity between the nucleic acid targeting sequence and the target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 consecutive nucleotides.
[0162] The Cas protein-binding segment of a directing nucleic acid can comprise two complementary nucleotide segments (e.g., crRNA and tracrRNA). These two complementary nucleotide segments (e.g., crRNA and tracrRNA) can be covalently linked via intercalated nucleotides (e.g., linkers in the case of a single directing nucleic acid). The two complementary nucleotide segments (e.g., crRNA and tracrRNA) can hybridize to form a double-stranded RNA double helix or hairpin of the Cas protein-binding segment, thus producing a stem-loop structure. crRNA and tracrRNA can be covalently linked via the 3' end of crRNA and the 5' end of tracrRNA. Alternatively, tracrRNA and crRNA can be covalently linked via the 5' end of tracrRNA and the 3' end of crRNA.
[0163] The length of the Cas protein-binding region that guides nucleic acids can range from about 10 nucleotides to about 100 nucleotides, for example, from about 10 nucleotides (nt) to about 20 nt, from about 20 nt to about 30 nt, from about 30 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt. For example, the length of the Cas protein-binding region that guides nucleic acids can range from about 15 nucleotides (nt) to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt, or from about 15 nt to about 25 nt.
[0164] The length of the dsRNA double strand of the Cas protein-binding region that guides nucleic acid binding can range from about 6 base pairs (bp) to about 50 bp. For example, the length of the dsRNA double strand of the protein-binding region can range from about 6 bp to about 40 bp, from about 6 bp to about 30 bp, from about 6 bp to about 25 bp, from about 6 bp to about 20 bp, from about 6 bp to about 15 bp, from about 8 bp to about 40 bp, from about 8 bp to about 30 bp, from about 8 bp to about 25 bp, from about 8 bp to about 20 bp, or from about 8 bp to about 15 bp. For example, the length of the dsRNA double strand of the Cas protein-binding region can range from about 8 bp to about 10 bp, from about 10 bp to about 15 bp, from about 15 bp to about 18 bp, from about 18 bp to about 20 bp, from about 20 bp to about 25 bp, from about 25 bp to about 30 bp, from about 30 bp to about 35 bp, from about 35 bp to about 40 bp, or from about 40 bp to about 50 bp.
[0165] In some embodiments, the length of the dsRNA duplex of the Cas protein-binding region can be 36 base pairs. The percentage of complementarity between the nucleotide sequences of the dsRNA duplex that hybridizes to form the protein-binding region can be at least about 60%. For example, the percentage of complementarity between the nucleotide sequences of the dsRNA duplex that hybridizes to form the protein-binding region can be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some cases, the percentage of complementarity between the nucleotide sequences of the dsRNA duplex that hybridizes to form the protein-binding region is 100%.
[0166] The length of the adapter (e.g., the sequence linking crRNA and tracrRNA in a single guide nucleic acid) can be from about 3 nucleotides to about 100 nucleotides. For example, the length of the adapter can be from about 3 nucleotides (nt) to about 90 nt, from about 3 nucleotides (nt) to about 80 nt, from about 3 nucleotides (nt) to about 70 nt, from about 3 nucleotides (nt) to about 60 nt, from about 3 nucleotides (nt) to about 50 nt, from about 3 nucleotides (nt) to about 40 nt, from about 3 nucleotides (nt) to about 30 nt, from about 3 nucleotides (nt) to about 20 nt, or from about 3 nucleotides (nt) to about 10 nt. For example, the length of the adapter can be from about 3 nt to about 5 nt, from about 5 nt to about 10 nt, from about 10 nt to about 15 nt, from about 15 nt to about 20 nt, from about 20 nt to about 25 nt, from about 25 nt to about 30 nt, from about 30 nt to about 35 nt, from about 35 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt. In some embodiments, the adapter for the DNA-targeting RNA is 4 nt.
[0167] The guiding nucleic acid of the system disclosed herein may include modifications or sequences that provide additional desired characteristics (e.g., stability of the modified or regulated protein; subcellular targeting; tracking with fluorescent labeling; binding site of the protein or protein complex, etc.). Examples of such modifications include, for example, a 5' cap (7-methylguanylate cap (m7G)); a 3' polyadenylated tail (3' poly(A) tail); a riboswitch sequence (e.g., to allow regulation of protein and / or protein complex stability and / or accessibility); a stability control sequence; a sequence that forms a dsRNA double helix (hairpin); modifications or sequences that target RNA to subcellular locations (e.g., the nucleus, mitochondria, chloroplasts, etc.); modifications or sequences that provide tracking (e.g., direct conjugation to fluorescent molecules, conjugation to a portion that promotes fluorescence detection, a sequence that allows fluorescence detection, etc.); and modifications or sequences that provide binding sites for proteins (e.g., proteins that act on DNA, including transcription activators, transcription repressors, DNA methyltransferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof).
[0168] The guiding nucleic acid can contain one or more modifications (e.g., base modifications, backbone modifications) to provide a nucleic acid with novel or enhanced characteristics (e.g., improved stability). The guiding nucleic acid can contain a nucleic acid affinity tag. The nucleoside can be a base-sugar combination. The base moiety of the nucleotide can be a heterocyclic base. The two most common types of such heterocyclic bases are purines and pyrimidines. The nucleotide can be a nucleoside that further includes a phosphate ester group covalently linked to the sugar moiety of the nucleoside. For those nucleosides containing pentofuranosyl sugars, the phosphate group can be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming the guiding nucleic acid, the phosphate ester group can covalently link adjacent nucleosides to form a linear polymer compound. Furthermore, the corresponding ends of this linear polymer compound can be further linked to form a cyclic compound; however, a linear compound may be suitable. Additionally, the linear compound can have internal nucleotide base complementarity and thus can be folded in a manner that facilitates the formation of fully or partially double-stranded compounds. Moreover, within the guiding nucleic acid, the phosphate ester group can often be involved in forming the inter-nucleoside backbone of the guiding nucleic acid. The binding or backbone of the nucleic acid can be 3' to 5' phosphodiester binding.
[0169] The guiding nucleic acid may include a modified backbone and / or modified internucleotide bonds. The modified backbone may include those that retain phosphorus atoms in the backbone and those that do not have phosphorus atoms in the backbone.
[0170] Suitable guide nucleic acid backbones containing phosphorus atoms may include, for example, thiophosphates, chiral thiophosphates, dithiophosphates, phosphate triesters, aminoalkyl phosphate triesters, methyl and other alkylphosphonates (such as 3'-alkylphosphonene esters, 5'-alkylphosphonene esters, chiral phosphonates), phosphonites, aminophosphates (including 3'-aminoaminophosphates and aminoalkylaminophosphates), diaminophosphates, thiocarbonylaminophosphates, thiocarbonylalkylphosphonates, thiocarbonylalkylphosphonate triesters, selenophosphates, and boron phosphates having normal 3'-5' linkages, 2'-5' linkage analogues, and those having reverse polarity where one or more of the internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages. Suitable guide nucleic acids with reverse polarity may contain a single 3'-3' linkage (such as a single reverse nucleotide residue with a nucleobase deletion or a hydroxyl group at its position) at the nucleotide linkage closest to the 3' end. It can also include various salts (e.g., potassium chloride or sodium chloride), mixed salts, and free acid forms.
[0171] The guiding nucleic acid may contain one or more thiophosphate esters and / or heteroatomic nucleosides linked together, particularly -CH2-NH-O-CH2-, -CH2-N(CH3)-O-CH2- (methylene (methylimino) or MMI backbone), -CH2-ON(CH3)-CH2-, -CH2-N(CH3)-N(CH3)-CH2- and -ON(CH3)-CH2-CH2- (where the natural phosphodiester nucleotide linking is represented as -OP(=O)(OH)-O-CH2-).
[0172] The nucleic acid may contain a morpholine backbone structure. For example, the nucleic acid may contain a 6-membered morpholine ring instead of a ribose ring. In some of these examples, the internucleotide bond between diaminophosphate or other non-phosphodiester bonds replaces the phosphodiester bond.
[0173] Guide nucleic acids can contain polynucleotide backbones formed by short-chain alkyl or cycloalkyl nucleosides, mixed heteroatoms, or one or more short-chain heteroatoms or heterocyclic nucleosides. These can include those having: morpholino linkages (partially formed by the sugar portion of the nucleoside); siloxane backbones; sulfide, sulfoxide, and sulfone backbones; formylacetyl and thioformylacetyl backbones; methyleneformylacetyl and thioformylacetyl backbones; riboacetyl backbones; olefin-containing backbones; aminosulfonate backbones; methyleneimino and methylenehydrazine backbones; sulfonate and sulfonamide backbones; amide backbones; and other backbones having mixed N, O, S, and CH2 components.
[0174] Guided nucleic acids can include nucleic acid mimics. The term "mimic" can be intended to include polynucleotides in which only the furanose ring or both the furanose ring and the intermolecular bonds are replaced by non-furanose groups; substitution of only the furanose ring can also be referred to as a sugar substitute. The heterocyclic base moiety or modified heterocyclic base moiety can be retained for hybridization with a suitable target nucleic acid. One such nucleic acid can be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of the polynucleotide can be replaced by an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotide can be retained and directly or indirectly bound to the aza-nitrogen atom of the amide moiety of the backbone. The backbone in a PNA compound can contain two or more linked aminoethylglycine units to the amide-containing backbone of the PNA. The heterocyclic base moiety can be directly or indirectly bound to the aza-nitrogen atom of the amide moiety of the backbone.
[0175] Guide nucleic acids can contain linked morpholino units (morpholinonucleotides) with heterocyclic bases attached to the morpholino ring. Linking groups can connect the morpholino monomer units within the morpholinonucleotide. Nonionic morpholino-based oligomers can exhibit undesirable interactions with cellular proteins. Morpholino-based polynucleotides can serve as nonionic mimics of guide nucleic acids. Various compounds within the morpholino group can be linked using different linking groups. Another type of polynucleotide mimic can be called cyclohexenyl nucleic acid (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced by the cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used in the synthesis of oligomers using phosphoramidite chemistry. Incorporating CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements, exhibiting stability similar to natural complexes. Other modifications can include locked nucleic acids (LNAs), in which the 2'-hydroxy group is linked to the 4' carbon atom of the sugar ring, forming a 2'-C,4'-C-oxymethylene bond, thus forming a bicyclic sugar moiety. This bond can be a methylene (-CH2-) group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2. LNAs and LNA analogs can exhibit very high duplex thermal stability (Tm = +3 °C to +10 °C) with complementary nucleic acids, stability against 3'-exonuclease degradation, and good solubility.
[0176] Guide nucleic acids may contain one or more substituted sugar moieties. Suitable polynucleotides may contain sugar substituents selected from the following: OH; F; O-, S-, or N-alkyl; O-, S-, or N-alkenyl; O-, S-, or N-ynyl; or O-alkyl-O-alkyl, wherein the alkyl, alkenyl, and ynyl groups may be substituted or unsubstituted C1 to C2 groups. 10 Alkyl or C2 to C 10 Alkenyl and ynyl groups. O((CH2)) is particularly suitable. n O) m CH3, O(CH2) n OCH3, O(CH2) n NH2, O(CH2) n CH3, O(CH2) n ONH2 and O(CH2) n ON((CH2) n CH3)2, where n and m range from 1 to approximately 10. The sugar substituents can be selected from C1 to C2. 10Lower alkyl groups, substituted lower alkyl groups, alkenyl groups, alkynyl groups, aryl groups, O-alkaneyl groups or O-aryl groups, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2, heterocyclic alkyl groups, heterocyclic alkaneyl groups, aminoalkylamino groups, polyalkylamino groups, substituted silyl groups, RNA cleaving groups, reporter groups, intercalators, groups used to improve the pharmacokinetic properties of guiding nucleic acids, or groups used to improve the pharmacodynamic properties of guiding nucleic acids, and other substituents with similar properties. Suitable modifications may include 2'-methoxyethoxy (2'-O-CH2CH2OCH3, also known as 2'-O-(2-methoxyethyl) or 2'-MOE, i.e., alkoxyalkoxy group). Other suitable modifications may include 2'-dimethylaminoethoxy(O(CH2)2ON(CH3)2 group, also known as 2'-DMAOE) and 2'-dimethylaminoethoxyethoxy (also known as 2'-O-dimethyl-amino-ethoxy-ethyl or 2'-DMAEOE), 2'-O-CH2-O-CH2-N(CH3)2.
[0177] Other suitable sugar substituents can include methoxy (-O-CH3), aminopropoxy (-OCH2CH2NH2), allyl (-CH2-CH=CH2), -O-allyl (-O--CH2—CH=CH2), and fluorine (F). The 2'-sugar substituent can be located at the arabinose (top) or ribose (bottom) position. A suitable 2'-arabinose modification is 2'-F. Similar modifications can also be made at other positions on the oligomer, particularly at the 3' terminal nucleotide or at the 3' position of the sugar and the 5' position of the 5' terminal nucleotide in the 2'-5' linked nucleotide. The oligomer can also have sugar mimics, such as a cyclobutyl moiety replacing the pentofuranosyl sugar.
[0178] Nucleic acids can also contain nucleobase (or “base”) modifications or substitutions. As used herein, “unmodified” or “natural” nucleobases can include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases can include other synthetic and natural nucleobases such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3)uracil and cytosine and other alkynyl derivatives of pyrimidine bases, 6-azo Uracil, cytosine, and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halogenated, 8-amino, 8-hydrothio, 8-thioalkyl, 8-hydroxy, and other 8-substituted adenine and guanine, 5-halogenated, especially 5-bromo, 5-trifluoromethyl, and other 5-substituted uracil and cytosine, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-nitroguanine and 8-nitroadenine, 7-denitroguanine and 7-denitroadenine, and 3-denitroguanine and 3-denitroadenine. The modified nucleobases may include tricyclic pyrimidines such as phenoxazincytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clasts such as substituted phenoxazincytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazolecytidine (2H-pyrimido(4,5-b)indol-2-one), and pyridoindolcytidine (H-pyrido(3',2':4,5)pyrrolo(2,3-d)pyrimido-2-one).
[0179] The heterocyclic base moiety may include those in which the purine or pyrimidine base is replaced by other heterocycles such as 7-deadenine, 7-deadenine, 2-aminopyridine, and 2-pyridone. Nucleobases can be used to increase the binding affinity of polynucleotide compounds. These nucleobases may include 5-substituted pyrimidines, 6-azapyrimidines, and N-2, N-6, and O-6 substituted purines, including 2-aminopropyladenine, 5-propynyluracil, and 5-propynylcytosine. 5-methylcytosine substitution can increase the stability of the nucleic acid duplex by 0.6°C–1.2°C and can be a suitable base substitution (e.g., when combined with 2'-O-methoxyethyl sugar modification).
[0180] Modification of the guiding nucleic acid can include one or more moieties or conjugates that are chemically linked to the guiding nucleic acid to enhance its activity, cellular distribution, or cellular uptake. These moieties or conjugates can include conjugate groups covalently bonded to functional groups such as primary or secondary hydroxyl groups. Conjugate groups can include, but are not limited to, intercalators, reporter molecules, polyamines, polyamides, polyethylene glycols, polyethers, groups that enhance the pharmacokinetic properties of oligomers, and groups that can enhance the pharmacokinetic properties of oligomers. Conjugate groups can include, but are not limited to, cholesterol, lipids, phospholipids, biotin, phenazine, folic acid, phenanthridine, anthraquinones, acridine, fluorescein, rhodamine, coumarin, and dyes. Groups that enhance pharmacokinetic properties include those that improve uptake, enhance degradation resistance, and / or strengthen sequence-specific hybridization with the target nucleic acid. Groups that can enhance pharmacokinetic properties include those that improve the uptake, distribution, metabolism, or excretion of nucleic acids. The conjugated portion may include, but is not limited to, lipid portions such as cholesterol portions, bile acids, thioethers (e.g., hexyl-S-triphenylmethylthiol), thiocholesterol, aliphatic chains (e.g., dodecyl glycol or undecyl residues), phospholipids (e.g., di-hexadecyl-racemic-glycerol or triethylammonium 1,2-di-O-hexadecyl-racemic-glycero-3-H-phosphonate), polyamines or polyethylene glycol chains or adamantaneacetic acid, palmityl portions or octadecylamine or hexylamino-carbonyl-hydroxycholesterol portions.
[0181] In some embodiments, at least one guide RNA polynucleotide of the system or method provided herein can bind to at least a portion of a genome (e.g., a plant genome) or a gene (e.g., a plant gene). In some cases, at least one guide RNA polynucleotide is capable of complexing with a site-directed nuclease to direct the site-directed nuclease to target a portion of a target nucleic acid (e.g., a site in the genome or gene).
[0182] In some embodiments, the system described herein includes at least one guide RNA polynucleotide capable of forming a complex with the site-directed nuclease portion of the fusion protein of the system. In some embodiments, the system described herein includes at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides capable of forming a complex with the site-directed nuclease portion of the fusion protein of the system.
[0183] In some embodiments, the guiding nucleic acid comprises a nucleotide sequence having at least 70% (e.g., at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100%) identity with SEQ ID NO:17 or 19.
[0184] This document also provides kits that include components of the systems described in this disclosure. In some embodiments, the kits include one or more of the fusion proteins and / or polynucleotides described herein. VI. Methods
[0185] On the other hand, this document provides methods for editing one or more nucleic acids using the fusion proteins and / or systems described herein. In some embodiments, these methods include contacting a nucleic acid containing a fusion protein binding site (i.e., the nucleic acid to be edited) with at least one fusion protein as described herein, wherein contacting the nucleic acid with at least one fusion protein results in editing of the nucleic acid. The nucleic acid (i.e., the nucleic acid to be edited) can be any suitable nucleic acid. In some embodiments, the nucleic acid is part of a chromosome. In some embodiments, the nucleic acid is part of a genome (e.g., a plant genome).
[0186] As described herein and illustrated in the examples below, the methods presented herein can increase the frequency of one or more desired nucleic acid editing outcomes (e.g., fragment replacement via HDR).
[0187] In some embodiments, the nucleic acid to be edited by this method comprises a target region. As used herein, "target region" refers to a portion of the nucleic acid targeted for editing. For example, the target region may be a portion of the gene to be edited. In some embodiments, at least a portion of the target region is replaced by at least a portion of a donor polynucleotide. In some embodiments, the target region comprises at least one (e.g., two, three, four, five, six, seven, eight, nine, ten or more) nuclease cleavage sites. In some embodiments, the nucleic acid comprises at least one binding site. In some embodiments, the target region is flanked by a nuclease cleavage site. In some embodiments, the nucleic acid comprises a first binding site adjacent to the 5' end of the target region and a second binding site adjacent to the 3' end of the target region.
[0188] In some embodiments, as detailed below, the nucleic acid to be edited includes a first binding site and a second binding site. In some embodiments, the first binding site and the second binding site are different sequences, and the method includes providing two different fusion proteins, one that binds to the first binding site and one that binds to the second binding site. In some embodiments, the first binding site and the second binding site are the same sequence, and the method includes providing a fusion protein capable of binding both the first binding site and the second binding site.
[0189] In some embodiments, the methods described herein include providing a donor polynucleotide. In some embodiments, the donor polynucleotide comprises a left homologous arm (i.e., a homologous arm complementary to a sequence upstream of the target region) and a right homologous arm (i.e., a homologous arm complementary to a sequence downstream of the target region). In such embodiments, the target region of the nucleic acid to be edited is flanked by a homologous arm, and the target region contains at least one fusion protein binding site (i.e., at least one cleavage site). Figure 2 An exemplary embodiment is shown.
[0190] In some embodiments of the methods provided herein, the recruitment domain of the fusion protein specifically binds to a recruitment sequence in the donor polynucleotide. Without being bound by any particular theory, the fusion protein, binding to a binding site in the nucleic acid target region (i.e., by binding of a site-directed nuclease) and the recruitment sequence in the donor polynucleotide (i.e., by specific binding of the recruitment domain), tethers the donor polynucleotide to a closely proximate cleavage site formed by an SDN. In some embodiments, this close proximity increases the likelihood that the cleaved nucleic acid will be repaired in a manner that results in the incorporation of at least a portion of the donor polynucleotide into the target region. In some embodiments, the donor polynucleotide contains at least one homologous arm, and repair occurs via HDR. In some embodiments, repair occurs via NHEJ or MMEJ, wherein the end of the donor polynucleotide is attached to the end of the cleaved nucleic acid in the target region.
[0191] In some embodiments of the methods provided herein, for example, in the examples herein, at least one site-directed nuclease of the fusion protein comprises a CRISPR-associated nuclease. In such embodiments, the method may further include providing guide RNA for targeting the fusion protein to a binding site. In some embodiments, the method includes providing at least one first guide RNA and at least one second guide RNA. In some embodiments, the at least one first guide RNA comprises a nucleotide sequence complementary to a first binding site of the nucleic acid to be edited. In some embodiments, the at least one second guide RNA comprises a nucleotide sequence complementary to a second binding site of the nucleic acid to be edited.
[0192] The methods described herein include providing at least one fusion protein, a nucleic acid to be edited, and a donor polynucleotide, and may also include providing at least one guide RNA. These various components can be provided using any suitable technique. For example, providing a fusion protein may include introducing the fusion protein into a cell or introducing a recombinant nucleic acid, construct, or vector encoding the fusion protein into a cell. Similarly, gRNA can be provided by introducing the gRNA itself or a nucleic acid sequence encoding the gRNA. In some embodiments, the fusion protein and gRNA may be encoded by the same DNA construct or vector. Example 1. Cro-Cas9 fusion induces allelic substitution at the ZmALS2 locus.
[0193] This example demonstrates that using donor DNA containing a Cro-binding sequence, the Cro-Cas9 fusion protein improves the efficiency of homology-directed repair (HDR)-mediated allele substitution. In this example, as... Figure 3 As shown in the small image above, the donor DNA contains a Cro operator site (O R 3) The Cro protein, acting as a homodimer, binds to the operator site. The Cro protein is tethered to Cas9 via the XTEN linker.
[0194] A schematic design of the Cro-Cas9 fusion is shown. Figure 4 In this study, a single-chain dimer of the N15 phage Cro protein (two monomers linked by a 15-amino acid linker: SEQ ID NO: 15) was fused to the N-terminus of Streptococcus pyogenes Cas9 (SpCas9) via a 32-amino acid linker (SEQ ID NO: 16). The N-terminus of the fusion protein was labeled with a 3X FLAG peptide and then with an SV40 nuclear localization signal (NLS), while the C-terminus of the fusion protein was labeled with another NLS.
[0195] Svitashev et al. (2015, Plant Physiol.) reported Cas9-induced allelic substitution at the ZmALS2 locus (a homolog of maize acetolactate synthase on chromosome 5). In this example, the same target site (gRNA represented by SEQ ID NO:17) was selected to evaluate the efficacy of the Cro-Cas9 fusion in improving the efficiency of allelic substitution.
[0196] DNA constructs were constructed to express the Cro-Cas9 fusion protein and gRNA in corn cells. The corn-codon-optimized coding sequence of the fusion protein was operatively linked to the sugarcane ubiquitin-4 promoter and the Agrobacterium tumefaciens carmine synthase terminator. A single guide RNA (sgRNA) targeting the ZmALS2 locus was operatively linked to the rice (Oryza sativa) U3 promoter, and a nine-base sequence of thymine followed the sgRNA spacer to terminate transcription. A control construct lacking the Cro domain was created, but otherwise identical to the Cro-Cas9 expression vector.
[0197] The donor DNA serving as the repair template for HDR was based on a 657-bp homologous fragment derived from the coding region of ZmALS2, beginning at codon 4 (GCT for alanine) and ending at codon 222 (CGG for arginine). The donor DNA sequence is represented by SEQ ID NO:20. Eight single nucleotide substitutions were introduced into the homologous fragment around the Cas9 cleavage site. Of these, only two conferred amino acid changes. The G-to-C substitution at codon 111 changed methionine to isoleucine, simultaneously disrupting the TGG PAM of the gRNA, protecting the donor DNA from Cas9 cleavage. The C-to-T substitution at codon 119 changed proline to serine, conferring resistance to chlorsulfuron-methyl herbicide in the gene product. The nucleotide substitutions also generated three restriction sites as molecular barcodes, which aid in the detection of alleles involved in HDR editing. A 20-bp phage N15O will be specifically recognized and bound by the Cro protein. R 3. An operon sequence (SEQ ID NO: 18) was added to each flanking segment containing the substituted homologous fragment to prepare donor DNA. The donor DNA was cloned into a high-copy cloning vector flanked by PmeI and AscI restriction sites for excision.
[0198] Linearized vector (0.2 pmol) and double-stranded donor DNA (2 pmol) were co-delivered into corn embryogenic callus cells using a gene gun. Transgenic callus was selected on PMI selective medium and then regenerated into transgenic plants using standard tissue culture procedures. Regenerated plants were sampled for DNA extraction, and a TaqMan assay was designed to distinguish between HDR-substituted resistance alleles and wild-type alleles. A pair of primers outside the homologous region was designed to amplify a 1.5 kb fragment covering the entire homology region via PCR, and the amplicon was subjected to restriction fragment length polymorphism (RFLP) analysis and Sanger sequencing for sequence verification.
[0199] The results are summarized in Table 1. In each of the three allelic substitution plants generated by the Cro-Cas9 fusion construct, one of the two alleles underwent HDR-mediated allelic substitution, with all eight nucleotide substitutions installed, while the other allele carried an insertion / deletion (indel) mutation at the Cas9 cleavage site. The control vector generated only one plant in which one of the two alleles was partially substituted with a subset of the installed eight substitutions. Table 1. Allelic substitution efficiency at the ZmALS2 locus. Example 2. Cro-Cas12a fusion induces allelic substitution at the ZmGL2 locus.
[0200] This example demonstrates that using donor DNA containing a Cro-binding sequence, the Cro-Cas12a fusion protein improves the efficiency of homology-directed repair (HDR)-mediated allele substitution. In this example, as... Figure 3 (See the smaller image below) The donor DNA contains a Cro operator site (O R 3) The Cro protein, acting as a homodimer, binds to the operator site. The Cro protein is tethered to Cas12a via the XTEN linker.
[0201] A schematic design of the Cro-Cas12a fusion is shown. Figure 5 In this study, the same single-chain N15 Cro dimer as described in Example 1 was fused to the N-terminus of the ND2006Cas12a (LbCas12a)D156R variant of Trichophyton spp. via a 30-amino acid (GGGGS)6 linker (SEQ ID NO:26). The fusion protein was flanked by one SV40 NLS at the N-terminus and two SV40 NLS at the C-terminus.
[0202] K. Lee et al., "Activities and specificities of CRISPR / Cas9 and Cas12a nucleases for targeted mutagenesis in maize," Plant Biotech.J., 17:362–372 (2019), reported that LbCas12a can effectively induce DNA double-strand breaks at the Glossy2 (ZmGL2) locus in maize. In this example, the same target site (gRNA is SEQ ID NO:19) was selected to evaluate the efficacy of the Cro-LbCas12a fusion in improving allele substitution efficiency.
[0203] A DNA construct was constructed to express the Cro-LbCas12a fusion protein and gRNA in corn cells. The corn-codon-optimized coding sequence of the fusion protein was operatively linked to the sucrose ubiquitin-4 promoter and the Agrobacterium tumefaciens carmine synthase terminator. A mature crRNA targeting the ZmGL2 locus (side-joined with a hammerhead (HH) structure and a hepatitis D virus (HDV) ribozyme for processing) was operatively linked to another copy of the sucrose ubiquitin-4 promoter and another copy of the carmine synthase terminator. An identical construct lacking the Cro domain was constructed as a control.
[0204] The donor DNA used as the repair template for HDR was based on an 859-bp homology fragment taken from the coding region of ZmGL2. The donor DNA sequence is represented by SEQ ID NO:21. Homology to the left (upstream) of the cleavage site is located in the first intron, and homology to the right (downstream) of the cleavage site is located in the second exon. Introducing eight consecutive nucleotide substitutions (ACAAACTT to TAGTGACC) into the homology fragment in the middle of the gRNA prototype spacer produces two consecutive early stop codons and protects the donor from Cas12a cleavage. The same N15 O as described above... R 3. Sequences were added to each flanking segment containing the substituted homologous fragment to prepare donor DNA. The donor DNA was cloned into a cloning vector and could be prepared in linear form by PCR using the vector as a template.
[0205] A mixture of 0.15 pmol of linearized vector and 2.25 pmol of double-stranded donor DNA was co-delivered into corn embryogenic callus cells using a gene gun. Transgenic callus was selected on PMI selective medium and then regenerated into transgenic plants using a tissue culture program. Regenerated plants were sampled for DNA extraction, and a TaqMan assay was designed to distinguish between functional knockout alleles and wild-type alleles with HDR substitution. A pair of primers outside the homologous region was designed to amplify a 1.4 kb fragment covering the entire homology region via PCR, and the amplicon was Sanger sequenced to verify the expected mutation and homologous genome junction.
[0206] The results are summarized in Table 2. In all plants, eight nucleotide mutations were inserted into the ZmGL2 locus. In those plants, six plants generated from the Cro-LbCas12a fusion and three plants generated from the LbCas12a control underwent perfect HDR at both genomic homologous junctions in the replaced alleles. In contrast, two plants generated from the Cro-LbCas12a fusion and seven plants generated from the LbCas12a control failed to undergo perfect HDR at at least one genomic homologous junction in the replaced alleles. Notably, one plant generated from Cro-LbCas12a underwent perfect HDR replacement in both alleles. Table 2. Allelic substitution efficiency at the ZmGL2 locus. *A plant with a homozygous perfect HDR displacement. Reference sequence list SEQ ID NO:1 N15 Cro monomeric amino acid sequence MKPEELVRHFGDVEKAAVGVGVTPGAVYQWLQAGEIPPLRQSDIEVRTAYKLKSDFTSQRMGKEGHNSGTK SEQ ID NO:2 λCro monomer amino acid sequence MEQRITLKDYAMRFGQTKTAKDLGVYQSAINKAIHAGRKIFLTINADGSVYAEEVKPFPSNKKTTA SEQ ID NO:3 P22 Cro monomeric amino acid sequence MYKKDVIDHFGTQRAVAKALGISDAAVSQWKEVIPEKDAYRLEIVTAGALKYQENAYRQAA SEQ ID NO:4 434Cro monomer amino acid sequence MQTLSERLKKRRIALKMTQTELATKAGVKQQSIQLIEAGVTKRPRFLFEIAMALNCDPVWLQYGTKRGKAA SEQ ID NO:5 N15 Cro O R 3 TTATAGCTGGCTATAA SEQ ID NO:6 λCro O R 3 (Natural) TATCACCGCAAGGGATA SEQ ID NO:7 λCro O R 3 (Synthesis) TATCACCGCGGGTGATA SEQ ID NO:8 P22 Cro O R 3 AGTTAAGTCATCTTAAAT SEQ ID NO:9 434Cro O R 3 (Natural) ACAAGAAAAACTGT SEQ ID NO:10 434 Cro O R 3 (Synthesis) ACAATATATATTGT SEQ ID NO:11 Cro-Cas9 amino acid sequence SEQ ID NO:12 Cro-Cas9 nucleotide sequence SEQ ID NO:13 Cro-Cas12a amino acid sequence SEQ ID NO:14 Cro-Cas12a nucleotide sequence SEQ ID NO:15 15 amino acid linkers. GGGSGGGSGGGSGGG SEQ ID NO:16 32 amino acid linkers. SGGSSGGSSGSETPGTSESATPESSGGSSGGS SEQ ID NO:17 ZmALS2 (corn acetyllactate synthase homolog) gRNA GCTGCTCGATTCCGTCCCCA SEQ ID NO:18 20bp bacteriophage N15 OR3 operon. CTTTATAGCTGGCTATAATT SEQ ID NO:19 Corn Glossy2(ZmGL2)gRNA GTCACAGATCACAAACTTCAAATG SEQ ID NO:20 The donor DNA used as a template for HDR repair is based on a 657-bp homology fragment taken from the coding region of ZmALS2, starting at the fourth codon (GCT stands for alanine) and ending at the 222nd codon (CGG stands for arginine). GCTCCCCCGGCCACCCCGCTCCGGCCGTGGGGCCCCACCGATCCCCGCAAGGGCGCCGACATCCTCGTCGAGTCCCTCGAGCGCTGCGGCGTCCGCGACGTCTTCGCCTACCCCGGCGGCGCGTCCATGGAGATCCACCAGGCACTCACCCGCTCCCCCGTCATCGCCAACCACCTCTTCCGCCACGAGCAAGGGGAGGCCTTTGCGGCCTCCGGCTACGCGCGCTCCTCGGGCCGCGTCGGCGTCTGCATCGCCACCTCCGGCCCCGGCGCCACCAACCTTGTCTCCGCGCTCGCCGACGCGCTGCTCGATTCTGTGCCGATCGTCGCTATCACCGGTCAGGTGTCGCGACGCATGATTGGCACCGACGCCTTCCAGGAGACGCCCATCGTCGAGGTCACCCGCTCCATCACCAAGCACAACTACCTGGTCCTCGACGTCGACGACATCCCCCGCGTCGTGCAGGAGGCTTTCTTCCTCGCCTCCTCTGGTCGACCGGGGCCGGTGCTTGTCGACATCCCCAAGGACATCCAGCAGCAGATGGCGGTGCCTGTCTGGGACAAGCCCATGAGTCTGCCTGGGTACATTGCGCGCCTTCCCAAGCCCCCTGCGACTGAGTTGCTTGAGCAGGTGCTGCGTCTTGTTGGTGAATCCCGG SEQ ID NO:21 The donor DNA used as a repair template for HDR is based on an 859-bp homology fragment taken from the ZmGL2 coding region. TACCCACATAGCACGTACGGAGTATGAAGTCAATCAAAGGTGCAAATCACGGCTTTGTATACTAGCTAGACTAGTCTATTACAGTACAGACAATTACTGCAAACTAATTGTGCTGGTCGATTACTTCTGTTACTAAGGGCATGTACAACTTAGATACAACTGCACGGTACTCCAAGTATAAGACACAACTAAAACACAATATAATACAGTGGTC ATGTCTAAAACATGTGTCTTACCATATTTATTGTACCAATCAGGGCATTCAATAAATTAAAGTGACCAATCAGATAGTCTCATGTCTCGAATATAGAGCTAAGACACCGTGTCTTCGTCAAAATACATGTCTTGAGATTTTTTACATTCACCCTCCTAGACACACTCTAAGACACAACTTAAGACACCCCACGGTACATGCCCTAACTACGTACT CCCTCCGTCCTTTTTTATTTATCGTTTCTTTGGTCACAGATCTAGTGACCCAAATGCGGTGGGCTGGCTGGGGTTCAGCTGGGCGCACCTCATCGGCGACATCCCGTCGGCCGCCACCTGCTTCAACAAGTGGGCGCAGATCCTCAGCGGCAAGAAGCCGGAAGCCACCGTCCTCACCCCGCCGAACCAGCCGCTGCAGGGCCAGTCCCCCGC GGCGCCGCGCTCCGTCAAGCAGGTCGGGCCCATGGAGGACCTCTGGCTGGTCCCCGCGGGCCGCGACATGGCGTGCTACTCCTTCCACGTCAGCGACGCGGTGCTCAAGAAGCTCCACCAGCAGCAGAATGGGCGCCAGGACGCCGCCGCTGGCACCTTCGAGCTCGTGTCGGCGCTGGTGTGGCAGGCGGTGGCCAAGATCAGGGGCGACGTGG
[0207] All patents, patent publications, patent applications, journal articles, books, technical references, etc., discussed in this disclosure are incorporated herein by reference in their entirety for all purposes.
[0208] It should be understood that the figures and descriptions in this disclosure have been simplified to illustrate and clearly understand the elements relevant to this disclosure. It should be understood that the figures are for illustrative purposes and not presented as structural diagrams. Omitted details and modifications or alternative embodiments are within the knowledge of those skilled in the art.
[0209] It is understood that, in certain aspects of this disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component to provide an element or structure or perform a given function or one or more functions. Such substitution is considered to be within the scope of this disclosure unless it would render certain embodiments of this disclosure inoperable.
[0210] The examples presented herein are intended to illustrate potential and specific implementations of this disclosure. It will be understood that these examples are primarily intended for the purpose of explaining this disclosure to those skilled in the art. Variations may be made to these figures or the operations described herein without departing from the spirit of this disclosure. For example, in some cases, method steps or operations may be performed or carried out in a different order, or operations may be added, deleted, or modified.
[0211] Where a numerical range is provided, it should be understood that each intermediate value between the upper and lower limits of the range (the smallest decimal place to the units digit of the lower limit, unless the context explicitly states otherwise) is also specifically disclosed. This covers any smaller range between any stated or non-statement intermediate value in the stated range and any other stated or intermediate value in the stated range. The upper and lower limits of these smaller ranges may be independently included or excluded from the range, and each range in which no one or both limits are included is also covered by this technique, depending on any limits specifically excluded from the stated range. Where the stated range includes one or both limits, it also includes ranges that exclude one or both of those included limits.
[0212] In the foregoing description, numerous specific details have been set forth to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention described herein can be practiced without one or more of these specific details. In other instances, features and procedures well-known to those skilled in the art have not been described to avoid obscuring the invention. Embodiments of this disclosure have been described for illustrative and not restrictive purposes. While the invention has been described primarily with reference to specific embodiments, other embodiments are contemplated that will become apparent to those skilled in the art upon reading this disclosure, and such embodiments are intended to be included within the methods of the invention. Therefore, this disclosure is not limited to the embodiments depicted above or in the accompanying drawings, and various embodiments and modifications may be made without departing from the scope of the following claims.
Claims
1. A fusion protein comprising a site-directed nuclease fused to a recruitment domain containing a site-specific DNA-binding domain.
2. The fusion protein of claim 1, wherein the site-directed nuclease comprises a CRISPR-associated nuclease.
3. The fusion protein of claim 2, wherein the CRISPR-associated nuclease is selected from the group consisting of: Cas5, Cas6, Cas7, Cas8, Cas9, Cas12a, Cas12b, Cas12i, Cas12j, Cas12L, Cas12e, Cas12c, Cas12d, Cas12g, Cas12h, TnpB, Cas13a, Cas13b, Cas14, and their nicking enzymes or inactivated forms.
4. The fusion protein of claim 3, wherein the CRISPR-associated nuclease is a Cas9 enzyme.
5. The fusion protein of claim 3, wherein the CRISPR-associated nuclease is a Cas12a enzyme.
6. The fusion protein according to any one of claims 1 to 5, wherein the recruitment domain is a Cro repressor family protein.
7. The fusion protein of claim 6, wherein the Cro repressor family protein comprises N15 Cro, λCro, P22Cro, 434Cro, or any combination thereof.
8. The fusion protein according to any one of claims 1 to 7, wherein the recruitment domain comprises an amino acid sequence having at least 90% identity with any one of SEQ ID NO: 1-4.
9. The fusion protein of any one of claims 1 to 8, wherein the recruitment domain comprises a dimerization domain.
10. The fusion protein of any one of claims 1 to 9, wherein the fusion protein comprises a linker located between the site-directed nuclease and the recruitment domain.
11. The fusion protein of claim 10, wherein the adapter comprises any one of SEQ ID NO: 6, 7, 15 or 16.
12. The fusion protein according to any one of claims 1 to 11, wherein the fusion protein contains a nuclear localization signal.
13. The fusion protein according to any one of claims 1 to 12, wherein the fusion protein comprises an amino acid sequence having at least 90% identity with SEQ ID NO: 11 or 13.
14. A recombinant nucleic acid encoding a fusion protein as described in any one of claims 1 to 13.
15. A DNA construct comprising a promoter operatively linked to the recombinant nucleic acid of claim 14.
16. The DNA construct of claim 15, wherein the promoter comprises at least one of an inducible promoter, a constitutive promoter, an egg cell-specific promoter, a pollen-specific promoter, or a apical meristem-specific promoter.
17. The DNA construct of claim 15 or 16, wherein the promoter is a ubiquitin 4 promoter, actin promoter, tubulin promoter, MADS box promoter, or plant virus promoter.
18. A vector comprising the recombinant nucleic acid as described in claim 14 or the DNA construct as described in any one of claims 15 to 17.
19. A cell comprising the recombinant nucleic acid as claimed in claim 14, the DNA construct as claimed in any one of claims 15 to 17, or the vector as claimed in claim 18.
20. The cell of claim 19, wherein the cell is a plant cell.
21. The cell of claim 20, wherein the plant cell is a corn plant cell, a soybean plant cell, a rice plant cell, a wheat plant cell, or a sunflower plant cell.
22. A method for editing nucleic acids, the method comprising: a. To provide at least one fusion protein as described in any one of claims 1 to 13; b. Provide the nucleic acid, wherein the nucleic acid contains a first binding site and a target region containing a portion of the nucleic acid, wherein the first binding site is within or adjacent to the target region; c. Providing a donor polynucleotide comprising a donor nucleotide region and at least one recruitment sequence that specifically binds to the recruitment domain of the at least one fusion protein; and d. Contact the nucleic acid and the donor polynucleotide with the at least one fusion protein, wherein the at least one fusion protein specifically binds to the first binding site of the nucleic acid and specifically binds to the recruitment sequence of the donor polynucleotide. This results in editing of the target region of the nucleic acid.
23. The method of claim 22, wherein the first binding site is adjacent to the 5′ end or 3′ end of the target region.
24. The method of claim 22 or 23, wherein the nucleic acid further comprises a second binding site, wherein the second binding site is within or adjacent to the target region, and wherein the at least one fusion protein specifically binds to the first binding site and the second binding site of the nucleic acid.
25. The method of claim 24, wherein the second binding site is adjacent to the 5′ end or 3′ end of the target region.
26. The method of any one of claims 22 to 25, wherein the recruitment domain of the fusion protein comprises a Cro repressor family protein, and the at least one recruitment sequence comprises Cro O R 3. Manipulator subsequences.
27. The method of claim 26, wherein the Cro O R 3. The operon sequence contains N15O R 3. Operator sequence (optionally, SEQ ID NO:18), λO R 3. Operator sequence, P22 O R 3 operon sequences, 434O R 3. Operator sequences or combinations thereof.
28. The method of any one of claims 22 to 27, wherein the donor polynucleotide comprises at least one homologous arm, wherein the at least one homologous arm comprises a nucleotide sequence complementary to a portion of the target region of the nucleic acid.
29. The method of any one of claims 22 to 28, wherein the donor polynucleotide comprises at least two recruitment sequences.
30. The method of claim 29, wherein the donor polynucleotide comprises a first recruitment sequence adjacent to the 5′ end of the donor nucleotide region and a second recruitment sequence adjacent to the 3′ end of the donor nucleotide region.
31. The method of claim 29 or 30, wherein the at least two recruited sequences are not within the donor nucleotide region.
32. The method of any one of claims 22 to 31, wherein the site-directed nuclease of the at least one fusion protein comprises a CRISPR-associated nuclease, and the method further comprises providing at least one guide RNA, wherein the at least one guide RNA comprises a nucleotide sequence complementary to the first binding site and / or the second binding site of the nucleic acid.
33. The method of any one of claims 22 to 32, wherein the editing of the target region of the nucleic acid is to replace at least a portion of the target region with at least a portion of the donor polynucleotide.
Citation Information
Patent Citations
Simultaneous gene editing and haploid induction
US10519456B2
Methods and compositions for using zinc finger endonucleases to enhance homologous recombination
US20030232410A1
Genomic editing in zebrafish using zinc finger nucleases
US20090203140A1
Methods and compositions for RNA-directed target DNA modification and for RNA-directed modulation of transcription
US20140068797A1
Optimized protein linkers and methods of use
US20210017506A1