Methods and compositions for generating dominant alleles using genome editing
Targeted editing techniques enable precise generation of dominant-negative and dominant-positive alleles by inducing specific genomic modifications, addressing the low frequency and specificity issues of natural mutagenesis methods, and facilitating controlled gene expression.
Patent Information
- Application Number
- JP2025143644
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-08-14
- Filing Date
- 2025-08-29
- Publication Date
- 2026-01-09
AI Technical Summary
Natural and random mutagenesis techniques for generating dominant alleles, such as dominant-negative or dominant-positive alleles, occur at low frequency and are difficult to target in specific genes of interest.
Utilizing targeted editing techniques, including RNA-guided nucleases and donor sequences, to induce specific genomic modifications such as inversions, deletions, and insertions to generate dominant-negative or dominant-positive alleles, thereby controlling gene expression.
This approach allows for precise and efficient generation of dominant alleles, enabling reduced or enhanced gene expression, and provides a method to create desired genetic modifications in various organisms.
Smart Images

Figure 2026003132000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No. 62 / 854,142, filed May 29, 2019, U.S. Provisional Application No. 62 / 886,726, filed August 14, 2019, and U.S. Provisional Application No. 62 / 886,732, filed August 14, 2019, all of which are incorporated herein by reference in their entireties.
[0002] Field The present disclosure relates to methods and compositions for generating dominant alleles through targeted editing of the genome.
[0003] Sequence Listing Reference The Sequence Listing contained in the file entitled "P34497WO00_SL.TXT", 172,842 bytes (measured in MS-Windows®), created on May 28, 2020, has been submitted electronically herewith and is incorporated by reference in its entirety. [Background technology]
[0004] background A dominant allele is an allele that suppresses the contribution of a second allele at the same locus. A dominant allele can be a dominant negative allele or a dominant positive allele. A dominant negative allele, or antimorph, is an allele that acts opposite to the normal allele function. For example, a dominant negative allele often prevents the normal function of an allele in the heterozygous or homozygous state. A dominant positive allele can increase normal gene function (e.g., a hypermorph) and / or confer a broader or new function to a gene (e.g., a neomorph).
[0005] Natural and random mutagenesis techniques (e.g., ethyl methyl sulfonate and T-DNA insertion) have been used to generate mutations in various cell types. However, dominant mutations occur at low frequency and are difficult to obtain in a given gene of interest. Therefore, methods and compositions for selectively editing genomes to create dominant-negative or dominant-positive alleles would be beneficial. Summary of the Invention
[0006] overview In one aspect, the disclosure provides a method for generating a dominant-negative allele of a gene in a cell, comprising using targeted editing techniques to invert a portion of the gene to generate an antisense RNA transcript that can induce suppression of the unmodified allele.
[0007] In one aspect, the present disclosure provides a method of generating a dominant-negative allele of a gene in a cell, comprising deleting a portion of a chromosome between a first gene region and a second gene region using targeted editing technology, wherein an antisense RNA transcript of the first gene region is generated after deletion of the portion of the chromosome.
[0008] In one aspect, the disclosure provides a method for generating a dominant negative allele of a gene in one or more cells, the method comprising: (a) inducing a first double-stranded break and a second double-stranded break flanking a targeted region of the gene; (b) identifying one or more cells containing an inversion of the targeted region of the gene, wherein the inversion results in production of an antisense RNA transcript from the targeted region; and (c) selecting one or more cells containing an inversion of the targeted region of the gene.
[0009] In one aspect, the disclosure provides a method for reducing expression of a protein in a cell, the method comprising: (a) inducing a first double-stranded break and a second double-stranded break flanking a targeted region of a chromosome; and (b) identifying one or more cells that contain an inversion in the targeted region of the chromosome, wherein expression of the protein is reduced compared to a control cell that does not contain an inversion within the targeted region.
[0010] In one aspect, the disclosure provides a method comprising: (a) identifying a chromosomal region comprising a first gene region comprising a first promoter and a first coding region and a second gene region comprising a second promoter and a second coding region, wherein the first coding region and the second coding region are separated by an intervening region and the first promoter and the second promoter are positioned in opposite orientations; (b) inducing first and second double-stranded breaks flanking the targeted region; (c) identifying one or more cells comprising a deletion of the targeted region of the chromosome; and (d) selecting one or more cells comprising a deletion of the targeted region of the chromosome.
[0011] In one aspect, the disclosure provides a method of reducing expression of a gene in at least one cell, the method comprising: (a) inducing a double-stranded break at a target site of the gene using targeted editing technology; and (b) inserting a donor sequence at the double-stranded break, wherein the donor sequence comprises a tissue-specific or tissue-preferred promoter, and the donor sequence is inserted at the target site such that the tissue-specific or tissue-preferred promoter is in a reverse orientation compared to the gene; and (b) identifying at least one cell containing an insertion of the reversed donor sequence, wherein expression of the gene is reduced compared to a control cell that does not contain the insertion of the donor sequence.
[0012] In one aspect, the disclosure provides a method of modifying gene expression, the method comprising: (a) inducing a double-stranded break at a target site using targeted editing technology; (b) inserting a donor sequence at the double-stranded break, where the donor sequence comprises an endogenous element (e.g., a promoter, enhancer, or promoter / enhancer fragment) or a designed element that can induce increased or ectopic expression of the gene; and (c) identifying at least one cell that contains the donor sequence, wherein expression of the target gene is increased in at least one tissue compared to a control cell that does not contain the donor sequence.
[0013] In one aspect, the disclosure provides a method of promoting gene expression, the method comprising: (a) inducing a double-stranded break at a target site using targeted editing technology; (b) inserting a donor sequence at the double-stranded break, where the donor sequence comprises an endogenous element (e.g., a promoter, enhancer, or promoter / enhancer fragment) or a designed element that can induce increased or ectopic expression of the gene; and (c) identifying at least one cell that contains the donor sequence, wherein expression of the target gene is increased in at least one tissue compared to a control cell that does not contain the donor sequence.
[0014] In one aspect, the disclosure provides a method of generating a dominant positive allele, the method comprising: (a) inducing a double-stranded break at a target site using targeted editing technology; (b) inserting a donor sequence at the double-stranded break, wherein the donor sequence comprises a sequence of an endogenous gene; and (c) identifying at least one cell that contains the donor sequence, wherein expression of the gene is increased in at least one tissue compared to a control cell that does not contain the donor sequence.
[0015] In one aspect, the disclosure provides a method of reducing expression of a gene in a cell, the method comprising: (a) identifying a chromosomal region comprising a first gene region comprising a first promoter and a first coding region and a second gene region comprising a second promoter and a second coding region, wherein the first coding region and the second coding region are separated by an intervening region and the first promoter and the second promoter are positioned in opposite orientations; (b) inducing first and second double-stranded breaks flanking the targeted region using targeted editing technology, wherein the targeted region comprises the second coding region and the intervening region; and (c) identifying one or more cells comprising a deletion of the targeted region, wherein the second promoter generates at least one antisense RNA of the first coding region and expression of the first coding region is reduced compared to a control cell that does not comprise the deletion of the targeted region.
[0016] In one aspect, the disclosure provides a method of reducing expression of a protein of interest in a cell, the method comprising: (a) identifying a chromosomal region comprising a gene region encoding the protein of interest, the gene region comprising a first promoter and a coding region for the protein, and a second chromosomal region comprising a second promoter and an intervening region, wherein the coding region for the protein of interest and the second promoter are separated by the intervening region, and the first promoter and the second promoter are positioned in opposite orientations; (b) inducing a first double-stranded break and a second double-stranded break flanking the intervening region using targeted editing technology; and (c) identifying one or more cells comprising a deletion of the intervening region, wherein expression of the protein of interest is reduced compared to a control cell that does not comprise the deletion of the intervening region.
[0017] In one aspect, the disclosure provides a method for generating an inversion in a targeted region of a gene, comprising: (a) providing at least one RNA-guided nuclease, or one or more vectors encoding at least one RNA-guided nuclease, to one or more cells, wherein the at least one RNA-guided nuclease is capable of inverting at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, at least 110, at least 111, at (b) providing a nucleic acid sequence encoding an RNA-guided nuclease capable of binding to a stretch of at least 25, or at least 26, nucleotides, wherein a first target site and a second target site are linked, and wherein at least one RNA-guided nuclease creates a double-stranded break at the first target site and the second target site in the gene; (b) identifying one or more cells containing an inversion in a targeted region of the gene, wherein the inversion results in production of an antisense RNA transcript from the targeted region; and (c) selecting one or more cells containing an inversion in the targeted region of the gene.
[0018] (a) identifying a chromosomal region comprising a first genetic region comprising a first promoter and a first coding region and a second genetic region comprising a second promoter and a second coding region, wherein the first coding region and the second coding region are separated by an intervening region and the first promoter and the second promoter are positioned in opposite orientations; and (b) providing at least one RNA-guided nuclease, or one or more vectors encoding at least one RNA-guided nuclease, to one or more cells, wherein the at least one RNA-guided nuclease is located at a target region of the chromosome. (c) providing a chromosomal RNA-guided nuclease capable of binding to at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 stretches of nucleotides of a first target site and a second target site flanking a targeting region, the targeting region comprising a second coding region and an intervening region, wherein the RNA-guided nuclease creates a double-stranded break at the first target site and the second target site in the chromosome; (c) identifying one or more cells comprising a deletion of the targeted region; and (d) selecting one or more cells comprising a deletion of the targeted region.
[0019] In one aspect, the present disclosure provides a method comprising: (a) providing one or more cells with at least one RNA-guided nuclease and at least one donor molecule, or one or more vectors encoding at least one RNA-guided nuclease and at least one donor molecule, wherein the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site of at least one gene, wherein the donor molecule comprises a designed element, wherein the RNA-guided nuclease creates a double-stranded break at the target site, and the designed element is inserted at the double-stranded break; (b) identifying one or more cells comprising insertion of the designed element at the target site; and (c) selecting one or more cells comprising insertion of the designed element at the target site.
[0020] In one aspect, the disclosure provides a method for mitochondrial cell proliferation and differentiation comprising (a) providing one or more vectors encoding at least one RNA-guided nuclease and at least one donor molecule, or at least one RNA-guided nuclease and at least one donor molecule, to one or more cells, wherein the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site of at least one gene, and wherein the donor molecule is tissue-specific or (b) identifying one or more cells that contain an insertion at the target site of the sequence encoding the tissue-specific or tissue-preferred promoter, such that the sequence encoding the tissue-specific or tissue-preferred promoter is in a reverse orientation relative to the gene; and (c) selecting one or more cells that contain an insertion at the target site of the sequence encoding the tissue-specific or tissue-preferred promoter.
[0021] In one aspect, the disclosure provides a method comprising: (a) providing one or more RNA-guided nucleases, or one or more vectors encoding one or more RNA nucleases, to one or more cells, wherein the one or more RNA-guided nucleases are capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site, and wherein the one or more RNA-guided nucleases create a double-stranded break at the target site; (b) identifying at least one cell comprising an insertion or deletion at the target site, wherein the insertion or deletion at the target site results in the generation of a dominant-negative allele of at least one gene; and (c) selecting the one or more cells comprising the dominant-negative allele of the at least one gene.
[0022] In one aspect, the disclosure provides a method comprising: (a) providing one or more RNA-guided nucleases, or one or more vectors encoding one or more RNA nucleases, to one or more cells, wherein the one or more RNA-guided nucleases are capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site, and wherein the one or more RNA-guided nucleases create a double-stranded break at the target site; (b) identifying at least one cell comprising an insertion or deletion at the target site, wherein the insertion or deletion at the target site results in the generation of a dominant positive allele of at least one gene; and (c) selecting the one or more cells comprising the dominant positive allele of the at least one gene.
[0023] In one aspect, the disclosure provides a method comprising: (a) using a targeted editing technique to create a first double-strand break (DSB) and a second DSB in a first allele of a gene in a cell; (b) using a targeted editing technique to create a third DSB in a second allele of the gene in the cell; and (c) identifying a cell that comprises an insertion of a region of the first allele in a reverse orientation at the site of the third DSB in the second allele, thereby generating a modified second allele.
[0024] In one aspect, the disclosure provides a method of generating a dominant-negative allele of a gene using targeted editing technology, comprising introducing at least one non-coding RNA target site into the gene.
[0025] In one embodiment, the present disclosure provides a method for generating the dominant allele of a gene, comprising using targeted editing technology to introduce nonsense mutation into gene to generate truncated protein or protein that changes the amino acid composition downstream of mutation.This technology can also be combined with a second targeted editing mutation to restore the normal amino acid sequence frame downstream of the first mutation, and generate a nonsense region within the polypeptide.
[0026] In one aspect, the disclosure provides a method comprising: (a) providing to a cell an engineered pentatricopeptide repeat (PPR) protein, or a vector encoding the engineered PPR protein operably linked to a promoter, wherein the engineered PPR protein is capable of binding to an RNA transcript of a target gene; (b) selecting one or more cells from step (a) that express the engineered PPR protein; and (c) identifying one or more cells selected in step (b) that comprise altered expression of the target gene.
[0027] In one aspect, the disclosure provides a method of generating a dominant-negative allele of a gene in a cell, comprising using targeted editing technology to insert an inverted copy of the gene or portion thereof adjacent to a native copy of the gene to generate an inverted repeat sequence capable of producing an antisense RNA transcript of the gene or portion thereof.
[0028] In one aspect, the present disclosure provides a method of generating a dominant negative or dominant positive allele of a gene in a cell, comprising deleting a portion of the gene using targeted editing technology, wherein a microprotein is produced after deletion of the portion of the gene.
[0029] In one aspect, the present disclosure provides a method for generating a dominant-negative or dominant-positive allele of a gene in a cell, comprising deleting a portion of an intergenic region using targeted editing technology, wherein the deletion places the gene under the control of an upstream promoter. In some embodiments, the upstream promoter drives increased expression of the gene. In some embodiments, the upstream promoter drives decreased expression of the gene. In some embodiments, the upstream promoter drives a change in temporal expression of the gene. In some embodiments, the upstream promoter drives a change in tissue-specific expression of the gene.
[0030] In one aspect, the disclosure provides a method of generating a dominant negative allele of at least one gene in at least one cell, the method comprising: (a) introducing into the at least one cell a genome editing system including: (i) a site-specific nuclease or a molecule encoding the site-specific nuclease; (ii) a single guide RNA (sgRNA) or a molecule encoding the sgRNA; and (iii) at least a first tethered guide oligo (tgOligo) and a second tgOligo, or one or more molecules encoding the first and second tgOligos, operably linked to at least one promoter; and (b) generating a first double-strand break (DSB) and a second DSB in the at least one gene. wherein the first tgOligo and the second tgOligo hybridize to the 3' free ends of opposing strands at the first DSB and the second DSB, resulting in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 2500, or at least 5000 nucleotides of at least one gene, thereby generating a dominant negative allele of the gene encoding the truncated protein; and (c) identifying and selecting at least one cell containing the truncated protein.
[0031] In one aspect, the disclosure provides a method of generating a dominant negative allele of at least one gene in at least one cell, the method comprising: (a) introducing into the at least one cell one or more vectors encoding at least a first tethered guide oligo (tgOligo) and a second tgOligo operably linked to (i) at least one site-specific nuclease, (ii) at least one single guide RNA (sgRNA), and (iii) at least one promoter; and (b) creating a first double-strand break (DS) in the gene. (B) generating a first DSB and a second DSB, wherein the first tgOligo and the second tgOligo hybridize to the 3' free ends of opposing strands at the first DSB and the second DSB, such that the region of at least one gene between the first DSB and the second DSB is in the opposite orientation, thereby generating a dominant negative allele of at least one gene that encodes an antisense RNA transcript of the gene; and (c) identifying and selecting at least one cell containing the antisense RNA transcript of at least one gene.
[0032] In one aspect, the disclosure provides a modified plant cell comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
[0033] In one aspect, the disclosure provides a modified chromosome comprising a non-transposon-mediated deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
[0034] In one aspect, the present disclosure provides a modified plant or portion thereof comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the endogenous locus of the gene, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
[0035] In one aspect, the disclosure provides a modified cell comprising: (a) a non-transposon-mediated genomic deletion of at least one gene or a portion thereof at the endogenous locus of the at least one gene; or (b) a non-transposon-mediated, non-T-DNA-mediated insertion of a polynucleotide sequence into at least one gene, wherein the deletion or insertion creates a dominant positive allele of the at least one gene.
[0036] In one aspect, the disclosure provides a modified cell comprising a non-transposon-mediated genomic deletion or inversion of at least one gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the gene's native transcript sequence.
[0037] In one aspect, the disclosure provides a modified cell comprising a targeted edit of at least one gene or portion thereof, wherein the targeted edit produces an RNA transcript that is complementary to the native transcript sequence of the gene.
[0038] In one aspect, the disclosure provides a modified cell comprising at least one dominant negative allele of at least one gene generated by targeted editing technology, wherein when the at least one dominant negative allele is transcribed, the allele produces an RNA transcript capable of forming a hairpin loop secondary structure.
[0039] In one aspect, the disclosure provides a modified cell comprising a non-transgenic dominant negative allele of a gene, wherein the dominant negative allele comprises a heterologous non-coding RNA target site at the endogenous locus of the gene.
[0040] In one aspect, the disclosure provides a modified cell comprising a non-transgenic dominant positive allele of a gene, wherein the dominant positive allele comprises a heterologous non-coding RNA target site at the endogenous locus of the gene.
[0041] In one aspect, the disclosure provides a modified cell comprising at least one insertion or deletion in the endogenous locus of at least one gene generated by targeted editing techniques, wherein the insertion or deletion results in expression of a truncated protein.
[0042] In one aspect, the disclosure provides modified cells comprising a dominant negative allele of at least one gene, comprising an inverted copy of the gene adjacent to a native copy of the gene at the endogenous locus of the gene. BRIEF DESCRIPTION OF THE DRAWINGS [Brief explanation of the drawings]
[0043] [Figure 1] Includes Panel A and Panel B. Panel A shows that two sRNAs coupled to two RNA-guided nucleases can hybridize to target DNA on the same or different strands, separating a region of DNA from the rest of the DNA strand(s) by generating two double-strand breaks (DSBs). Panel B shows various outcomes that can occur after two RNA-guided nucleases generate two DSBs in DNA. Native cellular machinery can repair the two DSBs by deleting the entire region between the DSBs, deleting a small portion of the region from the 5' or 3' end, or inverting the region between the DSBs. [Figure 2]Panels A and B are included. Panel A shows a representation of the maize genomic region near the GA20 oxidase_5 gene. As shown in panel A, the methyltransferase / SAMT gene is adjacent to the GA20 oxidase_5 gene but in the opposite orientation. The intervening region between the SAMT promoter and the GA20 oxidase_5-encoding gene was deleted after generating double-stranded breaks at each end. Panel B shows the structure of the region after deletion of the intervening region. By removing the intervening region, the SAMT promoter can drive expression of an antisense GA20 oxidase_5 RNA transcript, which can form a double-stranded RNA with the sense GA20 oxidase_5 RNA transcript generated by the native GA20 oxidase_5 promoter. [Figure 3] Panels A, B, C, and D are included. Panel A shows GUS staining in Arabidopsis thaliana plants, demonstrating that the native 3 promoter expresses GUS only in root tissue. Panel B shows the genomic structure surrounding the native 3 promoter and GUS transgene. Panel C shows expanded GUS staining after insertion of a native genomic expression element, such as an enhancer element or a designed expression element, upstream of the native 3 TATA box. Panel D shows the genomic structure surrounding the native 3 promoter after targeted insertion of a native enhancer or a designed element. [Figure 4] Panels A, B, C, and D are included. Panel A shows the structure of a gene of interest, demonstrating that a double-stranded break can be generated immediately upstream of the polyadenylation site within the 3'-UTR of the gene of interest. Panel B shows the insertion of an antisense promoter at the double-stranded break site. The promoter can be the gene's native promoter or any promoter of interest. Panel C shows the optional generation of two double-stranded breaks surrounding the gene's native promoter, resulting in its deletion (Panel D). [Figure 5]Includes panels A, B, C, and D. Panel A shows GUS staining in Arabidopsis thaliana plants containing a GUS transgene under the control of the native 3 promoter. GUS is expressed throughout the plant. Panel B shows the genomic region surrounding the native 3 promoter and GUS transgene. Panel C shows reduced GUS staining when a leaf-specific promoter is inserted downstream of the GUS transgene to transcribe antisense GUS RNA. Panel D shows the genomic region surrounding the native 3 promoter and GUS transgene after insertion of an antisense leaf-specific promoter. [Figure 6] The diagram illustrates the possible outcomes of protein truncation. The top flow shows normal protein interactions that bring all encoded components together in the correct conformation, allowing the protein complex to function. The middle flow shows truncation of one of the protein components, removing a functional unit of the protein but retaining an interaction domain that can still bind to interacting proteins and block the interaction site from the fully functional form of the truncated protein. In the middle case, this acts as a dominant-negative allele. The bottom flow shows the strategic deletion of a portion of the protein that encodes a regulatory element of activity, leaving the functional and interaction domains intact, resulting in a protein complex that is either constitutively active or repressed in activity. [Figure 7] Panels A, B, C, D, and E are included. Panel A shows a heterozygous genomic locus. Panel B shows that a nuclease (represented by scissors) generates one DSB in the first allele. Panel C shows that the nuclease generates two DSBs in the second allele. In one possible outcome (Panel D), the region between the two DSBs on the second allele (Panel C) is inverted and inserted into the single DSB in the first allele. Panel E shows the generation of a hairpin RNA transcript that transcribes the allele structure in Panel D. [Figure 8]A schematic diagram of the insertion of a non-coding RNA target site in a gene of interest is provided. Such insertion can lead to the production of transient secondary siRNAs that downregulate the gene of interest (GOI). [Figure 9] A schematic diagram of a Cas9-mediated double-strand break (DSB) and a tethered guide oligo (tgOligo) bound to a target DNA site is provided. The Cas9-PAM interaction occurs on the non-target strand, and sgRNA-DNA annealing occurs on the target strand. The blunt ends of the Cas9 cleavage site are held in place by Cas9 at the 5' end of the non-target strand (PAM position) and both cut ends (3' and 5') of the target strand. The 3' cut end of the non-target strand is free and "flaps." The 3' free "flap" end of the non-target strand can be up to 35 nucleotides long, which may be sufficient for specific complementary binding. A tgOligo (e.g., ssDNA template) can be included for incorporation of desired nucleotide modifications. The drawing scheme used here continues in subsequent figures. [Figure 10] Cas9 is shown conjugated to a homodimer domain (top) and a heterodimer domain (center and bottom right) to facilitate dimerization. Ligands for the homodimer and heterodimer domains are shown (bottom left). Drawing schemes for the ligand, homodimer or heterodimer domain, ssDNA-binding domain, and other components used herein are shown in subsequent figures. The components of the Cas9 / sgRNA complex and target DNA are shown as exemplified in Figure 9. Drawing schemes for the various dimerization domains used herein are shown in subsequent figures. [Figure 11] This figure shows the use of catalytically inactivated Cas9 (dCas9) to increase genome editing efficiency. Panel 1 illustrates that dCas9 binds to DNA at the target site specified by gRNA, creating a loop structure accessible for template-based editing. Panel 2 illustrates a modified scheme to further facilitate template-based editing via dCas9 conjugated with an ssDNA binding domain. The editing efficiency of this modified scheme is expected to be higher than that of Panel 1 because the ssDNA template binds to the dCas9 complex and is brought into close proximity to the gRNA target. [Figure 12] An example of a construct containing Cas9, gRNA, and tgOligo is provided. RZ stands for ribozyme, an enzyme that cleaves a 15-bp recognition site (RZ site) in RNA. [Figure 13] This figure provides examples of various approaches to improving genome editing efficiency. Using a dimerization domain (see Figure 10), tgOligo (see Figure 9), or a combination of both, can facilitate the restoration of a complete knockout (deletion) of a genomic region flanked by two gRNA target sites. Panel 1 shows a knockout (KO) event promoted by dimerization. Panel 2 shows a KO event promoted by tgOligo. Panel 3 shows a KO event promoted by the combination of dimerization and tgOligo. Panel 4 shows an inversion event promoted by tgOligo. Panel 5 shows an inversion event promoted by dimerization. Panel 6 shows an inversion event assisted by the combination of Cas9 dimerization / inactivation and tgOligo. Only configurations in which the two gRNAs recognize different strands of the target dsDNA are shown. The same concept is equally applicable to other configurations in which the two gRNAs recognize the same strand of the target dsDNA. [Figure 14]This example demonstrates the generation of a dominant knockout allele via genomic inversion by editing the maize BR2 gene. Two exemplary gRNAs are used. The first gRNA (shown on the left) targets the end of the first exon of BR2, while the second gRNA (shown on the right) recognizes the start codon region of the adjacent GRMZM2G491632 gene. Inversion of the genomic segment flanked by these two gRNAs can result in a BR2 antisense partial transcript (see transcript 1). This BR2 antisense transcript is produced via GRMZM2G491632 promoter activity. By adjusting the relative positions of the two gRNAs, a BR2 antisense full transcript (e.g., by shifting the first gRNA to the left to target the start codon region of the BR2 gene) or a BR2 antisense transcript under the control of the native BR2 promoter (e.g., by shifting the second gRNA to the right to target the stop codon region of the BR2 gene) can be achieved. [Figure 15] Examples are provided of template-based editing or site-directed integration (SDI) facilitated by dimerization at a single location (panels 1 and 2) or multiple locations (panel 3), and template-based editing or SDI facilitated by dimerization / tgOligo (panel 4). [Figure 16] Illustrative examples of template editing (Panel 1), site-directed integration (Panel 2), and / or recombination (Panel 3) using tgOligo are provided. [Figure 17] This paper provides an example of head-to-tail stacking of an inverted Y1 gene to produce antisense transcripts and silence gene expression. This approach can create a dominant mutant Y1 allele for a normally recessive trait. This dominant allele remains under the control of the native Y1 promoter. [Figure 18]We present an example of a microprotein. The target of a microprotein is often a transcriptional regulator that binds to DNA as an active homodimer. The microprotein interferes with its target by forming a non-functional heterodimeric complex that cannot bind to DNA. The DBD is the DNA-binding domain, and the PPI is the protein-protein interaction domain. [Figure 19] Panels A and B are included. Panel A shows a representation of the maize genomic region near the MIR1 gene. As shown in panel A, the GRMZM2G150302 gene is adjacent to and upstream of the GA20 oxidase_5 gene. The intervening region between the GRMZM2G150302 promoter and the MIR1-encoding gene was deleted after generating double-stranded breaks at each end. Panel B shows the structure of the region after deletion of the intervening region. By removing the intervening region, the GRMZM2G150302 promoter can drive expression of the MIR1 gene. [Figure 20] We provide a specific example for generating antisense RNA molecules targeting the Zm.GA20ox5 and Zm.GA20ox3 genes by deleting the genomic region between Zm.GA20ox5 and its neighboring gene Zm.SAMT, which faces in the opposite direction, via genome editing. [Figure 21] Illustrated are the genomic locations of various guide RNA target sites in three exemplary vectors for creating a genomic deletion between the Zm.GA20ox5 gene and its neighboring Zm.SAMT gene. [Figure 22] The average height of wild-type and homozygous edited plants is shown in inches (Y-axis). [Figure 23] The average height of wild-type plants and homozygous or heterozygous edited plants is shown in inches (Y-axis). [Figure 24] The concentrations of GA12 and GA9 in edited and control plants are shown in pmol / g (Y-axis). [Figure 25] The concentrations of GA20 and GA53 in edited and control plants are shown in pmol / g (Y-axis). [Figure 26]The concentrations of active gibberellic acids GA1, GA3, and GA4 in edited and control plants are shown in pmol / g (Y-axis). [Figure 27] A specific example is provided in which genomic modification of the Zm.GA20ox3 locus encodes an RNA transcript having an inverted sequence capable of hybridizing to the corresponding sequence of the RNA transcript to produce a stem-loop structure, in order to cause suppression of one or both copies or alleles of the endogenous Zm.GA20ox3 and Zm.GA20ox5 loci. [Figure 28] Average height of wild-type and heterozygous edited maize plants is shown. [Figure 29] The number of 21-mer small RNAs per million reads (Y-axis) detected in plant samples containing the edited GA20ox3 allele mapped to a region in the stem of the edited stem-loop containing the inverted sequence of GA20ox5 and the corresponding sequence of the edited GA20ox3 gene. [Figure 30] The concentrations of GA12 and GA9 in edited and control corn plants are shown in pmole / g (Y-axis). [Figure 31] The concentrations of GA20 and GA53 in edited and control corn plants are shown in pmole / g (Y-axis). [Figure 32] The concentrations of GA1, GA3, and GA4 in edited and control corn plants are shown in pmole / g (Y-axis). DETAILED DESCRIPTION OF THE INVENTION
[0044] Detailed Description Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. As will be recognized by those skilled in the art, many methods may be used in the practice of the present disclosure. Indeed, the present disclosure is in no way limited to the methods and materials described. In this disclosure, the following terms are defined as follows:
[0045] The present specification provides methods and compositions for generating dominant alleles using targeted editing techniques in a wide range of organisms, including plants, animals, fungi, and protozoa. A dominant allele is an allele that suppresses the contribution of a second allele at the same locus. A dominant allele can be a "dominant-negative allele" or a "dominant-positive allele." A dominant-negative allele, or antimorph, is an allele that acts opposite to the normal function of an allele. Dominant-negative alleles usually malfunction, either directly inhibiting the activity of the wild-type protein (e.g., via dimerization) or inhibiting the activity of a second protein (e.g., an activator or downstream component of a pathway) required for the normal function of the wild-type protein. For example, dominant-negative alleles prevent or reduce the normal function of an allele in heterozygous or homozygous states. Dominant-positive alleles can increase or expand normal gene function (e.g., hypermorphs) or confer new functions to a gene (e.g., neomorphs). A semidominant allele occurs when the penetrance of the associated phenotype in individuals heterozygous for the allele is less than the penetrance observed in individuals homozygous for the allele.
[0046] The practice of the present disclosure will employ, unless otherwise indicated, conventional techniques of biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics and biotechnology, which are within the skill of those in the art. Green and Sambrook, Molecular Cloning: A Laboratory Manual, 4th Edition (2012), Current Protocols In Molecular Biology (FMAusubel, et al. eds., (1987)), Series Methods In Enzymology (Academic Press, Inc.): PCR 2: A Practical Approach (MJMacPherson, BD Hames and GRTaylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual; Animal Cell Culture (RIFreshney, ed. (1987)), Recombinant Protein Purification: Principles And Methods, 18-1142-75, GE Healthcare Life Sciences, CNStewart, A. Touraev, V. Citovsky, T. Tzfira eds. (2011) Plant Transformation Technologies (Wiley-Blackwell), and RH Smith (2013) Plant Tissue Culture. Techniques and Experiments (Academic Press, Inc.).
[0047] All references cited herein, including, for example, all patents, published patent applications, and non-patent publications, are incorporated by reference in their entirety.
[0048] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" can include a plurality of compounds, including mixtures of compounds.
[0049] The term "and / or," when used in the context of a list of two or more items, means that any one of the listed items may be used alone or in combination with any one or more of the listed items. For example, the phrase "A and / or B" is intended to mean one or both of A and B, i.e., A alone, B alone, or A and B in combination. The phrase "A, B, and / or C" is intended to mean A alone, B alone, C alone, A and B in combination, A and C in combination, B and C in combination, or A, B, and C in combination.
[0050] Nucleic acid molecules provided herein include deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), as well as their functional analogs, such as complementary DNA (cDNA). Nucleic acid molecules provided herein may be single-stranded or double-stranded. Nucleic acid molecules contain the nucleotide bases adenine (A), guanine (G), thymine (T), and cytosine (C). In RNA molecules, uracil (U) substitutes for thymine. The symbol "N" can be used to represent any nucleotide base (e.g., A, G, C, T, or U). As used herein, "encoding" refers to a polynucleotide encoding the amino acids of a polypeptide. A sequence of three nucleotide bases encodes an amino acid. As used herein, "expressed," "expression," or "expressing" refers to the transcription of RNA from a DNA molecule. As used herein, the terms "polypeptide," "peptide," and "protein" are used interchangeably and refer to a polymer of amino acid residues. The term also applies to amino acid polymers in which one or more amino acids are chemical analogs or modified derivatives of corresponding naturally occurring amino acids. "Messenger RNA" or "mRNA" refers to an RNA transcript transcribed from a polynucleotide that can be translated into a protein. Typically, DNA encodes mRNA, which in turn encodes a protein. When DNA is transcribed by RNA polymerase to ultimately produce a protein, the RNA polymerase typically produces a sense mRNA strand from the antisense DNA strand. The sense strand of DNA or RNA extends from 5' to 3', and the antisense strand extends from 3' to 5'. The sense and antisense strands of the same polynucleotide are complementary to each other.
[0051] In one embodiment, the nucleic acid molecules provided herein include protein-encoding nucleic acid molecules that are codon-optimized for eukaryotic cells. In another embodiment, the protein-encoding nucleic acid molecule is codon-optimized for plant cells. In another embodiment, the protein-encoding nucleic acid molecule is codon-optimized for monocotyledonous plant species. In a further embodiment, the protein-encoding nucleic acid molecule is codon-optimized for corn or soybean cells.
[0052] The term "percent identity" or "percent identity" as used herein with respect to two or more nucleotide or protein sequences is calculated by (i) comparing two optimally aligned sequences (nucleotide or protein) over a comparison window, (ii) determining the number of positions where the same nucleic acid base (in the case of a nucleotide sequence) or amino acid residue (in the case of a protein) occurs in both sequences to determine the number of matching positions, (iii) dividing the number of matching positions by the total number of positions in the comparison window, and then (iv) multiplying this quotient by 100% to determine the percent identity. When "percent identity" is calculated relative to a reference sequence without specifying a specific comparison window, the percent identity is determined by dividing the number of matching positions in the region of alignment by the total length of the reference sequence. Thus, in this application, when two sequences (query and subject) are optimally aligned (allowing gaps in their alignment), the "percent identity" of the query sequence is equal to the number of identical positions between the two sequences divided by the total number of positions in the entire length of the query sequence (or comparison window), multiplied by 100%. When percentages of sequence identity are used with respect to proteins, it is recognized that non-identical residue positions often differ by conservative amino acid substitutions. Conservative amino acid substitutions are those in which an amino acid residue is replaced with another amino acid residue having similar chemical properties (e.g., charge or hydrophobicity), thereby not changing the functional properties of the molecule. When multiple sequences differ by conservative substitutions, the percent sequence identity may be adjusted upward to correct for the conservative nature of the substitution. Sequences that differ by such conservative substitutions are said to have "sequence similarity" or "similarity."
[0053] For optimally aligning sequences and calculating their percent identity, various pairwise or multiple sequence alignment algorithms and programs are known in the art, such as ClustalW or Basic Local Alignment Search Tool® (BLAST), which can be used to compare sequence identity or similarity between two or more nucleotide or protein sequences. Although other alignment and comparison methods are known in the art, the alignment and percent identity between two sequences (including the percent identity ranges mentioned above) can be determined by the ClustalW algorithm. For example, Chenna R. et al., “Multiple sequence alignment with the Clustal series of programs,” Nucleic Acids Research 31:3497-3500 (2003), Thompson JD et al., “Clustal W: Improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice,” Nucleic Acids Research 22:4673-4680 (1994), Larkin MA et al., “Clustal W and Clustal tool.” J. Mol. Biol. 215:403-410 (1990). The contents and disclosures of which are incorporated herein by reference in their entirety.
[0054] The term "percent complementarity" or "percent complementary," as used herein with respect to two nucleotide sequences, is similar to the concept of percent identity, but refers to the percentage of nucleotides in the query sequence that optimally base-pair or hybridize with the nucleotides of the subject sequence when the query and subject sequences are linearly aligned and optimally base-paired without secondary folding structures such as loops, stems, or hairpins. Such percent complementarity can be between two DNA strands, two RNA strands, or a DNA strand and an RNA strand. "Percent complementarity" is calculated by (i) optimally base-pairing or hybridizing the two nucleotide sequences in a linear, fully unfolded alignment (i.e., without folding or secondary structures) across the comparison window, (ii) determining the number of base-paired positions between the two sequences across the comparison window to determine the number of complementary positions, (iii) dividing the number of complementary positions by the total number of positions in the comparison window, and (iv) multiplying this quotient by 100% to determine the percent complementarity of the two sequences. The optimal base pairing of two sequences can be determined based on the known pairing of nucleotide bases such as GC, AT, and AU through hydrogen bonds. When "percent complementarity" is calculated relative to a reference sequence without specifying a specific comparison window, the percent identity is determined by dividing the number of complementary positions between two linear sequences by the total length of the reference sequence. Thus, in the present application, when two sequences (query and subject) are optimally base-paired (allowing for mismatched or non-base-paired nucleotides), the "percent complementarity" of the query sequence is equal to the number of base-paired positions between the two sequences divided by the total number of positions in the total length of the query sequence, multiplied by 100%.
[0055] The term "operably linked" refers to a functional relationship between a promoter or other regulatory element and an associated transcribable DNA sequence or coding sequence of a gene (or transgene), such that the promoter or the like operates to initiate, support, influence, cause, and / or enhance the transcription and expression of the associated transcribable DNA sequence or coding sequence, at least in particular tissue(s), developmental stage(s), and / or condition(s). Regulatory elements, in addition to promoters, include, but are not limited to, enhancers, leaders, transcription start sites (TSSs), linkers, 5' and 3' untranslated regions (UTRs), introns, polyadenylation signals, and termination regions or sequences, etc., that are suitable, necessary, or preferred for regulating or enabling expression of a gene or transcribable DNA sequence in a cell. Such additional regulatory element(s) may be optional and may be used to enhance or optimize expression of a gene or transcribable DNA sequence. In this application, an "enhancer" can be distinguished from a "promoter" in that an enhancer typically lacks a transcription initiation site, TATA box, or equivalent sequence, and therefore is insufficient to drive transcription alone. As used herein, a "leader" can be generally defined as the DNA sequence in the 5'-UTR of a gene (or transgene) that is between the transcription start site (TSS) and the 5' end of the transcribable DNA sequence or the start site of the gene's protein-coding sequence.
[0056] As commonly understood in the art, the term "promoter" refers to a DNA sequence that contains an RNA polymerase binding site, a transcription initiation site, and / or a TATA box and that supports or enhances the transcription and expression of an associated transcribable polynucleotide sequence and / or gene (or transgene). Promoters can be synthetically produced, altered, or derived from known or naturally occurring promoter sequences or other promoter sequences. Promoters can also include chimeric promoters, which contain a combination of two or more heterologous sequences. Thus, promoters of the present application can include variants of promoter sequences that are similar in composition to, but not identical to, other promoter sequence(s) known or provided herein. Promoters can be classified according to various criteria related to the expression pattern of the associated coding or transcribable sequence or gene (including transgene) operably linked to the promoter, such as constitutive, developmental, tissue-specific, inducible, etc. A promoter that drives expression in all or most tissues of a plant is referred to as a "constitutive" promoter. A promoter that drives expression during a specific period or stage of development is referred to as a "developmental" promoter. A promoter that drives enhanced expression in certain plant tissues relative to other plant tissues is referred to as a "tissue-promoting" or "tissue-preferred" promoter. Thus, a "tissue-preferred" promoter causes relatively higher or selective expression in specific plant tissue(s), with lower levels of expression in other plant tissue(s). A promoter that expresses in specific plant tissue(s) and has little or no expression in other plant tissues is referred to as a "tissue-specific" promoter. An "inducible" promoter is a promoter that initiates transcription in response to environmental stimuli such as cold, drought, or light, or other stimuli such as wounding or chemical application. Promoters can also be classified in terms of their origin, for example, heterologous, homologous, chimeric, synthetic, etc.A "heterologous" promoter is a promoter sequence that has a different origin from the associated transcribable sequence, coding sequence, or gene (or transgene) and / or does not naturally occur in the plant species being transformed.
[0057] Examples describing promoters that can be used herein include, but are not limited to, U.S. Pat. No. 6,437,217 (maize RS81 promoter), U.S. Pat. No. 5,641,876 (rice actin promoter), U.S. Pat. No. 6,426,446 (maize RS324 promoter), U.S. Pat. No. 6,429,362 (maize PR-1 promoter), U.S. Pat. No. 6,232,526 (maize A3 promoter), U.S. Pat. No. 6,177,611 (constitutive maize promoters), U.S. Pat. Nos. 5,322,938, 5,352,605, 5,359,142, and 5,530,196 (35S promoter). No. 6,433,252 (maize L3 oleosin promoter), U.S. Pat. No. 6,429,357 (rice actin 2 promoter and rice actin 2 intron), U.S. Pat. No. 5,837,848 (root-specific promoter), U.S. Pat. No. 6,294,714 (light-inducible promoter), U.S. Pat. No. 6,140,078 (salt-inducible promoter), U.S. Pat. No. 6,252,138 (pathogen-inducible promoter), U.S. Pat. No. 6,175,060 (phosphorus deficiency-inducible promoter), U.S. Pat. No. 6,635,806 (gamma-coixin promoter), and U.S. Patent Application No. 09 / 757,089 (maize chloroplast aldolase promoter).Additional promoters that can be used include the nopaline synthase (NOS) promoter (Ebert et al., 1987), the octopine synthase (OCS) promoter (carried on tumor-inducing plasmids of Agrobacterium tumefaciens), caulimovirus promoters, such as the cauliflower mosaic virus (CaMV) 19S promoter (Lawton et al., Plant Molecular Biology (1987) 9:315-324), the CaMV 35S promoter (Odell et al., Nature (1985) 313:810-812), the figwort mosaic virus 35S promoter (U.S. Patent Nos. 6,051,753 and 5,378,619), the sucrose synthase promoter (Yang and Russell, Proceedings of the National Academy of Sciences, 1999). Sciences, USA (1990) 87:4144-4148), the R gene complex promoter (Chandler et al., Plant Cell (1989) 1:1175-1183), and the chlorophyll a / b binding protein gene promoters, PC1SV (U.S. Patent No. 5,850,019), and AGRtu.nos (GenBank accession number V00087; Depicker et al., Journal of Molecular and Applied Genetics (1982) 1:561-573; Bevan et al., 1983) promoters.
[0058] Promoter hybrids can also be used and constructed to enhance transcriptional activity (see U.S. Pat. No. 5,106,739) or to combine desired transcriptional activity, inducibility, and tissue or developmental specificity. Promoters that function in plants include, but are not limited to, inducible, viral, synthetic, constitutive, temporally regulated, spatially regulated, and spatio-temporally regulated promoters. Other tissue-enhanced, tissue-specific, or developmentally regulated promoters are known in the art and are contemplated as having utility in practicing the present disclosure.
[0059] As used herein, "heterologous" with respect to a promoter is a promoter sequence that has a different origin from the associated transcribable DNA sequence, coding sequence, or gene (or transgene) and / or does not naturally occur in the plant species being transformed. The term "heterologous" more broadly refers to a combination of two or more DNA molecules or sequences, such as a promoter and an associated transcribable DNA sequence, coding sequence, or gene, that is artificial and not normally found in nature.
[0060] Furthermore, the term "recombinant" with respect to polynucleotide (DNA or RNA) molecules, proteins, constructs, vectors, etc., refers to polynucleotide or protein molecules or sequences that are man-made and not normally found in nature and / or exist in a context not normally found in nature, including polynucleotide (DNA or RNA) molecules, proteins, constructs, etc. that contain combinations of polynucleotide or protein sequences that do not naturally occur contiguous or adjacent to each other without human intervention, and / or polynucleotide molecules, proteins, constructs, etc. that contain at least two polynucleotide or protein sequences that are heterologous to each other. Recombinant polynucleotide or protein molecules, constructs, etc. can contain polynucleotide or protein sequence(s) that are (i) separated from other polynucleotide or protein sequence(s) that occur contiguous to each other in nature and / or (ii) adjacent to (or adjacent to) other polynucleotide or protein sequence(s) that do not occur contiguous to each other in nature. Such recombinant polynucleotide molecules, proteins, constructs, etc. may also refer to polynucleotide or protein molecules or sequences that have been genetically engineered and / or constructed extracellularly. For example, a recombinant DNA molecule can include any suitable plasmid, vector, etc., and can include linear or circular DNA molecules. Such plasmids, vectors, etc. can contain various maintenance elements, including a prokaryotic origin of replication and a selectable marker, possibly in addition to a plant selectable marker gene, etc., as well as one or more transgenes or expression cassettes.
[0061] As used herein, " adjacent " refers to a nucleic acid sequence that is close to or next to another nucleic acid sequence.In one embodiment, adjacent nucleic acid sequences are physically connected.In another embodiment, adjacent nucleic acid sequences or genes are located immediately adjacent to each other, so that there is no intervening nucleotide between the end of the first nucleic acid sequence and the start of the second nucleic acid sequence. In an embodiment, a first gene and a second gene are adjacent to each other if they are separated by fewer than 50,000, fewer than 25,000, fewer than 10,000, fewer than 9000, fewer than 8000, fewer than 7000, fewer than 6000, fewer than 5000, fewer than 4000, fewer than 3000, fewer than 2500, fewer than 2000, fewer than 1750, fewer than 1500, fewer than 1250, fewer than 1000, fewer than 900, fewer than 800, fewer than 700, fewer than 600, fewer than 500, fewer than 400, fewer than 300, fewer than 200, fewer than 100, fewer than 75, fewer than 50, fewer than 25, fewer than 20, fewer than 10, fewer than 5, fewer than 4, fewer than 3, fewer than 2, or fewer than 1 nucleotide.
[0062] In one aspect, the methods and compositions provided herein include a vector. As used herein, the terms "vector" and "plasmid" are used interchangeably and refer to a circular, double-stranded DNA molecule that is physically separate from chromosomal DNA. In one aspect, a plasmid or vector as used herein is capable of replicating in vivo. As used herein, a "transformation vector" is a plasmid capable of transforming plant cells. In certain aspects, the plasmids provided herein are bacterial plasmids. In another aspect, the plasmids provided herein are Agrobacterium Ti plasmids or are derived from Agrobacterium Ti plasmids.
[0063] In one embodiment, the plasmid or vector provided herein is a recombinant vector. As used herein, the term "recombinant vector" refers to a vector formed by genetic recombination experimental methods such as molecular cloning. In another embodiment, the plasmid provided herein is a synthetic plasmid. As used herein, a "synthetic plasmid" is an artificially created plasmid that can perform the same function (e.g., replication) as a natural plasmid (e.g., Ti plasmid). Without limitation, those skilled in the art can synthesize a plasmid from individual nucleotides or create a synthetic plasmid by splicing nucleic acid molecules together from different existing plasmids.
[0064] As used herein, "modified," in the context of plants, seeds, plant components, plant cells, and plant genomes, refers to a state that includes changes or variations from their natural or native state. By way of example, a "native transcript" of a gene refers to an RNA transcript produced from an unmodified gene. Typically, a native transcript is a sense transcript. A modified plant or seed contains molecular changes, such as genetic or epigenetic modifications, in its genetic material. Typically, the modified plant or seed, or its parent or progenitor lineage, has been subjected to mutagenesis, genome editing (e.g., but not limited to, by methods using site-specific nucleases), genetic transformation (e.g., but not limited to, by Agrobacterium transformation or biolistic methods), or a combination thereof. In one aspect, the modified plants provided herein do not contain non-plant genetic material or sequences. In yet another aspect, the modified plants provided herein do not contain interspecies genetic material or sequences. In one aspect, the present disclosure provides methods and compositions relating to modified plants, seeds, plant components, plant cells, and products made from the modified plants, seeds, plant parts, and plant cells. In one aspect, the modified seeds provided herein give rise to the modified plants provided herein. In one aspect, the modified plants, seeds, plant components, plant cells, or plant genomes provided herein comprise a recombinant DNA construct or vector provided herein. In another aspect, the products provided herein comprise the modified plants, plant components, plant cells, or plant chromosomes or genomes provided herein.The present disclosure provides modified plants having desirable or enhanced characteristics, such as, but not limited to, disease, insect, or pest resistance (e.g., viral resistance, bacterial resistance, fungal resistance, nematode resistance, arthropod resistance, gastropod resistance); herbicide resistance; environmental stress resistance; improved quality, such as yield, nutritional enhancement, environmental or stress tolerance; any desired change in plant physiology, growth, development, morphology, or plant product(s), including starch production, altered oil production, high oil production, altered fatty acid content, high protein production, fruit ripening, animal and human nutritional enhancement, biopolymer production, pharmaceutical peptides, and secretable peptide production; improved processing characteristics; improved digestibility; low raffinose; industrial enzyme production; improved flavor; nitrogen fixation; hybrid seed production; and fiber production.
[0065] As used herein, "genome editing" or editing refers to targeted mutagenesis, insertion, deletion, inversion, substitution, or translocation of a nucleotide sequence of interest in a genome using targeted editing techniques. The nucleotide sequence of interest can be of any length, for example, at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 75, at least 100, at least 250, at least 500, at least 1000, at least 2500, at least 5000, at least 10,000, or at least 25,000 nucleotides. As used herein, "targeted editing techniques" refers to any method, protocol, or technique that allows for precise and / or targeted editing of specific locations within a genome (e.g., editing is not random). Without limitation, the use of site-specific nucleases is one example of a targeted editing technique. Another non-limiting example of targeted editing technology is the use of one or more tethered guide oligos (tgOligo). As used herein, "targeted editing" refers to targeted mutagenesis, insertion, deletion, inversion, or substitution caused by targeted editing technology. The nucleotide sequence of interest can be an endogenous genomic sequence or a transgenic sequence.
[0066] In one embodiment, "targeted editing technology" refers to any method, protocol, or technology that allows for precise and / or targeted editing of specific locations within a genome (e.g., the editing is not random). Without limitation, the use of site-specific nucleases is an example of a targeted editing technology.
[0067] In one embodiment, targeted editing technology is used to edit an endogenous locus or endogenous gene. In another embodiment, targeted editing technology is used to edit a transgene. As used herein, an "endogenous gene" or a "native copy" of a gene refers to a gene that originates within a given organism, cell, tissue, genome, or chromosome. An "endogenous gene" or a "native copy" of a gene is a gene that has not previously been modified by human action.
[0068] As used herein, "locus" refers to a specific location on a chromosome or other nucleic acid molecule. Without limitation, a locus may include a polynucleotide that encodes a protein or RNA. A locus may also include non-coding RNA. A locus may include a gene. A locus may include a promoter, a 5' untranslated region (UTR), an exon, an intron, a 3'-UTR, or any combination thereof. A locus may include a coding region.
[0069] As used herein, "physically linked" refers to two or more loci located on the same nucleic acid molecule.
[0070] As used herein, a "coding region," "gene region," or "gene" refers to a polynucleotide capable of producing a functional unit (such as, but not limited to, a protein or a non-coding RNA molecule). A "coding region," "gene," or "gene region" may include a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription termination site, a polyadenylation site, one or more exons, one or more introns, a 5'-UTR, a 3'-UTR, or any combination thereof. A "coding region sequence," "gene sequence," or "gene region sequence" may include a polynucleotide sequence that encodes a promoter, an enhancer sequence, a leader sequence, a transcription start site, a transcription termination site, a polyadenylation site, one or more exons, one or more introns, a 5'-UTR, a 3'-UTR, or any combination thereof. In one embodiment, a "gene" encodes a non-coding RNA molecule or a precursor thereof. In another embodiment, a "gene" encodes a protein.
[0071] Non-limiting examples of non-coding RNA molecules include microRNAs (miRNAs), miRNA precursors (pre-miRNAs), small interfering RNAs (siRNAs), small RNAs (18-26 nt in length) and their encoding precursors, heterochromatin siRNAs (hc-siRNAs), Piwi-interacting RNAs (piRNAs), hairpin double-stranded RNAs (hairpin dsRNAs), trans-acting siRNAs (ta-siRNAs), naturally occurring antisense siRNAs (nat-siRNAs), CRISPR RNAs (crRNAs), tracer RNAs (tracrRNAs), guide RNAs (gRNAs), and single guide RNAs (sgRNAs). In one aspect, the non-coding RNA provided herein is selected from the group consisting of microRNAs, small interfering RNAs, secondary small interfering RNAs, transfer RNAs, ribosomal RNAs, trans-acting small interfering RNAs, naturally occurring antisense small interfering RNAs, heterochromatin small interfering RNAs, and precursors thereof. In another embodiment, the non-coding RNA provided herein is selected from the group consisting of miRNA, pre-miRNA, siRNA, hc-siRNA, piRNA, hairpin dsRNA, ta-siRNA, nat-siRNA, crRNA, tracrRNA, gRNA, and sgRNA. The non-coding RNA is often 100% complementary to the non-coding RNA target site, or at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% complementary. The non-coding RNA target site can be present in a DNA molecule or an RNA molecule. Hybridization (or binding) of the non-coding RNA to the non-coding RNA target site can have various consequences. By way of example, some non-coding RNAs (including but not limited to, miRNA, siRNA, ta-siRNA) support the cleavage of mRNA transcripts that contain complementary non-coding RNA target sites.Alternatively, non-coding RNAs (for example, but not limited to, miRNA, siRNA, ta-siRNA) can help inhibit protein translation of mRNA transcripts by binding to non-coding RNA target sites in mRNA. Some non-coding RNAs (for example, but not limited to, hc-siRNA) help induce epigenetic changes to DNA. In one embodiment, the non-coding RNA target site is an miRNA target site or an siRNA target site. In another embodiment, the gene provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 non-coding RNA target sites. In another embodiment, the gene provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 heterologous non-coding RNA target sites. In another embodiment, the dominant negative alleles provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 non-coding RNA target sites. In another embodiment, the dominant negative alleles provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 heterologous non-coding RNA target sites. In some embodiments, the endogenous genes provided herein are modified to comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 heterologous non-coding RNA target sites.
[0072] As used herein, " allele " refers to a variant of a given locus or gene in a genome. If the same allele exists in both chromosomes of a chromosome pair in a cell, the cell is considered to be homozygous at a given locus. If each member of a chromosome pair contains different alleles for a given locus, the cell is considered to be heterozygous for that locus. At least one allele is possible for a given locus, but typically multiple alleles are possible for any given locus in a genome.
[0073] As used herein, the terms "suppress," "inhibit," "inhibition," "inhibiting," and "downregulation" are defined as any method known in the art or described herein that reduces the expression or function of a gene product (e.g., mRNA, protein, non-coding RNA). "Inhibition" can relate to a comparison of two cells, e.g., a modified cell and a control cell. Inhibition of gene product expression or function can also relate to a comparison in a plant cell, organelle, organ, tissue, or plant component within the same plant or between different plants, including a comparison in a developmental or temporal stage within the same plant or plant component or between plants or plant components. "Inhibition" includes any relative decrease in function or production of a gene product of interest, up to and including complete elimination of function or production of that gene product. The term "inhibition" encompasses any method or composition that downregulates the translation and / or transcription of a target gene product or the functional activity of a target gene product. "Inhibition" does not necessarily include complete elimination of expression of a gene product. In certain embodiments, a gene product in a modified cell provided herein comprises expression that is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, or 100% less than expression of the gene product in a control cell. In another embodiment, a gene product in a modified cell provided herein comprises expression that is between 1% and 100%, between 1% and 95%, between 1% and 90%, between 1% and 80%, between 1% and 70%, between 1% and 60%, between 1% and 50%, between 1% and 40%, between 1% and 30%, between 1% and 25%, between 1% and 20%, between 1% and 15%, between 1% and 10%, between 1% and 5%, between 5% and 25%, between 5% and 50%, between 5% and 75%, between 5% and 100%, between 10% and 25%, between 10% and 50%, between 10% and 75%, between 10% and 100%, between 25% and 50%, between 25% and 75%, between 25% and 100%, or between 50% and 100% less than expression of the gene product in a control cell.
[0074] As used herein, "target site" refers to the position in a polynucleotide sequence where a site-specific nuclease binds and cleaves to introduce a double-strand break in the nucleic acid backbone.In another embodiment, the target site comprises at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 29, or at least 30 consecutive nucleotides.In another embodiment, the target site provided herein is at least 10, at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, or at least 500 nucleotides. In one embodiment, the site-specific nuclease binds to the target site. In another embodiment, the site-specific nuclease binds to the target site via a guided non-coding RNA (i.e., including but not limited to, CRISPR RNA or single guide RNA, both of which are described in detail below). In one embodiment, the non-coding RNA provided herein is complementary to the target site. As will be understood, perfect complementarity is not required for a non-coding RNA to bind to a target site. At least one, at least two, at least three, at least four, or at least five, at least six, at least seven, or at least eight mismatches between the target site and the non-coding RNA can be tolerated. As used herein, "target region" or "targeted region" refers to a polynucleotide sequence desired to be modified. In one embodiment, a "target region," "targeted region," or "target gene" is flanked by two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more target sites. "Target gene" refers to a polynucleotide sequence that encodes a gene that one desires to modify.In one embodiment, the polynucleotide sequence comprising the target gene further comprises one or more target sites.In another embodiment, the target region comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more target genes.Without being limited thereto, in one embodiment, the target region can be subject to deletion or inversion.As used herein, the term "flanking" used to describe a target region refers to two or more target sites that physically surround the target region, with one target site located on each side of the target region.
[0075] A target site may be located within a polynucleotide sequence encoding a leader, enhancer, transcription initiation site, promoter, 5'-UTR, exon, intron, 3'-UTR, polyadenylation site, or termination sequence. As will be understood, a target site may also be located upstream or downstream of a sequence encoding a leader, enhancer, transcription initiation site, promoter, 5'-UTR, exon, intron, 3'-UTR, polyadenylation site, or termination sequence. In one embodiment, the target site is located within 10, 20, 30, 40, 50, 75, 100, 125, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 2000, 2500, 5000, 10,000, or 25,000 nucleotides of a polynucleotide encoding a leader, enhancer, transcription start site, promoter, 5'-UTR, exon, intron, 3'-UTR, polyadenylation site, gene, or termination sequence.
[0076] As used herein, "upstream" refers to a nucleic acid sequence located before the 5' end of a linked nucleic acid sequence. As used herein, "downstream" refers to a nucleic acid sequence located after the 3' end of a linked nucleic acid sequence. As used herein, "5'" refers to the beginning of a coding DNA sequence or the beginning of an RNA molecule. As used herein, "3'" refers to the end of a coding DNA sequence or the end of an RNA molecule. As will be understood, "inversion" refers to reversing the orientation of a given polynucleotide sequence. For example, if the sample sequence 5'-ATGATC-3' is inverted, it becomes the inverted 5'-CTAGTA-3'. Furthermore, the sample sequence 5'-ATGATC-3' is considered to be in the "opposite orientation" relative to the sample sequence 5'-CTAGTA-3'.
[0077] As used herein, "donor molecule" is defined as a nucleic acid sequence selected for site-directed targeted insertion into a genome. In some embodiments, the donor molecule comprises a "donor sequence." In one embodiment, the targeted editing technology provided herein comprises the use of one or more, two or more, three or more, four or more, or five or more donor molecules or donor sequences. The donor molecules or donor sequences provided herein can be of any length. For example, donor molecules or donor sequences provided herein are between 2 and 50,000, 2 and 10,000, 2 and 5000, 2 and 1000, 2 and 500, 2 and 250, 2 and 100, 2 and 50, 2 and 30, 15 and 50, 15 and 100, 15 and 500, 15 and 1000, 15 and 5000, 18 and 30, 18 and 26, 20 and 26, 20 and 50, 20 and 100, 20 and 250, 20 and 500, 20 and 1000, 20 and 5000, or 20 and 10,000 nucleotides in length. Donor molecules or donor sequences can include one or more genes encoding genetic sequences that are actively transcribed and / or translated. Such transcribed sequences may encode proteins or non-coding RNA. In one embodiment, the donor molecule or donor sequence may comprise a polynucleotide sequence that does not contain a functional gene or entire gene (i.e., the donor molecule may simply comprise a regulatory sequence such as a promoter), or may not contain any identifiable gene expression elements or any actively transcribed gene sequence. Furthermore, the donor molecule or donor sequence may be linear or circular, single-stranded or double-stranded. It may be delivered to cells as a naked nucleic acid, complexed with one or more delivery agents (e.g., liposomes, poloxamers, protein-encapsulated T-strands, etc.), or contained in a bacterial or viral delivery vehicle, such as Agrobacterium tumefaciens or geminivirus, respectively. In another aspect, the donor molecule or donor sequence provided herein is operably linked to a promoter.In yet a further embodiment, the donor molecules or donor sequences provided herein are transcribed into RNA. In another embodiment, the donor molecules or donor sequences provided herein are not operably linked to a promoter.
[0078] In some embodiments, the donor molecules or donor sequences provided herein may contain at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten genes. In some embodiments, the donor molecules or donor sequences provided herein do not contain any genes. Without limitation, the genes provided herein may include insecticide resistance genes, herbicide resistance genes, nitrogen use efficiency genes, water use efficiency genes, nutritional quality genes, DNA binding genes, selectable marker genes, RNAi constructs, site-specific genome modification enzyme genes, single guide RNAs of CRISPR / Cas9 systems, geminivirus-based expression cassettes, or plant virus expression vector systems. In one embodiment, the donor molecules or donor sequences provided herein contain a polynucleotide encoding a promoter. In another embodiment, the donor molecules or donor sequences provided herein contain a polynucleotide encoding a tissue-specific or tissue-preferred promoter. In yet another embodiment, the donor molecules or donor sequences provided herein contain a polynucleotide encoding a constitutive promoter. In another embodiment, the donor molecule or donor sequence provided herein comprises a polynucleotide encoding an inducible promoter. In another embodiment, the donor molecule or donor sequence comprises a polynucleotide encoding a structure selected from the group consisting of a leader, an enhancer, a transcription initiation site, a 5'-UTR, an exon, an intron, a 3'-UTR, a polyadenylation site, a transcription termination site, a promoter, a full-length gene, a partial gene, a gene, or a non-coding RNA. In one embodiment, the donor molecule or donor sequence provided herein comprises one or more, two or more, three or more, four or more, or five or more designed elements.
[0079] As used herein, a "donor template" may be a recombinant DNA donor template, and is defined as a nucleic acid molecule having a nucleic acid template or insertion sequence for site-directed targeted insertion or recombination into the genome of a plant cell by repairing a nick or double-stranded DNA break in the genome of the plant cell. For example, a "donor template" may be used as a template for site-directed integration of a DNA segment encoding an antisense sequence of interest, or for introducing a mutation, such as an insertion or deletion, into a target site in the genome of a plant. The targeted genome editing technology provided herein may involve the use of one or more, two or more, three or more, four or more, or five or more donor templates. A "donor template" may be a single-stranded or double-stranded DNA molecule or RNA molecule or a plasmid. The "insertion sequence" of a donor template is a sequence designed for targeted insertion into the genome of a plant cell, and it may be of any suitable length. For example, the insert sequence of the donor template may be between 2 and 50,000, between 2 and 10,000, between 2 and 5000, between 2 and 1000, between 2 and 500, between 2 and 250, between 2 and 100, between 2 and 50, between 2 and 30, between 15 and 50, between 15 and 100, between 15 and 500, between 15 and 1000, between 15 and 5000, between 15 and 1000, between 15 and 5000, between 18 and 30, between 18 and 26, between 20 and 26, between 20 and 50, between 20 and 100, between 20 and 250, between 20 and 500, between 20 and 30, between 18 ...30, between 18 and 26, between 20 and 50, between 20 and 100, between 20 and 250, between 20 and 500, between 20 and 30, between 18 and 26 The length can be between 1000, 20 and 5000, 20 and 10,000, 50 and 250, 50 and 500, 50 and 1000, 50 and 5000, 50 and 10,000, 100 and 250, 100 and 500, 100 and 1000, 100 and 5000, 100 and 10,000, 250 and 500, 250 and 1000, 250 and 5000, or 250 and 10,000 nucleotides or base pairs. The donor template also has at least one homologous sequence or arm (such as two homologous arms) to direct integration of the mutation or insertion sequence into the target site in the plant's genome via homologous recombination, where the homologous sequence or arm(s) are identical or complementary or have a percent identity or percent complementarity to a sequence at or near the target site in the plant's genome.When the donor template comprises homology arm(s) and an insert sequence, the homology arm(s) flank or surround the insert sequence of the donor template.
[0080] The donor template may be linear or circular, and may be single-stranded or double-stranded. The donor template may be delivered to a cell as a naked nucleic acid (e.g., via a particle gun), as a complex with one or more delivery agents (e.g., liposomes, proteins, poloxamers, protein-encapsulated T-strands, etc.), or contained in a bacterial or viral delivery vehicle, such as Agrobacterium tumefaciens or geminivirus, respectively. The insert sequence of the donor template or the insert sequence provided herein may include a transcribable DNA sequence or segment that can be transcribed into all or a portion of an RNA molecule, such as an antisense sequence or portion of an RNA molecule.
[0081] As used herein, a "designed element" refers to a polynucleotide that is capable of directing a desired pattern of expression of an operably linked polynucleotide. In one embodiment, the designed element comprises at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, or at least 5000 nucleotides. In another embodiment, the designed element comprises between 20 and 50 nucleotides. In yet another embodiment, the designed element comprises 10 to 5,000, 10 to 2,500, 10 to 1,000, 10 to 500, 10 to 100, 20 to 50, 20 to 100, 20 to 500, 20 to 1,000, 20 to 2,000, 20 to 5,000, 50 to 100, 50 to 500, 50 to 1,000, or 50 to 5,000 nucleotides. In one embodiment, the designed element comprises a constitutive promoter. In another embodiment, the designed element comprises an inducible promoter. In another embodiment, the designed element comprises a tissue-specific or tissue-preferred promoter. In another embodiment, the designed element comprises a native promoter. In another embodiment, the designed element comprises a non-native promoter. In another embodiment, the designed element comprises a tissue-specific or tissue-preferred promoter element. In another embodiment, the designed element comprises a transcriptional enhancer element. In another embodiment, the designed element comprises a transcriptional repressor element.
[0082] One aspect of the present application relates to methods for screening and selecting cells for targeted editing, as well as methods for selecting cells containing targeted editing. Nucleic acids can be isolated using a variety of techniques. For example, nucleic acids can be isolated using any method, including, but not limited to, recombinant nucleic acid technology and / or polymerase chain reaction (PCR). General PCR techniques are described, for example, in "PCR Primer: A Laboratory Manual," Dieffenbach & Dveksler, Eds., Cold Spring Harbor Laboratory Press, 1995. Recombinant nucleic acid technology includes, for example, restriction enzyme digestion and ligation, which can be used to isolate nucleic acids. Isolated nucleic acids can also be chemically synthesized as single nucleic acid molecules or as a series of oligonucleotides. Polypeptides can be purified from natural sources (e.g., biological samples) by known methods, such as DEAE ion exchange, gel filtration, and hydroxyapatite chromatography. Polypeptides can also be purified, for example, by expressing nucleic acids in expression vectors. Additionally, purified polypeptides can be obtained by chemical synthesis. The degree of purity of a polypeptide can be measured using any appropriate method, for example, column chromatography, polyacrylamide gel electrophoresis, or HPLC analysis.
[0083] In one aspect, the present disclosure provides methods for detecting recombinant nucleic acids and polypeptides in modified and unmodified plant cells. Nucleic acids can also be detected using, but not limited to, hybridization. Hybridization between nucleic acids is described in detail in Sambrook et al. (1989, Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY).
[0084] Antibodies may be used to detect polypeptides. Techniques for detecting polypeptides using antibodies include enzyme-linked immunosorbent assay (ELISA), Western blot, immunoprecipitation, and immunofluorescence. The antibodies provided herein may be polyclonal or monoclonal. Antibodies having specific binding affinity for the polypeptides provided herein can be generated using methods well known in the art. The antibodies provided herein can be attached to a solid support, such as a microtiter plate, using methods well known in the art.
[0085] Detection (e.g., of an amplification product, hybridization complex, or polypeptide) can be achieved using a detectable label. The term "label" is intended to encompass the use of indirect as well as direct labeling. Detectable labels include enzymes, prosthetic groups, fluorescent materials, luminescent materials, bioluminescent materials, and radioactive materials.
[0086] Screening and selection of modified, engineered, or transgenic plants or plant cells can be performed by any methodology known to those skilled in the art. Examples of screening and selection methodologies include, but are not limited to, Southern analysis, PCR amplification for detecting polynucleotides, Northern blot, RNase protection, primer extension, RT-PCR amplification for detecting RNA transcripts, Sanger sequencing, next-generation sequencing technologies (e.g., Illumina, PacBio, Ion Torrent, 454), enzyme assays for detecting enzyme or ribozyme activity of polypeptides and polynucleotides, and protein gel electrophoresis, Western blot, immunoprecipitation, and enzyme-linked immunoassay for detecting polypeptides. Other techniques, such as in situ hybridization, enzyme staining, and immunostaining, can also be used to detect the presence or expression of polypeptides and / or polynucleotides. Methods for performing all of the referenced techniques are known.
[0087] Genome editing or targeted editing can be performed by using one or more site-specific nucleases. The site-specific nuclease can induce a double-strand break (DSB) at a target site in the genomic sequence, which is then repaired by either the natural process of homologous recombination (HR) or non-homologous end joining (NHEJ). Sequence modifications such as insertions and deletions can occur at the DSB site through NHEJ repair. If two DSBs flanking one target region are created, the break can be repaired via NHEJ by reversing the orientation of the targeted DNA (also called "inversion"). HR can be used to integrate a donor nucleic acid sequence into the target site. Without being limited to any theory, to integrate a donor nucleic acid sequence (or donor molecule) into a DSB, the donor molecule comprises a polynucleotide of interest flanked by first and second homologous regions, where the first and second homologous regions are homologous to each side of the DSB at the target site. Then, the homologous recombination mechanism in cells repairs DSB by incorporating donor molecule into target site.In one embodiment, the double-strand break provided herein is repaired by NHEJ.In another embodiment, the double-strand break provided herein is repaired by HR.
[0088] Although double-stranded breaks occur only between two nucleotides on each strand, the double-stranded break site can contain at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 75, at least 80, at least 90, or at least 100 nucleotides. As used herein, "double-stranded break site" refers to a polynucleotide sequence that a site-specific nuclease or guide RNA recognizes and binds to.
[0089] In some embodiments, the vectors or constructs provided herein comprise polynucleotides encoding at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten site-specific nucleases. In another embodiment, the cells provided herein already contain site-specific nucleases. In some embodiments, the polynucleotides encoding the site-specific nucleases provided herein are stably transformed into the cells. In another embodiment, the polynucleotides encoding the site-specific nucleases provided herein are transiently transformed into the cells. In another embodiment, the polynucleotides encoding the site-specific nucleases are under the control of a regulatable promoter, a constitutive promoter, a tissue-specific promoter, or any promoter useful for expressing site-specific nucleases.
[0090] In one embodiment, the vectors comprise a cassette encoding a site-specific nuclease and a donor molecule in cis, such that the site-specific nuclease enables site-specific integration of the donor molecule upon contacting the genome of a cell. In one embodiment, a first vector comprises a cassette encoding a site-specific nuclease and a second vector comprises a donor molecule, such that the site-specific nuclease provided in trans enables site-specific integration of the donor molecule upon contacting the genome of a cell.
[0091] The site-specific nucleases provided herein can be used as part of targeted editing technology. Non-limiting examples of site-specific nucleases used in the methods and / or compositions provided herein include meganucleases, zinc finger nucleases (ZFNs), transcription activator-like effector nucleases (TALENs), RNA-guided nucleases (e.g., Cas9 and Cpf1), recombinases (including but not limited to, for example, serine recombinases linked to DNA recognition motifs, tyrosine recombinases linked to DNA recognition motifs), transposases (including but not limited to, for example, DNA transposases linked to DNA-binding domains), or any combination thereof. In one embodiment, the methods provided herein comprise the use of one or more, two or more, three or more, four or more, or five or more site-specific nucleases to induce one, two, three, four, five, or six or more DSBs at one, two, three, four, five, or six or more target sites.
[0092] In one embodiment, a site-specific nuclease protein is provided to a cell. In another embodiment, a nucleic acid sequence (e.g., a vector) encoding the site-specific nuclease protein is provided to a cell. In another embodiment, a site-specific nuclease protein and a guide RNA are provided separately to a cell. In another embodiment, a site-specific nuclease protein and a guide RNA are provided to a cell as a complex. In some embodiments, the site-specific nuclease protein and the guide RNA are assembled into a complex in vitro, in vivo, or ex vivo.
[0093] In one embodiment, a genome editing system provided herein (e.g., meganucleases, ZFNs, TALENs, CRISPR / Cas9 systems, CRISPR / Cpf1 systems, recombinases, transposases), or a combination of genome editing systems provided herein, is used in a method to introduce one or more insertions, deletions, substitutions, or inversions into a locus in a cell to generate a dominant negative allele or a dominant positive allele.
[0094] Site-specific nucleases, such as meganucleases, ZFNs, TALENs, Argonaute proteins (non-limiting examples of Argonaute proteins include Thermus thermophilus Argonaute (TtAgo), Pyrococcus furiosus Argonaute (PfAgo), Natronobacterium gregoryi Argonaute (NgAgo), homologs thereof, or modified forms thereof), Cas9 nucleases (non-limiting examples of RNA-guided nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, , Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, their homologs, or modified forms thereof) induce double-stranded DNA breaks at target sites in genomic sequences, which are then repaired by the natural process of HR or NHEJ. Sequence modifications then occur at the break site. Sequence modifications can include inversions, deletions, or insertions, which result in gene disruption in the case of NHEJ or integration of a nucleic acid sequence in the case of HR.
[0095] In some embodiments, the site-specific nucleases provided herein are selected from the group consisting of zinc finger nucleases, meganucleases, RNA-guided nucleases, TALE-nucleases, recombinases, transposases, or any combination thereof. In another embodiment, the site-specific nucleases provided herein are selected from the group consisting of Cas9 or Cpf1. In another embodiment, the site-specific nucleases provided herein are selected from the group consisting of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm11, Csm12, Csm13, Csm14, Csm15, Csm16, Csm17, Csm18, Csm19, Csm10, Csm11, Csm12, Csm19, Csm16, Csm18, Csm19, Csm19, Csm19, Csm10, Csm11, Csm12, Csm13, Csm14, Csm15, Csm16, Csm17, Csm18, Csm19 ... In another embodiment, the RNA-guided nuclease provided herein is selected from the group consisting of Cas9 or Cpf1. In another embodiment, the RNA-guided nuclease provided herein is selected from the group consisting of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm11, Csm12, Csm13, Csm14, Csm15, Csm16, Csm17, Csm18, Csm19, Csm20, Csm21, Csm22, Csm23, Csm24, Csm25, Csm26, Csm27, Csm28, Csm29, C In another embodiment, the RNA-guided nuclease is selected from the group consisting of sm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, homologs thereof, or modified forms thereof. In another embodiment, the RNA-guided nuclease is Cas9 nuclease or a homolog or modified form thereof.In one embodiment, the RNA-guided nuclease is a Cas9 protein from Streptococcus pyogenes, Streptococcus thermophilius, Staphylococcus aureus, Neisseria meningitides, or Treponema denticola, or a modified version thereof. In another embodiment, the RNA-guided nuclease is Cpf1, or a homolog or modified version thereof.
[0096] In another embodiment, the methods and / or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 site-specific nucleases. In yet another embodiment, the methods and / or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 polynucleotides encoding at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 site-specific nucleases.
[0097] In one embodiment, the targeted editing technology described herein involves the use of a recombinase. In some embodiments, the tyrosine recombinase linked to the DNA recognition motif provided herein is selected from the group consisting of Cre recombinase, Gin recombinase, Flp recombinase, and Tnp1 recombinase. In some embodiments, the Cre recombinase or Gin recombinase provided herein is tethered to a zinc finger DNA binding domain. The Flp-FRT site-directed recombination system is derived from the 2μ plasmid of baker's yeast Saccharomyces cerevisiae. In this system, Flp recombinase (flippase) recombines sequences between flippase recognition target (FRT) sites. The FRT site contains 34 nucleotides. Flp binds to the "arms" of the FRT site (one arm is inverted) and cleaves the FRT sites at both ends of the intervening nucleic acid sequence. After cleavage, Flp recombines the nucleic acid sequence between the two FRT sites. Cre-lox is a site-directed recombination system derived from bacteriophage P1 similar to the Flp-FRT recombination system. Cre-lox can be used to invert, delete, or translocate nucleic acid sequences. In this system, Cre recombinase recombines a pair of lox nucleic acid sequences. Lox sites contain 34 nucleotides, with the first and last 13 nucleotides (arms) being palindromic. During recombination, the Cre recombinase protein binds to two lox sites on different nucleic acids and cleaves them at the lox sites. The cleaved nucleic acids are then spliced together (reciprocal translocation), completing the recombination. In another embodiment, the lox sites provided herein are loxP, lox2272, loxN, lox511, lox5171, lox71, lox66, M2, M3, M7, or M11 sites.
[0098] In another embodiment, the serine recombinase linked to the DNA recognition motif provided herein is selected from the group consisting of PhiC31 integrase, R4 integrase, and TP-901 integrase. In another embodiment, the DNA transposase linked to the DNA binding domain provided herein is selected from the group consisting of TALE-piggyBac and TALE-Mutator.
[0099] In one embodiment, the targeted editing technology described herein involves the use of zinc finger nucleases (ZFNs). ZFNs are synthetic proteins consisting of engineered zinc finger DNA binding domains fused to the cleavage domain of FokI restriction nuclease. Due to the modification of the zinc finger DNA binding domain, ZFNs can be designed to cleave almost any long stretch of double-stranded DNA. ZFNs form dimers from monomers composed of the non-specific DNA cleavage domain of FokI nuclease fused to a zinc finger array engineered to bind to the target DNA sequence.
[0100] The DNA-binding domain of a ZFN typically consists of an array of three or four zinc fingers. The amino acids at positions -1, +2, +3, and +6 relative to the start of the zinc finger ∞-helix, which contribute to site-specific binding to target DNA, can be altered and customized to accommodate specific target sequences. Other amino acids form a consensus backbone to generate ZFNs with different sequence specificities. Rules for selecting target sequences for ZFNs are known in the art.
[0101] Because the FokI nuclease domain requires dimerization to cleave DNA, two ZFNs with C-terminal regions are required to bind opposite DNA strands (5-7 nt apart) at the cleavage site. If the two ZF binding sites are in a batch structure, the ZFN monomers can cleave the target site. As used herein, the term ZFN is broad and includes monomeric ZFNs that can cleave double-stranded DNA without assistance from another ZFN. The term ZFN is also used to refer to one or both members of a pair of ZFNs engineered to function together to cleave DNA at the same site.
[0102] Without being limited to any scientific theory, the DNA binding specificity of zinc finger domains can, in principle, be re-engineered using one of a variety of methods, such that customized ZFNs could theoretically be constructed to target almost any gene sequence. Publicly available methods for engineering zinc finger domains include context-dependent assembly (CoDA), oligomerization pool engineering (OPEN), and modular assembly.
[0103] In one embodiment, the methods and / or compositions provided herein comprise one or more, two or more, three or more, four or more, or five or more ZFNs. In another embodiment, the ZFNs provided herein are capable of generating targeted DSBs. In one embodiment, a vector comprising a polynucleotide encoding one or more, two or more, three or more, four or more, or five or more ZFNs is provided to a cell by transformation methods known in the art (e.g., but not limited to, viral transfection, particle bombardment, PEG-mediated protoplast transfection, or Agrobacterium-mediated transformation).
[0104] In one aspect, the targeted editing technology described herein involves the use of meganucleases. Meganucleases are unique enzymes commonly identified in microorganisms with high activity and long recognition sequences (>14 nt) that result in site-specific digestion of target DNA. Engineered versions of naturally occurring meganucleases typically have extended DNA recognition sequences (e.g., 14-40 nt). Because the DNA recognition and cleavage functions of meganucleases are intertwined in a single domain, engineering meganucleases can be more challenging than engineering ZFNs and TALENs. Specialized methods of mutagenesis and high-throughput screening have been used to generate novel meganuclease variants that recognize unique sequences and have improved nuclease activity.
[0105] In one aspect, the methods and / or compositions provided herein comprise one or more, two or more, three or more, four or more, or five or more meganucleases. In another aspect, the meganucleases provided herein are capable of generating targeted DSBs. In one aspect, a vector comprising a polynucleotide encoding one or more, two or more, three or more, four or more, or five or more meganucleases is provided to a cell by transformation methods known in the art (e.g., but not limited to, viral transfection, particle bombardment, PEG-mediated protoplast transfection, or Agrobacterium-mediated transformation).
[0106] In one embodiment, the targeted editing technology described herein involves the use of transcription activator-like effector nucleases (TALENs). TALENs are artificial restriction enzymes created by fusing a transcription activator-like effector (TALE) DNA-binding domain to a FokI nuclease domain. When each member of a TALEN pair binds to a DNA site flanking the target site, the FokI monomers dimerize, causing double-stranded DNA breaks at the target site. In addition to the wild-type FokI cleavage domain, mutant variants of the FokI cleavage domain have been designed to improve cleavage specificity and activity. The FokI domain functions as a dimer, requiring two constructs with unique DNA-binding domains for the target genome site, appropriately oriented and spaced. Both the number of amino acid residues between the TALEN DNA-binding domain and the FokI cleavage domain, and the number of bases between the two individual TALEN binding sites, are parameters for achieving high levels of activity.
[0107] TALENs are artificial restriction enzymes generated by fusing a transcription activator-like effector (TALE) DNA-binding domain to a nuclease domain. In one embodiment, the nuclease is selected from the group consisting of PvuII, MutH, TevI, and FokI, AlwI, MlyI, SbfI, SdaI, StsI, CleDORF, Clo051, and Pept071. When each member of a TALEN pair binds to a DNA site flanking the target site, the FokI monomer dimerizes, causing a double-stranded DNA break at the target site.
[0108] As used herein, the term TALEN is broad and includes monomeric TALENs that can cut double-stranded DNA without the assistance of another TALEN.The term TALEN is also used to refer to one or both members of a pair of TALENs that work together to cut DNA at the same site.
[0109] Transcription activator-like effectors (TALEs) can be engineered to bind to virtually any DNA sequence. TALE proteins are DNA-binding domains derived from various plant bacterial pathogens in the genus Xanthomonas. Xanthomonas pathogens secrete TALEs into host plant cells during infection. TALEs translocate to the nucleus, where they recognize and bind to specific DNA sequences in the promoter regions of specific genes in the host genome. TALEs have a central DNA-binding domain composed of 13 to 28 repeating monomers of 33 to 34 amino acids. The amino acids in each monomer are highly conserved except for hypervariable amino acid residues at positions 12 and 13. The two variable amino acids are called repeating variable dimers (RVDs). The amino acid pairs NI, NG, HD, and NN in the RVD preferentially recognize adenine, thymine, cytosine, and guanine / adenine, respectively, and the RVD can be adjusted to recognize stretches of DNA bases. This simple relationship between amino acid sequence and DNA recognition made it possible to engineer specific DNA-binding domains by selecting combinations of repeat segments containing appropriate RVDs.
[0110] In addition to the wild-type FokI cleavage domain, mutant variants of the FokI cleavage domain have been designed to improve cleavage specificity and activity. The FokI domain functions as a dimer, requiring two constructs with unique DNA-binding domains for the target genome site, oriented and spaced appropriately. Both the number of amino acid residues between the TALEN DNA-binding domain and the FokI cleavage domain, and the number of bases between the two individual TALEN binding sites, are parameters for achieving high levels of activity. The cleavage domains of PvuII, MutH, and TevI are useful alternatives to FokI and FokI variants for use with TALEs. PvuII functions as a highly specific cleavage domain when coupled to a TALE (see Yank et al. 2013. PLoS One. 8:e82539). MutH is capable of introducing strand-specific nicks in DNA (see Gabsalilow et al. 2013. Nucleic Acids Research. 41:e83). TevI introduces a double-strand break in DNA at the targeted site (see Beurdeley et al., 2013. Nature Communications. 4:1762).
[0111] The amino acid sequence of the TALE binding domain and its relationship to DNA recognition allows for designable proteins. Software programs such as DNA Works can be used to design TALE constructs. Other methods for designing TALE constructs are known to those skilled in the art. See Doyle et al., Nucleic Acids Research (2012) 40:W117-122. Cermak et al., Nucleic Acids Research (2011). 39:e82, and tale-nt.cac.cornell.edu / about.
[0112] In one embodiment, the methods and / or compositions provided herein comprise one or more, two or more, three or more, four or more, or five or more TALENs. In another embodiment, the TALENs provided herein are capable of generating targeted DSBs. In one embodiment, a vector comprising a polynucleotide encoding one or more, two or more, three or more, four or more, or five or more TALENs is provided to a cell by a transformation method known in the art (for example, but not limited to, viral transfection, particle gun, PEG-mediated protoplast transfection, or Agrobacterium-mediated transformation).
[0113] In one embodiment, the targeted editing technology described herein comprises the use of RNA-guided nuclease.CRISPR / Cas9 system or CRISPR / Cpf1 system is an alternative to ZFN and TALEN of FokI-based method.CRISPR system is based on RNA-guided nuclease that recognizes the DNA sequence of target site by using complementary base pairing.
[0114] In certain embodiments, the vectors provided herein contain any combination of nucleic acid sequences encoding RNA-guided nucleases (non-limiting examples of RNA-guided nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, , Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, homologs thereof, or modified forms thereof), and optionally a guide RNA required to target the respective nuclease. As used herein, the term "guide RNA" or gRNA generally refers to an RNA molecule (or collectively a group of RNA molecules) that can bind to an RNA-guided endonuclease and help target the nuclease to a specific location within a target polynucleotide (e.g., DNA).
[0115] Without being limited to any particular scientific theory, CRISPR / Cas nucleases are part of the adaptive immune system of bacteria and archaea, protecting them from invading nucleic acids such as viruses by cleaving target DNA in a sequence-dependent manner. Immunity is achieved by incorporating short fragments of invading DNA, known as spacers, between approximately 20-nucleotide-long CRISPR repeats at the proximal end of the CRISPR locus (CRISPR array). A well-described Cas protein is the Cas9 nuclease (also known as Csn1), which is part of the class 2, type II CRISPR / Cas system in Streptococcus pyogenes. See Makarova et al., Nature Reviews Microbiology (2015) doi:10.1038 / nrmicro3569. Cas9 contains a RuvC-like nuclease domain at its amino terminus and an HNH-like nuclease domain located in the center of the protein. The Cas9 protein also contains a PAM-interacting (PI) domain, a recognition lobe (REC), and a BH domain. Another type II system, the Cpf1 nuclease, acts similarly to Cas9, but Cpf1 does not require tracrRNA. See Cong et al. Science (2013) 339:819-823, Zetsche et al., Cell (2015) doi:10.1016 / j.cell.2015.09.038, U.S. Patent Publication No. 2014 / 0068797, U.S. Patent Publication No. 2014 / 0273235, U.S. Patent Publication No. 2015 / 0067922, U.S. Patent No. 8,697,359, U.S. Patent No. 8,771,945, U.S. Patent No. 8,795,965, U.S. Patent No. 8,865,406, U.S. Patent No. 8,871,445, U.S. Patent No. 8,889,356, U.S. Patent No. 8,889,418, U.S. Patent No. 8,895,308, and U.S. Patent No. 8,906,616, each of which is incorporated herein by reference in its entirety.
[0116] When Cas9 or Cpf1 cleaves the targeted DNA, the endogenous double-strand break (DSB) repair mechanism is activated. DSBs can be repaired by non-homologous end joining, which can incorporate insertions or deletions (indels) into the targeted locus. If two DSBs flanking one target region are created, the breaks can be repaired by reversing the orientation of the targeted DNA. Alternatively, if a donor polynucleotide homologous to the target DNA sequence is provided, the DSB can be repaired by homology-directed repair. This repair mechanism allows the donor polynucleotide to be precisely integrated into the targeted DNA sequence.
[0117] Without being limited to any particular scientific theory, in class 2, type II CRISPR / Cas systems, a spacer-containing CRISPR array is transcribed upon encountering recognized invasive DNA and processed into a short interfering CRISPR RNA (crRNA) approximately 40 nucleotides in length. The crRNA hybridizes with a trans-activating crRNA (tracrRNA) to activate and guide the Cas9 nuclease to the target site. The nucleic acid molecules provided herein can combine the crRNA and tracrRNA into a single nucleic acid molecule, referred to herein as a "single-stranded guide RNA (sgRNA)." A prerequisite for cleavage of the target site by Cas9 is the presence of a conserved protospacer adjacent motif (PAM) downstream of the target DNA, typically containing the sequence 5-NGG-3 but rarely NAG. Specificity is provided by a so-called "seed sequence," located approximately 12 bases upstream of the PAM, which must match between the RNA and the target DNA. Cpf1 acts similarly to Cas9, but does not require tracrRNA. Therefore, in embodiments utilizing Cpf1, the sgRNA can be replaced with a crRNA. The Cpf1 PAM motif is located upstream of the target site. Furthermore, for the Cpf1 orthologs LbCpf1 and AsCpf1, the PAM sequence is 5-TTTV-3, where V can be A, C, or G. In one embodiment, when two or more sgRNAs are provided herein, the first sgRNA and the second sgRNA are complementary to different strands of a double-stranded DNA molecule. In another embodiment, when two or more sgRNAs are provided herein, the first sgRNA and the second sgRNA are complementary to the same strand of a double-stranded DNA molecule. As used herein, "protospacer adjacent motif" (PAM) refers to a 2-6 base pair DNA sequence immediately upstream or downstream of the target sequence of a CRISPR complex. In another embodiment, the first and second gRNAs target different PAM sequences. In another embodiment, the first and second gRNAs target the same PAM sequence.
[0118] In one embodiment, the methods and / or compositions provided herein comprise one or more, two or more, three or more, four or more, or five or more Cas9 nucleases. In one embodiment, the methods and / or compositions provided herein comprise one or more polynucleotides encoding one or more, two or more, three or more, four or more, or five or more Cas9 nucleases. In another embodiment, the Cas9 nucleases provided herein are capable of generating targeted DSBs. In one embodiment, the methods and / or compositions provided herein comprise one or more, two or more, three or more, four or more, or five or more Cpf1 nucleases. In one embodiment, the methods and / or compositions provided herein comprise one or more polynucleotides encoding one or more, two or more, three or more, four or more, or five or more Cpf1 nucleases. In another embodiment, the Cpf1 nucleases provided herein are capable of generating targeted DSBs.
[0119] When Cas9 nuclease hybridizes to the target site via the sgRNA, Cas9 generates two blunt-end cuts in the double-stranded DNA. The "target strand" of the double-stranded DNA is complementary to the sgRNA, and the "non-target strand" contains a PAM motif adjacent to the cut site on the non-target strand and at its 3' end. Cas9 retains the target strand and the PAM motif, but the 3' cut end of the non-target strand is free and referred to as the "3' flap." In one embodiment, the 3' flap contains at least 10, at least 15, at least 20, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, or at least 40 nucleotides.
[0120] In one embodiment, a vector comprising a polynucleotide encoding a site-specific nuclease and optionally one or more, two or more, three or more, or four or more sgRNAs is provided to a plant cell by transformation methods known in the art (e.g., but not limited to, particle bombardment, PEG-mediated protoplast transfection, or Agrobacterium-mediated transformation). In one embodiment, a vector comprising a polynucleotide encoding a Cas9 nuclease and optionally one or more, two or more, three or more, or four or more sgRNAs is provided to a plant cell by transformation methods known in the art (e.g., but not limited to, particle bombardment, PEG-mediated protoplast transfection, or Agrobacterium-mediated transformation). In another embodiment, a vector comprising a polynucleotide encoding Cpf1 and optionally one or more, two or more, three or more, or four or more crRNAs is provided to a cell by transformation methods known in the art (e.g., but not limited to, viral transfection, particle bombardment, PEG-mediated protoplast transfection, or Agrobacterium-mediated transformation).
[0121] In one aspect, the RNA-guided nucleases provided herein include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm11, Csm12, Csm13, Csm14, Csm15, Csm16, Csm17, Csm18, Csm19, Csm20, Csm21, Csm22, Csm23, Csm24, Csm25, Csm26, Csm27, Csm28, Csm29, Csm30, Csm31, Csm32, Csm33, Csm34, Csm35, Csm36, Csm37, Csm38, Csm39, Csm40, Csm41, Csm42, Csm43, Csm44, Csm45, Csm46, Csm47, Csm48, Csm49 ...9, Csm49, Csm41, Csm42, Csm43, Csm44, Csm45, Csm46, Csm47, Csm48, C The protein is selected from the group consisting of mr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, homologs thereof, or modified forms thereof, Argonaute (non-limiting examples of Argonaute proteins include Thermus thermophilus Argonaute (TtAgo), Pyrococcus furiosus Argonaute (PfAgo), Natronobacterium gregoryi Argonaute (NgAgo), homologs thereof, or modified forms thereof), a DNA guide of an Argonaute protein, and any combination thereof. In another embodiment, the RNA-guided nuclease provided herein is selected from the group consisting of Cas9 and Cpf1. The RNA-guided nuclease provided herein comprises Cas9. In one embodiment, the RNA-guided nuclease provided herein is selected from the group consisting of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4 , Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, homologs thereof, or modified forms thereof.In one embodiment, the site-specific nuclease is Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Csel, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cm r1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, TtAgo, PfAgo, and NgAgo. In another embodiment, the RNA-guided nuclease is selected from the group consisting of Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm11, Csm12, Csm13, Csm14, Csm15, Csm16, Csm17, Csm18, Csm19, Csm20, Csm21, Csm22, Csm23, Csm24, Csm25, Csm26, Csm27, Csm28, Csm29, Csm30, Csm31, Csm32, Csm33, Csm4, Csm5, Csm6, Csm7, Csm8, Csm9, Csm10, Csm11, Csm12, Csm13, Csm14, Csm15, Csm16, Csm17, Csm18, Csm19, Csm19, Csm20, Csm21, Csm22, Csm23, Csm24, Csm25, Csm26, Csm27, Csm28, Csm29, Csm31, Csm29, Csm32, Csm40, Csm41, Csm42, Csm43, Csm44, Csm45, Csm46, Csm47, Csm48, Csm49, C selected from the group consisting of mr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, CasX, CasY, TtAgo, PfAgo, and NgAgo.
[0122] Nucleases such as Cas9 can also be engineered to form catalytically inactive forms, such as catalytically inactivated Cas9 (dCas9). dCas9 binds to DNA at the target site specified by the gRNA, creating a loop structure accessible for template-based editing (Figure 11, panel 1). dCas9 may be further modified to form a fusion with an ssDNA-binding domain to further facilitate template-based editing (Figure 11, panel 2). The editing efficiency of this modified dCas9-ssDNA binding scheme is expected to be higher than that of a dCas9-only approach, since the ssDNA template binds to the dCas9 complex and is brought into proximity with the gRNA target. As used herein, "inactivated Cas nuclease" (dCas) refers to an enzymatically inactive form of a Cas nuclease protein that can bind to DNA but cannot cleave it. In one embodiment, the nuclease provided herein is dCas. In another embodiment, the site-specific nuclease provided herein is dCas.
[0123] In one aspect, the methods and compositions provided herein can be used to edit loci in eukaryotic cells. In one aspect, the eukaryotic cells provided herein are part of a multicellular eukaryotic organism. In another aspect, the eukaryotic cells provided herein are unicellular organisms. In another aspect, the eukaryotic cells provided herein are selected from the group consisting of animal cells, plant cells, fungal cells, and protozoan cells. In one aspect, the animal cells provided herein are selected from the group consisting of insect cells, arachnid cells, arthropod cells, crustacean cells, rotifer cells, cnidarian cells, flatworm cells, mollusk cells, gastropod cells, nematode cells, annelid cells, vertebrate cells, mammalian cells, avian cells, fish cells, reptile cells, and amphibian cells. In another aspect, the plant cells provided herein are monocotyledonous or dicotyledonous plant cells. In yet another aspect, the plant cells provided herein are algae cells. In yet another aspect, the plant cell provided herein is selected from the group consisting of a corn cell, a wheat cell, a sorghum cell, a canola cell, a soybean cell, an alfalfa cell, a cotton cell, and a rice cell.In yet another aspect, the plant cells provided herein are selected from the group consisting of acacia cells, alfalfa cells, dill cells, apple cells, apricot cells, artichoke cells, yellow jasmine cells, asparagus cells, avocado cells, banana cells, barley cells, bean cells, sugar beet cells, blackberry cells, blueberry cells, broccoli cells, Brussels sprout cells, cabbage cells, rapeseed cells, cantaloupe cells, carrot cells, cassava cells, cauliflower cells, and the like. Lawrence cell, celery cell, Chinese cabbage cell, cherry cell, coriander leaf cell, citrus cell, clementine cell, coffee cell, corn cell, cotton cell, cucumber cell, Douglas fir cell, eggplant cell, endive cell, arkberry cell, eucalyptus cell, fennel cell, fig cell, forest tree cell, gourd cell, grape cell, grapefruit cell, honeydew cell, jicama cell, kiwi cell, lettuce cell, leek cell, lemon cell, lime cell, ta Spruce cell, mango cell, maple cell, melon cell, mushroom cell, nectarine cell, nut cell, oat cell, okra cell, onion cell, orange cell, ornamental plant cell, papaya cell, parsley cell, pea cell, peach cell, peanut cell, pear cell, pepper cell, persimmon cell, pine cell, pineapple cell, plantain cell, plum cell, pomegranate cell, poplar cell, potato cell, pumpkin cell, quince cell, radiata pine cell, radicchio In another aspect, the plant cell provided herein is selected from the group consisting of a corn cell, a soybean cell, a rapeseed cell, a raspberry cell, a rice cell, a rye cell, a sorghum cell, a southern pine cell, a soybean cell, a spinach cell, a gourd cell, a strawberry cell, a sugar beet cell, a sugarcane cell, a sunflower cell, a sweet corn cell, a sweet potato cell, a sweetgum cell, a tangerine cell, a tea cell, a tobacco cell, a tomato cell, a grass cell, a vine cell, a watermelon cell, a wheat cell, a yam cell, and a zucchini cell.
[0124] In yet another aspect, the engineered plants provided herein are algae. In yet another aspect, the engineered plants or seeds provided herein are selected from the group consisting of corn plants, wheat plants, sorghum plants, rapeseed plants, soybean plants, alfalfa plants, cotton plants, and rice plants. In yet another aspect, the engineered plants or seeds provided herein are selected from the group consisting of acacia plants, alfalfa plants, dill plants, apple plants, apricot plants, artichoke plants, yellow bell plants, asparagus plants, avocado plants, banana plants, barley plants, bean plants, sugar beet plants, blackberry plants, blueberry plants, broccoli plants, Brussels sprout plants, cabbage plants, rapeseed plants, cantaloupe plants, carrot plants, cassava plants, or the like. Ba plants, cauliflower plants, celery plants, Chinese cabbage plants, cherry plants, coriander plants, citrus plants, clementine plants, coffee plants, corn plants, cotton plants, cucumber plants, Douglas fir plants, eggplant plants, endive plants, esculenta plants, eucalyptus plants, fennel plants, fig plants, forest tree plants, gourd plants, grape plants, grapefruit plants, honeydew plants, jicama plants, kiwi plants, lettuce plants, leek plants, lemon plants, la Imu plants, loblolly pine plants, mango plants, maple plants, melon plants, mushroom plants, nectarine plants, nut plants, oat plants, okra plants, onion plants, orange plants, ornamental plants, papaya plants, parsley plants, pea plants, peach plants, peanut plants, pear plants, pepper plants, persimmon plants, pine plants, pineapple plants, plantain plants, plum plants, pomegranate plants, poplar plants, potato plants, pumpkin plants, quince plants, radiata pine plants, rad The plant is selected from the group consisting of an icchio plant, a radish plant, a rapeseed plant, a raspberry plant, a rice plant, a rye plant, a sorghum plant, a southern pine plant, a soybean plant, a spinach plant, a gourd plant, a strawberry plant, a sugar beet plant, a sugar cane plant, a sunflower plant, a sweet corn plant, a sweet potato plant, a sweetgum plant, a tangerine plant, a tea plant, a tobacco plant, a tomato plant, a turf plant, a vine plant, a watermelon plant, a wheat plant, a yam plant, and a zucchini plant.In another aspect, the plant provided herein is selected from the group consisting of a corn plant, a soybean plant, a rapeseed plant, a cotton plant, a wheat plant, and a sugarcane plant.
[0125] In yet another aspect, the modified plants provided herein are algae. In yet another aspect, the modified plants provided herein are selected from the group consisting of corn plants, wheat plants, sorghum plants, rapeseed plants, soybean plants, alfalfa plants, cotton plants, and rice plants. In yet another aspect, the modified plants provided herein are acacia plants, alfalfa plants, dill plants, apple plants, apricot plants, artichoke plants, yellow bell plants, asparagus plants, avocado plants, banana plants, barley plants, bean plants, sugar beet plants, blackberry plants, blueberry plants, broccoli plants, Brussels sprout plants, cabbage plants, rapeseed plants, cantaloupe plants, carrot plants, and cassava plants. , cauliflower plant, celery plant, Chinese cabbage plant, cherry plant, coriander leaf plant, citrus plant, clementine plant, coffee plant, corn plant, cotton plant, cucumber plant, Douglas fir plant, eggplant plant, endive plant, esculenta plant, eucalyptus plant, fennel plant, fig plant, forest tree plant, gourd plant, grape plant, grapefruit plant, honeydew plant, jicama plant, kiwi plant, lettuce plant, leek plant, lemon plant, lime Plants, loblolly pine plants, mango plants, maple plants, melon plants, mushroom plants, nectarine plants, nut plants, oat plants, okra plants, onion plants, orange plants, ornamental plants, papaya plants, parsley plants, pea plants, peach plants, peanut plants, pear plants, pepper plants, persimmon plants, pine plants, pineapple plants, plantain plants, plum plants, pomegranate plants, poplar plants, potato plants, pumpkin plants, quince plants, radiata pine ... The plant is selected from the group consisting of a corn plant, a radish plant, a rapeseed plant, a raspberry plant, a rice plant, a rye plant, a sorghum plant, a southern pine plant, a soybean plant, a spinach plant, a gourd plant, a strawberry plant, a sugar beet plant, a sugar cane plant, a sunflower plant, a sweet corn plant, a sweet potato plant, a sweetgum plant, a tangerine plant, a tea plant, a tobacco plant, a tomato plant, a turf plant, a vine plant, a watermelon plant, a wheat plant, a yam plant, and a zucchini plant.
[0126] In yet another aspect, the modified seeds provided herein are selected from the group consisting of corn seeds, wheat seeds, sorghum seeds, rapeseed seeds, soybean seeds, alfalfa seeds, cotton seeds, and rice seeds. In yet another aspect, the modified seeds provided herein are selected from the group consisting of acacia seeds, alfalfa seeds, dill seeds, apple seeds, apricot seeds, artichoke seeds, yellow bell seeds, asparagus seeds, avocado seeds, banana seeds, barley seeds, bean seeds, sugar beet seeds, blackberry seeds, blueberry seeds, broccoli seeds, Brussels sprout seeds, cabbage seeds, rapeseed seeds, cantaloupe seeds, carrot seeds, cassava seeds, Cauliflower seeds, celery seeds, Chinese cabbage seeds, cherry seeds, coriander leaf seeds, citrus seeds, clementine seeds, coffee seeds, corn seeds, cotton seeds, cucumber seeds, Douglas fir seeds, eggplant seeds, endive seeds, esculenta seeds, eucalyptus seeds, fennel seeds, fig seeds, forest tree seeds, gourd seeds, grape seeds, grapefruit seeds, honeydew seeds, jicama seeds, kiwi seeds, lettuce seeds, leek seeds, lemon seeds, lime seeds , loblolly pine seeds, mango seeds, maple seeds, melon seeds, mushroom seeds, nectarine seeds, nut seeds, oat seeds, okra seeds, onion seeds, orange seeds, ornamental plant seeds, papaya seeds, parsley seeds, pea seeds, peach seeds, peanut seeds, pear seeds, pepper seeds, persimmon seeds, pine seeds, pineapple seeds, plantain seeds, plum seeds, pomegranate seeds, poplar seeds, potato seeds, pumpkin seeds, quince seeds, radiata pine seeds, radish The seed is selected from the group consisting of chio seeds, radish seeds, rapeseed seeds, raspberry seeds, rice seeds, rye seeds, sorghum seeds, southern pine seeds, soybean seeds, spinach seeds, gourd seeds, strawberry seeds, sugar beet seeds, sugarcane seeds, sunflower seeds, sweet corn seeds, sweet potato seeds, sweetgum seeds, tangerine seeds, tea seeds, tobacco seeds, tomato seeds, turf seeds, vine seeds, watermelon seeds, wheat seeds, yam seeds, and zucchini seeds.
[0127] In yet another aspect, the modified chromosomes provided herein are from algae. In yet another aspect, the modified chromosomes provided herein are selected from the group consisting of a maize chromosome, a wheat chromosome, a sorghum chromosome, a rapeseed chromosome, a soybean chromosome, an alfalfa chromosome, a cotton chromosome, and a rice chromosome.In yet another aspect, the modified chromosome provided herein is an acacia chromosome, an alfalfa chromosome, a dill chromosome, an apple chromosome, an apricot chromosome, an artichoke chromosome, a yellow bell chromosome, an asparagus chromosome, an avocado chromosome, a banana chromosome, a barley chromosome, a bean chromosome, a sugar beet chromosome, a blackberry chromosome, a blueberry chromosome, a broccoli chromosome, a Brussels sprout chromosome, a cabbage chromosome, a rapeseed chromosome, a cantaloupe chromosome, a carrot chromosome, a cassava chromosome, a cauliflower chromosome, a pea ... - chromosome, celery chromosome, Chinese cabbage chromosome, cherry chromosome, coriander leaf chromosome, citrus chromosome, clementine chromosome, coffee chromosome, corn chromosome, cotton chromosome, cucumber chromosome, Douglas fir chromosome, eggplant chromosome, endive chromosome, kerria chromosome, eucalyptus chromosome, fennel chromosome, fig chromosome, forest tree chromosome, gourd chromosome, grape chromosome, grapefruit chromosome, honeydew chromosome, jicama chromosome, kiwi chromosome, lettuce chromosome, leek chromosome, lemon chromosome, lime chromosome, taeda Pine chromosome, mango chromosome, maple chromosome, melon chromosome, mushroom chromosome, nectarine chromosome, nut chromosome, oat chromosome, okra chromosome, onion chromosome, orange chromosome, plant chromosome, papaya chromosome, parsley chromosome, pea chromosome, peach chromosome, peanut chromosome, pear chromosome, pepper chromosome, persimmon chromosome, pine chromosome, pineapple chromosome, plantain chromosome, plum chromosome, pomegranate chromosome, poplar chromosome, potato chromosome, pumpkin chromosome, quince chromosome, radiata pine chromosome, radish The chromosome is selected from the group consisting of a chio chromosome, a radish chromosome, a rapeseed chromosome, a raspberry chromosome, a rice chromosome, a rye chromosome, a sorghum chromosome, a southern pine chromosome, a soybean chromosome, a spinach chromosome, a gourd chromosome, a strawberry chromosome, a sugar beet chromosome, a sugarcane chromosome, a sunflower chromosome, a sweet corn chromosome, a sweet potato chromosome, a sweetgum chromosome, a tangerine chromosome, a tea chromosome, a tobacco chromosome, a tomato chromosome, a grass chromosome, a liana chromosome, a watermelon chromosome, a wheat chromosome, a yam chromosome, and a zucchini chromosome.
[0128] In one aspect, the cells provided herein are modified cells. In another aspect, the plants provided herein are modified plants. In yet another aspect, the plant cells provided herein are modified plant cells. In yet another aspect, the seeds provided herein are modified seeds. In a further aspect, the chromosomes provided herein are modified chromosomes.
[0129] According to one embodiment, the modified plants, plant cells, cells, seeds, or chromosomes provided herein comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten deletions generated by targeted editing technology. According to one embodiment, the modified plants, plant cells, cells, seeds, or chromosomes provided herein comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten insertions generated by targeted editing technology. According to one embodiment, the modified plants, plant cells, cells, seeds, or chromosomes provided herein comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten inversions generated by targeted editing technology. According to one aspect, the modified plants, plant cells, cells, seeds, or chromosomes provided herein include at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 deletions generated by targeted editing techniques; at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 insertions generated by targeted editing techniques; at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 inversions generated by targeted editing techniques, or any combination thereof. In yet another embodiment, the modified plants, plant cells, cells, seeds, or chromosomes provided herein contain at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 dominant negative alleles generated by targeted editing techniques.In yet another embodiment, the modified plants, plant cells, cells, seeds, or chromosomes provided herein comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten dominant positive alleles generated by targeted editing technology. In yet another embodiment, the modified plants, plant cells, cells, seeds, or chromosomes provided herein comprise at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten dominant negative alleles generated by targeted editing technology, at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten dominant positive alleles generated by targeted editing technology, or any combination thereof.
[0130] According to another aspect of the present application, there are provided modified plant(s), plant cell(s), seed(s), chromosome(s), and plant part(s) that comprise a genome editing event in the genome of at least one plant cell, the genome editing event comprising an insertion, deletion, substitution, or inversion of a targeted locus.
[0131] In one aspect, the present disclosure provides a modified plant cell produced by any one of the methods provided herein. In another aspect, the present disclosure provides a modified chromosome produced by any one of the methods provided herein. In yet another aspect, the present disclosure provides a modified cell comprising a modified chromosome provided herein. In yet a further aspect, the present disclosure provides a modified plant or modified plant tissue regenerated from a modified cell provided herein. In yet another aspect, the present disclosure provides a product comprising a modified chromosome provided herein. In an aspect, the present disclosure provides a product comprising a modified cell provided herein. As used herein, "product" refers to any article or substance intended for human use, human consumption, animal use, or animal consumption, including any component, part, or accessory comprising a modified cell or modified chromosome provided herein.
[0132] The methods and compositions provided herein can edit any locus within a genome. Also provided herein are chromosomes edited using the methods and compositions provided herein. In one embodiment, the genome provided herein is a nuclear genome, a mitochondrial genome, or a plastid genome. In another embodiment, the plastid genome provided herein comprises a chloroplast genome. In one embodiment, the method provided herein generates a double-strand break in a chromosome. In one embodiment, the chromosome provided herein is a nuclear chromosome, a mitochondrial chromosome, or a chloroplast chromosome. In another embodiment, the chromosome provided herein is an extra chromosome or an artificial chromosome. An extra chromosome, or B chromosome, is an extra chromosome found in addition to the normal diploid complement of chromosomes in a cell. The extra chromosome is dispensable and is not required for the normal development of a cell or organism. In one embodiment, the extra chromosome provided herein is a maize extra chromosome or a rye extra chromosome.
[0133] The targeted editing methods disclosed herein may involve transient transfection or stable transformation of cells of interest (e.g., plant cells). According to one aspect of the present application, a method is provided for transforming a cell, tissue, or explant with a recombinant DNA molecule or construct comprising a transcribable DNA sequence or transgene operably linked to a promoter to produce a transgenic or genome-edited cell. According to another aspect of the present application, a method is provided for transforming a plant cell, tissue, or explant with a recombinant DNA molecule or construct comprising a transcribable DNA sequence or transgene operably linked to a plant-expressible promoter to produce a transgenic or genome-edited plant or plant cell. As used herein, "transgene" refers to a polynucleotide introduced into a genome by any method known in the art.
[0134] Numerous methods for transforming chromosomes or plastids in plant cells with recombinant DNA molecules or constructs are known in the art and can be used in accordance with the methods of the present application to produce transgenic plant cells and plants. Any suitable method or technique for transforming plant cells known in the art can be used in accordance with the methods of the present invention. Effective methods for plant transformation include bacterial-mediated transformation, such as Agrobacterium- or Rhizobium-mediated transformation, and biolistic transformation. Various methods are known in the art for transforming explants with transformation vectors via bacterial-mediated transformation or biolistic transformation, followed by culturing these explants to regenerate or develop transgenic plants. Other plant transformation methods, such as microinjection, electroporation, vacuum infiltration, pressure, sonication, silicon carbide fiber agitation, and PEG-mediated transformation, are also known in the art. Transgenic plants produced by these transformation methods may be chimeric or non-chimeric with respect to the transformation event, depending on the method and explant used.
[0135] Methods for transforming plant cells are well known to those skilled in the art. For example, specific instructions for transforming plant cells by particle bombardment using particles coated with recombinant DNA can be found in U.S. Patent Nos. 5,550,318, 5,538,880, 6,160,208, 6,399,861 and 6,153,812, and Agrobacterium-mediated transformation can be found in U.S. Patent Nos. 5,159,135, 5,824,877, 5,591,616, 6,384,301, 5,750,871, 5,463,174 and 5,188,958, all of which are incorporated herein by reference. Further methods for transforming plants can be found, for example, in Compendium of Transgenic Crop Plants (2009) Blackwell Publishing. Any suitable method known to those of skill in the art can be used to transform plant cells with any of the nucleic acid molecules provided herein.
[0136] Recipient cells or explant targets for transformation include, but are not limited to, seed cells, fruit cells, leaf cells, cotyledon cells, hypocotyl cells, meristem cells, embryo cells, endosperm cells, root cells, shoot cells, stem cells, pod cells, flower cells, inflorescence cells, stem cells, base cells, style cells, stigma cells, receptacle cells, petal cells, sepal cells, pollen cells, anther cells, floral cells, ovary cells, ovule cells, pericarp cells, phloem cells, bud cells, or vascular tissue cells. In another aspect, the present disclosure provides plant chloroplasts. In a further aspect, the present disclosure provides epidermal cells, stomatal cells, trichome cells, root hair cells, storage root cells, or tuber cells. In another aspect, the present disclosure provides protoplasts. In another aspect, the present disclosure provides plant callus cells. Any cell from which a fertile plant can be regenerated is considered useful as a recipient cell in the practice of the present disclosure. Callus can be initiated from a variety of tissue sources, including, but not limited to, immature embryos or embryo parts, seedling apical meristems, microspores, etc. These cells capable of growing as callus can serve as recipient cells for transformation. Actual transformation methods and materials for producing the transgenic plants of the present disclosure (e.g., various media and recipient target cells, transformation of immature embryos, and subsequent regeneration of fertile transgenic plants) are disclosed, for example, in U.S. Pat. Nos. 6,194,636 and 6,232,526 and U.S. Patent Application Publication No. 2004 / 0216189, all of which are incorporated herein by reference. Transformed explants, cells, or tissues may be subjected to additional culture steps, such as callus induction, selection, regeneration, etc., as known in the art. Transformed cells, tissues, or explants containing recombinant DNA inserts can be grown, developed, or regenerated into transgenic plants in culture, plugs, or soil according to methods known in the art. In one aspect, the present disclosure provides plant cells that are not reproductive material and do not mediate the natural reproduction of plants. In another aspect, the present disclosure also provides plant cells that are reproductive material and mediate the natural reproduction of plants. In another aspect, the present disclosure provides plant cells that are incapable of sustaining themselves through photosynthesis.In another aspect, the present disclosure provides somatic plant cells. Somatic cells, unlike reproductive cells, do not mediate plant reproduction. In one aspect, the present disclosure provides non-reproductive plant cells.
[0137] The modified plant may be further crossed with itself or with other plants to produce modified seeds and progeny. Modified plants may also be prepared by crossing a first plant containing a recombinant DNA sequence insertion with a second plant lacking the insertion. For example, a recombinant DNA sequence can be introduced into a first plant line suitable for transformation, which can then be crossed with a second plant line to introgress the recombinant DNA sequence into the second plant line. Modified plants may also be prepared by crossing the modified plant with an unmodified plant. The progeny of these crosses may be further backcrossed to a more desirable line for 6-8 generations or multiple times, such as backcrossing, to produce progeny plants having substantially the same genotype as the original parent line but for the introduction of the recombinant DNA construct or modified sequence.
[0138] The modified plants, cells, or explants provided herein may be of elite varieties or lines. An elite variety or line refers to any variety resulting from breeding and selection for superior agronomic performance. The modified plants, cells, or explants provided herein may be hybrid plants, cells, or explants. As used herein, a "hybrid" is produced by crossing two plants from different varieties, lines, or species such that the offspring contain genetic material from each parent. Those skilled in the art will recognize that higher-order hybrids may also be produced. For example, a first hybrid can be produced by crossing variety C with variety D to produce a C×D hybrid, and a second hybrid can be produced by crossing variety E with variety F to produce an E×F hybrid. The first and second hybrids may be further crossed to produce a higher-order hybrid (C×D)×(E×F) containing the genetic information of all four parent varieties. The modified plants provided herein are fertile. The modified plants provided herein are male or female sterile modified plants and are unable to reproduce without human intervention. In one aspect, the modified plants provided herein reproduce asexually or via vegetative reproduction. In yet another aspect, the modified plants provided herein reproduce via sexual reproduction.
[0139] The recombinant DNA molecules or constructs of the present application may comprise or be contained within a DNA transformation vector for use in transforming target plant cells, tissues, or explants. Such transformation vectors of the present application may generally comprise sequences or elements necessary or beneficial for effective transformation, in addition to at least one selectable marker gene, at least one expression cassette, and / or a transcribable DNA sequence encoding one or more site-specific nucleases, and optionally one or more sgRNAs or crRNAs. In the case of Agrobacterium-mediated transformation, the transformation vector may comprise an engineered transfer DNA (or T-DNA) segment or region having at least two border sequences, a left border (LB) and a right border (RB), flanking the transcribable DNA sequence or transgene, such that insertion of the T-DNA into the plant genome results in a transformation event of the transcribable DNA sequence, transgene, or expression cassette. In other words, the transgene, transcribable DNA sequence, transgene, or expression cassette encoding site-specific nuclease(s), and / or sgRNA(s) or crRNA(s) will be located between the left and right borders of the T-DNA, possibly along with additional transgene(s) or expression cassette(s), such as a plant selectable marker transgene and / or other gene(s) of agronomic interest that can confer a trait or phenotype of agronomic interest to the plant. According to alternative embodiments, the transcribable DNA sequence, transgene, or expression cassette encoding at least one site-specific nuclease, any necessary sgRNA(s) or crRNA(s), and the plant selectable marker transgene (or other gene(s) of agronomic interest) may be present in separate T-DNA segments of the same or different recombinant DNA molecule(s), e.g., for co-transformation. The transformation vector or construct may further comprise prokaryotic maintenance elements, which, in the case of Agrobacterium-mediated transformation, may be located within the vector backbone outside the T-DNA region(s).
[0140] If the plant selectable marker transgene confers tolerance or resistance to a selective agent, the plant selectable marker transgene in the transformation vector or construct of the present application can be used to support the selection of transformed cells or tissues by the presence of the selective agent, such as an antibiotic or herbicide. Thus, the selective agent can bias or favor the survival, development, growth, proliferation, etc., of transformed cells expressing the plant selectable marker gene, e.g., increasing the proportion of transformed cells or tissues in an R0 plant. Commonly used plant selectable marker genes include those that confer tolerance or resistance to antibiotics, such as kanamycin and paromomycin (nptII), hygromycin B (aph IV), streptomycin or spectinomycin (aadA), and gentamicin (aac3 and aacC4), or those that confer tolerance or resistance to herbicides, such as glufosinate (bar or pat), dicamba (DMO), and glyphosate (aroA or Cp4-EPSPS). Plant screenable marker genes, such as luciferase or green fluorescent protein (GFP), which allow for visual screening of transformants, or genes expressing beta-glucuronidase or the uidA gene (GUS), for which various chromogenic substrates are known, may be used. In one aspect, the vectors or polynucleotides provided herein contain at least one marker gene selected from the group consisting of nptII, aph IV, aadA, aac3, aacC4, bar, pat, DMO, EPSPS, aroA, GFP, and GUS.
[0141] According to certain embodiments of the present application, methods for transforming plant cells, tissues, or explants with recombinant DNA molecules or constructs may further include site-directed or targeted integration using site-specific nucleases. These methods allow a portion of a recombinant DNA donor molecule (i.e., an insertion sequence) to be inserted or integrated at a desired site or locus within a genome. The insertion sequence of the donor template may include a transgene or construct, such as a designed element or a tissue-specific promoter. The donor molecule may also have one or two homologous arms flanking the insertion sequence to enhance targeted insertion events via homologous recombination and / or homology-directed repair. Thus, the recombinant DNA molecules of the present application may further include a donor template for site-directed or targeted integration of a transgene or construct, such as a transgene or transcribable DNA sequence encoding a designed element or a tissue-specific promoter, into a genome.
[0142] As used herein, a "portion" of a nucleic acid sequence or molecule refers to any number of nucleotides less than the full length of the nucleic acid sequence. For example, a portion of a 100-nucleotide nucleic acid sequence can be any number of nucleotides from 1 to 99 nucleotides. Alternatively, a "portion" of a nucleic acid sequence can refer to any range from 0.01% to 99.99% of the full length of a given nucleic acid sequence.
[0143] Provided herein are methods for generating dominant alleles of gene regions using targeted editing technology. Also provided herein are cells produced by such methods and compositions used in such methods. Further provided herein are modified plants regenerated from cells subjected to the methods provided herein. In one aspect, the dominant negative alleles provided herein are capable of suppressing transcription of a locus or gene in a heterozygous state. In another aspect, the dominant negative alleles provided herein are capable of suppressing transcription of a locus or gene in a homozygous state.
[0144] A dominant-negative allele of a gene region can reduce or eliminate the function of the gene region product in the heterozygous state. Without limitation, a dominant-negative allele can be generated by editing an allele of a gene region such that the orientation of at least a portion of the polynucleotide encoding the gene region is reversed (e.g., a portion of the gene is flipped to a 3' to 5' orientation, while the remainder of the gene remains in a 5' to 3' orientation). Expression of the edited allele of the gene region will include an antisense RNA segment complementary to the sense RNA expressed by the unedited gene region. Without being bound by any scientific theory, processing of the complementary segment between the sense and antisense portions of the gene region RNA by the cell's intrinsic RNA silencing mechanism can reduce expression of the edited and unedited gene region alleles in a dominant-negative manner. In one embodiment, the antisense RNA transcripts provided herein are capable of suppressing complementary sense RNA transcripts. In another embodiment, the antisense RNA transcripts provided herein suppress complementary sense RNA transcripts.
[0145] In one embodiment, the present disclosure provides a method for generating a dominant-negative allele of a gene in a cell, comprising using targeted editing technology to invert a portion of the gene to generate an antisense RNA transcript that can induce the suppression of the unmodified allele of the gene.In one embodiment, the expression of the unmodified allele is reduced compared to a control cell that does not contain the antisense RNA transcript.In another embodiment, the targeted editing technology provided herein comprises the use of at least one site-specific nuclease.In one embodiment, the antisense RNA transcript provided herein is a partial antisense RNA transcript.In another embodiment, the antisense RNA transcript provided herein is a complete antisense RNA transcript.Without being limited thereto, a partial antisense RNA transcript can be generated by inverting only one region of a gene, rather than inverting the entire gene.For example, if an mRNA transcript is coded by three exons, only the second exon can be inverted to generate a partial antisense RNA transcript. As will be understood, inverting any number of nucleotides of a gene region that is less than the entire length of the gene region can produce a partial antisense RNA transcript. For example, if a gene region contains 500 nucleotides, inverting a 200-nucleotide region will produce a partial antisense RNA transcript. Inverting all 500 nucleotides will produce a complete antisense RNA transcript. In some embodiments, the antisense RNA transcripts provided herein are capable of suppressing the expression of complementary nucleic acid sequences. In some embodiments, the antisense RNA transcripts provided herein are capable of suppressing the expression of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 genes. In some embodiments, the antisense RNA transcripts provided herein are capable of suppressing the expression of a first gene region. In some embodiments, the antisense RNA transcripts provided herein are capable of suppressing the expression of a protein encoded by a complementary nucleic acid sequence.As will be appreciated by those skilled in the art, 100% complementarity between the antisense RNA transcript and the second nucleic acid is not required to induce the suppressed expression of the second nucleic acid.For example, an antisense RNA transcript that comprises at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5% or 100% complementarity to the second nucleic acid sequence may be able to suppress the expression of the second nucleic acid sequence.
[0146] In one embodiment, the antisense RNA transcript transcribed by the dominant negative allele provided herein can down-regulate its own expression.In another embodiment, the antisense RNA transcript transcribed by the dominant negative allele provided herein can down-regulate the expression of the unmodified allele at the same locus.In one embodiment, the expression of unmodified allele is reduced compared with the control cell that does not contain antisense RNA transcript.
[0147] In one aspect, the disclosure provides a method of generating a dominant negative allele of a gene in one or more cells, the method comprising: a) inducing a first double-stranded break and a second double-stranded break flanking a targeted region of the gene; b) identifying one or more cells containing an inversion of the targeted region of the gene, wherein the inversion results in the production of an antisense RNA transcript from the targeted region of the gene; and c) selecting one or more cells containing an inversion of the targeted region of the gene.
[0148] In another aspect, the disclosure provides a method for reducing expression of a protein in a cell, the method comprising: a) inducing a first double-stranded break and a second double-stranded break flanking a targeted region of a chromosome; and b) identifying one or more cells that contain an inversion in the targeted region of the chromosome, wherein expression of the protein is reduced compared to a control cell that does not contain an inversion within the targeted region.
[0149] In a further aspect, the present disclosure provides a method for generating an inversion in a targeted region of a gene, the method comprising: a) providing at least one RNA-guided nuclease, or one or more vectors encoding the RNA-guided nuclease, to one or more cells, wherein the at least one RNA-guided nuclease is capable of binding to at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 consecutive nucleotides of a first target site and a second target site flanking the targeted region of the gene, wherein the first target site and the second target site are linked, and wherein the at least one RNA-guided nuclease creates a double-stranded break at the first target site and the second target site of the gene; b) identifying one or more cells comprising an inversion in the targeted region of the gene, wherein the inversion results in the production of an antisense RNA transcript from the targeted region; and c) selecting one or more cells comprising an inversion in the targeted region of the gene.
[0150] In one embodiment, the methods or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 site-specific nucleases. In another embodiment, the methods or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 sgRNAs. In a further embodiment, the methods or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 tgOligos. In another embodiment, the methods or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 donor molecules. In another embodiment, the methods or compositions provided herein comprise at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 donor sequences.
[0151] In another embodiment, a method or composition provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 6, at least 7, at least 8, at least 9, or at least 10 vectors encoding at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 site-specific nucleases. In another embodiment, a method or composition provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 vectors encoding at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 sgRNAs. In another embodiment, a method or composition provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 6, at least 7, at least 8, at least 9, or at least 10 vectors encoding at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 tgOligos. In another embodiment, a method or composition provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 vectors encoding at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 donor molecules.In another embodiment, a method or composition provided herein comprises at least 1, at least 2, at least 3, at least 4, at least 6, at least 7, at least 8, at least 9, or at least 10 vectors encoding at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 donor sequences.
[0152] In one embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five site-specific nucleases. In one embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five RNA-guided nucleases. In another embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five sgRNAs. In one embodiment, the methods or compositions provided herein include one or more vectors comprising a first sgRNA and a second sgRNA. In a further embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five donor molecules.
[0153] In certain embodiments, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five site-specific nucleases and at least one, at least two, at least three, at least four, or at least five sgRNAs. In another embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five RNA-guided nucleases and at least one, at least two, at least three, at least four, or at least five sgRNAs. In certain embodiments, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five sgRNAs and at least one, at least two, at least three, at least four, or at least five donor molecules.
[0154] In another embodiment, the vector provided herein encodes at least one RNA-guided nuclease, a first sgRNA, and a second sgRNA. In a further embodiment, the at least one RNA-guided nuclease, a first sgRNA, and a second sgRNA are encoded by two or more or three or more vectors. In another embodiment, the vector provided herein encodes at least one RNA-guided nuclease, an sgRNA, and a donor molecule. In a further embodiment, the at least one RNA-guided nuclease, an sgRNA, and a donor molecule are encoded by two or more or three or more vectors.
[0155] In another embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five site-specific nucleases and at least one, at least two, at least three, at least four, or at least five donor molecules. In one embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five RNA-guided nucleases and at least one, at least two, at least three, at least four, or at least five donor molecules. In another embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five site-specific nucleases, at least one, at least two, at least three, at least four, or at least five sgRNAs, and at least one, at least two, at least three, at least four, or at least five donor molecules. In another embodiment, the vectors provided herein encode at least one, at least two, at least three, at least four, or at least five RNA-guided nucleases, at least one, at least two, at least three, at least four, or at least five sgRNAs, and at least one, at least two, at least three, at least four, or at least five donor molecules. In another embodiment, the vectors provided herein encode at least one site-specific nuclease, at least one donor molecule, and at least one sgRNA. In another embodiment, the vectors provided herein encode at least one RNA-guided nuclease, at least one donor molecule, and at least one sgRNA.
[0156] In one embodiment, one or more site-specific nucleases, one or more sgRNAs, and one or more donor molecules provided herein are encoded by a single vector. In one embodiment, one or more site-specific nucleases, one or more sgRNAs, and one or more donor molecules provided herein are encoded by two or more or three or more vectors. In yet another embodiment, one or more sgRNAs and one or more donor molecules provided herein are encoded by two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more vectors. In one embodiment, at least one RNA-guided nuclease, a first sgRNA, and a second sgRNA are encoded by a single vector. In another embodiment, at least one RNA-guided nuclease, a first sgRNA, and a second sgRNA are encoded by two or more or three or more vectors. In one embodiment, at least one site-specific nuclease, a first sgRNA, and a second sgRNA are encoded by a single vector. In another embodiment, at least one site-specific nuclease, a first sgRNA, and a second sgRNA are encoded by two or more or three or more vectors. In one embodiment, at least one RNA-guided nuclease, at least one sgRNA, and at least one donor molecule are encoded by a single vector. In one embodiment, at least one RNA-guided nuclease, at least one sgRNA, and at least one donor molecule are encoded by two or more or three or more vectors. In one embodiment, at least one site-specific nuclease, at least one sgRNA, and at least one donor molecule are encoded by a single vector. In one embodiment, at least one site-specific nuclease, at least one sgRNA, and at least one donor molecule are encoded by two or more or three or more vectors.
[0157] In one embodiment, one or more Cas9 nucleases, one or more sgRNAs, and one or more donor molecules provided herein are encoded by a single vector. In one embodiment, one or more Cas9 nucleases, one or more sgRNAs, and one or more donor molecules provided herein are encoded by two or more or three or more vectors. In one embodiment, at least one Cas9 nuclease, a first sgRNA, and a second sgRNA are encoded by a single vector. In another embodiment, at least one Cas9 nuclease, a first sgRNA, and a second sgRNA are encoded by two or more or three or more vectors. In one embodiment, at least one Cas9 nuclease, at least one sgRNA, and at least one donor molecule are encoded by a single vector. In one embodiment, at least one Cas9 nuclease, at least one sgRNA, and at least one donor molecule are encoded by two or more or three or more vectors.
[0158] In yet another embodiment, any vector described herein further encodes at least one, at least two, at least three, at least four, or at least five marker genes. In one embodiment, the marker genes provided herein are selected from the group consisting of nptII, aph IV, aadA, aac3, aacC4, bar, pat, DMO, EPSPS, aroA, GFP, and GUS.
[0159] Using targeted editing technology, a genomic locus can be converted into a locus capable of generating an RNAi-inducing hairpin when the edited locus is transcribed into RNA. In cells heterozygous at a locus of interest (e.g., two polymorphic alleles are present), two or more nucleases are used to generate two double-stranded breaks (e.g., a first double-stranded break and a second double-stranded break) in the first allele and one double-stranded break (e.g., a third double-stranded break) in the second allele. When the nucleases cleave the alleles, portions of the first allele flanking the first and second double-stranded breaks are released from the genomic DNA. In one result, the released portion of the first allele is inverted and incorporated into the third double-stranded break in the second allele, thereby creating an edited locus capable of generating an RNAi-inducing hairpin when the edited locus is transcribed.
[0160] The present disclosure provides a method comprising: a) using a targeted editing technique to create a first double-stranded break and a second double-stranded break in a first allele of a gene in a cell; and using a targeted editing technique to create a third double-stranded break in a second allele of the gene in the cell; and c) identifying a cell containing an insertion of a region of the first allele in the opposite orientation at the site of the third double-stranded break in the second allele, thereby creating a modified second allele. In one embodiment, the modified second allele is a dominant negative allele. In another embodiment, the modified second allele is a dominant positive allele. In some embodiments, the first double-stranded break and the second double-stranded break are at the same nucleotide sequence or the same nucleotide position in the first allele and the second allele. In some embodiments, the first double-stranded break and the second double-stranded break are at the same nucleotide sequence in the first allele and the second allele. In some embodiments, the first double-stranded break and the second double-stranded break are at the same nucleotide position in the first allele and the second allele. In some embodiments, the nucleotide sequence of the first allele is not identical to the nucleotide sequence of the second allele (e.g., the cell is heterozygous for the locus). In some embodiments, the nucleotide sequence of the third double-stranded break site in the second allele is not present in the first allele. In some embodiments, the modified second allele provided herein transcribes RNA that can form a hairpin loop secondary structure. In some embodiments, the region of the first allele can include any number of nucleotides, including up to the full length of the first allele.In an embodiment, the region of the first allele is at least 10, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 40, at least 50, at least At least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 3000, at least 4000, at least 5000, or at least 10,000 nucleotides. In another embodiment, the region of the first allele comprises 18 to 5,000, 18 to 4,000, 18 to 3,000, 18 to 2,000, 18 to 1,000, 18 to 500, 18 to 400, 18 to 300, 18 to 200, 18 to 100, 18 to 50, 18 to 30, 50 to 500, 50 to 1,000, 100 to 500, 100 to 1,000, or 500 to 5,000 nucleotides.
[0161] In one aspect, the disclosure provides a modified cell comprising at least one dominant negative allele of at least one gene generated by targeted editing technology, wherein when the at least one dominant negative allele is transcribed, the allele produces an RNA transcript capable of forming a hairpin loop secondary structure.
[0162] The present disclosure provides a method for generating a dominant-negative allele of a gene in a cell, comprising using targeted editing technology to insert an inverted copy of the gene or a portion thereof adjacent to a native copy of the gene to generate an inverted repeat sequence capable of producing an antisense RNA transcript of the gene or a portion thereof. In one embodiment, the inverted repeat sequence is capable of forming a hairpin loop secondary structure. In another embodiment, the dominant-negative allele generates at least one RNA transcript capable of forming a hairpin loop secondary structure. In one embodiment, the inverted copy of the gene and the native copy of the gene are separated by a spacer sequence. In some embodiments, the spacer sequence comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 250, at least 500, or at least 1000 nucleotides. In yet another embodiment, the dominant-negative allele is operably linked to a promoter of the native copy of the gene.
[0163] In another aspect, the disclosure provides modified cells comprising a dominant negative allele of at least one gene, comprising an inverted copy of the gene adjacent to a native copy of the gene at the endogenous locus of the gene.
[0164] Dominant-negative alleles can also be generated by deleting a region between a first gene region and a second adjacent gene region of DNA, editing the genome so that the first gene region and the second gene region are in opposite orientations on the chromosome (e.g., the first gene region is in a 5' to 3' orientation, and the second gene region is in a 3' to 5' orientation on the same DNA strand of the chromosome). Without being bound by any scientific theory, such deletion allows the promoter of the second gene region to express an antisense RNA transcript of the first gene region, while the native promoter of the first gene region expresses a sense RNA transcript. The sense and antisense RNA transcripts are complementary to each other, and when processed by the cell's intrinsic RNA silencing mechanism, they can reduce the expression of the edited and unedited first gene region alleles in a dominant-negative manner. Without being bound by any scientific theory, it is also believed that antisense RNA molecules transcribed from mutant or edited alleles of endogenous genes or loci may affect the expression level(s) of genes through a variety of mechanisms, including nonsense-mediated decay, nonstop decay, no-go decay, DNA or histone methylation or other epigenetic changes, inhibiting or reducing the efficiency of transcription and / or translation, ribosomal interference, interference with mRNA processing or splicing, and / or ubiquitin-mediated protein degradation by the proteasome.See, for example, Nickless, A. et al., "Control of gene expression through the nonsense-mediated RNA decay pathway," Cell Biosci 7:26 (2017); Karamyshev, A. et al., "Lost in Translation: Ribosome-Associated mRNA and Protein Quality Controls," Frontiers in Genetics 9:431 (2018); Inada, T., "Quality controls induced by aberrant translation," Nucleic Acids Res 48:3 (2020); and Szadeczky-Kardoss, I. et al., "The nonstop decay and the RNA silencing systems operate cooperatively in plants," Nucleic Acids Res 46:9 (2018), the contents and disclosures of which are incorporated herein by reference in their entirety. These different mechanisms may act instead of, or in addition to, RNA interference (RNAi), transcriptional gene silencing (PGS), and / or post-transcriptional gene silencing (PTGS) mechanisms. See, for example, Wilson, R.C. et al., "Molecular Mechanisms of RNA Interference," Annu Rev Biophysics 42:217-39 (2013), and Guo, Q. et al., "RNA Silencing in Plants: Mechanism, Technologies, and Applications in Horticulture Crops," Current Genomics 17:476-489 (2016), the entire contents and disclosures of which are incorporated herein by reference. Some of the above mechanisms may reduce expression of the edited allele itself, while others may reduce expression of other copies or alleles of the endogenous locus(s) or gene(s).Such dominant or semi-dominant effect(s) on gene(s) may act through non-canonical repression mechanisms that do not involve significant or detectable levels of RNAi and / or formation of targeted small RNAs.
[0165] In one embodiment, the present disclosure provides a method for generating a dominant negative allele of a gene in a cell, comprising using targeted editing technology to delete a portion of chromosome between a first gene region and a second gene region, wherein the antisense RNA transcript of the first gene region is generated after the deletion of the portion of chromosome.In another embodiment, the targeted editing technology provided herein comprises the use of at least one site-specific nuclease.In some embodiments, the antisense RNA transcript provided herein is a partial antisense RNA transcript.In some embodiments, the partial antisense RNA transcript is shorter than the corresponding sense RNA transcript. In certain embodiments, the partial antisense RNA transcript is at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, or at least 2500 nucleotides shorter than the corresponding sense RNA transcript. In another embodiment, partial antisense RNA transcript is at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, or at least 95% shorter than corresponding sense RNA transcript.In another embodiment, the antisense RNA transcript provided herein is a complete antisense RNA transcript.In some embodiments, the complete antisense RNA transcript is the same length as corresponding sense RNA transcript.In some embodiments, the antisense RNA transcript provided herein suppresses the expression of a first gene region.In certain embodiments, the antisense RNA transcripts provided herein are capable of suppressing expression of a first gene region.
[0166] In another aspect, the disclosure provides a method comprising: a) identifying a chromosomal region comprising a first gene region comprising a first promoter and a first coding region and a second gene region comprising a second promoter and a second coding region, wherein the first coding region and the second coding region are separated by an intervening region and the first promoter and the second promoter are positioned in opposite orientations; b) inducing first and second double-stranded breaks flanking the targeted regions; c) identifying one or more cells comprising a deletion of the targeted region of the chromosome; and d) selecting one or more cells comprising a deletion of the targeted region of the chromosome.
[0167] As used herein, an "intervening region" or "intervening sequence" refers to a polynucleotide sequence between a first polynucleotide sequence and a second polynucleotide sequence that are physically linked. In one embodiment, the intervening region or intervening sequence is between a first gene and a second gene. In an embodiment, the intervening region or intervening sequence is between a first gene region and a second gene region. In one embodiment, the intervening region or intervening sequence is between a first coding region and a second coding region. In another embodiment, the intervening region or intervening sequence is between a first target site and a second target site. In one embodiment, the intervening region or intervening sequence is between a first target gene and a second target gene. In one embodiment, all or a portion of the intervening region or intervening sequence is inverted by targeted editing technology. In another embodiment, all or a portion of the intervening region or intervening sequence is deleted by targeted editing technology. In one embodiment, the intervening region or sequence comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 1250, at least 1500, at least 1750, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, at least 9000, at least 10,000, at least 15,000, at least 20,000, at least 25,000, or at least 50,000 nucleotides. In some embodiments, the intervening region or intervening sequence comprises at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, or at least 10 genes. In one embodiment, the intervening region or intervening sequence is located on a chromosome. In one embodiment, the intervening region or intervening sequence is located on a vector. In one embodiment, the intervening region or intervening sequence comprises a DNA sequence. In one embodiment, the intervening region or intervening sequence comprises an RNA sequence.In one embodiment, the intervening region or intervening sequence comprises an endogenous nucleic acid sequence. In another embodiment, the intervening region or intervening sequence comprises a transgenic nucleic acid sequence. In one embodiment, the intervening region or intervening sequence comprises an endogenous nucleic acid sequence and a transgenic nucleic acid sequence.
[0168] In one embodiment, the first gene region is selected from the group consisting of the GA20 oxidase gene region, the GA3 oxidase gene region, the brachytic2 gene region, and the Y1 gene region. In another embodiment, the first gene region is the GA20 oxidase gene region or the GA3 oxidase gene region. In a further embodiment, the first gene region is the GA20 oxidase gene region. In one embodiment, the first gene region is the GA3 oxidase gene region. In yet another embodiment, the first gene region is the brachytic2 gene region. In another embodiment, the first gene region is the Y1 gene region.
[0169] GA oxidase in cereal plants consists of a family of related GA oxidase genes. For example, in maize, there is a family of at least nine GA20 oxidase genes, including GA20 oxidase_1, GA20 oxidase_2, GA20 oxidase_3, GA20 oxidase_4, GA20 oxidase_5, GA20 oxidase_6, GA20 oxidase_7, GA20 oxidase_8, and GA20 oxidase_9. The DNA and protein sequences of GA20 oxidase_3 and GA20 oxidase_5, respectively, are provided in Table 1. [Table 1]
[0170] The wild-type genomic DNA sequence of the GA20 oxidase_3 locus from the reference genome is provided in SEQ ID NO:27, and the wild-type genomic DNA sequence of the GA20 oxidase_5 locus from the reference genome is provided in SEQ ID NO:31.
[0171] For the maize GA20 oxidase_3 gene (also known as Zm.GA20ox3), SEQ ID NO:27 provides 3,000 nucleotides upstream (5') of the 5'-UTR of GA20 oxidase_3, with nucleotides 3001 to 3096 corresponding to the 5'-UTR, nucleotides 3097 to 3665 corresponding to the first exon, nucleotides 3666 to 3775 corresponding to the first intron, nucleotides 3776 to 4097 corresponding to the second exon, nucleotides 4098 to 5314 corresponding to the second intron, nucleotides 5315 to 5584 corresponding to the third exon, and nucleotides 5585 to 5800 corresponding to the 3'-UTR. SEQ ID NO:27 also provides 3,000 nucleotides downstream (3') of the end of the 3'-UTR (nucleotides 5801 to 8800).
[0172] For the maize GA20 oxidase_5 gene (also known as Zm.GA20ox5), SEQ ID NO:31 provides 3,000 nucleotides upstream of the GA20 oxidase_5 start codon (nucleotides 1-3,000), with nucleotides 3001-3,791 corresponding to the first exon, nucleotides 3,792-3,906 corresponding to the first intron, nucleotides 3,907-4,475 corresponding to the second exon, nucleotides 4,476-5,197 corresponding to the second intron, nucleotides 5,198-5,473 corresponding to the third exon, and nucleotides 5,474-5,859 corresponding to the 3'-UTR. SEQ ID NO:31 also provides 3,000 nucleotides downstream (3') of the end of the 3'-UTR (nucleotides 5,860-8,859).
[0173] In the maize genome, the Zm.GA20ox5 gene is located next to the Zm.SAMT gene. These two genes are separated by an approximately 550-bp intergenic region, with the Zm.SAMT gene located downstream and facing in the opposite direction to the Zm.GA20ox5 gene. Reference genome sequences for the region encompassing the Zm.GA20ox5 and Zm.SAMT genes are provided in SEQ ID NOs: 35 and 36. SEQ ID NO: 35 represents the sequence of the sense strand of the Zm.GA20ox5 gene, encompassing both the Zm.GA20ox5 and Zm.SAMT genes ("GA20ox5_SAMT genomic sequence" in Table 2). SEQ ID NO: 35 partially overlaps with SEQ ID NO: 31, and has a shorter Zm.GA20ox5 upstream sequence and a longer Zm.GA20ox5 downstream sequence compared to SEQ ID NO: 31. SEQ ID NO: 36 represents the sequence of the sense strand of the Zm.SAMT gene (i.e., the antisense strand of the Zm.GA20ox5 gene), encompassing both the Zm.GA20ox5 gene and the Zm.SAMT gene ("SAMT_GA20ox5 genome sequence" in Table 2). In Table 2 below, elements or regions of the reference genome Zm.GA20ox5 / Zm.SAMT sequence are annotated by reference to the nucleotide coordinates of these elements or regions in SEQ ID NO: 35 or 36.
[0174] It has previously been shown that suppression of GA20 oxidase gene(s) via transgenic suppression and / or targeting a subset of one or more GA oxidase genes (e.g., artificial microRNA-mediated suppression of both the GA20 oxidase_3 and GA20 oxidase_5 genes) can be effective in achieving a low-height, semi-dwarf phenotype with increased resistance to lodging but without reproductive heterogeneity in the panicle. See PCT Application No. PCT / US2017 / 047405 and U.S. Application No. 15 / 679,699, both of which were filed August 17, 2017, and published as WO / 2018 / 035354 and US20180051295, respectively. Furthermore, knocking out the GA20 oxidase_3, GA20 oxidase_5, or both genes via genome editing can also result in reduced plant height and increased lodging resistance, affecting GA hormone levels. See PCT Application Nos. PCT / US2019 / 018128, PCT / US2019 / 018131, and PCT / US2019 / 018133, all filed February 15, 2019.
[0175] In one embodiment, the first gene region comprises a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity or complementarity to a sequence selected from the group consisting of SEQ ID NO: (insert GA20 cDNA sequence).
[0176] In another embodiment, the first gene region comprises a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity or complementarity to a sequence selected from the group consisting of SEQ ID NO: (insert BR2 cDNA sequence).
[0177] In one embodiment, the first gene region comprises a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity or complementarity to a sequence selected from the group consisting of SEQ ID NO: (insert GA3 cDNA sequence).
[0178] In one embodiment, the first gene region comprises a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity or complementarity to a sequence selected from the group consisting of SEQ ID NO: (insert Y1 cDNA sequence).
[0179] In one embodiment, the deletions provided herein include all or a portion of the second gene region. In another embodiment, the deletions provided herein include all of the second gene region. In yet another embodiment, the deletions provided herein include a portion of the second gene region.
[0180] In yet another aspect, the present disclosure provides a method for reducing expression of a gene in a cell, the method comprising: a) identifying a chromosomal region comprising a first gene region comprising a first promoter and a first coding region and a second gene region comprising a second promoter and a second coding region, wherein the first coding region and the second coding region are separated by an intervening region and the first promoter and the second promoter are positioned in opposite directions; b) using targeted editing technology to induce first and second double-strand breaks flanking the targeted region, wherein the targeted region comprises the second coding region and the intervening region; and c) identifying one or more cells comprising a deletion of the targeted region of the chromosome, wherein the second promoter generates at least one antisense RNA of the first coding region, and expression of the first coding region is reduced compared to a control cell that does not comprise a deletion of the targeted region. In one aspect, the deletion results in a portion of the first coding region being transcribed in the opposite direction.
[0181] In a further aspect, the present disclosure provides a method for detecting a chromosomal region comprising: a) identifying a chromosomal region comprising a first gene region comprising a first promoter and a first coding region and a second gene region comprising a second promoter and a second coding region, wherein the first coding region and the second coding region are separated by an intervening region and the first promoter and the second promoter are positioned in opposite orientations; and b) providing at least one RNA-guided nuclease, or one or more vectors encoding at least one RNA-guided nuclease, to one or more cells, wherein the at least one RNA-guided nuclease is a) providing a nuclease capable of binding to at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 stretches of nucleotides of a first target site and a second target site flanking a targeted region of a chromosome, wherein the targeted region comprises a second coding region and an intervening region, wherein the RNA-guided nuclease creates a double-stranded break at the first target site and the second target site of the chromosome; b) identifying one or more cells comprising a deletion of the targeted region; and c) selecting one or more cells comprising a deletion of the targeted region.
[0182] In one aspect, the present disclosure provides a modified plant or portion thereof comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof. In one aspect, the present disclosure provides a modified plant cell comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof. In another aspect, the present disclosure provides a modified plant or modified plant tissue comprising a modified plant cell comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
[0183] It is known in the art that transposons, or transposable elements, are DNA sequences that can change position within a genome. Transposons can create insertions, deletions, or inversions in a genome. In some embodiments, the methods, compositions, and cells provided herein do not involve the use of transposons (e.g., "non-transposon-mediated").
[0184] In one aspect, the present disclosure provides a modified chromosome comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof. In one aspect, the present disclosure provides a modified cell comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof. In another aspect, the present disclosure provides a modified cell comprising a modified chromosome comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
[0185] In one aspect, the present disclosure provides a product comprising a modified chromosome comprising a non-transposon-mediated genomic deletion or inversion of a gene or a portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or a portion thereof. In one aspect, the present disclosure provides a product comprising a modified plant or a portion thereof comprising a non-transposon-mediated genomic deletion or inversion of a gene or a portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or a portion thereof. In one aspect, the present disclosure provides a product comprising a modified plant cell comprising a non-transposon-mediated genomic deletion or inversion of a gene or a portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or a portion thereof. In one aspect, the disclosure provides a product comprising a modified cell containing a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof. In certain embodiments, the product comprises silage, flour, cellulose, sugar, starch, fat, syrup, or protein derived from the plant, plant part, or plant cell.
[0186] In one aspect, the present disclosure provides a modified cell comprising: a) a non-transposon-mediated genomic deletion of at least one gene or a portion thereof at the endogenous locus of the at least one gene; or b) a non-transposon-mediated, non-T-DNA-mediated insertion of a polynucleotide sequence into at least one gene, wherein the deletion or insertion creates a dominant positive allele of the at least one gene. In one aspect, the insertion comprises a regulatory element. In another aspect, the regulatory element is selected from the group consisting of a promoter sequence, a transcription start site sequence, a transcription termination site sequence, an enhancer sequence, and a designed element.
[0187] In another embodiment, the present disclosure provides a modified cell comprising a non-transposon-mediated genomic deletion or inversion of at least one gene or a portion thereof at the endogenous locus of at least one gene, wherein the deletion or inversion creates a dominant-negative allele of at least one gene. In yet another embodiment, the present disclosure provides a modified cell comprising a non-transposon-mediated genomic deletion or inversion of at least one gene or a portion thereof at the endogenous locus of at least one gene, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene. In one embodiment, the present disclosure provides a modified cell comprising targeted editing of at least one gene or a portion thereof, wherein the targeted editing produces an RNA transcript complementary to the native transcript sequence of the gene. In one embodiment, the RNA transcript is a complete antisense transcript. In another embodiment, the RNA transcript is a partial antisense transcript. In a further embodiment, the RNA transcript is a partial sense transcript. In yet another embodiment, the RNA transcript is a complete sense transcript. In another embodiment, the RNA transcript is a native transcript of the gene. In a further embodiment, the native transcript of the gene is a partial or complete sense transcript.
[0188] Dominant alleles of gene regions can also be created by inserting designed elements into the promoter of the gene region to induce constitutive expression of the gene region.
[0189] In one aspect, the disclosure provides a method of modifying gene expression, the method comprising: a) inducing a double-stranded break at a target site in a gene using targeted editing technology; b) inserting a donor sequence at the double-stranded break, wherein the donor sequence comprises a designed element that can induce increased or ectopic expression of the gene; and c) identifying at least one cell that contains the insertion of the donor sequence, wherein expression of the gene is increased in at least one tissue compared to a control cell that does not contain the insertion of the donor sequence.
[0190] In another aspect, the present disclosure provides a method comprising: a) providing one or more cells with at least one RNA-guided nuclease and at least one donor molecule, or one or more vectors encoding at least one RNA-guided nuclease and at least one donor molecule, wherein the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site of at least one gene, wherein the donor molecule comprises a designed element, wherein the RNA-guided nuclease creates a double-stranded break at the target site, and the donor molecule is inserted at the double-stranded break; b) identifying one or more cells comprising insertion of the donor molecule at the target site; and c) selecting one or more cells comprising insertion of the donor molecule at the target site.
[0191] In some embodiments, the target site is located downstream of the TATA box upstream of the gene. In some embodiments, the target site is located upstream of the TATA box upstream of the gene. In some embodiments, the target site is located upstream of the TATA box operably linked to at least one gene. In other embodiments, the target site is located downstream of the TATA box operably linked to at least one gene. In some embodiments, the target site is located within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, or 5000 nucleotides of the TATA box operably linked to at least one gene. In yet other embodiments, the target site is located between 10 and 5000, 10 and 2500, 10 and 1500, 10 and 1000, 10 and 750, 10 and 500, 10 and 250, 10 and 100, 20 and 100, 20 and 250, 20 and 500, or 50 and 500 nucleotides from a TATA box operably linked to at least one gene. In yet another embodiment, the target site is located within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, or 5000 nucleotides from the promoter of the gene. In yet another embodiment, the target site is located upstream of the upstream initiator element of the gene. In yet another embodiment, the target site is located downstream of the upstream initiator element of the gene.In yet another embodiment, the target site is located within 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, or 5000 nucleotides of the initiator element of the gene.
[0192] In one embodiment, a "TATA box" comprises the core DNA sequence 5'-TATAAA-3' or a variant thereof and is frequently associated with the promoters of eukaryotic genes. Typically, but not always, the TATA box is located approximately 25-35 nucleotides upstream from the transcription start site of the gene. The TATA box often functions as a binding site for transcription factors to enable expression of an operably linked gene or for histones to inhibit expression of an operably linked gene. In one embodiment, the TATA box is an initiator element. An initiator element is a core promoter that facilitates binding of transcription factors to enhance expression of an operably linked gene. In one embodiment, the initiator sequence provided herein comprises the sequence 5'-[C / T][C / T]AN[A / T][C / T][C / T]-3'.
[0193] Dominant negative alleles can also be created by editing the genome to include a tissue-specific or tissue-preferential promoter of a gene region, so that the tissue-specific or tissue-preferential promoter is in the opposite direction to the target gene.For example, by placing a tissue-specific promoter in the reverse direction downstream of the 3'-UTR of a gene region, the tissue-specific promoter can generate a complete antisense gene region RNA transcript.Without being bound by any theory, the antisense gene region RNA transcript expressed by the antisense tissue-specific promoter can suppress the expression of the gene region in a tissue-specific manner.
[0194] In one embodiment, the present disclosure provides a method for reducing the expression of a gene in at least one cell, comprising: a) using targeted editing technology to induce a double-strand break at a target site of the gene; b) inserting a donor sequence at the double-strand break, wherein the donor sequence comprises a tissue-specific or tissue-preferred promoter, and the donor sequence is inserted at the target site so that the tissue-specific or tissue-preferred promoter is in the reverse orientation compared to the gene; and c) identifying at least one cell containing the insertion of the donor sequence in the reverse orientation, wherein the expression of the gene is reduced compared to a control cell that does not contain the insertion of the donor sequence. In one embodiment, the method provided herein further comprises using targeted editing technology to remove the native promoter of the gene. As used herein, "native promoter" refers to a promoter that generates the sense mRNA transcript of the operably linked gene.
[0195] In another aspect, the present disclosure provides a method comprising: a) providing one or more cells with at least one RNA-guided nuclease and at least one donor molecule, or one or more vectors encoding at least one RNA-guided nuclease and at least one donor molecule, wherein the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site of at least one gene, wherein the donor molecule comprises a sequence encoding a tissue-specific or tissue-preferred promoter, wherein the RNA-guided nuclease creates a double-stranded break at the target site, and the donor molecule is inserted at the double-stranded break; b) identifying one or more cells comprising an insertion of the donor molecule at the target site, such that the tissue-specific or tissue-preferred promoter is in a reversed orientation compared to the gene; and c) selecting one or more cells comprising an insertion of the donor molecule at the target site.
[0196] In one embodiment, the target site is located downstream of the 3'-UTR of the gene. In another embodiment, the target site is located within the 3'-UTR of the gene. In another embodiment, the target site is located within an intron of the gene. In a further embodiment, the target site is located within an exon of the gene. In one embodiment, the target site is located with the 5'-UTR of the gene. In another embodiment, the target site is located upstream of the 5'-UTR of the gene. In yet a further embodiment, the target site is located within the promoter of the gene.
[0197] In one embodiment, the donor molecule comprises a polynucleotide encoding a promoter. In one embodiment, the donor molecule comprises a polynucleotide encoding a promoter selected from the group consisting of a tissue-specific promoter, a tissue-preferred promoter, a constitutive promoter, and an inducible promoter. In another embodiment, the donor molecule provided herein comprises a polynucleotide encoding a tissue-specific or tissue-preferred promoter. In yet another embodiment, the donor molecule provided herein comprises a polynucleotide encoding a constitutive promoter. In another embodiment, the donor molecule provided herein comprises a polynucleotide encoding an inducible promoter.
[0198] In one embodiment, the tissue-specific or tissue-preferred promoter is selected from the group consisting of a leaf-specific promoter, a leaf-preferred promoter, a stem-specific promoter, a stem-preferred promoter, a vascular-specific promoter, a vascular-preferred promoter, a root-specific promoter, a root-preferred promoter, an inflorescence-specific promoter, an inflorescence-preferred promoter, a pollen-specific promoter, a pollen-preferred promoter, an anther-specific promoter, an anther-preferred promoter, an ovule-specific promoter, an ovule-preferred promoter, a seed-specific promoter, a seed-preferred promoter, an embryo-specific promoter, an embryo-preferred promoter, an endosperm-specific promoter, an endosperm-preferred promoter, a pericarp-specific promoter, a pericarp-preferred promoter, an aleurone-specific promoter, an aleurone-preferred promoter, a meristem-specific promoter, a meristem-preferred promoter, a fruit-specific promoter, a fruit-preferred promoter, a pod-specific promoter, a pod-preferred promoter, an epidermis-specific promoter, an epidermis-preferred promoter, a mitochondrion-specific promoter, a mitochondrion-preferred promoter, a chloroplast-specific promoter, and a chloroplast-preferred promoter. In another embodiment, the tissue-specific or tissue-preferred promoter provided herein is an RTBV promoter. In some embodiments, the tissue-specific or tissue-preferred promoters provided herein express antisense mRNA transcripts of genes. In some embodiments, the tissue-specific or tissue-preferred promoters provided herein express antisense mRNA transcripts of genes.
[0199] Targeted editing technology can be used to insert a donor molecule into a target site in a genomic locus. When a donor molecule containing a non-coding RNA target site is inserted into the 5'-UTR, exon, intron, or 3'-UTR of a gene of interest, the RNA transcription or protein translation of the gene of interest can be suppressed by a complementary non-coding RNA. When the gene of interest is targeted by a non-coding RNA (e.g., miRNA or siRNA), the cleaved mRNA of the gene of interest can generate a secondary siRNA, which can further suppress the transcription or translation of the gene of interest. Because the secondary siRNA is complementary to the allele regardless of whether the non-coding RNA target site is inserted, such secondary suppression can act in a dominant manner.
[0200] In one embodiment, engineered or artificial miRNAs are created to target native gene regions. In another embodiment, gene regions are edited to be complementary to native miRNAs. Engineered miRNAs are useful for suppressing target genes with increased specificity. See, for example, Parizotto et al., Genes Dev. 18:2237-2242 (2004), and U.S. Patent Application Publication Nos. 2004 / 0053411, 2004 / 0268441, 2005 / 0144669, and 2005 / 0037988, the contents and disclosures of which are incorporated herein by reference. miRNAs are non-protein-coding RNAs. Cleavage of the miRNA precursor molecule results in the formation of a mature miRNA, typically about 19 to about 25 nucleotides in length (generally about 20 to about 24 nucleotides in length in plants), e.g., 19, 20, 21, 22, 23, 24, or 25 nucleotides in length, with a sequence corresponding to the gene targeted for repression and / or its complement. The mature miRNA hybridizes to the target mRNA transcript and guides the binding of a protein complex to the target transcript, which functions to inhibit translation and / or result in transcript degradation, thereby negatively regulating or repressing the expression of the targeted gene. In plants, miRNA precursors are also useful for directing the in-phase production of siRNAs, trans-acting siRNAs (ta-siRNAs), in processes that require RNA-dependent RNA polymerase to effect target gene repression. See, for example, Allen et al., Cell 121:207-221 (2005), Vaucheret Science STKE, 2005:pe43 (2005), and Yoshikawa et al. Genes Dev., 19:2164-2175 (2005), the contents and disclosures of which are incorporated herein by reference.
[0201] Plant miRNAs regulate their target genes by recognizing and binding to nearly perfectly complementary sequences (miRNA recognition sites) in target transcripts, followed by cleavage of the transcripts by RNase III enzymes such as Argonaute 1. Plants do not tolerate certain mismatches between a given miRNA recognition site and the corresponding mature miRNA, particularly mismatched nucleotides at positions 10 and 11 of the mature miRNA. Positions within the mature miRNA are shown in the 5' to 3' direction. Typically, perfect complementarity between a given miRNA recognition site and the corresponding mature miRNA is required at positions 10 and 11 of the mature miRNA. See, for example, Franco-Zorrilla et al. (2007) Nature Genetics, 39:1033-1037 and Axtell et al. (2006) Cell, 127:565-577.
[0202] Many microRNA genes (MIR genes) have been identified and are publicly available in databases ("miRBase," available online at microrna.sanger.ac.uk / sequences; see also Griffiths-Jones et al. (2003) Nucleic Acids Res., 31:439-441). MIR genes have been reported to occur in both isolated and clustered intergenic regions within the genome, but may also be located wholly or partially within introns of other genes (both protein-coding and non-protein-coding). For a recent review of miRNA biogenesis, see Kim (2005) Nature Rev. Mol. Cell. Biol., 6:376-385. Transcription of MIR genes can, at least in some cases, occur under the enhanced control of the MIR gene's own promoter. The primary transcript, termed the "pri-miRNA," can be quite large (several kilobases) and may be polycistronic, containing one or more pre-miRNAs (foldback structures containing stem-loop sequences that are processed into mature miRNAs) as well as the usual 5' "cap" and polyadenylation tail of an mRNA. See, e.g., Figure 1 in Kim (2005) Nature Rev. Mol. Cell. Biol., 6:376-385.
[0203] Transgenic expression of miRNAs (whether naturally occurring or artificial) can be used to regulate the expression of miRNA target gene(s). MiRNA recognition sites have been identified in all regions of mRNAs, including the 5' untranslated, coding, and 3' untranslated regions, demonstrating that the location of miRNA target sites relative to the coding sequence does not necessarily affect repression (see, e.g., Jones-Rhoades and Bartel (2004). Mol. Cell, 14:787-799; Rhoades et al. (2002) Cell, 110:513-520; Allen et al. (2004) Nat. Genet., 36:1282-1290; Sunkar and Zhu (2004) Plant Cell, 16:2001-2019). Because miRNAs are key regulatory elements in eukaryotes, transgenic repression of miRNAs is useful for manipulating biological pathways and responses. Promoters of miR genes can have highly specific expression patterns (e.g., cell-specific, tissue-specific, temporal-specific, or inducible) and are therefore useful in recombinant constructs that direct such specific transcription of operably linked DNA sequences. Various utilities of miRNAs, their precursors, their recognition sites, and their promoters are described in detail in U.S. Patent Application Publication No. 2006 / 0200878 A1, incorporated herein by reference. Non-limiting examples of these utilities include: (1) expression of native miRNAs or miRNA precursor sequences that repress target genes; (2) expression of artificial miRNAs or miRNA precursor sequences that repress target genes; (3) expression of transgenes containing miRNA recognition sites, where the transgene is repressed upon expression of the mature miRNA; and (4) expression of transgenes driven by miRNA promoters.
[0204] As demonstrated by Zeng et al. (2002) Mol. Cell, 9:1327-1333, the design of an artificial miRNA sequence can be as simple as substituting sequences complementary to the intended target for nucleotides in the miRNA stem region of a miRNA precursor. One non-limiting example of a general method for determining nucleotide changes in a native miRNA sequence to produce an engineered miRNA precursor includes the following steps: (a) selecting a unique target sequence of at least 18 nucleotides specific to the target gene by using sequence alignment tools such as BLAST (see, e.g., Altschul et al. (1990) J. Mol. Biol., 215:403-410; Altschul et al. (1997) Nucleic Acids Res., 25:3389-3402) of both tobacco cDNA and genomic DNA databases to identify any potential matches with the target transcript orthologs and unrelated genes, thereby avoiding unintended silencing of non-target sequences; (b) analyzing the target gene for unwanted sequences (e.g., matches with sequences from non-target species) and determining the GC content, Reynolds score (Reynolds et al. (2004) Nature 106:101-102) of each potential 19-mer segment; Biotechnol., 22:326-330), and scoring for functional asymmetry characterized by a negative free energy difference (“.DELTA..DELTA.G” or “ΔΔG”) (see Khvorova et al. (2003) Cell, 115:209-216) [Preferably, 19-mers are selected that have all or most of the following characteristics: (1) a Reynolds score >4, (2) about 40% to about 60% GC content, (3) a negative ΔΔG, (4) a terminal adenosine, (5) lack of stretches of four or more identical nucleotides, (6) location near the 3′ end of the target gene, and (7) minimal differences from the miRNA precursor transcript.The position of every third nucleotide in the siRNA has been reported to be particularly important in influencing the efficacy of RNAi, and an algorithm called "siExplorer" is publicly available at rna.chem.tu-tokyo.ac.jp / siexplorer.htm (see Katoh and Suzuki (2007) Nucleic Acids Res., 10.1093 / nar / gkl1120)]; (c) determining the reverse complement of the selected 19-mer for use in generating a modified mature miRNA (the additional nucleotide at position 20 preferably matches the selected target sequence, and the nucleotide at position 21 is preferably selected to be unpaired to prevent spread of silencing to the target transcript or paired with the target sequence to enhance spread of silencing to the target transcript); and (d) transforming the artificial miRNA into plants.
[0205] The siRNA pathway involves non-stepwise cleavage of a longer double-stranded RNA intermediate (RNA duplex) into small interfering RNA (siRNA). The size or length of siRNAs ranges from about 19 to about 25 nucleotides or base pairs, although general classes of siRNAs include those containing 21 or 24 base pairs. Thus, a transcribable DNA sequence or repression element of the present application can encode an RNA molecule that is at least about 19 to about 25 nucleotides in length, e.g., 19, 20, 21, 22, 23, 24, or 25 nucleotides in length.
[0206] In the ta-siRNA pathway, miRNAs help guide the in-phase processing of siRNA primary transcripts in a process that requires RNA-dependent RNA polymerase to produce double-stranded RNA precursors. ta-siRNAs are defined by the absence of secondary structure, the miRNA target site that initiates double-stranded RNA production, the requirement for DCL4 and RNA-dependent RNA polymerase (RDR6), and the production of multiple approximately 21-nt small RNAs that are identical in phase and have perfectly matched duplexes with two-nucleotide 3' overhangs (see Allen et al. (2005) Cell, 121:207-221). The size or length of ta-siRNAs ranges from about 20 to about 22 nucleotides or base pairs, but is most often 21 base pairs. Thus, donor molecules or vectors of the present application can encode RNA molecules that are at least about 20 to about 22 nucleotides in length, e.g., 20, 21, or 22 nucleotides in length. The donor molecules and vectors provided herein may also comprise a ta-siRNA scaffold. For methods of constructing suitable ta-siRNA scaffolds, see US Pat. No. 9,309,512, which is incorporated herein by reference in its entirety.
[0207] The present disclosure provides a method for generating a dominant-negative allele of a gene, comprising using targeted editing technology to introduce at least one non-coding RNA target site into the gene. In one embodiment, the dominant-negative allele of the gene is downregulated compared to an allele of the gene that does not include the at least one non-coding RNA target site. In another embodiment, a secondary siRNA complementary to the gene is generated. In another embodiment, the at least one non-coding RNA target site is an miRNA target site or an siRNA target site. In a further embodiment, the at least one non-coding RNA target site is introduced into a region of the gene selected from the group consisting of a 5'-UTR, an exon, an intron, and a 3'-UTR. In another embodiment, the at least one non-coding RNA target site is introduced into an exon of the gene. In another embodiment, the at least one non-coding RNA target site is introduced into an intron of the gene. In another embodiment, the at least one non-coding RNA target site is introduced into the 5'-UTR of the gene. In yet another embodiment, the at least one non-coding RNA target site is introduced into the 3'-UTR of the gene.
[0208] In another aspect, the disclosure provides a modified cell comprising a non-transgenic dominant negative allele of a gene, wherein the dominant negative allele comprises a heterologous non-coding RNA target site at the endogenous locus of the gene.
[0209] Dominant alleles can also be created by editing an allele of a gene region encoding a protein to produce a truncated protein, where the edited truncated protein interferes with the activity of the wild-type protein and exerts a dominant effect. In one embodiment, a dominant-positive allele is created by introducing targeted editing into a gene encoding a protein to produce a truncated protein. In one embodiment, a dominant-negative allele is created by introducing targeted editing into a gene encoding a protein to produce a truncated protein. In some embodiments, the truncated proteins provided herein interfere with protein-protein binding, DNA-protein binding, or RNA-protein binding. In one embodiment, the truncated proteins provided herein are microproteins. As used herein, a microprotein refers to a protein of approximately 100 to 200 amino acids in length that encodes only a protein-protein interaction or binding domain (see, e.g., Seo et al., Trends in Plant Sciences, 2011, 10:541-549). Microproteins often arise from functional genes that have undergone mutations that eliminate functional protein domains. In one embodiment, the microprotein is at least 50, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, or at least 225 amino acids in length. In one embodiment, the microprotein inhibits the activity of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten proteins in a cell. In another embodiment, the microprotein promotes the activity of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten proteins in a cell. Without limitation, the microprotein may compete with a second protein for a binding site in a third protein.In one embodiment, the microprotein may block the binding of a second protein to a third protein, resulting in inhibition of activity, or in another embodiment, the microprotein may bind to the third protein in place of the second protein, promoting the activity of the third protein.
[0210] In one aspect, a plant comprising a dominant negative allele encoding a microprotein comprises an improved trait selected from the group consisting of flowering time, meristem size, insect resistance, herbicide tolerance, and shade avoidance. In one aspect, a plant comprising a dominant positive allele encoding a microprotein comprises an improved trait selected from the group consisting of flowering time, meristem size, insect resistance, herbicide tolerance, and shade avoidance.
[0211] In one aspect, the truncated protein provided herein is selected from the group consisting of a truncated CLAVATA protein, a truncated CORYNE protein, a truncated BAM receptor, a truncated receptor-like protein kinase 2 (RPK2) protein, and a truncated G protein beta subunit 1 (AGB1) protein. In another aspect, the CLAVATA protein provided herein is a CLAVATA1 protein, a CLAVATA2 protein, or a CLAVATA3 protein.
[0212] In one aspect, the disclosure provides a method of generating a dominant negative allele of a gene, the method comprising: a) inducing a double-stranded break in the genome of at least one cell using a targeted editing technique at a target site of the gene, where the double-stranded break is repaired by non-homologous end joining; and b) identifying at least one cell containing an insertion or deletion at the target site, where the insertion or deletion at the target site results in the generation of a dominant negative allele of the gene.
[0213] In yet another aspect, the disclosure provides a modified cell comprising at least one insertion or deletion in the endogenous locus of at least one gene generated by targeted editing techniques, wherein the insertion or deletion results in expression of a truncated protein.
[0214] In another aspect, the disclosure provides a method comprising: a) providing at least one RNA-guided nuclease, or one or more vectors encoding at least one RNA-guided nuclease, to one or more cells, wherein the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site of at least one gene, and wherein the RNA-guided nuclease creates a double-stranded break at the target site; b) identifying at least one cell comprising an insertion or deletion at the target site, wherein the insertion or deletion at the target site results in the generation of a dominant-negative allele of the at least one gene; and c) selecting the one or more cells comprising the dominant-negative allele of the at least one gene.
[0215] The present disclosure also provides a method for generating a dominant allele of a gene, comprising using targeted editing technology to introduce nonsense mutation into gene to create truncated protein.In some embodiments, the truncated protein is a microprotein.In some embodiments, the targeted editing technology comprises the deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 2500 or at least 5000 nucleotides. In some embodiments, the targeted editing technique comprises an insertion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 2500, or at least 5000 nucleotides. In some embodiments, the targeted editing technique comprises an inversion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 25, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 500, at least 1000, at least 2500, or at least 5000 nucleotides.
[0216] In one embodiment, a dominant positive allele is created by introducing targeted editing into a gene encoding a protein to produce a truncated protein. In one embodiment, the present disclosure provides a method for generating a dominant positive allele of a gene, the method comprising: a) using targeted editing technology to induce a double-strand break in the genome of at least one cell at a target site of the gene, wherein the double-strand break is repaired by non-homologous end joining; and b) identifying at least one cell that contains an insertion or deletion at the target site, wherein the insertion or deletion at the target site results in the generation of a dominant positive allele of the gene.
[0217] In another aspect, the disclosure provides a method comprising: a) providing one or more vectors to one or more cells, wherein the one or more vectors comprise at least one polynucleotide encoding at least one RNA-guided nuclease, wherein the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a target site of at least one gene, wherein the RNA-guided nuclease creates a double-stranded break at the target site, and the double-stranded break is repaired by non-homologous end joining; b) identifying at least one cell comprising an insertion or deletion at the target site, wherein the insertion or deletion at the target site results in generation of a dominant positive allele of the at least one gene; and c) selecting one or more cells comprising the dominant positive allele of the at least one gene.
[0218] In one embodiment, the insertion or deletion provided herein disables intron / exon splice sites.Intron / exon splice sites refer to the boundary between introns and exons within a gene.In eukaryotes, introns are typically processed from RNA transcripts by spliceosomes so that mRNA transcripts containing only exon sequences are produced, but this is not always the case.If intron / exon splice sites are disrupted, spliceosomes may not properly remove intron sequences, resulting in proteins with one or more nonsense mutations that generate premature stop codons.In some embodiments, nonsense mutations generate truncated proteins. In certain embodiments, the truncated protein contains at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 250, at least 275, at least 300, at least 350, at least 400, at least 450, or at least 500 fewer amino acids than the endogenous protein encoded by the gene lacking the nonsense mutation.
[0219] In one embodiment, the nonsense mutation is a mutation that results in a premature stop codon in the transcribed mRNA. In another embodiment, the insertion or deletion provided herein is located within an exon. In another embodiment, the insertion or deletion provided herein is located within an intron. In another embodiment, the insertion or deletion provided herein is located within a 5'-UTR or 3'-UTR. In some embodiments, the insertion or deletion provided herein is located within a structure selected from the group consisting of an intron / exon splice site, an exon, an intron, a 5'-UTR, and a 3'-UTR. In yet another embodiment, the dominant negative allele provided herein comprises one or more, two or more, three or more, four or more, or five or more insertions and / or deletions. In yet another embodiment, the dominant positive allele provided herein comprises one or more, two or more, three or more, four or more, or five or more insertions and / or deletions.
[0220] In another embodiment, the nonsense mutations provided herein are located within an exon. In one embodiment, the insertions or deletions provided herein are located within a structure selected from the group consisting of an intron / exon splice site and an exon. The insertions or deletions provided herein can produce proteins with one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more nonsense mutations.
[0221] In some embodiments, the dominant negative alleles provided herein comprise polynucleotides that contain premature stop codons compared to the polynucleotides of control alleles. A premature stop codon is a stop codon located upstream of the normal stop codon of a gene. A premature stop codon produces a truncated protein. A stop codon is a nucleotide triplet in mRNA that signals the end of protein translation from mRNA. In one embodiment, the dominant negative alleles provided herein comprise polynucleotides that encode truncated proteins. In one embodiment, the truncated proteins provided herein are at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 250, at least 275, at least 300, at least 400, or at least 500 amino acids shorter than the full-length protein. In certain embodiments, the truncated proteins provided herein are generated by insertion or deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 1000, at least 1500, or at least 2000 nucleotides.
[0222] In some embodiments, the dominant positive alleles provided herein comprise polynucleotides that contain premature stop codons compared to the polynucleotides of control alleles. A premature stop codon is a stop codon located upstream of the normal stop codon of a gene. A premature stop codon produces a truncated protein. A stop codon is a nucleotide triplet in mRNA that signals the end of protein translation from mRNA. In one embodiment, the dominant positive alleles provided herein comprise polynucleotides that encode truncated proteins. In one embodiment, the truncated proteins provided herein are at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 125, at least 150, at least 175, at least 200, at least 225, at least 250, at least 275, at least 300, at least 400, or at least 500 amino acids shorter than the full-length protein. In certain embodiments, the truncated proteins provided herein are generated by insertion or deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 400, at least 500, at least 1000, at least 1500, or at least 2000 nucleotides.
[0223] The present disclosure provides a method for generating a dominant-negative allele of a gene in a cell, comprising deleting a portion of the gene using targeted editing technology, wherein a microprotein is generated after the deletion of the portion of the gene. In one embodiment, the truncated protein is a microprotein. In another embodiment, the dominant-negative allele provided herein encodes a microprotein. In a further embodiment, the dominant-positive allele provided herein encodes a microprotein. As used herein, "microprotein" refers to a short single-domain protein that has the ability to interfere with a larger multidomain protein. In one embodiment, the microprotein provided herein interferes with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten other proteins. In one embodiment, the microprotein provided herein can prevent a second protein from binding to a nucleic acid molecule. In another embodiment, the microprotein provided herein can prevent a second protein from binding to a third protein. The third protein may or may not be identical to the second protein. In another embodiment, the microproteins provided herein are capable of binding to at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten other proteins. In one embodiment, the microproteins provided herein are capable of forming heterodimers with at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten other proteins. In another embodiment, the microproteins provided herein are capable of forming homodimers.As used herein, "homodimer" refers to the hybridization or association of two identical molecules (e.g., protein A and protein A), and "heterodimer" refers to the hybridization or association of two different macromolecules (e.g., protein A and protein B; protein A and DNA; protein A and RNA).
[0224] Members of the pentatricopeptide repeat (PPR) gene family are common in plant genomes. Many PPR proteins can bind to RNA molecules in a sequence-specific manner. PPR proteins contain 2 to 30 PPR motifs, each of which aligns with a single nucleotide in the RNA molecule. Within a PPR motif, amino acids present at two or three specific positions confer nucleotide specificity. For example, but not limited to, when threonine is at position 6 and asparagine is at position 1, the PPR motif binds to adenine nucleotides; when threonine is at position 6 and aspartic acid is at position 1, the PPR motif binds to guanine nucleotides; when asparagine is at position 6 and aspartic acid is at position 1, the PPR motif binds to uracil (or thymine) nucleotides; and when asparagine is at position 6 and asparagine or serine is at position 1, the PPR motif binds to cytosine nucleotides.
[0225] Engineered PPR proteins can be generated by at least two construction strategies, but these are not limiting. In the first strategy, PPR proteins are constructed by treating each PPR motif as a separate block, such that a PPR protein is constructed by arranging multiple desired motifs. The resulting engineered PPR protein can then bind to a target RNA molecule. However, this strategy does not always work because each PPR motif contains an internal scaffold between positions 1' and 6, and this intra-motif scaffold is not shared between different PPR proteins. The second strategy utilizes an existing intra-motif scaffold. In the second strategy, site-directed mutagenesis of positions 1' and 6 is used to edit an existing PPR protein to become specific for a new target RNA molecule.
[0226] As used herein, "engineered PPR protein" and "engineered PPR motif" refer to a synthetically produced PPR protein or PPR motif that does not occur in nature and is capable of site-specifically binding to an RNA sequence.
[0227] The present disclosure provides a method comprising: a) providing a cell with an engineered PPR protein or a vector encoding the engineered PPR protein operably linked to a promoter, wherein the engineered PPR protein is capable of binding to an RNA transcript of a target gene; b) selecting one or more cells from step (a) that express the engineered PPR protein; and c) identifying one or more cells selected in step (b) that exhibit altered expression of the target gene. In one embodiment, the engineered PPR protein is capable of binding to at least one non-coding RNA target site of the RNA transcript. In one embodiment, the engineered PPR protein binds to at least one non-coding RNA target site of the RNA transcript. In one embodiment, the altered expression is increased expression. In another embodiment, the altered expression is decreased expression. In one embodiment, the promoter is the native promoter of the target gene. In another embodiment, the promoter is selected from the group consisting of a constitutive promoter, a tissue-specific promoter, a tissue-preferred promoter, and an inducible promoter.
[0228] In one embodiment, the engineered PPR protein provided herein binds to a non-coding RNA target site of a target RNA molecule and prevents the non-coding RNA from cleaving the target RNA or inhibiting the translation of the target RNA. In another embodiment, the engineered PPR protein provided herein directs the degradation of the target RNA molecule. In one embodiment, the engineered PPR protein provided herein comprises at least one RNA nuclease domain. In another embodiment, the RNA nuclease domain provided herein is an NYN nuclease domain or an SMR (small MutS-related) domain.
[0229] In certain embodiments, the engineered PPR proteins or engineered PPR motifs provided herein bind to at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, or at least 35 nucleotides of an RNA molecule. In another embodiment, an engineered PPR protein or engineered PPR motif provided herein binds to 5-35, 5-30, 5-25, 5-20, 5-15, 5-14, 5-13, 5-12, 5-11, 5-10, 10-35, 10-30, 10-25, 10-20, 10-15, 10-14, 10-13, 10-12, or 15-30 nucleotides of an RNA molecule.
[0230] In some embodiments, the engineered PPR proteins provided herein can act as dominant-negative alleles. In some embodiments, the engineered PPR proteins provided herein can act as dominant-positive alleles.
[0231] In one embodiment, the engineered PPR protein is targeted to mitochondria or chloroplasts.In another embodiment, the engineered PPR protein is targeted to the nucleus.In yet another embodiment, the engineered PPR protein is targeted to the cytoplasm of cells.Without being limited to any theory, the protein can be targeted to a specific cellular structure by adding or editing a transport peptide at the N-terminus of the protein.
[0232] In one embodiment, the genome editing system provided herein comprises tgOligo as a tether molecule. In another embodiment, the tether molecule is a crosslinker coupled to a nuclease or a DNA targeting guide molecule. In a further embodiment, the tether molecule is a dimerization domain coupled to a nuclease.
[0233] In one embodiment, the tether molecule can connect two or more DNA binding mechanisms bound to two genomic loci. In another embodiment, the tether molecule can connect two or more DNA binding mechanisms bound to two genomic loci located within a single chromosome flanking the target genomic region. In another embodiment, the tether molecule can connect two or more DNA binding mechanisms bound to two genomic loci on separate chromosomes.
[0234] In one aspect, the disclosure provides a method of generating a dominant negative allele of at least one gene in at least one cell, comprising: a) introducing into the at least one cell a genome editing system including: i) a site-specific nuclease or a molecule encoding the site-specific nuclease; ii) an sgRNA or a molecule encoding the sgRNA; and iii) one or more molecules encoding at least a first tethered guide oligo (tgOligo) and a second tgOligo, or the first and second tgOligos, operably linked to at least one promoter; and b) generating a first double-stranded break and a second double-stranded break in the at least one gene, wherein the first tgOligo is a tethered guide oligo (tgOligo) and a second tgOligo are operably linked to at least one promoter. wherein the tgOligo and the second tgOligo hybridize to the 3' free ends of opposing strands in the first double-stranded break and the second double-stranded break, resulting in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, at least 100, at least 250, at least 500, at least 750, at least 1000, at least 2500, or at least 5000 nucleotides of at least one gene, thereby generating a dominant negative allele of the gene encoding the truncated protein; and c) identifying and selecting at least one cell containing the truncated protein.
[0235] In one aspect, the disclosure provides a method of generating a dominant negative allele of at least one gene in at least one cell, comprising: a) introducing into the at least one cell a genome editing system including: i) a site-specific nuclease or a molecule encoding the site-specific nuclease; ii) an sgRNA or a molecule encoding the sgRNA; and iii) one or more molecules encoding at least a first tethered guide oligo (tgOligo) and a second tgOligo, or the first and second tgOligos, operably linked to at least one promoter; and b) generating a first double-stranded break and a second double-stranded break in the at least one gene. wherein the first tgOligo and the second tgOligo hybridize to the 3' free ends of opposing strands in the first double-stranded break and the second double-stranded break, resulting in a deletion of 1 to 5,000, 5 to 5,000, 10 to 5,000, 25 to 2,500, 25 to 1,000, 25 to 750, 25 to 500, 25 to 100, 50 to 5,000, 50 to 1,000, 50 to 500, 100 to 1,000, or 1,000 to 5,000 nucleotides in at least one gene, thereby generating a dominant-negative allele of the gene encoding the truncated protein; and c) identifying and selecting at least one cell containing the truncated protein.
[0236] In another aspect, the disclosure provides a method of generating a dominant negative allele of at least one gene in at least one cell, comprising: a) introducing into the at least one cell one or more vectors encoding at least a first tgOligo and a second tgOligo operably linked to: i) at least one site-specific nuclease; ii) at least one sgRNA; and iii) at least one promoter; and b) generating a first double-strand break and a second double-strand break in the gene. and c) generating a dominant-negative allele of at least one gene encoding an antisense RNA transcript of the gene, wherein the first tgOligo and the second tgOligo hybridize to the 3' free ends of opposing strands in the first double-stranded break and the second double-stranded break, such that the region of the at least one gene between the first double-stranded break and the second double-stranded break is in the opposite orientation, and thereby generating a dominant-negative allele of at least one gene encoding an antisense RNA transcript of the gene; and c) identifying and selecting at least one cell containing the antisense RNA transcript of the at least one gene.
[0237] As used herein, a "tethered guide oligo" (tgOligo) refers to an oligonucleotide containing a sequence segment that can hybridize to the 3' free end (this 3' free end is also referred to as a 3' free flap) of the non-target strand of a double-stranded DNA molecule that is recognized and cleaved by a CRISPR gRNA-Cas complex. When a tgOligo recognizes and hybridizes to the 3' free end of the non-target strand of the target site of a gRNA, the tgOligo corresponds to that gRNA. A tgOligo can be a DNA molecule, an RNA molecule, or a mixture of nucleotides. A hybrid tgOligo is a tgOligo that can recognize and hybridize to the non-target 3' free ends generated by two separate CRISPR gRNA-Cas complexes.
[0238] As used herein, "tethered guide RNA" (tgRNA) refers to an RNA molecule that contains both a guide RNA (gRNA) sequence and a tether RNA sequence, where the tether RNA sequence is capable of hybridizing to a desired genomic site (a site referred to as the "tether site").
[0239] In one embodiment, the method provided herein includes the use of one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more tgOligos. In one embodiment, tgOligos are DNA molecules. In another embodiment, tgOligos are RNA molecules. In yet a further embodiment, tgOligos are a mixture of DNA molecules and RNA molecules. In one embodiment, tgOligos are single-stranded. In another embodiment, tgOligos are double-stranded. In one embodiment, at least one or at least two tgOligos are used simultaneously with at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten site-specific nucleases. In another embodiment, at least one tgOligo is not used simultaneously with a site-specific nuclease. In some embodiments, at least one or at least two tgOligos are tethered to at least one or at least two Cas9 proteins. In one embodiment, a first tgOligo is tethered to a first Cas9 protein and a second tgOligo is tethered to a second Cas9 protein. In another embodiment, at least one or at least two tgOligos are tethered to at least one or at least two inactivated Cas9 proteins.
[0240] In yet another embodiment, the tgOligo provided herein comprises at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 500, at least 1000, at least 2500, at least 5000, at least 10,000, or at least 25,000 nucleotides. In further embodiments, the tgOligo provided herein contains 5 to 25,000 nucleotides, 5 to 10,000 nucleotides, 5 to 5000 nucleotides, 20 to 10,000 nucleotides, 20 to 5000 nucleotides, 20 to 1000 nucleotides, 20 to 500 nucleotides, 20 to 250 nucleotides, 50 to 2500 nucleotides, 50 to 1000 nucleotides, 50 to 500 nucleotides, 50 to 250 nucleotides, 100 to 2500 nucleotides, 100 to 1000 nucleotides, 100 to 500 nucleotides, or 1000 to 10,000 nucleotides.
[0241] In one embodiment, the first tgOligo and the second tgOligo are at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% complementary to each other. In one embodiment, the first tgOligo and the second tgOligo are at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least At least 125, at least 150, at least 175, at least 200, at least 250, at least 500, at least 1000, at least 2500, or at least 5000 nucleotides are at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% complementary to each other. In one embodiment, the first tgOligo comprises a sense strand, and the second tgOligo comprises an antisense strand.
[0242] In one embodiment, tgOligo is provided to a cell. In another embodiment, tgOligo is encoded by a vector. In another embodiment, the site-specific nuclease and tgOligo are encoded by a single vector. In yet another embodiment, the site-specific nuclease and tgOligo are encoded by two or more vectors.
[0243] The methods provided herein are suitable for generating dominant alleles of protein-coding genes and non-coding RNA. Non-limiting examples of target genes in plant genomes contemplated by the present disclosure include genes for disease, insect, or pest resistance; herbicide resistance; quality improvement, such as yield, nutritional enhancement, environmental tolerance, or stress tolerance; or starch production (see U.S. Patent Nos. 6,538,181, 6,538,179, 6,538,178, 5,750,876, 6,477, 6,538,189 ...89, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538,189, 6,538, 6,295); altered oil yield (U.S. Patent Nos. 6,444,876, 6,426,447, 6,380,462); high oil yield (U.S. Patent Nos. 6,495,739, 5,608,149, 6,483,008, 6,476,295); altered fatty acid content (U.S. Patent Nos. 6,828,475, 6,822,141, 6,770,465 , 6,706,950, 6,660,849, 6,596,538, 6,589,767, 6,537,750, 6,489,461, 6,459,018); high protein yield (U.S. Patent No. 6,380,466); fruit ripening (U.S. Patent No. 5,512,466); animal and human nutrition promotion (U.S. Patent No. 6,723,837, U.S. Patent No. 6,659,018). 3,530, 6,5412,59, 5,985,605, 6,171,640); or biopolymers (U.S. Patent Nos. RE37,543, 6,228,623, 5,958,745, and U.S. Patent Publication No. US20030028917).Also, environmental stress resistance (U.S. Patent No. 6,072,103); pharmaceutical peptides and secretable peptides (U.S. Patent Nos. 6,812,379, 6,774,283, 6,140,075, 6,080,560); improved processing properties (U.S. Patent No. 6,476,295); improved digestibility (U.S. Patent No. 6,531,648); low raffinose (U.S. Patent No. 6,166,292); industrial improved enzyme production (U.S. Patent No. 5,543,576); improved flavor (U.S. Patent No. 6,011,199); nitrogen fixation (U.S. Patent No. 5,229,114); hybrid seed production (U.S. Patent No. 5,689,041); fiber production (U.S. Patent Nos. 6,576,818, 6,271,443, 5,981,834, 5,869,720); and biofuel production (U.S. Patent No. 5,998,700).
[0244] In one embodiment, the gene edited by the methods provided herein is selected from the group consisting of the Y1 gene, the brachytic2 gene, the GA3 oxidase gene, and the GA20 oxidase gene. In another embodiment, the gene edited by the methods provided herein encodes a non-coding RNA. In certain embodiments, the non-coding RNA edited by the methods provided herein is selected from the group consisting of microRNA, small interfering RNA, transfer RNA, ribosomal RNA, trans-acting small interfering RNA, naturally occurring antisense small interfering RNA, heterochromatin small interfering RNA, and precursors thereof. In yet another embodiment, the gene edited by the methods provided herein encodes a miRNA. In a further embodiment, the gene edited by the methods provided herein encodes a precursor miRNA (pre-miRNA).
[0245] In one embodiment, the GA20 oxidase gene provided herein is encoded by an mRNA that encodes a protein having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity to a sequence selected from the group consisting of SEQ ID NO: (insert protein sequence of GA20). In another embodiment, the brachytic2 gene provided herein is encoded by an mRNA that encodes a protein having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity to a sequence selected from the group consisting of SEQ ID NO: (insert protein sequence of BR2).
[0246] In one aspect, the unmodified alleles provided herein comprise a polynucleotide sequence having at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% identity or complementarity to a sequence selected from the group consisting of SEQ ID NOs: (listing the sequences of GA and BR2).
[0247] In another embodiment, the non-coding RNA edited by the methods provided herein is selected from the group consisting of microRNA, small interfering RNA, transfer RNA, ribosomal RNA, trans-acting small interfering RNA, naturally occurring antisense small interfering RNA, heterochromatin small interfering RNA, and precursors thereof. [Example]
[0248] Example 1. Generation of a dominant allele of GA20 oxidase via targeted genomic inversion Two functional guide RNAs (gRNAs) for the CRISPR / RNA-guided nuclease system were engineered to target the flanking regions (left and right target sites) of the GA20 oxidase_5 gene in the maize genome. See Figure 1, panel A. Each of the two target sites is unique within the maize genome. The plant hormone gibberellin plays an important role in several plant developmental processes, including germination, cell elongation, flowering, embryogenesis, and seed development. Specific biosynthetic enzymes (e.g., GA20 oxidase and GA3 oxidase) and catabolic enzymes (e.g., GA2 oxidase) in the GA pathway are important for influencing GA levels within plant tissues.
[0249] A transfer DNA (T-DNA) vector suitable for Agrobacterium transformation is used. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains an expression cassette containing a promoter operable in plant cells operably linked to polynucleotides encoding the two gRNAs described above.
[0250] Immature corn embryos are co-cultured with Agrobacterium containing the T-DNA vector for three days. The polynucleotide between the LB and RB sequences is integrated into the nuclear genome of the immature corn embryos. Upon expression of the integrated polynucleotide, the gRNA guides a nuclease to each of two target sites within the GA20 oxidase_5 gene, where the nuclease creates a double-strand break at each target site.
[0251] In most events, the region between the target sites is deleted, and non-homologous end-joining repair mechanisms join the flanking regions. Less frequently, some events create insertion / deletion mutations at the left target site, the right target site, or both. In still other events, the entire targeted region is inverted, referred to as a "full inversion." See Figure 1, panel B. Transformation events containing full inversions are identified using suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing). Transformed embryos containing a full inversion targeted to the 5' end of the GA20 oxidase_5 gene are selected and used to regenerate modified plants using techniques standard in the art.
[0252] Without being bound by any scientific theory, the presence of a complete inversion in one allele of GA20 oxidase_5 produces a population of antisense mRNA under the control of the native GA20 oxidase_5 promoter. The inverted region of the edited GA20 oxidase_5 allele mRNA and the corresponding region in the unedited GA20 oxidase_5 allele mRNA are complementary to each other and can form dsRNA. Thus, even when the modified corn plant is heterozygous for the edited GA20 oxidase_5 allele, the edited allele can reduce expression of the GA20 oxidase_5 gene in the modified corn plant.
[0253] RNA is extracted from modified corn plants identified as containing a complete inversion of the targeted region of the GA20 oxidase_5 gene. RNA is also extracted from control corn plants lacking the complete inversion. Using suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing), downregulation of GA20 oxidase_5 is confirmed in modified corn plants containing the complete inversion of the targeted region.
[0254] Example 2. Generation of a dominant allele of BR2 via targeted genomic inversion Two functional guide RNAs (gRNAs) for the RNA-guided nuclease system were generated to target the flanking regions (left and right target sites) of the BRACHYTIC2 (BR2) gene in the maize genome. Each of the two target sites is unique within the maize genome.
[0255] A transfer DNA (T-DNA) vector suitable for use in Agrobacterium tumefaciens transformation is used. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains an expression cassette containing a promoter operable in plant cells operably linked to polynucleotides encoding the two gRNAs described above.
[0256] Immature corn embryos were co-cultured with Agrobacterium tumefaciens containing the T-DNA vector for three days. The polynucleotide between the LB and RB sequences was integrated into the nuclear genome of the immature corn embryos. Upon expression of the integrated polynucleotide, the gRNA guided the CRISPR endonuclease to each of the two target sites within the BR2 gene, where the CRISPR endonuclease created a double-strand break at each target site.
[0257] In most cases, the region between the target sites is deleted, and non-homologous end joining repair mechanisms join the flanking regions. Less frequently, some events create insertion / deletion mutations at the left target site, the right target site, or both. In still other events, the entire targeted region is inverted, a phenomenon known as a "full inversion." Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing full inversions. Transformed embryos containing a full inversion targeted to the 5' end of the BR2 gene are selected and used to regenerate plants using techniques standard in the art.
[0258] Without being bound by any scientific theory, the presence of a complete inversion in one allele of the BR2 gene produces a population of antisense mRNA under the control of the native BR2 promoter. The inverted region of the edited BR2 allele mRNA and the corresponding region in the unedited BR2 allele mRNA are complementary to each other and can form dsRNA. Therefore, even if the corn plant is heterozygous for the edited BR2 allele, the edited allele can reduce the expression of both alleles of the BR2 gene in the corn plant, thereby resulting in a short-stemmed phenotype.
[0259] Example 3. Generation of a dominant allele of GA20 oxidase via targeted genomic deletion placing the GA20 oxidase gene under the control of two inverted promoters The gene encoding GA20 oxidase_5 (also called GA20ox5) is located on maize chromosome 8. It is adjacent to the maize gene GRMZM2G049269, which encodes an S-adenosyl-L-methionine-dependent methyltransferase superfamily protein (hereafter referred to as "SAMT"). SAMT is a member of a large overlapping gene family, and no phenotypes associated with mutations in this gene have been reported in maize or Arabidopsis. The SAMT gene is located in the opposite orientation compared to the GA20 oxidase_5 gene (i.e., the SAMT gene is oriented 5' to 3', while the GA20 oxidase_5 gene is oriented 3' to 5' on the same DNA strand). See Figure 2, Panel A.
[0260] Two functional guide RNAs (gRNAs) for the RNA-guided nuclease system were created to target the genomic DNA region between the GA20 oxidase_5 gene and the SAMT gene. The first gRNA targeted an area near the transcription start site of the SAMT gene, and the second gRNA targeted a region near the transcription end site of the GA20 oxidase_5 gene. Each of the two target sites was unique within the maize genome. A transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation was used. The T-DNA construct contained several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contained a promoter operably linked to a polynucleotide encoding the RNA-guided nuclease and operable in plant cells. The second expression cassette contained a promoter operably linked to the CP4-EPSPS marker gene and operable in plant cells. The construct also contained an expression cassette operably linked to a polynucleotide encoding the two gRNAs described above and operable in plant cells.
[0261] Immature corn embryos were co-cultured with Agrobacterium containing the T-DNA vector for three days. The polynucleotide between the LB and RB sequences was integrated into the nuclear genome of the immature corn embryos. Upon expression of the integrated polynucleotide, the gRNA guided a nuclease to each of two target sites within the genomic DNA region between the GA20 oxidase_5 gene and the SAMT gene, where the nuclease created a double-strand break at each target site.
[0262] In most cases, the region between the target sites is deleted, and non-homologous end joining repair mechanism joins the flanking region. Using suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing), transformation events containing complete deletion are identified. Transformed embryos containing the targeted deletion between the GA20 oxidase_5 gene and the SAMT gene are selected and used to regenerate modified plants using standard techniques in the art.
[0263] Without being bound by any scientific theory, we propose that deleting the genomic DNA between the GA20 oxidase_5 gene and the SAMT promoter would enable the native SAMT promoter to generate an antisense mRNA transcript of the GA20 oxidase_5 gene, while the native GA20 oxidase_5 promoter would generate a sense mRNA transcript of the GA20 oxidase_5 gene. The complementary sense and antisense mRNA transcripts of the GA20 oxidase_5 gene can form dsRNA that can be processed by the RNA silencing mechanism inherent to maize cells (see Figure 2, panel B). The processed dsRNA can then suppress the expression of both alleles of GA20 oxidase_5. Furthermore, because the mRNA encoding GA20 oxidase_3 is highly similar to the sense GA20 oxidase_5 transcript, we predict that the deletion between the SAMT promoter and the GA20 oxidase_5 gene would also downregulate the expression of GA20 oxidase_3. It is also contemplated that suppression or silencing of the GA20 oxidase_5 gene may occur by other mechanisms provided herein (e.g., nonsense-mediated decay) instead of, or in addition to, any RNAi or PTGS form of suppression.
[0264] RNA is extracted from modified corn plants identified as containing the targeted deletion between the SAMT promoter and the GA20 oxidase_5 gene. RNA is also extracted from control corn plants lacking the deletion. Using suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing), downregulation of GA20 oxidase_5 and / or GA20 oxidase_3 is confirmed in modified corn plants containing the targeted deletion.
[0265] Example 4. Generation of a dominant allele of BR2 via targeted genomic deletion placing the BR2 gene under the control of two inverted promoters The gene encoding BR2 is located on maize chromosome 1. It is adjacent to the maize gene GRMZM2G491632, which is expressed in the opposite direction to BR2. Two functional guide RNAs (gRNAs) for the RNA-guided nuclease system were generated to target the genomic DNA region between the BR2 and GRMZM2G491632 genes. The first gRNA targets an area near the end of exon 1 of the BR2 gene, and the second gRNA targets a region near the start of the coding sequence of exon 1 of the GRMZM2G491632 gene.
[0266] A transfer DNA (T-DNA) vector suitable for Agrobacterium transformation is used. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains an expression cassette containing a promoter operable in plant cells operably linked to polynucleotides encoding the two gRNAs described above.
[0267] Immature corn embryos are co-cultured with Agrobacterium containing the T-DNA vector for three days. The polynucleotide between the LB and RB sequences is integrated into the nuclear genome of the immature corn embryos. Upon expression of the integrated polynucleotide, the gRNA guides the endonuclease to each of two target sites within the genomic DNA region between the BR2 gene and the GRMZM2G491632 gene, where the endonuclease creates a double-strand break at each target site.
[0268] In most cases, the region between the target sites is deleted, and non-homologous end joining repair mechanism joins the flanking region.Use suitable methods known in the art (for example, PCR, DNA hybridization (Southern) blot, sequencing) to identify the transformation event that contains deletion.Select the transformed embryo that contains the targeted deletion between BR2 gene and GRMZM2G491632 gene, and use standard techniques in the art to regenerate plants.
[0269] Without being bound by any scientific theory, it is believed that by removing the genomic DNA between the BR2 gene and the GRMZM2G491632 promoter, the native GRMZM2G491632 promoter can generate antisense mRNA transcripts of the BR2 gene, and the native BR2 promoter can generate sense mRNA transcripts of the BR2 gene. The complementary sense and antisense mRNA transcripts of the BR2 gene can form dsRNA that can be processed by the RNAi machinery inherent in maize cells. The processed dsRNA can suppress the expression of both BR2 alleles, thereby resulting in a short-stem phenotype.
[0270] Example 5. Creating a dominant allele by inserting elements designed using genome editing techniques into the native promoter. A gene containing a root-specific promoter has been identified in the Arabidopsis thaliana genome. See Figure 3, panel A. This promoter is used to drive GUS expression in plant roots. See Figure 3, panel B. A functional gRNA is designed to target a region upstream of the TATA box (the "target site") of the root-specific promoter. The gRNA is introduced into Arabidopsis using a transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operably linked to a polynucleotide encoding an RNA-guided nuclease and operable in plant cells. The second expression cassette contains a promoter operably linked to the CP4-EPSPS marker gene and operable in plant cells. The construct also contains an expression cassette containing a promoter operably linked to a polynucleotide encoding the above-mentioned gRNA and operable in plant cells. The second T-DNA construct comprises a polynucleotide encoding a donor molecule containing a designed element to be inserted at the target site between the LB and RB sequences. In one embodiment, the donor molecule contains a designed element flanked by homologous regions that are homologous to sequences present on either side of the target site. The designed element allows a previously root-specific gene to be constitutively expressed in all tissues of the plant when inserted into the gene's promoter region. In another embodiment, the donor molecule contains a designed sequence flanked by a target site targeted by the gRNA of T-DNA vector 1.
[0271] The floral dip method is used to transform Arabidopsis using the above-mentioned vector. See Clough and Bent, 1998, Plant J, 16:735-743, which is incorporated herein by reference in its entirety. When the integrated polynucleotide is expressed, the gRNA guides the nuclease to the target site, creating a double-strand break at the target site. In the case of a donor molecule that contains a designed sequence flanked by homologous arms, the homologous repair mechanism inherent in Arabidopsis cells inserts the designed element into the site of the double-strand break. In the case of a donor molecule that contains a designed sequence flanked by gRNA target sites, the gRNA guides the nuclease to create a double-strand break in the second T-DNA, thereby releasing the designed sequence, which can then be integrated into the genome target site via the NHEJ (non-homologous end joining) repair mechanism. In some insertion events, the promoter is inserted in the desired orientation. Without being bound to any particular theory, the presence of the designed element upstream of the TATA box induces constitutive expression of the gene throughout the modified plant, thereby creating a dominant allele of the gene.
[0272] Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing targeted insertion of the designed element in the desired orientation. Transformed Arabidopsis plants containing the designed element at the target site are selected and further tested.
[0273] Plants identified as containing the designed element upstream of the TATA box (see Figure 3, panel D) are tested for GUS expression using techniques standard in the art. Plants containing the designed element exhibit GUS expression throughout the plant. See Figure 3, panel C. RNA is also extracted from various tissues (e.g., roots, stems, leaves, inflorescences) of modified Arabidopsis plants identified as containing the designed element. RNA is also extracted from control Arabidopsis plants lacking the designed element. Suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing) are used to confirm that GUS is more widely and / or more strongly expressed in the modified Arabidopsis plants.
[0274] Example 6. Creating dominant alleles by inserting tissue-specific suppressor elements using genome editing techniques. Using standard techniques in the art, Arabidopsis thaliana plants containing a GUS transgene under the control of a promoter functional in leaf, vascular, and root tissues are generated. See Figure 4 and Figure 5, panels A and B, for an overview of the concepts provided in this example. A functional gRNA is designed to target a region downstream of the GUS gene (the "target site"). The gRNA is introduced into Arabidopsis using a transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operably linked to a polynucleotide encoding an RNA-guided nuclease and operable in plant cells. The second expression cassette contains a promoter operably linked to the CP4-EPSPS marker gene and operable in plant cells. The construct also contains an expression cassette containing a promoter operably linked to a polynucleotide encoding the gRNA described above and operable in plant cells. The second T-DNA construct includes a donor molecule containing a leaf-specific promoter, such as the COOLAIR promoter (see Chen and Penfield, Science, 2018, 360:6392), between the LB and RB sequences. In one embodiment, the donor molecule contains a promoter in antisense orientation flanked by homologous regions that are homologous to sequences on either side of the target site. The antisense leaf-specific promoter allows expression of the GUS antisense mRNA in tissues where the promoter is expressed (e.g., leaf tissue). Without being bound by any particular theory, antisense RNA transcripts of the GUS gene cause GUS silencing in leaf tissue but not in root tissue. In another embodiment, the donor molecule contains a promoter sequence flanked by the target site targeted by the gRNA of T-DNA vector 1.
[0275] The floral dip method is used to transform Arabidopsis using the above-mentioned vector. See Clough and Bent, 1998, Plant J, 16:735-743, which is incorporated herein by reference in its entirety. Upon polynucleotide expression, the gRNA guides the nuclease to the target site, creating a double-strand break at the target site. In the case of a donor molecule containing a promoter in antisense orientation flanked by homologous arms, the homologous repair mechanism inherent in Arabidopsis cells inserts a leaf-specific promoter at the site of the double-strand break downstream of the GUS gene, thereby placing the promoter in antisense orientation relative to the GUS gene. In the case of a donor molecule containing a promoter flanked by gRNA target sites, the gRNA guides the nuclease to create a double-strand break in the second T-DNA, thereby liberating the promoter sequence, which can then be integrated into the genomic target site via the NHEJ (non-homologous end joining) repair mechanism. In some insertion events, the promoter is inserted in antisense orientation. Without being bound by any particular theory, the presence of an antisense leaf-specific promoter induces a reduction in expression of GUS throughout the leaf tissue of the plant, thereby creating a dominant allele of the gene.
[0276] Transformation events containing targeted insertion of the promoter in the desired orientation are identified using suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing). Transformed Arabidopsis plants containing the leaf-specific promoter at the target site are selected and further tested.
[0277] Plants identified as containing a leaf-specific promoter downstream of the GUS gene in an antisense orientation relative to the GUS gene (see Figure 5, panel D) are tested for GUS expression using techniques standard in the art. Plants containing a leaf-specific promoter in an antisense orientation relative to the GUS gene exhibit GUS expression only in root tissue. See Figure 5, panel C. RNA is also extracted from various tissues (e.g., roots, stems, leaves, inflorescences) of modified Arabidopsis plants identified as containing a leaf-specific promoter. RNA is also extracted from control Arabidopsis plants lacking a leaf-specific promoter in an antisense orientation relative to the GUS gene. Reduced GUS expression in leaf tissue is confirmed using suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing).
[0278] Example 7. Creation of a dominant GA20 oxidase_5 allele by inserting a tissue-specific suppressor element using genome editing techniques. A functional gRNA is designed to target a region downstream of the 3'-UTR of the GA20 oxidase_5 gene (the "target site"). A transfer DNA (T-DNA) vector suitable for Agrobacterium transformation is used to introduce the gRNA into maize cells. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains an expression cassette containing a promoter operable in plant cells operably linked to a polynucleotide encoding the above-mentioned gRNA. The second T-DNA construct contains a donor molecule containing an RTBV promoter between the LB and RB sequences. In one embodiment, the donor molecule comprises an RTBV promoter in antisense orientation flanked by homologous regions homologous to sequences present on either side of the target site. In another embodiment, the donor molecule comprises a promoter sequence flanked by a target site targeted by the gRNA of T-DNA vector 1. The RTBV promoter in antisense orientation allows expression of GA20 oxidase_5 antisense mRNA in tissues where RTBV is expressed (e.g., stem and vascular tissues). Without being bound to any particular theory, antisense RNA transcripts of the GA20 oxidase_5 gene cause silencing of both GA20 oxidase_5 and GA20 oxidase_3 in stem and vascular tissues.
[0279] Immature corn embryos are co-cultured with Agrobacterium containing the T-DNA vector for three days. Upon polynucleotide expression, the gRNA guides the nuclease to the target site, creating a double-stranded break at the target site. In the case of donor molecules containing an RTBV promoter in antisense orientation flanked by homologous arms, the homology repair mechanism inherent to corn cells inserts the antisense RTBV promoter into the target site downstream of the 3' end of the GA20 oxidase_5 gene. In the case of donor molecules containing a promoter flanked by gRNA target sites, the gRNA guides the nuclease to create a double-stranded break in the second T-DNA, thereby liberating the promoter sequence, which can then be integrated into the genomic target site via the NHEJ (non-homologous end joining) repair mechanism. In some insertion events, the promoter is inserted in the antisense orientation.
[0280] Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing targeted insertion of the antisense RTBV promoter. Transformed embryos containing the antisense RTBV promoter at the target site are selected and used to regenerate modified plants using techniques standard in the art.
[0281] RNA is extracted from various tissues (e.g., roots, stems, leaves, inflorescences) of modified corn plants identified as containing the antisense RTBV promoter. RNA is also extracted from control corn plants lacking the antisense RTBV promoter at the target site. Using suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing), reduced expression of GA20 oxidase_5 and / or GA20 oxidase_3 is confirmed in the stem and vascular tissues of modified corn plants containing the antisense RTBV promoter compared to control corn plants.
[0282] Example 8. Creating dominant alleles by engineering truncated proteins using genome editing techniques. A) Engineering GA20 oxidase truncated proteins: Dominant alleles can be generated by targeted editing of the gene, resulting in truncated proteins or nonsense mutations of the protein. See Figure 6. The maize GA20 oxidase_5 and GA20 oxidase_3 genes are highly similar in sequence and structure. Both genes contain three exons. gRNAs are designed to introduce edits into the exons of both genes.
[0283] The gRNA is introduced into maize cells using a transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains an expression cassette containing a promoter operable in plant cells operably linked to a polynucleotide encoding the above-mentioned gRNA.
[0284] Immature corn embryos are co-cultured with Agrobacterium containing the T-DNA vector for three days. The polynucleotide between the LB and RB sequences is integrated into the nuclear genome of the immature corn embryo. Upon expression of the integrated polynucleotide, the gRNA guides the nuclease to the target site, creating a double-stranded break at the target site. Without being bound by theory, the cell's inherent non-homologous end-joining repair mechanism frequently incompletely repairs such breaks, which can lead to the insertion or deletion of one or more nucleotides. Such insertions or deletions in exons can result in premature stop codons (producing truncated proteins) or nonsense mutations. Premature stop codons have the potential to generate dominant alleles of GA20 oxidase_5 and GA20 oxidase_3.
[0285] Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing targeted insertions or deletions in the GA20 oxidase_5 and / or GA20 oxidase_3 genes. Transformed embryos containing insertions or deletions at the target site that can introduce premature stop codons or nonsense mutations resulting in truncated proteins are selected and used to regenerate modified plants using techniques standard in the art.
[0286] Proteins are extracted from modified corn plants identified as containing the identified insertions / deletions. Using suitable methods known in the art (e.g., Western blot, HPLC, LC / MS, ELISA, immunoprecipitation), the introduction of truncated proteins or nonsense mutations into the GA20 oxidase_5 and / or GA20 oxidase_3 genes is confirmed. Additional experiments to confirm the reduction of gibberellic acid in stem and / or vascular tissues are performed as described in Bensen et al., Plant Physiol. 1990, 94:77-84, which is incorporated herein by reference in its entirety.
[0287] B) Engineering a truncated Brachytic 2 (Br2) protein: gRNAs are designed to introduce edits within exons of the Brachytic 2 gene from maize.
[0288] The gRNA is introduced into maize cells using a transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains an expression cassette containing a promoter operable in plant cells operably linked to a polynucleotide encoding the above-mentioned gRNA.
[0289] Immature corn embryos are co-cultured with Agrobacterium containing a T-DNA vector for three days. The polynucleotide between the LB and RB sequences is integrated into the nuclear genome of the immature corn embryo. Upon expression of the integrated polynucleotide, the gRNA guides a nuclease to the target site, where the nuclease creates a double-strand break. Without being bound by theory, the cell's inherent non-homologous end-joining repair mechanism frequently incompletely repairs such breaks, which can lead to the insertion or deletion of one or more nucleotides. Such insertions or deletions in exons can result in premature stop codons (producing truncated proteins) or nonsense mutations. Premature stop codons have the potential to generate dominant alleles of Brachytic 2 (Br2).
[0290] Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing targeted insertions or deletions in the Br2 gene. Transformed embryos containing insertions or deletions at the target site that can introduce premature stop codons (resulting in truncated proteins) or nonsense mutations are selected and used to regenerate modified plants using techniques standard in the art.
[0291] Protein is extracted from modified corn plants identified as containing the identified insertion / deletion, and production of the Br2 truncated protein is confirmed using a suitable method known in the art (e.g., Western blot, HPLC, LC / MS, ELISA, immunoprecipitation).
[0292] Example 9. Creation of inverted repeats in target genes Using targeted editing techniques, genomic loci can be converted into loci capable of generating an RNAi-inducing hairpin when the edited locus is transcribed into RNA. See Figure 7. In cells heterozygous at a locus of interest (e.g., two polymorphic alleles are present), one or more nucleases are used to generate two double-stranded breaks in a first allele (e.g., a first double-stranded break and a second double-stranded break) and one double-stranded break in a second allele (e.g., a third double-stranded break). When the nucleases cleave the first and second alleles, portions of the first allele flanking the first and second double-stranded breaks are released from the genomic DNA. In one result, the released portion of the first allele becomes inverted and integrates into the third double-stranded break in the second allele, thereby creating an edited locus capable of generating an RNAi-inducing hairpin when the edited locus is transcribed.
[0293] First and second functional guide RNAs (gRNAs) for the RNA-guided nuclease system are generated. The first and second gRNAs are complementary to a first target site and a second target site, respectively, flanking a portion of a first allele of the GA20 oxidase_5 gene in the maize genome. The first gRNA is also complementary to a second allele of the GA20 oxidase_5 gene at a third target site (homologous to the first target site), but the second gRNA is not complementary to the second allele due to a polymorphism between the first and second GA20 oxidase_5 alleles at the second target site.
[0294] A transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation is constructed. The T-DNA construct contains several expression cassettes between the left border (LB) and right border (RB) sequences. The first expression cassette contains a promoter operable in plant cells operably linked to a polynucleotide encoding an RNA-guided nuclease. The second expression cassette contains a promoter operable in plant cells operably linked to the CP4-EPSPS marker gene. The construct also contains expression cassettes containing promoters operable in plant cells operably linked to polynucleotides encoding the first and second gRNAs described above.
[0295] Immature maize embryos are co-cultured with Agrobacterium containing the T-DNA vector for three days. Upon expression of the polynucleotide, the gRNA guides a nuclease to each of three target sites within the GA20 oxidase_5 allele, where the nuclease creates a double-strand break at each target site.
[0296] In most events, the region between the first and second target sites in the first GA20 oxidase_5 allele is deleted, and non-homologous end joining repair mechanisms join the flanking regions. Less frequently, in some events, insertion / deletion mutations are created at the first target site, the second target site, or both. In still other events, the entire targeted region of the first allele is integrated in the reverse orientation into a double-stranded break at the third target site. See Figure 7, panel D. Transformation events containing the inversion in the second GA20 oxidase_5 allele are identified using suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing). Transformed embryos containing the inversion in the second GA20 oxidase_5 allele are selected and used to regenerate modified plants using techniques standard in the art.
[0297] Without being bound by any scientific theory, the presence of an inversion in one allele of GA20 oxidase_5 creates a population of RNA transcripts that can form hairpin structures. These hairpins induce the RNAi machinery in the cell, resulting in the downregulation of GA20 oxidase_5 RNA transcripts in a dominant manner (e.g., both the edited and unedited alleles are downregulated).
[0298] RNA is extracted from modified corn plants identified as containing an inversion in a second allele capable of producing a hairpin RNA transcript. RNA is also extracted from control corn plants lacking the edited GA20 oxidase_5 allele. Downregulation of GA20 oxidase_5 is confirmed in modified corn plants containing an inversion in a second allele capable of producing a hairpin RNA transcript using suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing). Furthermore, because GA20 oxidase_5 and GA20 oxidase_3 share sequence similarity, the inversion in the second allele of GA20 oxidase_5 also results in downregulation of GA20 oxidase_3 RNA transcripts.
[0299] Example 10. Insertion of miRNA target sites into desired genomic loci Targeted editing techniques can be used to insert a donor molecule into a target site in a genomic locus. See Figure 8. When a donor molecule containing a non-coding RNA target site is inserted into the 5'-UTR, exon, intron, or 3'-UTR of a gene of interest, RNA transcription or protein translation of the gene of interest can be suppressed by a complementary non-coding RNA. When the gene of interest is targeted by a non-coding RNA (e.g., miRNA or siRNA), the cleaved mRNA of the gene of interest can generate secondary siRNAs, which can further suppress the transcription or translation of the gene of interest. Because the secondary siRNA is complementary to the allele regardless of whether the non-coding RNA target site is inserted, such secondary suppression can act in a dominant manner.
[0300] A functional guide RNA (gRNA) complementary to a target site within the 3'-UTR of the GA20 oxidase_5 gene was engineered. A transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation was constructed. The T-DNA construct contained a) an RNA guide / nuclease, b) a CP4-EPSPS marker gene, c) the gRNA described above, and d) a promoter operable in plant cells operably linked to a polynucleotide encoding a donor molecule between the left border (LB) and right border (RB) sequences. The donor molecule contained a 21-nucleotide sequence homologous to miR166 and first and second homologous regions to the 3'-UTR of the GA20 oxidase_5 gene on either side of the target site.
[0301] Immature maize embryos were co-cultured with Agrobacterium containing the T-DNA vector for three days. Upon expression of the polynucleotide, the gRNA guided a nuclease to the target site within the 3'-UTR of GA20 oxidase_5, where the nuclease created a double-strand break. Homologous recombination repair then inserted the donor molecule into the target site, thereby integrating the miR166 target site into the 3'-UTR of the GA20 oxidase_5 gene.
[0302] Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing the miR166 target site insertion into the 3'-UTR of the GA20 oxidase_5 gene. Transformed embryos containing the miR166 target site insertion into the 3'-UTR of the GA20 oxidase_5 gene are selected and used to regenerate modified plants using techniques standard in the art.
[0303] Without being bound by any scientific theory, the presence of miRNA binding sites in GA20 oxidase_5 generates a population of secondary siRNA transcripts that can suppress GA20 oxidase_5 RNA transcription in a dominant-negative manner.
[0304] RNA is extracted from modified corn plants identified as containing miR166 target site insertions in the 3'-UTR of the GA20 oxidase_5 gene. RNA is also extracted from control corn plants lacking the edited GA20 oxidase_5 gene. Downregulation of GA20 oxidase_5 is confirmed in the modified corn plants using suitable methods known in the art (e.g., quantitative reverse transcriptase PCR, reverse transcriptase PCR, RNA sequencing). Furthermore, due to the sequence similarity between GA20 oxidase_5 and GA20 oxidase_3, the miRNA target site in the GA20 oxidase_5 gene can also cause downregulation of GA20 oxidase_3 RNA transcripts.
[0305] Example 11. Generation of dominant-negative alleles by creating truncated proteins Targeted editing of a gene (e.g., nonsense mutation) resulting in a truncated protein can create a dominant-negative allele. In one embodiment, the targeted gene encodes a protein with a protein:protein interaction domain. See Figure 6. Examples of dominant truncated proteins are known in various plant species. For example, dominant mutant phenotypes caused by truncated proteins are known for FAZ1 in rice and AGAMOUS and SOC1 in Arabidopsis thaliana. The peptide CLAVATA3 (CLV3) is processed into a signal peptide (CLE), which is bound by receptor-like kinases CLV1 and CORYNE (CRN) and the receptor-like protein CLV2 to regulate the expression of WUSCHEL (WUS) in the meristem. Maize plants containing mutant alleles of CLV2 often exhibit ears with an increased number of grain rows. Without being limited to a particular theory, mutation of CLV2 may increase WUS expression, which increases the size of the meristem and result in an increased number of grain rows in the ear. CLV2 contains an extracellular domain and a transmembrane domain and forms a complex with CRN. Truncated CLV2 proteins can function in a dominant manner to increase maize meristem size.
[0306] A gRNA is designed to introduce a stop codon into the extracellular domain of maize CLV2. The gRNA is introduced into maize cells using a transfer DNA (T-DNA) vector suitable for use in Agrobacterium transformation. The T-DNA construct contains a) an RNA-guided nuclease, b) a CP4-EPSPS marker gene, and c) a promoter operable in plant cells operably linked to a polynucleotide encoding the gRNA, between the left border (LB) and right border (RB) sequences.
[0307] Immature corn embryos are co-cultured with Agrobacterium containing the T-DNA vector for three days. The polynucleotide between the LB and RB sequences is integrated into the nuclear genome of the immature corn embryo. Upon expression of the integrated polynucleotide, the gRNA guides the nuclease to the target site, creating a double-stranded break at the target site. Without being bound by any theory, the cell's intrinsic non-homologous end-joining repair mechanism frequently incompletely repairs such breaks, which can lead to the insertion or deletion of one or more nucleotides. Such insertions or deletions in exons can result in premature stop codons (producing truncated proteins) or nonsense mutations. Such mutations can generate dominant-negative alleles of CLV2.
[0308] Suitable methods known in the art (e.g., PCR, DNA hybridization (Southern) blot, sequencing) are used to identify transformation events containing targeted insertions or deletions in the CLV2 gene. Transformed embryos containing insertions or deletions at the target site that result in truncated proteins are selected and used to regenerate modified plants using techniques standard in the art.
[0309] Proteins are extracted from modified corn plants identified as containing the identified insertion / deletion. The introduction of the nonsense mutation into the CLV2 gene is confirmed using suitable methods known in the art (e.g., Western blot, HPLC, LC / MS, ELISA, immunoprecipitation). Additional phenotypic screening of the grain rows of the panicle is performed for increased meristem size. Light microscopy of meristem sections from modified and control plants is also performed to quantify the increase in meristem size.
[0310] Example 12. Prevention of target gene cleavage using engine...
Claims
1. 1. A method of generating a dominant negative allele of a gene in a cell, comprising using targeted editing technology to invert a portion of the gene to generate an antisense RNA transcript capable of inducing suppression of an unaltered allele of the gene.
2. 1. A method for generating a dominant negative allele of a gene in one or more cells, comprising: a) inducing a first double-stranded break and a second double-stranded break flanking the targeted region of the gene; b) identifying one or more cells containing an inversion of the targeted region of the gene, wherein the inversion results in the production of an antisense RNA transcript from the targeted region; c) selecting one or more cells containing said inversion of said targeted region of said gene.
3. 1. A method for reducing expression of a protein in a cell, comprising: a) inducing a first double-stranded break and a second double-stranded break at locations flanking a targeted region of a chromosome; and b) identifying one or more cells containing an inversion in the targeted region of the chromosome, wherein expression of the protein is reduced compared to a control cell that does not contain the inversion within the targeted region.
4. 1. A method for generating an inversion in a targeted region of a gene, comprising: a) providing at least one RNA-guided nuclease or one or more vectors encoding at least one RNA-guided nuclease to one or more cells, the at least one RNA-guided nuclease is capable of binding to a stretch of at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, or at least 26 nucleotides of a first target site and a second target site flanking the targeted region of the gene; the first target site and the second target site are linked; providing, wherein the at least one RNA-guided nuclease creates a double-stranded break at the first target site and the second target site of the gene; b) identifying one or more cells containing an inversion in the targeted region of the gene, wherein the inversion results in the production of an antisense RNA transcript from the targeted region; c) selecting one or more cells containing said inversion in said targeted region of said gene.
5. a) using a targeted editing technique to generate a first double-strand break (DSB) and a second DSB in a first allele of a gene in a cell; b) using a targeted editing technique to generate a third DSB in a second allele of the gene in the cell; and c) identifying cells that contain an insertion of a region of the first allele in a reverse orientation at the site of the third DSB in the second allele, thereby generating a modified second allele.
6. a) providing to a cell an engineered pentatricopeptide repeat (PPR) protein or a vector encoding the engineered PPR protein operably linked to a promoter, wherein the engineered PPR protein is capable of binding to an RNA transcript of a target gene; b) selecting one or more cells from step (a) that express the engineered PPR protein; and and c) identifying one or more cells selected in step (b) that contain altered expression of said target gene.
7. 1. A method of generating a dominant negative allele of a gene in a cell, comprising using targeted editing technology to insert an inverted copy of a gene or portion thereof adjacent to a native copy of the gene to generate an inverted repeat sequence capable of producing an antisense RNA transcript of the gene or portion thereof.
8. A modified plant cell comprising a non-transposon-mediated genomic deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
9. A modified chromosome comprising a non-transposon-mediated deletion or inversion of a gene or portion thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or portion thereof.
10. A modified cell comprising the modified chromosome of claim 9.
11. A modified plant or modified plant tissue regenerated from the modified plant cell of claim 8.
12. A product comprising the modified chromosome of claim 9.
13. A product comprising the modified cell of claim 10.
14. A modified plant or part thereof comprising a non-transposon-mediated genomic deletion or inversion of a gene or part thereof at the gene's endogenous locus, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene or part thereof.
15. 15. A product comprising the modified plant or part thereof of claim 14.
16. A modified cell comprising a non-transposon-mediated genomic deletion or inversion of at least one gene or a portion thereof at the endogenous locus of the gene, wherein the deletion or inversion results in the production of an RNA transcript comprising a sequence complementary to the native transcript sequence of the gene.
17. 1. A modified cell comprising a targeted edit of at least one gene or portion thereof, wherein the targeted edit produces an RNA transcript that is complementary to the native transcript sequence of the gene.
18. 1. A modified cell comprising at least one dominant negative allele of at least one gene generated by targeted editing technology, wherein when the at least one dominant negative allele is transcribed, the allele produces an RNA transcript capable of forming a hairpin loop secondary structure.
19. A modified cell comprising a dominant negative allele of at least one gene, the cell comprising an inverted copy of the gene adjacent to a native copy of the gene at the endogenous locus of the gene.
Citation Information
Patent Citations
Gene modification-mediated methods and compositions for generating dominant traits in eukaryotic systems
WO2015117041A1
Compositions and methods for treatment of cystic fibrosis
WO2017143061A1
Genome editing method
WO2018097257A1
EXON deletion correction of duchenne muscular dystrophy mutations in the dystrophin actin binding domain 1 using crispr genome editing
WO2019036599A1
Methods, compositions and components for crispr-CAS9 editing of tgfbr2 in t cells for immunotherapy
WO2019089884A2