Inhibition of the shade avoidance response in plants
By introducing non-natural mutations into plant HD-Zip transcription factors and utilizing CRISPR technology to suppress the shade avoidance response, the problem of growth inhibition caused by insufficient light in plants was solved, resulting in growth promotion and yield increase.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PAIRWISE PLANTS SERVICES INC
- Filing Date
- 2021-01-29
- Publication Date
- 2026-05-15
AI Technical Summary
Plants experience growth inhibition and yield reduction due to insufficient light during their shade response, and current technologies struggle to effectively suppress this survival mechanism.
By introducing non-natural mutations into the HD-Zip transcription factor in plants, and utilizing CRISPR-related effector proteins and guide nucleic acids, the binding of the HD-Zip transcription factor to DNA is disrupted or reduced, thereby decreasing its activity and inhibiting the shade avoidance response.
It effectively reduces the shade avoidance response of plants, promotes growth and increases yield, avoids resource waste, and enhances the survival ability of plants in competitive environments.
Smart Images

Figure CN115335392B_ABST
Abstract
Description
[0001] Declaration regarding the electronic submission of the sequence list
[0002] A sequence of ASCII text files, titled 1499.17.WO_ST25.txt, of size 465,868 bytes, was generated on January 28, 2021, and submitted via the EFS website, in lieu of a paper copy, as per 37 CFR § 1.821. This sequence is incorporated herein by reference to the specification incorporating its public content.
[0003] Priority Statement
[0004] This application claims the benefit of U.S. Provisional Application No. 62 / 968,596, filed January 31, 2020, pursuant to 35 USC §119(e), the entire contents of which are incorporated herein by reference. Technical Field
[0005] This invention relates to compositions and methods for modifying homologous domain-leucine zipper (HD-Zip) transcription factors to suppress shade avoidance responses in plants. The invention also relates to plants produced using the methods and compositions of this invention. Technical Background
[0006] Shade avoidance response (SAR) is a response to a decrease in the quality or quantity of available light (Kebrom and Brutnell, J ExpBot 58:3079–3089 (2007)), in which plants attempt to outcompete neighboring plants by growing toward a resource (primarily light). Overcrowding can lead to shade avoidance syndrome (SAS), in which plants become less vigorous and produce lower yields. Shade avoidance involves the relative proportions of red to far-red light in a plant's environment (Ballare et al., Science, 247:329–332 (1990)). Plants absorb most of the red light they have access to but reflect far-red light, including reflecting it toward nearby plants. When plants detect a persistent far-red light in their environment, they undergo morphological and physiological responses. These responses include reduced branching, increased plant height, reduced leaf area, auxin redistribution, increased ethylene production, and accelerated flowering. SAS is characterized by an increased root-to-shoot ratio, increased plant height, and reduced yield per plant. In typical monoculture environments, interplant competition through shade avoidance is considered a wasteful survival mechanism. Summary of the Invention
[0007] One aspect of the invention provides a plant or plant part thereof containing at least one non-naturally mutated endogenous homologous domain-leucine zipper (HD-Zip) transcription factor, wherein the mutation disrupts the binding of the HD-Zip transcription factor to DNA.
[0008] Another aspect of the invention provides a plant cell comprising an editing system, the editing system comprising: (a) a CRISPR-related effector protein; and (b) a guide nucleic acid (gRNA, gDNA, crRNA, crDNA) having a spacer sequence complementary to an endogenous target gene encoding an HD-Zip transcription factor.
[0009] Another aspect of the invention provides a plant cell containing a non-natural mutation within the DNA binding site of an HD-Zip transcription factor gene that prevents or reduces the binding of the HD-Zip transcription factor to DNA, wherein the mutation is a substitution, insertion, and / or deletion introduced using an editing system comprising a nucleic acid binding domain that binds to a target site in the HD-Zip transcription factor gene, wherein the HD-Zip transcription factor gene encodes: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the following amino acid sequence: RKKLRLSKDQSAVLEDSFREHPTLNPRQKAALAQQLGLRPRQVEVWFQNRRARTKLKQTEVDCEYLKRCCETLTEENRRLQKEVQELRALKLVSPHLYMHMSPPTTLTMCPSCERV (SEQ ID NO:38) NO:1)(Corn HB53) or RKKLRLSKDQAAVLEESFKEHNTLNPKQKAALAKQLNLKPRQVEVWFQNRRARTKLKQTEVDCEFLKRCCETLTEENRRLQREVAELRVLKLVAPHHYARMPPPTTLTMCPSCERL(SEQ ID NO:2)(Corn HB78); (c) a polypeptide containing a sequence having at least 80% sequence identity with the amino acid sequence LAKQLNLKPRQVEVWFQNRRARTKLKQTEVDCEFLKRCCETLTEENRRLQREV(SEQ ID NO:3); (d) a polypeptide containing a sequence having at least 95% sequence identity with the nucleotide sequence RQVEVWFQNRRARTKLKQTEVDCE(SEQ ID No:4); (e) a polypeptide containing the amino acid sequence RQVEVWFQNRRARTKXKQTEVDCE(SEQ ID No:4). (f) A polypeptide comprising: (i) a sequence having the amino acid sequence KKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; and (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L;(iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide comprising a sequence having the amino acid sequence shown in VWFQNRRA (SEQ ID NO:9).
[0010] Another aspect of the invention provides a plant or a portion thereof containing a mutation (e.g., at least one mutation) in an endogenous HD-Zip transcription factor that reduces the binding of the endogenous HD-Zip transcription factor to DNA, wherein the endogenous HD-Zip transcription factor comprises a polypeptide containing a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; wherein the mutation is a deletion, substitution, and / or insertion of at least one amino acid residue of amino acid residues 45-52 (VWFQNRRA (SEQ ID NO:9)) of the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2, optionally wherein the mutation is a non-natural mutation.
[0011] Another aspect of the present invention provides a plant or a portion thereof comprising an HD-Zip transcription factor gene, said gene encoding the amino acid sequence of SEQ ID NO:201.
[0012] Another aspect of the present invention provides a maize plant containing an HD-Zip transcription factor gene, said gene containing the amino acid sequence of SEQ ID NO:201.
[0013] Another aspect of the present invention provides a plant or a portion thereof comprising an HD-Zip transcription factor gene, said gene comprising the nucleotide sequence of SEQ ID NO:202.
[0014] Another aspect of the present invention provides a maize plant containing an HD-Zip transcription factor gene, said gene containing the nucleotide sequence of SEQ ID NO:202.
[0015] The present invention also provides a method for producing / cultivating non-transgenic genome-edited plants, comprising: (a) hybridizing the plant of the present invention with a non-transgenic plant to introduce a mutation present in the plant of the present invention into the non-transgenic plant; and (b) selecting progeny plants containing the mutation but without transgenes to produce non-transgenic genome-edited plants.
[0016] Another aspect of the present invention provides a method for editing a specific site in the genome of a plant cell, the method comprising: cutting a target site within an endogenous HD-Zip transcription factor gene in a site-specific manner, the endogenous HD-Zip transcription factor gene encoding: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (f) a polypeptide comprising: (i) having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:38). (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide containing the amino acid sequence VWFQNRRA (SEQ ID NO:9), thereby producing editing in the endogenous HD-Zip transcription factor gene of plant cells.
[0017] Another aspect of the present invention provides a method for manufacturing a plant, comprising: (a) contacting a plant cell population containing a wild-type endogenous gene encoding an HD-Zip transcription factor with a nuclease targeting the wild-type endogenous gene, wherein the nuclease is linked to a DNA-binding domain, said binding domain being bound to a nucleic acid sequence encoding the following sequences: (i) a polypeptide comprising a sequence having at least 80% sequence identity with an amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (ii) a polypeptide comprising a sequence having at least 80% sequence identity with an amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (iii) a polypeptide comprising a sequence having at least 80% sequence identity with an amino acid sequence of SEQ ID NO:3; (iv) a polypeptide comprising a sequence having at least 95% sequence identity with a nucleotide sequence of SEQ ID NO:4; (v) a polypeptide comprising a sequence having an amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (vi) a polypeptide comprising: (1) having an amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:38) (1) A sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (2) A sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (3) A sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (4) A sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9); and / or (vii) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 6). (a) a polypeptide of the sequence NO:9); (b) selecting from the population plant cells containing a mutant wild-type endogenous gene encoding the HD-Zip transcription factor, wherein the mutation is a substitution and / or deletion of at least one amino acid residue in the polypeptide of any one of (i)-(v), wherein the mutation reduces or eliminates the ability of the HD-Zip transcription factor to bind DNA; and (c) growing the selected plant cells into plants.
[0018] In some aspects, the present invention provides a method for reducing shade avoidance response in plants, comprising (a) contacting a plant cell containing a wild-type endogenous gene encoding an HD-Zip transcription factor with a nuclease targeting the wild-type endogenous gene, wherein the nuclease is linked to a DNA-binding domain that binds to a target site in the wild-type endogenous gene, said wild-type endogenous gene encoding: (i) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (ii) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (iii) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (iv) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (v) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (vi) a polypeptide comprising: (1) having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:5). (1) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (2) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (3) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (4) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9); and / or (vii) a polypeptide containing the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9), thereby producing mutant plant cells containing a wild-type endogenous gene encoding the HD-Zip transcription factor; and (b) causing plant cells to grow into plants, thereby reducing the shade avoidance response.
[0019] On the other hand, a method is provided for producing a plant or a portion thereof containing a cell having at least one mutated endogenous HD-Zip transcription factor gene, the method comprising contacting a target site in the HD-Zip transcription factor gene in the plant or plant portion with a nuclease comprising a cleavage domain and a DNA-binding domain, wherein the DNA-binding domain binds to the target site in the HD-Zip transcription factor gene, wherein the HD-Zip transcription factor gene encodes: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38; (f) A polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID No:9); and / or (g) a sequence having the amino acid sequence VWFQNRRA (SEQ ID No:9). A polypeptide with the sequence NO:9 is used to produce a plant or part thereof containing at least one cell with a mutation in the endogenous HD-Zip transcription factor gene.
[0020] In another aspect, a method is provided for producing a plant or a portion thereof comprising a mutant endogenous HD-Zip transcription factor having reduced DNA binding, the method comprising contacting a target site in an endogenous HD-Zip transcription factor gene in the plant or plant portion with a nuclease comprising a cleavage domain and a DNA-binding domain, wherein the DNA-binding domain binds to the target site in the HD-Zip transcription factor gene, wherein the HD-Zip transcription factor gene encodes: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38; (f) A polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9). A polypeptide with the sequence NO:9 is used to produce plants or parts thereof with a mutant endogenous HD-Zip transcription factor that has reduced DNA binding.
[0021] Another aspect of the invention provides a guide nucleic acid that binds to a target site in an HD-Zip transcription factor gene, the target site comprising a nucleotide sequence encoding the following sequences: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (f) a polypeptide comprising the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9).
[0022] Another aspect of the invention provides a system comprising the guide nucleic acid of the invention and a CRISPR-Cas effector protein associated with the guide nucleic acid.
[0023] Another aspect of the invention provides a gene editing system comprising a CRISPR-Cas effector protein associated with a guide nucleic acid, wherein the guide nucleic acid comprises a spacer sequence that binds to an HD-Zip transcription factor gene.
[0024] Another aspect of the present invention provides a complex comprising a CRISPR-Cas effector protein comprising a cleavage domain and a guide nucleic acid, wherein the guide nucleic acid binds to a target site in an HD-Zip transcription factor gene encoding (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9) and / or (f) a polypeptide containing the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9), wherein the cleavage domain cleaves the target strand in the HD-Zip transcription factor gene.
[0025] Another aspect of the invention provides a nucleic acid encoding an HD-Zip transcription factor having a mutated DNA binding site, wherein the mutated DNA binding site contains a mutation that disrupts DNA binding.
[0026] Plants containing one or more mutated HD-Zip transcription factors in their genomes are also provided, as well as polypeptides, polynucleotides, nucleic acid constructs, expression cassettes, and vectors for preparing the plants of the present invention, wherein the transcription factors have reduced ability to bind to DNA produced by the methods of the present invention.
[0027] These and other aspects of the invention will be set forth in more detail in the following description of the invention.
[0028] Brief description of the sequence
[0029] SEQ ID NO:1 shows a portion of the corn HB53 transcription factor (amino acid residues 173-288 of SEQ ID NO:38).
[0030] SEQ ID NO:2 shows a portion of the corn HB 78 transcription factor (amino acid residues 76-191 of SEQ ID NO:83).
[0031] SEQ ID NO:3-9 is a partial sequence of the HD-zip transcription factor.
[0032] SEQ ID NO:10-54 is an example of HB53 transcription factors from a variety of different plant species.
[0033] SEQ ID NO:55-98 and 303 are examples of HB78 transcription factors from a variety of different plant species.
[0034] SEQ ID NO:99-134 is a cDNA sequence of the HB53 transcription factor from a variety of different plant species.
[0035] SEQ ID NO:135-174 is the cDNA sequence of HB78 transcription factor from a variety of different plant species.
[0036] SEQ ID NO:175-182 shows exemplary spacer sequences for guiding nucleic acid targeting of HB78 and HB53 transcription factors.
[0037] SEQ ID NO:183-194 are exemplary cytosine deaminase amino acid sequences that can be used in this invention.
[0038] SEQ ID NO:195 is an exemplary uracil-DNA glycosylation inhibitor (UGI) that can be used in this invention.
[0039] SEQ ID NO:196-197 are exemplary regulatory sequences encoding promoters and introns.
[0040] SEQ ID NO:198-200 provides an example of the motif position adjacent to the prespacer of a type V CRISPR-Cas12a nuclease.
[0041] SEQ ID NO:201 provides an exemplary HB78 mutant amino acid sequence generated using the methods and compositions of the present invention.
[0042] SEQ ID NO:202 provides an exemplary HB78 mutant nucleotide sequence generated using the methods and compositions of the present invention. SEQ ID NO:202 encodes the amino acid sequence of SEQ ID NO:201.
[0043] SEQ ID NO:203-246 and 302 are Figure 1A-1BThe comparison shows partial HB78 amino acid sequences of different plant species.
[0044] SEQ ID No:247-291 is Figures 2A-2D The comparison shows partial HB53 amino acid sequences of different plant species.
[0045] SEQ ID NO:292-294 is Figure 6 The HB53 sequence shown.
[0046] SEQ ID NO:295-297 is Figure 8 The HB78 sequence shown.
[0047] SEQ ID NO:298-301 is Figure 9 The HB78 sequence shown.
[0048] SEQ ID NO:304 and SEQ ID NO:305 provide the genomic sequence and cDNA of wild-type HB78 from corn (corresponding to the WT HB78 amino acid sequence of SEQ ID NO:83 and the WT HB78 coding sequence of SEQ ID NO:171, respectively).
[0049] SEQ ID NO:306-309 provide the protein sequence, genomic sequence, coding sequence, and cDNA of HB78 with a 17-base-pair deletion, respectively.
[0050] SEQ ID NO:310 and SEQ ID NO:311 provide the genomic sequence and cDNA of wild-type HB53 from corn (corresponding to the WT HB53 amino acid sequence of SEQ ID NO:38 and the WT HB53 coding sequence of SEQ ID NO:132, respectively).
[0051] SEQ ID NO:312-315 provide the protein sequence, genomic sequence, coding sequence, and cDNA of HB53 with an 11-base-pair deletion, respectively.
[0052] SEQ ID NO:316-319 provide the protein sequence, genomic sequence, coding sequence, and cDNA of HB53 with an 8-base-pair deletion, respectively.
[0053] SEQ ID No:320-329 is Figures 10-12 the sequence shown in . Attached Figure Description
[0054] Figure 1 shows the alignment of the HD-Zip (HB78) amino acid sequences from 44 different plant species. The sequences are a portion of the full-length HB78 sequence consisting of consecutive amino acids (53 amino acid residues).
[0055] Figure 2 provides an alignment of the HD-Zip (HB53) amino acid sequences from 45 different plant species. The sequences are a portion of the full-length HB53 sequence consisting of consecutive amino acid residues (e.g., approximately 116 amino acid residues).
[0056] Figure 3 Examples illustrating the relationship between maize planting density and yield are provided. Increasing planting density increases plant yield until an inflection point is reached (arrow - optimal economic planting rate). Varieties with reduced shading also reach an inflection point at higher planting densities.
[0057] Figure 4 An example of a dominant negative strategy is provided. By removing the DNA-binding ability of a bifunctional protein, the dimerized complex will not activate gene expression.
[0058] Figure 5 The relationship between HDLZ class II proteins is shown.
[0059] Figure 6 Examples of editing the DNA-binding domain of HB53 in corn are provided, and exemplary target amino acid residues for modification are shown in boxes. The coding strand (SEQ ID NO:292) and the non-coding strand (SEQ ID NO:293) and the amino acid sequence of HB53 (SEQ ID NO:294) are shown.
[0060] Figure 7 The HB78 gene with annotations of exemplary guide nucleic acids is provided.
[0061] Figure 8 Examples of deletions in the HB78 gene (SEQ ID No:295) are provided, showing that deletions, for example, in exon 2 (SEQ ID No:297) result in truncation of exon 3, exon 4, and the DNA-binding domain. The protein sequence at the deletion site (SEQ ID NO:296) is also shown.
[0062] Figure 9Representative genome sequences of the edited plant (coding strand (SEQ ID NO:298) and non-coding strand (SEQ ID NO:299)) are provided, showing premature termination upstream of the HB78 DNA-binding domain. Protein sequences resulting from this premature termination are shown (SEQ ID NO:300 and SEQ ID NO:301).
[0063] Figure 10 The alignment of the positions of the bootstrap PWsp227 and the edits obtained in corn HB53 (co-occurrence sequence SEQ ID NO:320; corn SEQ ID NO:321; CE28392, CE28403, CE28409, CE28382, CE28390 SEQ ID NO:322; CE28492, CE28505, CE28514, CE28517, CE28522, CE28534, CE28544, CE28547 SEQ ID NO:320); CE28456, CE28477, CE28483 SEQ ID NO:323) is provided.
[0064] Figure 11 The comparison of the positions of the bootstrap PWsp227 and PWsp225 and the edits obtained in corn HB53 (common (common sequence SEQ ID NO:324; corn SEQ ID NO:325; CE28330…8D SEQ ID NO:326; CE28350…11D SEQ ID NO:327) is provided.
[0065] Figure 12 The comparisons provided show the positions of the bootstrap PWsp230 and PWsp232 and the edits obtained in corn HB78 (common sequences SEQ ID NO:328; Z. mays SEQ ID NO:329; CE28330 SEQ ID NO:328; CE28350 SEQ ID No:328).
[0066] Figure 13 Examples targeting HB53 (top) and HB78 (bottom) are provided. Plasmids pWISE443 and pWISE444 are shown with corresponding spacers. Plasmid pWISE445 contains all four spacers shown for HB53. Plasmids pWISE446 and pWISE447 are shown with corresponding spacers. Plasmid pWISE448 contains all four spacers shown for HB78. Plasmid pWISE451 contains all eight spacers (four for HB53, four for HB78, etc.). Figure 13 (As shown).
[0067] Figure 14 Exemplary E2 shade test results for HB53 knockout and HB53 / HB78 knockout are provided, showing no stem elongation in the edited lines when grown in shade.
[0068] Figure 15 Off-type analysis was provided, and no evidence of morphological heterosis or developmental delay was found in the edited plants. Detailed Implementation
[0069] The invention will now be described below with reference to the accompanying drawings and examples, which illustrate embodiments of the invention. This description is not intended to be a detailed list of all the different ways in which the invention may be practiced or all the features that may be added to the invention. For example, features described with respect to one embodiment may be incorporated into other embodiments, and features described with respect to a particular embodiment may be removed from that embodiment. Therefore, the invention contemplates that in some embodiments of the invention, any features or combinations of features set forth herein may be excluded or omitted. Furthermore, many variations and additions to the various embodiments presented herein will be apparent to those skilled in the art based on this disclosure, without departing from the invention. Therefore, the following description is intended to illustrate some specific embodiments of the invention, rather than to exhaustively describe all permutations, combinations, and variations thereof.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in the description of this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention.
[0071] All publications, patent applications, patents and other references cited in this article are incorporated in their entirety by reference for the purpose of teaching in relation to the sentences and / or paragraphs in which the citations are presented.
[0072] Unless otherwise stated in the context, the various features of the invention described herein can be used in any combination. Furthermore, the invention contemplates that in some embodiments, any feature or combination of features set forth herein may be excluded or omitted. For illustration, if the specification states that a composition comprises components A, B, and C, it is particularly intended that any one of A, B, or C, or any combination thereof, may be omitted and abandoned individually or in any combination.
[0073] As used in the description of the invention and the appended claims, unless the context clearly indicates otherwise, the singular forms “a”, “an”, and “the” are also intended to include the plural forms.
[0074] Furthermore, as used herein, “and / or” means and includes any and all possible combinations of one or more of the related listed items, as well as combinations that do not exist when interpreted in the alternative (“or”).
[0075] As used herein, the term “about” when referring to a measurable value such as amount or concentration means including a variation of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1% of the specified value, as well as the specified value itself. For example, “about X” (where X is a measurable value) means including X as well as variations of X of ±10%, ±5%, ±1%, ±0.5%, or even ±0.1%. The ranges of measurable values provided herein may include any other ranges and / or individual values thereof.
[0076] As used in this article, phrases such as “between X and Y” and “between about X and Y” should be interpreted as including both X and Y. As used in this article, phrases such as “between about X and Y” mean “between about X and about Y”, while phrases such as “from about X to Y” mean “from about X to about Y”.
[0077] Unless otherwise stated herein, the descriptions of numerical ranges herein are intended only as a shorthand for individually referring to each individual value falling within that range, and each individual value is incorporated into the specification as if it were described separately herein. For example, if the range 10 to 15 is disclosed, then 11, 12, 13, and 14 are also disclosed.
[0078] As used herein, the terms “comprise”, “comprises”, and “comprising” specify the presence of the said feature, integer, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0079] As used herein, the transitional phrase “consistently of…” means that the scope of the claim should be interpreted to include the specific materials or steps described in the claim, as well as those materials or steps that do not materially affect one or more essential and novel features of the claimed invention. Therefore, when used in the claims of this invention, the term “consistently of…” is not intended to be equivalent to “comprising”.
[0080] As used herein, the terms “increase,” “increasing,” “increased,” “enhance,” “enhanced,” “enhancing,” and “enhancement” (and their grammatical variations) describe an increase of at least about 15%, 20%, 25%, 50%, 75%, 100%, 150%, 200%, 300%, 400%, 500%, or more compared to a control.
[0081] As used herein, the terms “reduce,” “reduced,” “reducing,” “reduction,” “diminish,” and “decrease” (and their grammatical variations) describe, for example, a reduction of at least about 5%, 10%, 15%, 20%, 25%, 35%, 50%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or 100% compared to a control. In certain embodiments, the reduction may result in no or substantially no (i.e., a negligible amount, e.g., less than about 10% or even 5%) detectable activity or amount.
[0082] As used herein, the terms “express,” “expresses,” “expressed,” or “expression,” etc., relating to nucleic acid molecules and / or nucleotide sequences (e.g., RNA or DNA) indicate that the nucleic acid molecules and / or nucleotide sequences are transcribed and, optionally, translated. Thus, nucleic acid molecules and / or nucleotide sequences can express target polypeptides, or, for example, functional untranslated RNA.
[0083] "Heterologous" or "recombinant" nucleotide sequences are nucleotide sequences that are not naturally associated with the host cell to which they are introduced, including multiple non-natural copies of naturally occurring nucleotide sequences.
[0084] "Natural" or "wild-type" nucleic acid, nucleotide, polypeptide, or amino acid sequence refers to a naturally occurring or endogenous nucleic acid, nucleotide, polypeptide, or amino acid sequence. Therefore, for example, "wild-type mRNA" is mRNA naturally present in a reference organism or endogenous mRNA of a reference organism.
[0085] As used in this article, the term "heterozygous" refers to the genetic state in which different alleles are located at corresponding loci on homologous chromosomes.
[0086] As used in this article, the term "homozygous" refers to the genetic state in which the same alleles are located at corresponding loci on homologous chromosomes.
[0087] As used in this article, the term "allelic gene" refers to one of two or more different nucleotides or nucleotide sequences that are present at a particular locus.
[0088] A "locus" is the location on a chromosome where a gene, marker, or allele is located. In some implementations, a locus may contain one or more nucleotides.
[0089] As used herein, the terms “required allele,” “target allele,” and / or “goal allele” are used interchangeably and refer to an allele associated with a desired trait. In some embodiments, depending on the nature of the desired phenotype, the desired allele may be associated with an increase or decrease (relative to a control) in or in a given trait. In some embodiments of the invention, the phrases “required allele,” “target allele,” or “goal allele” refer to one or more alleles associated with increased plant yield under non-water stress conditions, relative to a control plant that does not have one or more target alleles.
[0090] A marker is "associated" with a trait when it is linked to that trait, and when the presence of that marker is an indicator of whether and / or to what extent the desired trait or form will appear in the plant / germplasm containing that marker. Similarly, a marker is "associated" with an allele or chromosomal spacer when it is linked to that allele or chromosomal spacer, and when the presence of that marker is an indicator of whether the allele or chromosomal spacer is present in the plant / germplasm containing that marker.
[0091] As used herein, the terms “backcross” and “backcrossing” refer to the process of backcrossing a progeny plant with one of its parents once or multiple times (e.g., 1, 2, 3, 4, 5, 6, 7, 8, etc.). In a backcross scheme, the “donor” parent is the parent plant that possesses the desired gene or locus to be introgressed. The “recipient” parent (used once or multiple times) or “recurrent” parent (used twice or more) refers to the parent plant in which the gene or locus is gradually introgressed. For example, see Ragot, M. et al., Marker-assisted Backcrossing: A Practical Example, in Techniques et Utilisations des Marqueurs Moleculaires Les Colloques, Vol. 72, pp. 45-56 (1995); and Openshaw et al., Marker-assisted Selection in Backcross Breeding, in Proceedings of the Symposium "Analysis of Molecular Marker Data," pp. 41-43 (1994). The initial cross produces the F1 generation. The term "BC1" refers to the second use of the cyclic parent, "BC2" refers to the third use of the cyclic parent, and so on.
[0092] As used herein, the term "cross" or "crossed" refers to the production of offspring (e.g., cells, seeds, or plants) through the fusion of gametes via pollination. This term includes sexual hybridization (pollination of one plant to another) and self-pollination (self-pollination, such as when pollen and ovules come from the same plant). The term "cross" refers to the act of producing offspring through the fusion of gametes via pollination.
[0093] As used herein, the terms “introgression,” “introgressing,” and “introgressed” refer to the natural and artificial transfer of a desired allele or combination of desired alleles at one or more loci from one genetic background to another. For example, a desired allele at a particular locus can be transferred to at least one offspring through sexual hybridization between two parents of the same species, where at least one parent carries the desired allele in its genome. Alternatively, for example, allele transfer can occur, for instance, in fused protoplasts via recombination between two donor genomes, where at least one donor protoplast carries the desired allele in its genome. The desired allele can be a selected allele, such as a marker, QTL, transgene, etc. Offspring containing the desired allele can be backcrossed once or multiple times (e.g., 1, 2, 3, 4, or more times) with a line having the desired genetic background, selecting for the desired allele, resulting in the desired allele becoming fixed in the desired genetic background. For example, a marker associated with increased yield under non-water stress conditions can be introgressed from a donor into a recurrent parent that does not contain the marker and does not show increased yield under non-water stress conditions. The resulting offspring can then be backcrossed once or multiple times and selected until the offspring possess the genetic marker associated with increased yield under non-water stress conditions in the context of the recurrent parent.
[0094] A genetic map is a description of the genetic linkages between loci on one or more chromosomes in a given species, typically presented as a graph or table. For each genetic map, the distance between loci is measured by the frequency of recombination between them. Various markers can be used to detect recombination between loci. A genetic map is the product of the mapping population, the types of markers used, and the polymorphic potential of each marker across different populations. The order and genetic distance between loci can vary from genetic map to genetic map.
[0095] As used herein, the term "genotype" refers to the genetic makeup of an individual (or population of individuals) at one or more loci, as opposed to an observable and / or detectable and / or expressed trait (phenotype). A genotype is defined by one or more alleles at one or more known loci inherited by an individual from its parents. The term genotype can be used to refer to the genetic makeup of an individual at a single locus, multiple loci, or more generally, to the individual genetic makeup of all genes in an individual's genome. Genotypes can be characterized, for example, indirectly using markers and / or directly by nucleic acid sequencing.
[0096] As used herein, the term "germplasm" refers to the genetic material of an individual (e.g., a plant), a group of individuals (e.g., a plant strain, variety, or family), or a clone derived from, or derived from, a strain, variety, species, or culture. Germplasm can be part of an organism or cell, or can be separated from an organism or cell. Generally, germplasm provides genetic material with a specific genetic composition that forms the basis for some or all of the genetic properties of an organism or cell culture. As used herein, germplasm includes cells, seeds, or tissues from which new plants can grow, as well as plant parts (e.g., leaves, stems, buds, roots, pollen, cells, etc.) that can be cultured into complete plants.
[0097] As used in this article, the terms “cultivar” and “variety” refer to a group of similar plants that can be distinguished from other varieties within the same species by structural or genetic characteristics and / or properties.
[0098] As used herein, the terms “exotic,” “exotic strain,” and “exotic germplasm” refer to any non-superior plant, strain, or germplasm. Generally, exotic plants / germplasm are not derived from any known superior plant or germplasm, but are selected to introduce one or more desired genetic elements into a breeding program (e.g., to introduce novel alleles into a breeding program).
[0099] As used in this article, the term "hybrid" in the context of plant breeding refers to the offspring of genetically distinct parents produced by hybridization of different lines, varieties, or species of plants (including but not limited to hybridization between two inbred lines).
[0100] As used herein, the term “inbreeding” refers to a plant or variety that is substantially homozygous. This term can refer to a plant or variety that is substantially homozygous throughout its genome, or to a plant or variety that is substantially homozygous with respect to a portion of the genome of particular interest.
[0101] A haplotype is an individual's genotype at multiple loci, that is, a combination of alleles. Typically, the genetic loci defining a haplotype are physically and genetically linked, meaning they are located on the same segment of a chromosome. The term "haplotype" can refer to polymorphism at a specific locus, such as a single marker locus, or polymorphism at multiple loci along a segment of a chromosome.
[0102] As used in this article, the term "heterogeneous" refers to nucleotides / peptides derived from a foreign species, or, if derived from the same species, to nucleotides whose natural form has been substantially modified in terms of composition and / or genomic loci through intentional human intervention.
[0103] As used herein, “shade response” is defined as the growth of a plant in response to a low red:far-red (R:FR) light ratio. Inhibition of shade response refers to the suppression of growth changes in response to a low R:FR light ratio. On one hand, inhibition of the shade response can be demonstrated by measuring the height of plants containing the trait of the present invention (e.g., the HD-Zip mutation as described herein) and syngeneic plants without the trait in a controlled environment with a low R:FR light ratio. When grown under the same conditions with an R:FR ratio of 0.16, plants containing the trait of the present invention will be at least 5% shorter (e.g., in height measured at the coleoptile, V1 sheath, or V2 sheath) than isogenetic plants without the trait (e.g., shorter by approximately 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%). %, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 110%, 120%, 130%, 140%, 150% or more, or any range or value thereof;For example, approximately 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25% to approximately 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%. 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90% 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 110%, 120%, 130%, 140%, 150% or more) (e.g., about 5% to about 10%, about 5% to about 15%, about 5% to about 20%, about 5% to about 25%, about 5% to about 30%, about 5% to about 40%, about 5% to about 50%, about 10% to about 20%, about 10% to about 30%, about 10% to about 50%, about 10% to Approximately 70%, 15% to approximately 20%, 15% to approximately 30%, 15% to approximately 50%, 20% to approximately 30%, 20% to approximately 50%, 20% to approximately 70%, 40% to approximately 50%, 40% to approximately 60%, 40% to approximately 80%, 40% to approximately 100%, 50% to approximately 70%, 50% to approximately 100%, 50% to approximately 125%, 75% to approximately 100%, 75% to approximately 120%, 75% to approximately 140%, etc.
[0104] Plants exhibiting SAR show excessive elongation of the hypocotyl and internodes, longer leaves, impaired root growth, premature flowering and reduced fruit set, low photosynthetic efficiency, enhanced green shoot growth, high lodging rate, accelerated senescence, reduced grain filling, as well as active disease suppression and herbivorous response mechanisms.
[0105] Compared to plants with reduced SAR, where SAR is reduced as described herein, plants with reduced SAR may have increased yield. As used herein, “increased yield” refers to any plant trait associated with growth, such as biomass, yield, nitrogen use efficiency (NUE), inflorescence size / weight, fruit yield, fruit weight, fruit size, seed size, seed number, leaf tissue weight, number of nodules, nodule weight, nodule activity, number of seed heads, number of tillers, number of flowers, number of tubers, tuber weight, corm weight, number of seeds, total seed weight, leaf emergence rate, rate of tiller emergence, emergence rate, root length, number of roots, root mass size and / or weight, or any combination thereof. Therefore, in some respects, “increased yield” may include, but is not limited to, increased inflorescence production, increased fruit production (e.g., increased number, weight and / or size of fruits; e.g., increased number, weight and / or size of ears of, for example, maize), increased fruit quality, increased number, size and / or weight of roots, increased meristem size, increased seed size, increased biomass, increased leaf size, increased nitrogen use efficiency, increased height and / or increased internode length, compared to control plants or portions thereof (e.g., plants that do not contain mutant endogenous nucleic acids encoding the HD-Zip transcription factor as described herein, grown in environments with a low R:FR light ratio (e.g., shaded environments; e.g., an R:FR ratio of about 0.16), including when grown in close proximity to other plants).
[0106] Seed weight is determined by grain morphological traits such as seed length, seed width, and seed thickness, as well as grain filling, all of which are controlled by quantitative genetics.
[0107] As used in this article, “reduced height” refers to the inhibition of stem elongation in response to enriched far-red light.
[0108] As used in this article, “reduced stem:root ratio” refers to a decrease in the ratio of aboveground biomass to belowground biomass.
[0109] As used herein, the terms “nucleic acid,” “nucleic acid molecule,” “nucleotide sequence,” and “polynucleotide” refer to linear or branched, single-stranded or double-stranded RNA or DNA, or hybrids thereof. The term also includes RNA / DNA hybrids. Less common bases, such as inosine, 5-methylcytosine, 6-methyladenine, and hypoxanthine, can also be used for antisense, dsRNA, and ribozyme pairing when synthesizing dsRNA. For example, polynucleotides containing C-5 propyne analogs of uridine and cytidine have been shown to bind RNA with high affinity and are potent antisense inhibitors of gene expression. Other modifications can also be made, such as modifications to the phosphodiester backbone or the 2'-hydroxyl group in the ribose group of RNA.
[0110] As used herein, the term "nucleotide sequence" refers to a heteropolymer of nucleotides or the sequence of these nucleotides from the 5' end to the 3' end of a nucleic acid molecule, including DNA or RNA molecules, including cDNA, DNA fragments or portions, genomic DNA, synthetic (e.g., chemically synthesized) DNA, plasmid DNA, mRNA, and antisense RNA, any of which may be single-stranded or double-stranded. The terms "nucleotide sequence," "nucleic acid," "nucleic acid molecule," "nucleic acid construct," "oligonucleotide," and "polynucleotide" are used interchangeably herein and refer to heteropolymers of nucleotides. The nucleic acid molecules and / or nucleotide sequences provided herein are presented in a left-to-right 5' to 3' orientation and are represented using the standard codes for representing nucleotide characteristics as set forth in US Sequencing Rules 37 CFR §§1.821-1.825 and WIPO Standard ST.25. As used herein, "5' region" may refer to the region of a polynucleotide closest to its 5' end. Therefore, for example, elements in the 5' region of a polynucleotide can be located anywhere from the first nucleotide at the 5' end of the polynucleotide to the nucleotide in the middle of the polynucleotide. As used herein, "3' region" can refer to the region of a polynucleotide closest to the 3' end of the polynucleotide. Therefore, for example, elements in the 3' region of a polynucleotide can be located anywhere from the first nucleotide at the 3' end of the polynucleotide to the nucleotide in the middle of the polynucleotide.
[0111] As used herein with respect to nucleic acids, the term "fragment" or "part" refers to a nucleic acid that is shortened (e.g., shortened by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more nucleotides) relative to a reference nucleic acid, and that contains the same or nearly the same portion as the corresponding part of the reference nucleic acid (e.g., 70%, 70%). A nucleotide sequence of consecutive nucleotides (1%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% identical), substantially composed of said nucleotide sequence and / or composed of said nucleotide sequence. If appropriate, such nucleic acid fragments may be contained within larger polynucleotides in which they are components. For example, the repeat sequence of the guide nucleic acid of the present invention may include a portion of a wild-type CRISPR-Cas repeat sequence (e.g., a wild-type CRISPR-Cas repeat sequence; for example, a repeat from a CRISPR-Cas system such as Cas9, Cas12a (Cpf1), Cas12b, Cas12c (C2c3), Cas12d (CasY), Cas12e (CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b and / or Cas14c, etc.).In some implementations, the nucleic acid fragment may contain approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 300, 350, 400, 450, 500, 550, 600, 660, or 700 nucleotides encoding the HD-Zip transcription factor. 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1550, 1600, 1650, 1700, 1750, 1800, 1850, 1900, 1950, 2000 or more consecutive nucleotides, substantially composed of or consisting of said consecutive nucleotides, the reduced activity of said HD-Zip transcription factor (e.g., reduced DNA binding) can lead to a reduced shade avoidance response in plants.
[0112] In some implementations, the fragment or portion may be a fragment or portion of the HD-Zip transcription factor. In some embodiments, the nucleic acid fragment or portion may be a nucleic acid fragment or portion thereof encoding any of the following amino acid sequences: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (f) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO:9).The fragment or portion thereof comprises about 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, or 32 nucleic acids encoding any one of the polypeptides (a)-(g) above. 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 10 A series of 5, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 130, 140, 150, 175, 200, 225, 250, 300, 350 or more consecutive nucleotides, or any range or value thereof. In some embodiments, "part" may be related to the number of amino acids deleted from the polypeptide. Therefore, for example, a deletion of a portion of the HD-Zip transcription factor may comprise the deletion of at least two consecutive nucleotides (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21) of the nucleotide sequence encoding the HD-Zip transcription factor polypeptide containing the amino acid sequence of SEQ ID NO:9 (VWFQNRRA). In some embodiments, the deletion may comprise a portion of the HD-Zip transcription factor containing exon 3 and exon 4, wherein exon 3 encodes the HD-Zip DNA-binding region. In some embodiments, the deletion may comprise a portion of the HD-Zip transcription factor containing both exon 3 and exon 4, and optionally, a portion of exon 2. In some embodiments, the portion of the HD-Zip polynucleotide that may contain the last 96 to 125 consecutive amino acid residues encoding the C-terminal portion of the HD-Zip polypeptide is omitted.
[0113] In some implementations, the "sequence-specific DNA binding domain" can bind to one or more segments or portions of the nucleotide sequence encoding the HD-Zip transcription factor described herein.
[0114] As used herein with respect to polypeptides, the terms "fragment" or "part" may refer to a polypeptide shortened relative to a reference polypeptide that comprises, or is substantially identical to (e.g., 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%) a continuous amino acid sequence that is identical to or substantially composed of the corresponding portion of the reference polypeptide. Where appropriate, such polypeptide fragments may be contained within a larger polypeptide for which they constitute a part. In some embodiments, the polypeptide fragment comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 300, 350, 400 or more consecutive amino acids of a reference polypeptide, and is substantially composed of or consisting of said consecutive amino acids.In some embodiments, the polypeptide fragment may comprise about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 300, 350, 400, 450, 500, 550, 600, 660, 700 or more consecutive amino acid residues of the HD-Zip transcription factor, substantially composed of or consisting of said consecutive amino acid residues, said consecutive amino acid residues being, for example (a) comprising SEQ ID NO:38 or SEQ ID NO:700. (b) A polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) A polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) A polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) A polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (f) A polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide comprising the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9).
[0115] In some embodiments, the fragment or portion may be a fragment or portion of the HD-Zip transcription factor. In some embodiments, the fragment or portion may be a fragment or portion of any of the following amino acid sequences: (a) a polypeptide containing a sequence having at least 80% sequence identity with the amino acid sequences of (a) to (g) above, wherein the fragment or portion comprises about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, ... 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 140, 150, 175, 200, 225, 250, 300, 350 or more consecutive amino acids, or any range or value of consecutive amino acids therein. In some embodiments, "part" may be associated with the number of amino acids deleted from the polypeptide. Therefore, for example, the deleted “part” of an HD-Zip transcription factor may include at least one amino acid residue (e.g., at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues) of the amino acid sequence (VWFQNRRA) of any HD-Zip transcription factor described herein, and / or at least two amino acid residues (e.g., at least 2, 3, 4, 5, 6, 7, or 8 amino acid residues).In some embodiments, the deletion of a portion of the HD-Zip transcription factor may comprise a portion of consecutive amino acid residues of SEQ ID NO:9 (e.g., at least 2, 3, 4, 5, 6, 7, or 8 consecutive amino acid residues). In some embodiments, the deletion comprises at least a portion of consecutive amino acid residues of SEQ ID No:9 (e.g., at least 1, 2, 3, 4, 5, 6, 7, or 8 consecutive amino acid residues), wherein the deletion length may be from at least 1 amino acid residue to about 120 amino acid residues of the HD-Zip transcription factor (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 2...). 0, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62 1, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103 The deletion may consist of 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120 or more consecutive amino acid residues, up to the full length of the HD-Zip transcription factor, and any range or value of consecutive amino acid residues therein, wherein at least one deleted amino acid residue originates from the DNA-binding region of the HD-Zip transcription factor. In some embodiments, the deletion may be a truncation comprising at least a portion of the consecutive amino acid residues of SEQ ID NO:9 (e.g., at least 1, 2, 3, 4, 5, 6, 7 or 8 consecutive amino acid residues). In some embodiments, the truncation may be a C-terminal truncation and comprise a length of at least 96 consecutive amino acid residues.In some embodiments, the truncation may be a C-terminal truncation and comprises a length of about 96 amino acid residues to about 125 amino acid residues (e.g., at least 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125 or more consecutive amino acid residues up to the full length, and any range or value of consecutive amino acid residues thereof), and wherein at least one truncated amino acid residue originates from the DNA-binding region of the HD-Zip transcription factor. In some embodiments, the deletion may include a deletion of exon 3, which contains the DNA-binding region of the HD-Zip transcription factor (e.g., the deleted portion of the HD-Zip transcription factor may be about 27 amino acid residues in length). In some embodiments, truncation may result from a deletion in exon 2 (e.g., a portion of exon 2) that results in the truncation of a portion of the amino acid residues encoded by exon 2 and all remaining amino acids after the deletion, thereby, for example, truncation of all amino acids encoded by exons 3 and 4. Therefore, in some embodiments, the deletion may result in the truncation of the C-terminal region of the HD-Zip transcription factor polypeptide, the C-terminal region comprising a DNA-binding region or at least a portion of a DNA-binding region.
[0116] In some embodiments, the deletion can cause a frameshift mutation, resulting in a truncation of the stop codon and the C-terminus of the HD-Zip transcription factor peptide. In some embodiments, the C-terminal truncation can result in a peptide containing 207 amino acids (e.g., the deleted or truncated portion includes all amino acid residues after amino acid residue 207; see, for example, the maize HD-Zip edited peptide SEQ ID NO:201).
[0117] In some implementations, the "sequence-specific DNA binding domain" can bind to one or more segments or portions of the nucleotide sequence encoding the HD-Zip transcription factor described herein.
[0118] As used in this article regarding nucleic acids, the term "functional fragment" refers to a nucleic acid that encodes a functional fragment of a polypeptide.
[0119] As used herein, the term "gene" refers to a nucleic acid molecule capable of producing mRNA, antisense RNA, miRNA, antimicroRNA antisense oligodeoxyribonucleotides (AMO), etc. Genes may or may not be used to produce functional proteins or gene products. Genes may include coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and / or 5' and 3' untranslated regions). Genes may be "isolated," meaning that the nucleic acid is substantially or essentially free of components typically found to bind to the nucleic acid in its native state. Such components include other cellular materials from recombinant production, culture media, and / or various chemicals used for the chemical synthesis of nucleic acids.
[0120] The term "mutation" refers to point mutations (e.g., missense or nonsense, or insertions or deletions of a single base pair that result in a frameshift), insertions, deletions, and / or truncations. When a mutation is the substitution of a residue in an amino acid sequence by another residue, or the deletion or insertion of one or more residues in the sequence, it is typically described by identifying the original residue, followed by identifying its position in the sequence and the identity of the newly substituted residue. In some embodiments, deletions can result in frameshift mutations, producing premature stop codons that truncate the protein.
[0121] As used herein, the terms “complementary” or “complementarity” refer to the natural binding of polynucleotides through base pairing under permissible salt and temperature conditions. For example, the sequence “AGT” (5' to 3') binds to the complementary sequence “TCA” (3' to 5'). Complementarity between two single-stranded molecules can be “partial,” where only some nucleotides bind, or it can be complete when there is perfect complementarity between the single-stranded molecules. The degree of complementarity between nucleic acid strands has a significant impact on the efficiency and strength of hybridization between nucleic acid strands.
[0122] As used herein, “complementary” can refer to 100% complementarity with a comparative nucleotide sequence, or it can refer to less than 100% complementarity with a comparative nucleotide sequence (e.g., complementarity of approximately 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, etc.).
[0123] Different nucleic acids or proteins that are homologous are referred to herein as “homologs.” The term homolog includes homologous sequences from the same species and other species, as well as orthologous sequences from the same species and other species. “Homology” refers to the level of similarity between two or more nucleic acid and / or amino acid sequences, expressed as a percentage of positional identity (i.e., sequence similarity or identity). Homology also refers to the concept of similar functional properties between different nucleic acids or proteins. Therefore, the compositions and methods of the present invention also comprise homologs to the nucleotide and polypeptide sequences of the present invention. As used herein, “orthologous homology” refers to homologous nucleotide and / or amino acid sequences in different species that originated from a common ancestral gene during speciation. Homologous products of the nucleotide sequence of the present invention have basic sequence identity with the nucleotides of the present invention (e.g., at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%).
[0124] As used herein, “sequence identity” refers to the degree to which two best-aligned polynucleotide or polypeptide sequences remain unchanged throughout the component (e.g., nucleotide or amino acid) alignment window. “Identity” can be readily calculated by known methods, including but not limited to those described in the following publications: Computational Molecular Biology (Lesk, AM, ed.) Oxford University Press, New York (1988); Biocomputing: Informatics and Genome Projects (Smith, DW, ed.) Academic Press, New York (1993); Computer Analysis of Sequence Data, Part I (Griffin, AM and Griffin, HG, ed.) Humana Press, New Jersey (1994); Sequence Analysis in Molecular Biology (von Heinje, G., ed.) Academic Press (1987); and Sequence Analysis Primer (Gribskov, M. and Devereux, J., ed.) Stockton Press, New York (1991).
[0125] As used herein, the term "sequence identity percentage" or "identity percentage" refers to the percentage of identical nucleotides in the linear polynucleotide sequence of a reference ("query") polynucleotide molecule (or its complementary strand) and the linear polynucleotide sequence of a test ("test") polynucleotide molecule (or its complementary strand) when two sequences are optimally aligned. In some embodiments, "identity percentage" may refer to the percentage of identical amino acids in the amino acid sequence compared to a reference polypeptide.
[0126] As used herein, in the context of two nucleic acid molecules, nucleotide sequences, or polypeptide sequences, the phrase “substantially identical” or “substantially identical” means that two or more sequences or subsequences, when compared and aligned for maximum correspondence, have at least about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% nucleotide or amino acid residue identity, as measured by using one of the following sequence comparison algorithms or by visual inspection. In some embodiments of the invention, substantial identity exists in continuous nucleotide regions of the nucleotide sequence of the invention, said continuous nucleotide regions being of length from about 10 nucleotides to about 20 nucleotides, from about 10 nucleotides to about 25 nucleotides, from about 10 nucleotides to about 30 nucleotides, from about 15 nucleotides to about 25 nucleotides, from about 30 nucleotides to about 40 nucleotides, from about 50 nucleotides to about 60 nucleotides, from about 70 nucleotides to about 80 nucleotides, from about 90 nucleotides to about 100 nucleotides or more nucleotides, and any range therewith, up to the full length of the sequence. In some embodiments, the nucleotide sequence may be substantially identical in at least about 20 nucleotides (e.g., about 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, or 40 nucleotides). In some embodiments of the invention, substantial identity exists in continuous amino acid residue regions of the polypeptide of the invention, which are of lengths of about 3 amino acid residues to about 20 amino acid residues, about 5 amino acid residues to about 25 amino acid residues, about 7 amino acid residues to about 30 amino acid residues, about 10 amino acid residues to about 25 amino acid residues, about 15 amino acid residues to about 30 amino acid residues, about 20 amino acid residues to about 40 amino acid residues, about 25 amino acid residues to about 40 amino acid residues, about 25 amino acid residues to about 50 amino acid residues, about 30 amino acid residues to about 50 amino acid residues, about 40 amino acid residues to about 50 amino acid residues, about 40 amino acid residues to about 70 amino acid residues, about 50 amino acid residues to about 70 amino acid residues, about 60 amino acid residues to about 80 amino acid residues, about 70 amino acid residues to about 80 amino acid residues, about 90 amino acid residues to about 100 amino acid residues, or more amino acid residues, and any range thereof, up to the full length of the sequence.In some embodiments, the polypeptide sequence can be in the form of at least about 8 consecutive amino acid residues (e.g., lengths of about 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 3...). 6, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107 The amino acids (108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 130, 140, 150, 175, 200, 225, 250, 275, 300, 325, 350 or more amino acids, or more consecutive amino acid residues) are generally the same as each other. In some embodiments, two or more HD-Zip transcription factors may be substantially identical to each other in at least about 8 consecutive amino acid residues (e.g., SEQ ID NO: 9), at least about 9 consecutive amino acid residues (e.g., SEQ ID NO: 7), at least about 11 consecutive amino acid residues (e.g., SEQ ID NO: 6), at least about 13 consecutive amino acid residues (e.g., SEQ ID NO: 8), at least about 24 consecutive amino acid residues (e.g., SEQ ID NO: 4-5), at least about 53 consecutive amino acid residues (e.g., SEQ ID NO: 3), at least about 116 consecutive amino acid residues (e.g., SEQ ID NO: 1-2), or combinations thereof. In some embodiments, substantially identical nucleotide or protein sequences perform substantially the same function as substantially identical nucleotide (or encoded protein sequences).
[0127] For sequence comparisons, a reference sequence is typically used as the comparison target. When using a sequence comparison algorithm, the test and reference sequences are input into the computer, and subsequence coordinates are specified if necessary, along with the sequence algorithm program parameters. The sequence comparison algorithm then calculates the percentage of sequence identity between one or more test sequences and the reference sequence based on the specified program parameters.
[0128] The optimal alignment of sequences for the comparison window is well known to those skilled in the art and can be performed using tools such as Smith and Waterman's local homology algorithm, Needleman and Wunsch's homology alignment algorithm, Pearson and Lipman's similarity search method, and optionally computerized implementations of these algorithms (such as GAP, BESTFIT, FASTA, and TFASTA). Wisconsin (A portion of the sample was obtained from Accelrys Inc., San Diego, California). The “identity score” of the aligned segments of the test and reference sequences is the number of common components shared by the two aligned sequences divided by the total number of components in the reference sequence segment (e.g., the entire reference sequence or a smaller defined portion of the reference sequence). The sequence identity percentage is expressed as the identity score multiplied by 100. The comparison of one or more polynucleotide sequences can be with a full-length polynucleotide sequence or a portion thereof, or with a longer polynucleotide sequence. For the purposes of this invention, BLASTX version 2.0 may also be used for translated nucleotide sequences, and BLASTN version 2.0 may be used for polynucleotide sequences to determine the “identity percentage”.
[0129] Two nucleotide sequences can be considered substantially complementary when they hybridize under stringent conditions. In some implementations, two nucleotide sequences considered substantially complementary hybridize under highly stringent conditions.
[0130] Nucleic acid hybridization experiments, such as the "strict hybridization conditions" and "strict hybridization washing conditions" in the context of Southern and Northern hybridization, are sequence-dependent and vary under different environmental parameters. Detailed guidelines for nucleic acid hybridization can be found in Laboratory Techniques in Biochemistry and Molecular Biology—Hybridization with Nucleic Acid Probes, Part I, Chapter 2, "Overview of principles of hybridization and the strategy of nucleic acid probe assays," Elsevier, New York (1993). Typically, highly stringent hybridization and washing conditions are chosen at specific ionic strengths and pH values, above the melting point (T0) of the particular sequence. m It is about 5°C lower.
[0131] T m This is the temperature at which 50% of the target sequence hybridizes with a perfectly matched probe (at a given ionic strength and pH). Very stringent conditions are chosen so that they are equal to the Ta of the specific probe. mIn Southern or Northern blotting, an example of stringent hybridization conditions for hybridizing complementary nucleotide sequences with more than 100 complementary residues on a filter membrane is 50% formamide and 1 mg heparin at 42°C, with hybridization performed overnight. An example of highly stringent washing conditions is washing with 0.1 5M NaCl for approximately 15 minutes at 72°C. An example of stringent washing conditions is washing with 0.2x SSC for 15 minutes at 65°C (see Sambrook, ibid., for a description of SSC buffer). Typically, a low-stringent wash precedes a high-stringent wash to remove background probe signal. For example, for duplexes exceeding 100 nucleotides, an example of a moderately stringent wash is washing with 1x SSC for 15 minutes at 45°C. For example, for duplexes exceeding 100 nucleotides, an example of a low-stringent wash is washing with 4–6x SSC for 15 minutes at 40°C. For short probes (e.g., about 10 to 50 nucleotides), stringent conditions typically involve a salt concentration of less than about 1.0 M Na ions, typically about 0.01 to 1.0 M Na ion concentration (or other salt) at pH 7.0 to 8.3, and a temperature typically at least about 30°C. Stringent conditions can also be achieved by adding a destabilizing agent such as formamide. Generally, in a specific hybridization assay, a signal-to-noise ratio that is twice (or higher) than that observed for irrelevant probes indicates that specific hybridization has been detected. If nucleotide sequences that do not hybridize under stringent conditions encode substantially identical proteins, then the nucleotide sequences remain substantially identical. This can occur, for example, when copies of nucleotide sequences are generated using the maximum codon degeneracy allowed by the genetic code.
[0132] Codon optimization can be performed to express the polynucleotide and / or recombinant nucleic acid constructs (e.g., expression cassettes and / or vectors) of the present invention. In some embodiments, polynucleotides, nucleic acid constructs, expression cassettes and / or vectors (e.g., those containing / encoding sequence-specific DNA-binding domains (e.g., from polynucleotide-guided endonucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), Argonaute proteins and / or CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins) (e.g., type I CRISPR-Cas effector proteins, type II CRISPR-Cas effector proteins, type III CRISPR-Cas effector proteins, type IV CRISPR-Cas effector proteins, type V CRISPR-Cas effector proteins or type VI CRISPR-Cas effector proteins), nucleases (e.g., endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins)) and nucleases (e.g., endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins)) can be expressed in plants. The present invention utilizes codon optimization for zinc finger nucleases and / or transcription activator-like effector nucleases (TALENs), deaminase proteins / domains (e.g., adenine deaminase, cytosine deaminase), polynucleotides encoding reverse transcriptases or domains, polynucleotides and / or affinity peptides encoding 5'-3' exonuclease polypeptides, peptide tags, etc. In some embodiments, the codon-optimized nucleic acids, polynucleotides, expression cassettes, and / or vectors of the present invention are compared with uncodon-optimized reference nucleic acids, polynucleotides, expression cassettes, and / or vectors. The delivery box and / or carrier have approximately 70% to approximately 99.9% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.9%, or 100%) identity or higher.
[0133] In any of the embodiments described herein, the polynucleotide or nucleic acid constructs of the present invention can be operatively linked to a variety of promoters and / or other regulatory elements for expression in plants and / or plant cells. Therefore, in some embodiments, the polynucleotide or nucleic acid constructs of the present invention may also comprise one or more promoters, introns, enhancers, and / or terminators operatively linked to one or more nucleotide sequences. In some embodiments, the promoter may be operatively associated with an intron (e.g., the Ubi1 promoter and introns). In some embodiments, the promoter associated with an intron may be referred to as a “promoter region” (e.g., the Ubi1 promoter and introns).
[0134] As used herein when referring to polynucleotides, “operably linked” or “operably associated” means that the elements are functionally related to each other and generally also physically related. Therefore, as used herein, the terms “operably linked” or “operably associated” refer to functionally related nucleotide sequences on a single nucleic acid molecule. Thus, a first nucleotide sequence operably linked to a second nucleotide sequence means the first nucleotide sequence is positioned in a functional relationship with the second nucleotide sequence. For example, if a promoter influences the transcription or expression of the nucleotide sequence, then the promoter is operably associated with the nucleotide sequence. Those skilled in the art will understand that a control sequence (e.g., a promoter) does not need to be sequential with the nucleotide sequence to which it is operably associated, as long as the control sequence can direct its expression. Therefore, for example, an intermediate, untranslated but still transcribed nucleic acid sequence may be present between the promoter and the nucleotide sequence, and the promoter can still be considered “operably linked” to the nucleotide sequence.
[0135] As used herein, when referring to polypeptides, the term "linked" refers to the attachment of one polypeptide to another. Polypeptides can be linked to another polypeptide directly (e.g., via peptide bonds) or via a linker (at the N-terminus or C-terminus).
[0136] The term "connector" is recognized in the art and refers to a chemical group or molecule that connects two molecules or parts (e.g., two domains of a fusion protein, such as a DNA-binding polypeptide or domain and peptide tag and / or reverse transcriptase and affinity polypeptide bound to the peptide tag; or a DNA endonuclease polypeptide or domain and peptide tag and / or reverse transcriptase and affinity polypeptide bound to the peptide tag). A connector may consist of a single linker molecule or may contain more than one linker molecule. In some embodiments, the connector may be an organic molecule, group, polymer, or chemical part, such as a divalent organic part. In some embodiments, the connector may be an amino acid or may be a peptide. In some embodiments, the connector is a peptide.
[0137] In some embodiments, the length of the peptide linker used in this invention can be from about 2 to about 100 or more amino acids, for example, about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, etc. 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids (e.g., lengths of about 2 to about 40, about 2 to about 50, about 2 to about 60, about 4 to about 40, about 4 to about 50, about 4...). From about 60, about 5 to about 40, about 5 to about 50, about 5 to about 60, about 9 to about 40, about 9 to about 50, about 9 to about 60, about 10 to about 40, about 10 to about 50, about 10 to about 60, or with lengths of about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 amino acids to about 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100 or more amino acids (e.g.,(The length is approximately 105, 110, 115, 120, 130, 140, 150 or more amino acids). In some embodiments, the peptide linker may be a GS linker.
[0138] As used herein, the terms “linked” or “fused” with respect to polynucleotides refer to the attachment of one polynucleotide to another. In some embodiments, two or more polynucleotide molecules may be linked by a linker, which may be an organic molecule, group, polymer, or chemical motif, such as a divalent organic motif. Polynucleotides may be linked or fused to another polynucleotide (at the 5' or 3' end) by covalent or non-covalent bonding or binding (including, for example, Watson-Crick base pairing) or by one or more linking nucleotides. In some embodiments, a polynucleotide motif of one structure may be inserted into another polynucleotide sequence (e.g., an extension of a hairpin structure in guide RNA). In some embodiments, the linking nucleotide may be a naturally occurring nucleotide. In some embodiments, the linking nucleotide may be a non-naturally occurring nucleotide.
[0139] A “promoter” is a nucleotide sequence that controls or regulates transcription of a nucleotide sequence (e.g., a coding sequence) that is operatively associated with a promoter. The coding sequence controlled or regulated by a promoter may encode a polypeptide and / or functional RNA. Generally, a “promoter” refers to a nucleotide sequence containing an RNA polymerase II binding site that directs the initiation of transcription. Typically, a promoter is located at 5' or upstream of the coding region of the corresponding coding sequence. Promoters may contain other elements that act as regulators of gene expression; for example, promoter regions. These include TATA box concordance sequences, and often CAAT box concordance sequences (Breathnach and Chambon, (1981) Annu. Rev. Biochem. 50:349). In plants, the CAAT box can be replaced by the AGGA box (Messing et al., (1983) in Genetic Engineering of Plants, T. Kosuge, C. Meredith and A. Hollaender (eds.), Plenum Press, pp. 211-227).
[0140] Promoters that can be used in this invention may include, for example, constitutive, inducible, time-regulated, developmentally regulated, chemically regulated, tissue-preferred, and / or tissue-specific promoters for the preparation of recombinant nucleic acid molecules, such as “synthetic nucleic acid constructs” or “protein-RNA complexes.” These different types of promoters are known in the art.
[0141] The choice of promoter can vary depending on the temporal and spatial requirements of expression, as well as the host cell to be transformed. Promoters for many different organisms are well-known in the art. Based on this extensive knowledge, suitable promoters can be selected for specific target host organisms. Therefore, for example, much is known about promoters upstream of highly constitutively expressed genes in model organisms, and this knowledge is readily accessible and can be applied to other systems where appropriate.
[0142] In some embodiments, promoters that are functional in plants can be used with the constructs of the present invention. Non-limiting examples of promoters for driving expression in plants include the promoter of the RubisCo small subunit gene 1 (PrbcS1), the promoter of the actin gene (Pactin), the promoter of the nitrate reductase gene (Pnr), and the promoter of the carbonic anhydrase replication gene 1 (Pdca1) (see Walker et al., Plant Cell Rep. 23:727-735 (2005); Li et al., Gene 403:132-142 (2007); Li et al., Mol Biol. Rep. 37:1143-1154 (2010)). PrbcS1 and Pactin are constitutive promoters, while Pnr and Pdca1 are inducible promoters. Pnr is nitrate-induced and ammonium-inhibited (Li et al., Gene 403:132-142 (2007)) and Pdca1 is salt-induced (Li et al., Mol Biol. Rep. 37:1143-1154 (2010)). In some embodiments, the promoter used in this invention is the RNA polymerase II (Pol II) promoter. In some embodiments, the U6 promoter or 7SL promoter from maize is used in the constructs of this invention. In some embodiments, the U6c promoter and / or 7SL promoter from maize is used to drive the expression of the guide nucleic acid. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean is used in the constructs of this invention. In some embodiments, the U6c promoter, U6i promoter, and / or 7SL promoter from soybean is used to drive the expression of the guide nucleic acid.
[0143] Examples of constitutive promoters for plants include, but are not limited to, the cestrum virus promoter (CMP) (US Patent No. 7,166,770), the rice actin 1 promoter (Wang et al. (1992) Mol. Cell. Biol. 12:3399-3406; and US Patent No. 5,641,876), the CaMV 35S promoter (Odell et al. (1985) Nature 313:810-812), the CaMV 19S promoter (Lawton et al. (1987) Plant Mol. Biol. 9:315-324), the nos promoter (Ebert et al. (1987) Proc. Natl. Acad. Sci. USA 84:5745-5749), and the Adh promoter (Walker et al. (1987) Proc. Natl. Acad. Sci. USA). 84:6624-6629), sucrose synthase promoter (Yang & Russell (1990) Proc. Natl. Acad. Sci. USA 87:4144-4148), and ubiquitin promoter. Constitutive promoters derived from ubiquitin accumulate in many cell types. Ubiquitin promoters have been cloned from several plant species for transgenic plants, such as sunflower (Binet et al., 1991. Plant Science 79:87-94), maize (Christensen et al., 1989. Plant Molec. Biol. 12:619-632), and Arabidopsis (Norris et al., 1993. Plant Molec. Biol. 21:895-906). The maize ubiquitin promoter (UbiP) has been developed in transgenic monocotyledonous plant systems, and its sequence and the construction of vectors for monocotyledonous plant transformation have been disclosed in patent publication EP 0 342 926. The ubiquitous protein promoter is suitable for expressing the nucleotide sequences of the present invention in transgenic plants, especially monocots. In addition, the promoter expression cassette described by McElroy et al. (Mol. Gen. Genet. 231:150-160 (1991)) can be readily modified to express the nucleotide sequences of the present invention and is particularly suitable for monocot hosts.
[0144] In some embodiments, tissue-specific / tissue-preferred promoters can be used to express heterologous polynucleotides in plant cells. Tissue-specific or preferential expression patterns include, but are not limited to, green tissue-specific or preferential, root-specific or preferential, stem-specific or preferential, flower-specific or preferential, or pollen-specific or preferential expression patterns. Promoters suitable for expression in green tissues include promoters of a number of genes regulating those involved in photosynthesis, many of which have been cloned from monocotyledonous and dicotyledonous plants. In one embodiment, the promoter that can be used in this invention is the maize PEPC promoter from the phosphoenol carboxylase gene (Hudspeth & Grula, Plant Molec. Biol. 12:579-589 (1989)). Non-limiting examples of tissue-specific promoters include those associated with genes encoding seed storage proteins (such as β-conglycinin, cruciferin, napin, and bean protein), maize proteins or oil body proteins (such as olein), or proteins involved in fatty acid biosynthesis (including acyl carrier proteins, stearoyl-ACP desaturases, and fatty acid desaturases (fad 2-1)), as well as other nucleic acids expressed during embryonic development (such as Bce4, see, for example, Kridl et al. (1991) Seed Sci. Res. 1:209-219; and European Patent No. 255378). Tissue-specific or tissue-preferred promoters for expressing the nucleotide sequences of the present invention in plants, particularly maize, include, but are not limited to, promoters that direct expression in roots, pith, leaves, or pollen. Such promoters are disclosed, for example, in WO 93 / 07278 (which is incorporated herein by reference in its entirety).Other non-limiting examples of tissue-specific or tissue-preferred promoters that can be used in this invention include: the cotton rubisco promoter disclosed in U.S. Patent 6,040,504; the rice sucrose synthase promoter disclosed in U.S. Patent 5,604,121; the root-specific promoter described in de Framond (FEBS 290:103-106 (1991); EP 0 452 269 belonging to Ciba-Geigy); the stem-specific promoter described in U.S. Patent 5,625,136 (belonging to Ciba-Geigy), which drives the expression of the maize trpA gene; the lilac yellow leaf curl virus promoter disclosed in WO 01 / 73087; and pollen-specific or preferential promoters, including but not limited to ProOsLPS10 and ProOsLPS11 from rice (Nguyen et al. Plant Biotechnol. Reports). 9(5):297-306(2015)), ZmSTK2_USP from maize (Wang et al. Genome 60(6):485-495(2017)), LAT52 and LAT59 from tomato (Twell et al. Development 109(3):705-713(1990)), Zm13 (US Patent No. 10,421,972), PLA2-δ promoter from Arabidopsis thaliana (US Patent No. 7,141,424) and / or ZmC5 promoter from maize (International PCT Publication No. WO1999 / 042587).
[0145] Other examples of plant tissue-specific / tissue-preferred promoters include, but are not limited to, root hair-specific cis-elements (RHE) (Kim et al., The Plant Cell 18:2958-2970 (2006)), root-specific promoters RCc3 (Jeong et al., Plant Physiol. 153:185-197 (2010)) and RB7 (US Patent No. 5,459,252), plant lectin promoters (Lindstrom et al. (1990) Der. Genet. 11:160-167; and Vodkin (1983) Prog. Clin. Biol. Res. 138:87-98), zeatol dehydrogenase 1 promoter (Dennis et al. (1984) Nucleic Acids Res. 12:3983-4000), and S-adenosyl-L-methionine synthase (SAMS) (Vander Mijnsbrugge et al. (1996) Plant and Cell). Physiology, 37(8):1108-1115), maize light harvesting complex promoter (Bansal et al. (1992) Proc. Natl. Acad. Sci. USA 89:3654-3658), maize heat shock protein promoter (O'Dell et al. (1985) EMBO J. 5:451-458; and Rochester et al. (1986) EMBO J. 5:451-458), pea small subunit RuBP carboxylase promoter (Cashmore, "Nuclear genes encoding the small subunit of ribulose-1,5-bisphosphate carboxylase" pp. 29-39 In: Genetic Engineering of Plants (Hollaender, editor, Plenum Press) 1983; and Poulsen et al. (1986) Mol. Gen. Genet. 205: 193-200), Ti plasmid mannitol synthase promoter (Langridge et al. (1989) Proc. Natl. Acad. Sci. USA 86: 3219-3223), Ti plasmid lipoic acid synthase promoter (Langridge et al. (1989), ibid.), petunia chalcone isomerase promoter (van Tunen et al. (1988) EMBO J. 7: 1257-1263), soybean glycine-rich protein 1 promoter (Keller et al. (1989) Genes Dev.3:1639-1646), truncated CaMV 35S promoter (O'Dell et al. (1985) Nature 313:810-812), potato tuber storage protein (patatin) promoter (Wenzler et al. (1989) Plant Mol. Biol. 13:347-354), root cell promoter (Yamamoto et al. (1990) Nucleic Acids Res. 18:7449), zein promoter (Kriz et al. (1987) Mol. Gen. Genet. 207:90-98; Langridge et al. (1983) Cell 34:1015-1022; Reina et al. (1990) Nucleic Acids Res. 18:6425; Reina et al. (1990) Nucleic Acids Res. 18:7449; and Wandelt et al. (1989) Nucleic Acids Res. 17:2354, globulin-1 promoter (Belanger et al. (1991) Genetics 129:863-872), α-tubulin cab promoter (Sullivan et al. (1989) Mol. Gen. Genet. 215:431-440), PEPCase promoter (Hudspeth & Grula (1989) Plant Mol. Biol. 12:579-589), R gene complex-related promoter (Chandler et al. (1989) Plant Cell 1:1175-1183), and chalcone synthase promoter (Franken et al. (1991) EMBO J. 10:2605-2612).
[0146] Useful promoters for seed-specific expression include the pea globulin promoter from peas (Czako et al. (1992) Mol. Gen. Genet. 235:33-40); and the seed-specific promoter disclosed in U.S. Patent No. 5,625,136. Useful promoters for expression in mature leaves are those that are switched at the onset of senescence, such as the SAG promoter from Arabidopsis thaliana (Gan et al. (1995) Science 270:1986-1988).
[0147] Alternatively, promoters that are functional in chloroplasts can be used. Non-limiting examples of such promoters include the 5'UTR of phage T3 gene 9 and other promoters disclosed in U.S. Patent No. 7,579,516. Other promoters that can be used in this invention include, but are not limited to, the S-E9 small subunit RuBP carboxylase promoter and the Kunitz trypsin inhibitor gene promoter (Kti3).
[0148] Other control elements that can be used in this invention include, but are not limited to, introns, enhancers, termination sequences and / or 5' and 3' untranslated regions.
[0149] Introns that can be used in this invention can be introns identified and isolated from plants and then inserted into expression cassettes for plant transformation. As those skilled in the art will understand, introns can contain sequences required for self-excision and are integrated into the nucleic acid construct / expression cassette within a frame. Introns can be used as spacers to separate multiple protein-coding sequences within a nucleic acid construct, or introns can be used within a protein-coding sequence to, for example, stabilize mRNA. If they are used in a protein-coding sequence, they are inserted into a “frame” that includes an excision site. Introns can also be combined with promoters to improve or modify expression. For example, promoter / intron combinations that can be used in this invention include, but are not limited to, combinations of the maize Ubi1 promoter and introns (see, for example, SEQ ID NO:196 and SEQ ID NO:197).
[0150] Non-limiting examples of introns that can be used in this invention include introns from the ADHI gene (e.g., Adh1-S introns 1, 2, and 6), the ubiquitin gene (Ubi1), the RuBisCO small subunit (rbcS) gene, the RuBisCO large subunit (rbcL) gene, the actin gene (e.g., actin-1 intron), the pyruvate dehydrogenase kinase gene (pdk), the nitrate reductase gene (nr), the replicating carbonic anhydrase gene 1 (Tdca1), the psbA gene, the atpA gene, or any combination thereof.
[0151] In some embodiments, the polynucleotide and / or nucleic acid constructs of the present invention may be "expression cassettes" or may be contained within an expression cassette. As used herein, "expression cassette" means a recombinant nucleic acid molecule containing, for example, one or more polynucleotides of the present invention (e.g., polynucleotides encoding sequence-specific DNA-binding domains, polynucleotides encoding deaminase proteins or domains, polynucleotides encoding reverse transcriptase proteins or domains, polynucleotides encoding 5'-3' exonuclease polypeptides or domains, guide nucleic acids and / or reverse transcriptase (RT) templates), wherein one or more polynucleotides are operatively associated with one or more control sequences (e.g., promoters, terminators, etc.). Therefore, in some embodiments, one or more expression cassettes may be provided, which are designed to express, for example, the nucleic acid constructs of the present invention (e.g., polynucleotides encoding sequence-specific DNA-binding domains, polynucleotides encoding nuclease peptide / domains, polynucleotides encoding deaminase protein / domains, polynucleotides encoding reverse transcriptase protein / domains, polynucleotides encoding 5'-3' exonuclease peptide / domains, polynucleotides encoding peptide tags and / or polynucleotides encoding affinity peptides, etc., or containing guide nucleic acids, extended guide nucleic acids and / or RT templates, etc.). When the expression cassette of the present invention contains more than one polynucleotide, the polynucleotide may be operatively linked to a single promoter driving the expression of all polynucleotides, or the polynucleotide may be operatively linked to one or more separate promoters (e.g., three polynucleotides may be driven by one, two or three promoters in any combination). When two or more different promoters are used, the promoters may be the same promoter or different promoters. Therefore, when contained in a single expression cassette, the polynucleotide encoding a sequence-specific DNA-binding domain, the polynucleotide encoding a nuclease protein / domain, the polynucleotide encoding a CRISPR-Cas effector protein / domain, the polynucleotide encoding a deaminase protein / domain, the polynucleotide encoding a reverse transcriptase polypeptide / domain (e.g., RNA-dependent DNA polymerase), and / or the polynucleotide encoding a 5'-3' exonuclease polypeptide / domain, the guide nucleic acid, the extended guide nucleic acid, and / or the RT template can each be operatively linked to a single promoter or an independent promoter in any combination.
[0152] Expression cassettes containing nucleic acid constructs of the present invention may be chimeric, meaning that at least one of its components is heterologous relative to at least one of its other components (e.g., a promoter from a host organism is operatively linked to a target polynucleotide to be expressed in the host organism, wherein the target polynucleotide is derived from an organism different from the host, or is generally found not to be associated with the promoter). Expression cassettes may also be naturally occurring, but already obtained in a recombinant form for heterologous expression.
[0153] The expression cassette may optionally include a transcription and / or translation termination region (i.e., a termination region) and / or an enhancer region that functions in the selected host cell. Various transcription terminators and enhancers are known in the art and are available for use in the expression cassette. The transcription terminator is responsible for the termination of transcription and proper polyadenylation of mRNA. The termination region and / or enhancer region may be native to the transcription initiation region, or to genes encoding sequence-specific DNA-binding proteins, nucleases, reverse transcriptases, deaminases, etc., or to the host cell, or to another source (e.g., foreign or heterologous to, for example, promoters, genes encoding sequence-specific DNA-binding proteins, genes encoding nucleases, genes encoding reverse transcriptases, genes encoding deaminases, etc., or to the host cell or any combination thereof).
[0154] The expression cassette of the present invention may also include a polynucleotide encoding a selectable marker, which can be used to select transformed host cells. As used herein, a “selectable marker” refers to a polynucleotide sequence that, when expressed, confers a unique phenotype on host cells expressing that marker, thereby distinguishing such transformed cells from those that are not tagged. Such a polynucleotide sequence may encode a selectable marker or a screenable marker, depending on whether the marker confers a trait selectable by chemical means, such as by using a selector (e.g., antibiotics), or depending on whether the marker is merely a trait identifiable by observation or testing, such as by screening (e.g., fluorescence). Numerous examples of suitable selectable markers are known in the art and can be used in the expression cassette described herein.
[0155] In addition to expression cassettes, the nucleic acid molecules / constructs and polynucleotide sequences described herein can also be used in conjunction with vectors. The term "vector" refers to a composition used to transfer, deliver, or introduce nucleic acids (or multiple nucleic acids) into cells. A vector contains a nucleic acid construct (e.g., one or more expression cassettes) containing one or more nucleotide sequences to be transferred, delivered, or introduced. Vectors used to transform host organisms are well known in the art. Non-limiting examples of vectors in the general class include viral vectors, plasmid vectors, phage vectors, phage particle vectors, granular vectors, fosmid vectors, phages, artificial chromosomes, microcircles, or Agrobacterium binary vectors, which may or may not be self-transmitting or mobile. In some embodiments, viral vectors may include, but are not limited to, retroviruses, lentiviruses, adenoviruses, adeno-associated viruses, or herpes simplex virus vectors. Vectors as defined herein can transform prokaryotic or eukaryotic hosts by integration into the cellular genome or by being present outside the chromosome (e.g., autonomously replicating plasmids with origins of replication). Additionally, shuttle vectors are included, which are DNA mediators capable of replicating naturally or by design in two different host organisms, selectable from actinomycetes and related species, bacteria, and eukaryotes (e.g., higher plants, mammals, yeast, or fungal cells). In some embodiments, the nucleic acid in the vector is under the control of a suitable promoter or other regulatory element for transcription in the host cell and is operatively linked to it. The vector can be a bifunctional expression vector that functions in multiple hosts. In the case of genomic DNA, this may contain its own promoter and / or other regulatory elements, while in the case of cDNA, this may be under the control of a suitable promoter and / or other regulatory elements for expression in a host cell. Therefore, the nucleic acids or polynucleotides of the present invention and / or expression cassettes containing them can be contained in vectors described herein and known in the art.
[0156] As used in this article, “contact,” “contacting,” “contacted,” and their grammatical variations refer to placing the components of a desired reaction together under conditions suitable for carrying out the desired reaction (e.g., transformation, transcriptional control, genome editing, creating notches and / or cutting). For example, under conditions where sequence-specific DNA-binding proteins, reverse transcriptases, and deaminases are expressed and the sequence-specific DNA-binding proteins are bound to the target nucleic acid, the target nucleic acid can be contacted with sequence-specific DNA-binding proteins (e.g., polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, transcription activator-like effector nucleases (TALEN), and / or Argonaute proteins) and deaminases or nucleic acid constructs encoding them. The reverse transcriptases and / or deaminases can be fused to the sequence-specific DNA-binding proteins or recruited to the sequence-specific DNA-binding proteins (e.g., via peptide tags fused to the sequence-specific DNA-binding proteins and affinity tags fused to the reverse transcriptases and / or deaminases), thus placing the deaminases and / or reverse transcriptases near the target nucleic acid, thereby modifying the target nucleic acid. Other methods utilizing other protein-protein interactions can be used to recruit reverse transcriptases and / or deaminases, and RNA-protein interactions and chemical interactions can also be used for protein-protein and protein-nucleic acid recruitment.
[0157] As used herein, “modifying,” “mutating,” or “mutation” (the terms are used interchangeably herein) of a target nucleic acid includes editing (e.g., mutation), covalent modification, exchanging / substituting nucleic acid / nucleotide bases, deletion, cleavage, nicking, and / or altering the transcriptional control of the target nucleic acid. In some embodiments, modification may include one or more single-base alterations (SNPs) of any type.
[0158] In the context of a target polynucleotide, “introducing,” “introduce,” “introduced” (and their grammatical variations) refers to presenting a target nucleotide sequence (e.g., a polynucleotide, an RT template, a nucleic acid construct, and / or a guide nucleic acid) to a plant, its plant parts, or its cells in a manner that allows the nucleotide sequence to enter the cell.
[0159] The terms “transformation” and “transfection” are used interchangeably and, as used herein, refer to the introduction of a heterologous nucleic acid into a cell. Cellular transformation can be stable or transient. Therefore, in some embodiments, the polynucleotide / nucleic acid molecules of the present invention can be used to stably transform host cells or host organisms (e.g., plants). In some embodiments, the polynucleotide / nucleic acid molecules of the present invention can be used to transiently transform host cells or host organisms.
[0160] In the context of polynucleotides, "transient conversion" means that polynucleotides are introduced into the cell but not integrated into the cell's genome.
[0161] In the context of introducing polynucleotides into cells, "stably introducing" or "stably introduced" means that the introduced polynucleotides are stably integrated into the cell's genome, thus the cell is stably transformed by the polynucleotides.
[0162] As used herein, "stable transformation" or "stable conversion" means the introduction of nucleic acid molecules into a cell and their integration into the cell's genome. Therefore, the integrated nucleic acid molecules can be inherited by their offspring, and more specifically, by offspring across multiple successive generations. As used herein, "genome" includes both the nuclear and plasmid genomes, and therefore includes the integration of nucleic acids into, for example, the chloroplast or mitochondrial genome. As used herein, stable transformation can also refer to transgenes that remain outside the chromosome (e.g., as microchromosomes or plasmids).
[0163] Transient transformation can be detected, for example, by enzyme-linked immunosorbent assay (ELISA) or Western blotting, which can detect the presence of peptides or polypeptides encoded by one or more transgenes introduced into the organism. Stable transformation of cells can be detected, for example, by Southern blot hybridization analysis of cellular genomic DNA and nucleic acid sequences, wherein the nucleic acid sequences specifically hybridize to the nucleotide sequences of the transgene introduced into the organism (e.g., a plant). Stable transformation of cells can also be detected, for example, by Northern blot hybridization assay of cellular RNA and nucleic acid sequences, wherein the nucleic acid sequences specifically hybridize to the nucleotide sequences of the transgene introduced into the host organism. Stable transformation of cells can also be detected, for example, by polymerase chain reaction (PCR) or other amplification reactions known in the art, which use specific primer sequences that hybridize to one or more target sequences of the transgene, resulting in the amplification of the transgene sequence, which can be detected according to standard methods. Transformation can also be detected by direct sequencing and / or hybridization protocols known in the art.
[0164] Therefore, in some embodiments, the nucleotide sequences, polynucleotides, nucleic acid constructs, and / or expression cassettes of the present invention can be transiently expressed and / or they can be stably integrated into the genome of a host organism. Thus, in some embodiments, the nucleic acid constructs of the present invention (e.g., one or more expression cassettes containing polynucleotides as described herein for editing) can be transiently introduced into cells with guide nucleic acids, thereby leaving no DNA in the cells.
[0165] The nucleic acid constructs of the present invention can be introduced into plant cells by any method known to those skilled in the art. Non-limiting examples of transformation methods include nucleic acid delivery mediated by bacteria (e.g., by *Bacillus subtilis*), nucleic acid delivery mediated by viruses, nucleic acid delivery mediated by silicon carbide or nucleic acid whiskers, nucleic acid delivery mediated by liposomes, microinjection, microparticle bombardment, calcium phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation, sonication, infiltration, PEG-mediated nucleic acid uptake, and any other electrochemical, chemical, physical (mechanical), and / or biological mechanism (including any combination thereof) leading to the introduction of nucleic acids into plant cells. Methods for transforming eukaryotes and prokaryotes are conventional methods known in the art and described throughout the literature (see, for example, Jiang et al. 2013. Nat. Biotechnol. 31:233-239; Ran et al. Nature Protocols 8:2281–2308 (2013)). General guidelines for various plant transformation methods known in the art include Miki et al. ("Procedures for Introducing Foreign DNA into Plants" in Methods in Plant Molecular Biology and Biotechnology, Glick, BR and Thompson, JE, eds. (CRC Press, Inc., Boca Raton, 1993), pp. 67-88) and Rakowoczy-Trojanowska (Cell. Mol. Biol. Lett. 7: 849-858 (2002)).
[0166] In some embodiments of the invention, cell transformation may include nuclear transformation. In other embodiments, cell transformation may include plastid transformation (e.g., chloroplast transformation). In yet another embodiment, the nucleic acids of the invention can be introduced into cells using conventional breeding techniques. In some embodiments, one or more of polynucleotides, expression cassettes, and / or vectors can be introduced into plant cells via Agrobacterium transformation.
[0167] Therefore, polynucleotides can be introduced into plants, plant parts, and plant cells in many ways known in the art. The method of the present invention does not depend on a specific method of introducing one or more nucleotide sequences into a plant, but only on their entry into the cell. When introducing a polynucleotide, it can be assembled as part of a single nucleic acid construct or as a separate nucleic acid construct, and can be located on the same or different nucleic acid constructs. Thus, polynucleotides can be introduced into target cells in a single transformation event or in separate transformation events, or polynucleotides can be integrated into plants as part of a breeding program.
[0168] For example, maize yields (bushels per acre) have steadily increased through intensive breeding. However, incremental yield increases have recently begun to plateau, and significant investments in field evaluation and breeding are needed to clearly demonstrate genetic gains. New genetic modification approaches are required to significantly increase yields in ways that are not possible through conventional methods. Crop yields can be increased in two distinct ways: 1) by increasing yield itself, where engineered plants gain an advantage, such as improved photosynthesis or optimized carbohydrate allocation, or 2) by eliminating residual survival mechanisms inconsistent with high-yield agriculture. Shade avoidance response (SAR) or shade avoidance syndrome (SAS) is one such survival mechanism. SAS / SAR is characterized by an increased root-to-shoot ratio, increased plant height, and decreased yield per plant; in a typical monoculture environment, this response to competition is a wasteful survival mechanism.
[0169] Therefore, the present invention solves the problem of tolerance to increased planting density (see, Figure 3 This invention describes the use of gene editing modifications to trigger key regulators of crop shade avoidance (e.g., dominant inactivation mutations) (see, e.g., ...). Figure 4 Plants with such edited genomes will have reduced shade-avoidance abilities. An example of a mutation used to address (e.g., reduce / weaken) SAR / SAS could be a mutation that removes the DNA-binding function of a dimerizing transcription factor. In some cases, this mutation could be a dominant inactivation mutation (see, for example,...). Figure 4 ).
[0170] HD (homogeneous domain)-LZ (leucine zipper) transcription factors have multiple functions in plants. Type II species are associated with photosensitivity and shade avoidance. A specific HD-LZ type II member (HB53) has been shown to be induced by shading treatment. In maize, ZmHB53 is the closest maize homolog of ATHB2 (an HDLZ closely associated with the shade avoidance response in Arabidopsis) (Carabelli et al., 1996; Steindler et al., 1999) (see, for example, Figure 5 The closely related HDLZ protein ZmHB78 has been identified as another target for attenuating shade-avoidance. A method for attenuating the DNA-binding ability of transcription factors (e.g., HB53, HB78) that can be used in this invention may include modifying a single amino acid (deletion, insertion, or substitution) or removing all or part of the DNA-binding domain by in-frame deletion.
[0171] Figure 1 provides an alignment of the amino acid sequences of HD-Zip (HB78) from 44 different plant species. The sequence shown in Figure 1 is a fraction of the full-length HB78 sequence, consisting of 53 amino acid residues. An alignment of the amino acid sequences of HD-Zip (HB53) from 45 different plant species is also provided (Figure 2). The sequence shown in Figure 2 is a fraction of the full-length HB53 sequence, consisting of approximately 116 amino acid residues. These alignments demonstrate substantial conservation in the target regions of both genes and prove that the present invention targeting the endogenous homologous domain-leucine zipper (HD-Zip) transcription factor (where mutations disrupt the binding of the HD-Zip transcription factor to DNA) will be predicted to function in different plant species to produce plants with a reduced shade avoidance response.
[0172] Examples of possible gene editing are provided in Figure 6 , Figure 7 , Figure 8 and Figure 9 middle. Figure 6 An example of editing the DNA-binding domain of HB53 in corn is provided, and exemplary target amino acid residues for modification are shown in boxes. Figure 7 The HB78 gene is provided with annotations featuring example guide nucleic acids. Figure 8 Examples of deletions in the HB78 gene (SEQ ID NO:295) are provided, showing that deletions, for example, in exon 2 (SEQ ID NO:297) result in deletions of exon 3, exon 4 and the DNA-binding domain, leading to deletions in the protein sequence (SEQ ID NO:296). Figure 9Representative genome sequences of the edited plant (coding strand (SEQ ID NO:298) and non-coding strand (SEQ ID NO:299)) are provided, showing premature termination upstream of the HB78 DNA-binding domain. Figure 13 Schematic diagrams are provided illustrating the use of plasmids pWISE443, pWISE444, pWISE446, and pWISE447 and their corresponding spacers to target HB53 (top) and HB78 (bottom). Plasmid pWISE448 contains all four spacers shown for HB78, plasmid pWISE445 contains all four spacers shown for HB53, and plasmid pWISE451 contains all eight spacers (four for HB53 and four for HB78, as shown below). Figure 13 (As shown).
[0173] In some embodiments, the present invention provides a plant or plant part thereof containing at least one non-naturally mutated endogenous homologous domain-leucine zipper (HD-Zip) transcription factor, wherein the mutation disrupts the binding of the HD-Zip transcription factor to DNA. In some embodiments, the HD-Zip transcription factor may be an HD-Zip type II (HD-Zip II) transcription factor, wherein the HD-Zip II transcription factor is capable of regulating the plant's response to light (e.g., regulating the shade avoidance response (SAR)). In some embodiments, the HD-Zip II transcription factor usable in the present invention may include, but is not limited to, orthologs of AtHB2, HB53, and / or HB78. In some embodiments, the HD-Zip II transcription factor usable in the present invention may be homeobox protein 53 (HB53) or homeobox protein 78 (HB78). The HD-Zip transcription factor usable in the present invention is an HD-Zip transcription factor comprising the following sequence: (a) comprising the sequence of SEQ ID NO:38 or SEQ ID NO:38. (a) A polypeptide having at least 80% sequence identity (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity) of the amino acid sequence NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 89%, or 100% sequence identity) of the amino acid sequence NO:83; Peptides with sequences of 7%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity: RKKLRLSKDQSAVLEDSFREHPTLNPRQKAALAQQLGLRPRQVEVWFQNRRARTKLKQTEVDCEYLKRCCETLTEENRRLQKEVQELRALKLVSPHLYMHMSPPTTLTMCPSCERV (SEQ ID NO:1) (Corn HB53) or RKKLRLSKDQAAVLEESFKEHNTLNPKQKAALAKQLNLKPRQVEVWFQNRRARTKLKQTEVDCEFLKRCCETLTEENRRLQREVAELRVLKLVAPHHYARMPPPTTLTMCPSCERL (SEQ ID NO:2) (Corn HB78);(c) A polypeptide comprising a sequence having at least 80% sequence identity (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity) with the amino acid sequence LAKQLNLKPRQVEVWFQNRRARTKLKQTEVDCEFLKRCCETLTEENRRLQREV (SEQ ID NO:3); (d) A polypeptide comprising a nucleotide sequence RQVEVWFQNRRARTKLKQTEVDCE (SEQ ID NO:3). NO:4) A polypeptide having a sequence having at least 95% sequence identity (e.g., at least about 95%, 96%, 97%, 99%, 99.5%, or 100% sequence identity); (e) A polypeptide comprising the amino acid sequence RQVEVWFQNRRARTKXKQTEVDCE (SEQ ID NO:5), wherein X is L or S; (f) A polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:5). (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and / or (g) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO: 9).
[0174] In some embodiments, the plant or plant part of the present invention comprises an HD-Zip transcription factor, said HD-Zip transcription factor comprising: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83 (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity); and (b) a polypeptide comprising a sequence having at least 80% sequence identity with the following amino acid sequences (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 89%, or 100% sequence identity). The polypeptide sequence with 7%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity is: RKKLRLSKDQSAVLEDSFREHPTLNPRQKAALAQQLGLRPRQVEVWFQNRRARTKLKQTEVDCEYLKRCCETLTEENRRLQKEVQELRALKLVSPHLYMHMSPPTTLTMCPSCERV (SEQ ID NO: 7%) SEQ ID NO:1)(Corn HB53) or RKKLRLSKDQAAVLEESFKEHNTLNPKQKAALAKQLNLKPRQVEVWFQNRRARTKLKQTEVDCEFLKRCCETLTEENRRLQREVAELRVLKLVAPHHYARMPPPTTLTMCPSCERL SEQ ID NO:2)(Corn HB78); (c) Contains amino acid sequences LAKQLNLKPRQVEVWFQNRRARTKLKQTEVDCEFLKRCCETLTEENRRLQREV (SEQ ID NO:2)(Corn HB78)(C) NO:3) A polypeptide having at least 80% sequence identity (e.g., at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 99%, or 100% sequence identity); (d) A polypeptide containing at least 95% sequence identity (e.g., at least about 95%, 96%, 97%, 99%, or 100%) of the nucleotide sequence RQVEVWFQNRRARTKLKQTEVDCE (SEQ ID NO:4).(e) a polypeptide comprising the amino acid sequence RQVEVWFQNRRARTKXKQTEVDCE (SEQ ID NO:5), wherein X is L or S; (f) a polypeptide comprising: (i) the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) the amino acid sequence VWFQNRRA (SEQ ID NO:5). The sequence of SEQ ID NO:9; and / or (g) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO:9). Therefore, the HD-Zip transcription factor used in this invention may comprise the amino acid sequence of any one of SEQ ID NO:1-8 or 10-98, which contains a DNA-binding domain containing the amino acid sequence VWFQNRRA (SEQ ID NO:9). For example, residues 173-288 of the HD-Zip transcription factor comprising the amino acid sequence of SEQ ID NO:38 comprise the amino acid sequence of SEQ ID NO:9. As another example, residues 76-191 of the HD-Zip transcription factor comprising the amino acid sequence of SEQ ID NO:83 comprise the amino acid sequence of SEQ ID NO:9. In addition to SEQ ID NO:9, other polypeptide domains identified in the HD-Zip transcription factor used in this invention include polypeptides containing at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 and / or SEQ ID NO:2, polypeptides containing a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3, polypeptides containing a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4, polypeptides containing a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S, polypeptides containing a sequence having the amino acid sequence shown in SEQ ID NO:6, wherein X1 is S or T, X2 is D or E, X3 is S or A, polypeptides containing the sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, X3 is V or L, and / or polypeptides containing the sequence having SEQ ID NO:9. The polypeptide with the amino acid sequence shown in NO:8, where X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N.
[0175] In some embodiments, at least one non-natural mutation in the endogenous homologous domain-leucine zipper (HD-Zip) transcription factor in plants can be a substitution, deletion, and / or insertion that disrupts the binding of the HD-Zip transcription factor to DNA. For example, the mutation can be a substitution, deletion, and / or insertion of one or more amino acid residues of the transcription factor. The at least one non-natural mutation may include a base substitution for A, T, G, or C, which results in an amino acid substitution that disrupts the binding of the HD-Zip transcription factor to DNA. In some embodiments, at least one non-natural mutation in the endogenous gene encoding the HD-Zip transcription factor may include a deletion. Such a deletion may include, for example, the deletion of all or part of the DNA-binding domain of the HD-Zip transcription factor (e.g., the deletion of at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues of SEQ ID NO:9 (VWFQNRRA)). In some embodiments, the deletion may be a truncation of a portion comprising consecutive amino acid residues of SEQ ID NO:9 (e.g., at least 2, 3, 4, 5, 6, 7, or 8 consecutive amino acid residues). In some embodiments, the deletion may be a truncation of at least a portion comprising consecutive amino acid residues of SEQ ID NO:9 (e.g., at least 2, 3, 4, 5, 6, 7, or 8 consecutive amino acid residues).Therefore, the length of the deletion can be from about 1 amino acid residue to about 120 amino acid residues of the HD-Zip peptide, or the length (e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29). 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 The deletion may consist of 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120 or more consecutive amino acid residues, up to the full length of the HD-Zip polypeptide, wherein the deletion includes at least a portion of the consecutive amino acid residues of SEQ ID NO:9. In some embodiments, the deletion produces a truncated HD-Zip transcription factor comprising the deletion of at least a portion of the consecutive amino acid residues of SEQ ID NO:9. In some embodiments, the deletion in the HD-Zip transcription factor polynucleotide may result in an early stop codon producing a truncated HD-Zip transcription factor, optionally truncated at the C-terminus of the HD-Zip transcription factor, wherein at least a portion of the consecutive amino acid residues of SEQ ID NO:9 are deleted.
[0176] The non-natural mutation in the endogenous gene encoding the HD-Zip transcription factor mutation used in this invention can be a dominant-recessive mutation. Dominant inactivation removes the DNA-binding function of the dimerized transcription factor. The transcription factor can still dimerize, but will lose its ability to bind to downstream gene regulatory regions, and therefore will lose its function. Thus, by removing the DNA-binding ability of the bifunctional protein, the dimerized complex will not activate gene expression (see, for example, Figure 5 ).
[0177] In some embodiments, a plant cell comprising an editing system is provided, the editing system comprising: (a) a CRISPR-related effector protein; and (c) a guide nucleic acid (gRNA, gDNA, crRNA, crDNA) having a spacer sequence complementary to an endogenous target gene encoding a wild-type HD-Zip transcription factor. The wild-type HD-Zip transcription factor can be any HD-Zip transcription factor involved in the shade avoidance response. In some embodiments, the HD-Zip transcription factor can be an HD-Zip type II transcription factor, optionally HB53 or HB78. In some embodiments, the HD-Zip transcription factor gene, which shares complementarity with the spacer sequence of the guide nucleic acid, can encode (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (f) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9); and / or (g) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO: 9). In some embodiments, the spacer sequence of the guide nucleic acid of the editing system of the present invention may comprise the nucleotide sequence of any one of SEQ ID NO: 175-182. In some embodiments, the nucleic acid binding domain that can be used in the editing system of the present invention may be derived from a polynucleotide-guided endonuclease, a CRISPR-Cas endonuclease (e.g., a CRISPR-Cas effector protein), a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein.In some implementations, the edited plant cells, as described herein, can be regenerated into plants, thereby providing plants with mutations in the HD-Zip transcription factor involved in the shade avoidance response and with a weakened shade avoidance response.
[0178] In some embodiments, the present invention provides plant cells comprising at least one non-naturally occurring mutation (e.g., one, two, three, four, five or more mutations) at the DNA binding site of an HD-Zip transcription factor gene, said mutation preventing or reducing the binding of the encoded HD-Zip transcription factor to DNA, wherein the mutation is a substitution, insertion and / or deletion introduced using an editing system comprising a nucleic acid binding domain that binds to a target site in the HD-Zip transcription factor gene, and wherein the HD-Zip transcription factor gene encodes: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having SEQ ID NO:38; (f) A polypeptide comprising: (i) the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO:9).
[0179] In some embodiments, a plant or a portion thereof is provided, the plant or portion thereof comprising a mutation of an endogenous HD-Zip transcription factor that reduces DNA binding of the endogenous HD-Zip transcription factor, wherein the endogenous HD-Zip transcription factor comprises a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; wherein the mutation is a deletion, substitution, and / or insertion of at least one amino acid residue of amino acid residues 45-52 (VWFQNRRA) (SEQ ID NO:9) of the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2. In some embodiments, the mutation of at least one amino acid residue of amino acid residues 45-52 of the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2 is generated after nuclease cleavage, the nuclease comprising a DNA-binding domain that binds to a target site within a target nucleic acid, the target nucleic acid encoding a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 (maize HB53) or SEQ ID NO:2 (maize HB78).
[0180] Mutations in endogenous HD-Zip transcription factors in plants or their parts can be insertions, substitutions, and / or deletions of at least one amino acid. In some embodiments, mutations may include the deletion of all or part of the DNA-binding domain within the endogenous HD-Zip transcription factor (e.g., the deletion of at least 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues of SEQ ID NO:9 (VWFQNRRA)).
[0181] In some implementations, non-limiting examples of plants or parts thereof include maize, soybean, rapeseed, wheat, rice, cotton, sugarcane, sugar beet, barley, oats, alfalfa, sunflower, safflower, oil palm, sesame, coconut, tobacco, potato, sweet potato, cassava, coffee, apple, plum, apricot, peach, cherry, pear, fig, banana, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, grape, tomato, cucumber, blackberry, raspberry, black raspberry, or certain species of the genus Brassica. In some embodiments, the plant part may be derived from plant cells, including but not limited to corn, soybean, rapeseed, wheat, rice, cotton, sugarcane, sugar beet, barley, oats, alfalfa, sunflower, safflower, oil palm, sesame, coconut, tobacco, potato, sweet potato, cassava, coffee, apple, plum, apricot, peach, cherry, pear, fig, banana, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, grape, tomato, cucumber, blackberry, raspberry, black raspberry, or certain species of Brassica. In some embodiments, the plant can be regenerated from the cells or plant parts of the present invention. Plants of the present invention containing at least one mutation in the HD-Zip transcription factor exhibit a weakened shade avoidance response (SAR).
[0182] In some embodiments, the present invention provides a plant or a plant portion thereof comprising an HD-Zip transcription factor gene, said HD-Zip transcription factor gene comprising the nucleotide sequence of SEQ ID NO:202 and / or encoding the amino acid sequence of any one of SEQ ID NO:201. In some embodiments, the present invention provides a maize plant or a plant portion thereof comprising an HD-Zip transcription factor gene, said HD-Zip transcription factor gene comprising the nucleotide sequence of SEQ ID No:202 and / or encoding the amino acid sequence of any one of SEQ ID No:201.
[0183] The present invention also provides a method for producing / breeding non-transgenic genome-edited (e.g., base-edited) plants, comprising: (a) hybridizing the plant of the present invention with a non-transgenic plant to introduce a mutation or modification from the plant of the present invention into the non-transgenic plant; and (b) selecting offspring plants containing the mutation or modification but without transgenes to produce non-transgenic genome-edited (e.g., base-edited) plants.
[0184] In some embodiments, a method is provided for providing multiple plants with increased yield when each of the multiple plants is planted adjacent to each other, the method comprising planting two or more plants of the present invention adjacent to each other, thereby providing multiple plants with increased yield compared to multiple control plants planted adjacent to each other (e.g., plants without edited HD-Zip transcription factor genes and reduced SAR).
[0185] "Closely adjacent" refers to a high planting density of any particular plant species that can lead to SAR (Specific Absorption Rate). For example, in some implementations, "closely adjacent" includes spacing of approximately 6.1 inches or less (e.g., spacing of approximately 6.1 inches, 6 inches, 5.9 inches, 5.8 inches, 5.7 inches, 5.6 inches, 5.5 inches, 5.4 inches, 5.3 inches, 5.2 inches, 5.2 inches, 5.1 inches, 5 inches, 4.9 inches, 4.8 inches, 4.7 inches, 4.6 inches, 4.5 inches, 4.4 inches, 4.3 inches, 4.2 inches, 4.1 inches, 4...). The density of plants produced by planting seeds in rows of 3.9 inches, 3.8 inches, 3.7 inches, 3.6 inches, 3.5 inches, 3.4 inches, 3.3 inches, 3.2 inches, 3.1 inches, 3 inches, 2.9 inches, 2.8 inches, 2.7 inches, 2.6 inches, 2.5 inches, 2.4 inches, 2.3 inches, 2.2 inches, 2.1 inches, 2 inches, 1.9 inches, 1.8 inches, 1.7 inches, 1.6 inches, 1.5 inches, 1.4 inches, 1.3 inches, 1.2 inches, 1.1 inches, 1 inch, 0.9 inches, 0.8 inches, 0.7 inches, 0.6 inches, 0.5 inches, etc., or any range or value thereof. In some embodiments, high-density planting includes approximately 35,000 seeds per acre at row spacings of 36 inches and 38 inches; or any density exceeding 35,000 seeds per acre at row spacings of 30 inches or greater. As those skilled in the art will understand, the number of seeds planted per acre will vary depending on the plant species in order to achieve high-density planting.
[0186] In some embodiments, a method for editing specific sites in the genome of a plant cell is provided, the method comprising: cutting a target site within an endogenous HD-Zip transcription factor gene in a site-specific manner, the endogenous HD-Zip transcription factor gene encoding: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (f) a polypeptide comprising: (i) having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:38). (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide containing the amino acid sequence VWFQNRRA (SEQ ID NO:9), thereby producing editing in the endogenous HD-Zip transcription factor gene of plant cells. Methods for editing plants may also include regenerating plants from plant cells containing edits in endogenous HD-Zip transcription factor genes to produce plants containing edits in endogenous HD-Zip transcription factor genes. In some embodiments, the editing results in a non-natural mutation in the endogenous HD-Zip transcription factor gene, said mutation producing an HD-Zip transcription factor with reduced DNA binding.
[0187] When compared with control plants that do not contain edited endogenous HD-Zip transcription factor genes, plants containing edited endogenous HD-Zip transcription factor genes as described herein, to provide HD-Zip transcription factors with reduced DNA binding, exhibit a weakened shade avoidance response. Plants containing endogenous HD-Zip transcription factor genes edited as described herein can be compared with plants grown under the same environmental conditions that were not so edited, such as environments with a low R:FR light ratio, for example, shading conditions (e.g., an R:FR ratio of about 0.16; or an R:FR ratio range of about 0.09 to about 0.7 (e.g., about 0.09, 0.10, 0.11, 0.12, 0.13, 0.14, 0.15, 0.16, 0.17, 0.18, 0.19, 0.2, 0.21, 0.23, 0.24, 0.25 to about 0.26, 0.27, 0.28, 0.29, 0.3, 0.35, 0.4, 0.45, 0.5, 0.55, 0.6, 0.65, 0.7 or any range or value thereof)).
[0188] In some embodiments, a method for manufacturing a plant is provided, the method comprising: (a) contacting a plant cell population containing a wild-type endogenous gene encoding an HD-Zip transcription factor with a nuclease targeting the wild-type endogenous gene, wherein the nuclease is linked to a DNA-binding domain, the binding domain being bound to a nucleic acid sequence encoding the following sequences: (i) a polypeptide comprising a sequence having at least 80% sequence identity with an amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (ii) a polypeptide comprising a sequence having at least 80% sequence identity with an amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (iii) a polypeptide comprising a sequence having at least 80% sequence identity with an amino acid sequence of SEQ ID NO:3; (iv) a polypeptide comprising a sequence having at least 95% sequence identity with a nucleotide sequence of SEQ ID NO:4; (v) a polypeptide comprising a sequence having an amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (vi) a polypeptide comprising: (1) having an amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:38) (1) A sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (2) A sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (3) A sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (4) A sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9); and / or (vii) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 6). (a) a polypeptide of the sequence NO:9); (b) selecting from the population plant cells containing a mutant wild-type endogenous gene encoding the HD-Zip transcription factor, wherein the mutation is a substitution and / or deletion of at least one amino acid residue in the polypeptide of any one of (i)-(v), wherein the mutation reduces or eliminates the ability of the HD-Zip transcription factor to bind DNA; and (c) growing the selected plant cells into plants.
[0189] In some embodiments, a method for reducing shade avoidance response in plants is provided, the method comprising (a) contacting a plant cell containing a wild-type endogenous gene encoding an HD-Zip transcription factor with a nuclease targeting the wild-type endogenous gene, wherein the nuclease is linked to a DNA-binding domain that binds to a target site in the wild-type endogenous gene, the wild-type endogenous gene encoding: (i) a polypeptide containing a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (ii) a polypeptide containing a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (iii) a polypeptide containing a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (iv) a polypeptide containing a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (v) a polypeptide containing a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (vi) a polypeptide comprising: (1) having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:5). (1) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (2) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (3) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (4) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9); and / or (vii) a polypeptide containing the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9), thereby producing mutant plant cells containing a wild-type endogenous gene encoding the HD-Zip transcription factor; and (b) causing plant cells to grow into plants, thereby reducing the shade avoidance response in plants.
[0190] In some embodiments, a method is provided for generating a plant or a portion thereof comprising a cell containing at least one mutated endogenous HD-Zip transcription factor gene, the method comprising contacting a target site in the endogenous HD-Zip transcription factor gene in the plant or plant portion with a nuclease comprising a cleavage domain and a DNA-binding domain, wherein the DNA-binding domain binds to the target site in the endogenous HD-Zip transcription factor gene, wherein the endogenous HD-Zip transcription factor gene encodes: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38; (f) A polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9). A polypeptide containing the sequence NO:9 is used to generate a plant or a portion thereof comprising at least one cell having a mutation in an endogenous HD-Zip transcription factor gene. In some embodiments, at least one cell of a plant or a portion thereof having a mutated endogenous HD-Zip transcription factor gene produces an HD-Zip transcription factor with reduced DNA binding.
[0191] In some embodiments, a method for producing a plant or a portion thereof comprising an endogenous HD-Zip transcription factor with a mutation exhibiting reduced DNA binding is described, the method comprising contacting a target site in an endogenous HD-Zip transcription factor gene in the plant or plant portion with a nuclease comprising a cleavage domain and a DNA-binding domain, wherein the DNA-binding domain binds to the target site in the HD-Zip transcription factor gene, wherein the HD-Zip transcription factor gene encodes:
[0192] (a) A polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83; (b) A polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (c) A polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (d) A polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (e) A polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (f) A polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9); and / or (g) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO:9), thereby producing a plant or a portion thereof having an endogenous HD-Zip transcription factor with a mutation that reduces DNA binding. In some embodiments, the endogenous HD-Zip transcription factor gene encodes an endogenous HD-Zip transcription factor comprising the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83, wherein the amino acid sequence of SEQ ID NO:38 or SEQ ID NO:83 comprises the amino acid sequence VWFQNRRA (SEQ ID NO:9), and the mutated endogenous HD-Zip transcription factor contains a mutation in the amino acid sequence VWFQNRRA (SEQ ID NO:9). In some embodiments, the endogenous HD-Zip transcription factor gene encodes an endogenous HD-Zip transcription factor comprising the amino acid sequence of any one of SEQ ID NO:10-98, wherein the amino acid sequence of any one of SEQ ID NO:10-98 comprises the amino acid sequence VWFQNRRA (SEQ ID NO:9), and the mutated endogenous HD-Zip transcription factor contains a mutation in the amino acid sequence VWFQNRRA (SEQ ID NO:9).
[0193] In some embodiments, plants containing the mutations described herein in the endogenous HD-Zip transcription factor, or portions thereof, exhibit a weakened / reduced shade-avoidance response compared to control plants (e.g., plants or plant parts that have not been exposed to the editing system) that do not contain the mutations in the endogenous HD-Zip transcription factor gene. In some embodiments, comparisons with control plants can be made between edited plants and control plants when grown under the same environmental conditions (e.g., shaded environments, such as low R:FR ratio environments). Plants containing endogenous HD-Zip transcription factors with mutations that cause a weakened / reduced shade avoidance response exhibit phenotypes including, but not limited to, the following: compared to plants without endogenous HD-Zip transcription factors with mutations that cause a weakened / reduced shade avoidance response, when planted in close proximity to one or more other plants, they have increased yield, reduced height, reduced stem:root ratio, shortened leaf length; enhanced stem mechanical strength; reduced lodging rate; delayed senescence; improved photosynthetic efficiency and grain filling; and / or enhanced defense responses against pathogens and herbivores. In some embodiments, plants with reduced SAR are at least about 5% shorter than control plants grown under the same environmental conditions (e.g., shaded environments, such as low R:FR ratio environments) (e.g., or about 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%). 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 110%, 120%, 130%, 140%, 150% or less, or any range or value thereof.
[0194] In some implementations, nucleases that contact plant cells, plant cell populations, and / or target sites cleave the endogenous HD-Zip transcription factor gene, and mutations are introduced into the DNA binding site of the endogenous HD-Zip transcription factor encoded by the endogenous HD-Zip transcription factor gene.
[0195] In some embodiments, the mutation in the endogenous HD-Zip transcription factor gene can be a non-naturally occurring mutation. In some embodiments, the non-naturally occurring mutation can be a substitution, insertion, and / or deletion. In some embodiments, a non-naturally occurring mutation as a substitution, insertion, and / or deletion can result in the substitution, insertion, and / or deletion of one or more amino acids in the endogenous HD-Zip transcription factor encoded by the endogenous HD-Zip transcription factor gene. Nucleases that can be used in this invention include, but are not limited to, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), endonucleases (e.g., Fok1), or CRISPR-Cas effector proteins.
[0196] In some embodiments, the HD-Zip transcription factor used in this invention may be an HD-Zip type II (HD-ZipII) transcription factor, wherein the HD-Zip II transcription factor is capable of regulating the plant's response to light (e.g., shade avoidance response (SAR)). In some embodiments, the HD-Zip II transcription factor may include, but is not limited to, orthologs of AtHB2, HB53, and / or HB78. In some embodiments, the HD-Zip II transcription factor may be homeobox protein 53 (HB53) or homeobox protein 78 (HB78).
[0197] In some embodiments, the present invention provides a guide nucleic acid (e.g., gRNA, gDNA, crRNA, crDNA) that binds to a target site in an HD-Zip transcription factor gene, the target site comprising a nucleotide sequence encoding: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9) and / or (f) a polypeptide containing the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9).
[0198] The spacer sequence of the guiding nucleic acid of the present invention may be complementary to a fragment or part of a nucleotide sequence encoding (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9) and / or (f) a polypeptide containing the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9).
[0199] In some embodiments, the target nucleic acid is an endogenous HD-Zip transcription factor gene that regulates the plant's response to light. In some embodiments, the target site in the target nucleic acid may encode (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; and (f) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9) and / or (f) a polypeptide containing the amino acid sequence VWFQNRRA (SEQ ID NO: 9).
[0200] In some embodiments, the guiding nucleic acid comprises a spacer having a nucleotide sequence having any one of SEQ ID NO:175-182. In some embodiments, the HD-Zip transcription factor may be an HD-Zip type II (HD-Zip II) transcription factor, optionally wherein the HD-Zip II transcription factor may be HB53 or HB78.
[0201] In some embodiments, a system is provided comprising the guide nucleic acid of the present invention and a CRISPR-Cas effector protein associated with the guide nucleic acid. In some embodiments, the system may further comprise a tracr nucleic acid associated with the guide nucleic acid and the CRISPR-Cas effector protein, optionally wherein the tracr nucleic acid and the guide nucleic acid are covalently linked.
[0202] In some embodiments, a gene editing system is provided comprising a CRISPR-Cas effector protein associated with a guide nucleic acid, wherein the guide nucleic acid comprises a spacer sequence that binds to the HD-Zip transcription factor gene. In some embodiments, the HD-Zip transcription factor gene used in the gene editing system encodes: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9) and / or (f) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO:9). In some embodiments, the HD-Zip transcription factor may be an HD-Zip type II (HD-Zip II) transcription factor, optionally wherein the HD-Zip II transcription factor may be HB53 or HB78.
[0203] In some embodiments, the guide nucleic acid of the gene editing system may comprise a spacer sequence having a nucleotide sequence complementary to the nucleotide sequence encoding the following polypeptides: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and (iv) a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9) and / or (f) a polypeptide comprising the amino acid sequence VWFQNRRA (SEQ ID NO: 9). In some embodiments, the guide nucleic acid of the gene editing system may comprise a spacer sequence having a nucleotide sequence having any one of SEQ ID NO: 175-182. In some embodiments, the gene editing system may also comprise a tracr nucleic acid associated with the guide nucleic acid and a CRISPR-Cas effector protein, optionally wherein the tracr nucleic acid is covalently linked to the guide nucleic acid.
[0204] The present invention also provides a complex comprising a CRISPR-Cas effector protein comprising a cleavage domain and a guide nucleic acid, wherein the guide nucleic acid binds to a target site in an HD-Zip transcription factor gene encoding (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO:6). (ii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S or N; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9) and / or (f) a polypeptide containing the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9), wherein the cleavage domain cleaves the target strand in the HD-Zip transcription factor gene.
[0205] This document also provides an expression cassette comprising (a) a polynucleotide encoding a CRISPR-Cas effector protein containing a cleavage domain and (b) a guide nucleic acid binding to a target site in an HD-Zip transcription factor gene, wherein the guide nucleic acid comprises a spacer sequence complementary to and binding to a nucleotide sequence encoding a polypeptide of: (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) a sequence having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (ii) a sequence having the amino acid sequence PX1X2 The sequence X2LTX3CPX4CER (SEQ ID NO:8), wherein X1 is P or A, X2 is T or A, X3 is V or M, and X4 is Q, S, or N; (iii) the sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO:7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9) and / or (f) a polypeptide comprising the sequence having the amino acid sequence VWFQNRRA (SEQ ID NO:9). In some embodiments, the HD-Zip transcription factor is an HD-Zip type II (HD-Zip II) transcription factor, optionally wherein the HD-Zip II transcription factor is HB53 or HB78.
[0206] In some embodiments, a nucleic acid is provided encoding an HD-Zip transcription factor (e.g., HD-Zip II, such as HB53 or HB78) having a mutated DNA binding site, wherein the mutated DNA binding site of the HD-Zip transcription factor contains a mutation that disrupts DNA binding of the HD-Zip transcription factor. In some embodiments, the mutation may eliminate the binding of the HD-Zip transcription factor to DNA, or may reduce the ability of the HD-Zip transcription factor to bind to DNA by at least 75% (e.g., at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%). The invention also provides plants or portions thereof comprising the nucleic acids of the invention. In some embodiments, when the plant is a maize plant, the maize plant may contain a dwarf / semi-dwarf phenotype. In some embodiments, the plant may be a wheat plant or a portion thereof, optionally wherein the nucleic acid may be contained in genome A, genome B, genome D, or any combination thereof. In some embodiments, compared with control plants planted adjacent to one or more other plants that do not contain nucleic acids encoding HD-Zip transcription factors (e.g., HD-Zip II, such as HB53 or HB78) with mutated DNA binding sites, and thus do not contain reduced shade avoidance responses, the plants of the present invention, maize plants, and / or wheat plants with reduced SAR when planted adjacent to one or more other plants may exhibit increased yield, reduced height, reduced stem:root ratio, shortened leaf length; enhanced stem mechanical strength; reduced lodging rate; delayed senescence; improved photosynthetic efficiency and grain filling; and / or enhanced defense responses against pathogens and herbivores. In some embodiments, when planted close to each other, the plants of the present invention containing reduced SAR are at least about 5% shorter than control plants grown under the same conditions (e.g., shaded environments, such as low R:FR ratio environments).
[0207] In some embodiments, the method of the present invention may further include regenerating plants from plant cells or plant parts containing at least one non-natural mutation in the endogenous homologous domain-leucine zipper (HD-Zip) transcription factor, wherein the mutation disrupts the binding of the HD-Zip transcription factor to DNA. In some embodiments, compared to control plants planted adjacent to one or more other plants that do not contain at least one non-natural mutation in the HD-Zip transcription factor and thus do not exhibit reduced shade avoidance response, plants containing at least one non-natural mutation in the HD-Zip transcription factor may have increased yield, reduced height, reduced stem-to-root ratio, shortened leaf length; enhanced stem mechanical strength; reduced lodging rate; delayed senescence; improved photosynthetic efficiency and grain filling; and / or enhanced defense responses against pathogens and herbivores. In some embodiments, the mutation is a non-naturally occurring mutation. In some embodiments, the mutation is a deletion. In some embodiments, the deletion is a dominant-recessive mutation.
[0208] The editing systems that can be used in this invention can be any site-specific (sequence-specific) genome editing system now known or developed in the future, which can introduce mutations in a target-specific manner. For example, editing systems (e.g., site-specific or sequence-specific editing systems) can include, but are not limited to, CRISPR-Cas editing systems, broad-spectrum nuclease editing systems, zinc finger nuclease (ZFN) editing systems, transcription activator-like effector nuclease (TALEN) editing systems, base editing systems, and / or primer editing systems, each of which can contain one or more polypeptides and / or one or more polynucleotides that, when expressed as a system in cells, can modify (mutate) target nucleic acids in a sequence-specific manner. In some embodiments, the editing system (e.g., site-specific or sequence-specific editing systems) can contain one or more polynucleotides and / or one or more polypeptides, including but not limited to nucleic acid-binding domains (DNA-binding domains), nucleases, and / or other polypeptides and / or polynucleotides.
[0209] In some embodiments, the editing system may include one or more sequence-specific nucleic acid-binding domains (DNA-binding domains), which may be derived from, for example, polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, transcription activator-like effector nucleases (TALENs), and / or Argonaute proteins. In some embodiments, the editing system may include one or more cleavage domains (e.g., nucleases), including but not limited to endonucleases (e.g., Fok1), polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, and / or transcription activator-like effector nucleases (TALENs). In some embodiments, the editing system may include one or more polypeptides, including but not limited to deaminases (e.g., cytosine deaminase, adenine deaminase), reverse transcriptases, Dna2 polypeptides, and / or 5'-flap endonucleases (FENs). In some implementations, the editing system may contain one or more polynucleotides, including but not limited to CRISPR array (CRISPR-guided) nucleic acids, extended guide nucleic acids, and / or reverse transcriptase templates.
[0210] In some embodiments, methods for modifying or editing HD-Zip transcription factors may include contacting a target nucleic acid (e.g., a nucleic acid encoding an HD-Zip transcription factor) with a base-editing fusion protein (e.g., a sequence-specific DNA-binding protein (e.g., a CRISPR-Cas effector protein or domain) fused with a deaminase domain (e.g., adenine deaminase and / or cytosine deaminase) and a guide nucleic acid, wherein the guide nucleic acid is capable of guiding / targeting the base-editing fusion protein to the target nucleic acid to encode a locus within the target nucleic acid. In some embodiments, the base-editing fusion protein and the guide nucleic acid may be contained in one or more expression cassettes. In some embodiments, the target nucleic acid may be contacted with the base-editing fusion protein and an expression cassette containing the guide nucleic acid. In some embodiments, the sequence-specific DNA-binding fusion protein and the guide nucleic acid may be provided in the form of a ribonucleoprotein (RNP). In some embodiments, cells may be contacted with more than one base-editing fusion protein and / or one or more guide nucleic acids, said guide nucleic acids being capable of targeting one or more target nucleic acids in the cell.
[0211] In some embodiments, methods for modifying or editing HD-Zip transcription factors may include contacting a target nucleic acid (e.g., a nucleic acid encoding an HD-Zip transcription factor) with a sequence-specific DNA-binding fusion protein fused to a peptide tag (e.g., a sequence-specific DNA-binding protein (e.g., a CRISPR-Cas effector protein or domain), a deaminase fusion protein (containing a deaminase domain fused to an affinity polypeptide capable of binding to a peptide tag (e.g., adenine deaminase and / or cytosine deaminase)) and a guide nucleic acid, wherein the guide nucleic acid is capable of directing the sequence-specific DNA-binding fusion protein to / targeting the target nucleic acid, and the sequence-specific DNA-binding fusion protein is capable of directing the deaminase fusion protein to / targeting the target nucleic acid via peptide tag-affinity polypeptide interactions. The target nucleic acid is recruited, thereby editing loci within the target nucleic acid. In some embodiments, a sequence-specific DNA-binding fusion protein can be fused with an affinity peptide of a peptide tag, and a deaminase can be fused with a peptide tag, thereby recruiting the deaminase to the sequence-specific DNA-binding fusion protein and the target nucleic acid. In some embodiments, the sequence-specific binding fusion protein, the deaminase fusion protein, and the guide nucleic acid can be contained in one or more expression cassettes. In some embodiments, the target nucleic acid can be contacted with the sequence-specific binding fusion protein, the deaminase fusion protein, and the expression cassette containing the guide nucleic acid. In some embodiments, the sequence-specific DNA-binding fusion protein, the deaminase fusion protein, and the guide nucleic acid can be provided in the form of a ribonucleoprotein (RNP).
[0212] In some implementations, methods such as primer editing can be used to generate mutations in endogenous HD-Zip transcription factor genes. In primer editing, an RNA-dependent DNA polymerase (reverse transcriptase, RT) and a reverse transcriptase template (RT template) are combined with a sequence-specific DNA-binding domain that confers the ability to recognize and bind to a target in a sequence-specific manner and also induces a cleavage containing a PAM chain within the target. The DNA-binding domain can be a CRISPR-Cas effector protein, and in this case, the CRISPR array or guide RNA can be an extended guide nucleic acid containing an extension portion comprising a primer binding site (PSB) and the edit to be integrated into the genome (template). Similar to base editing, primer editing can utilize various methods for recruiting proteins for editing to the target site, including non-covalent and covalent interactions between proteins and nucleic acids used in the selection process of genome editing.
[0213] In some embodiments, the mutation of the HD-Zip transcription factor gene can be an insertion, deletion, and / or point mutation, which produces an HD-Zip transcription factor with reduced DNA binding (e.g., a mutated HD-Zip transcription factor). In some embodiments, the plant part can be a cell. In some embodiments, the plant or its plant part can be any plant or part thereof described herein. In some embodiments, the plant that can be used in the present invention can be maize, soybean, rapeseed, wheat, rice, cotton, sugarcane, sugar beet, barley, oats, alfalfa, sunflower, safflower, oil palm, sesame, coconut, tobacco, potato, sweet potato, cassava, coffee, apple, plum, apricot, peach, cherry, pear, fig, banana, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, grape, tomato, cucumber, or certain species of the genus Brassica. In some embodiments, compared to control plants planted adjacent to one or more other plants that do not contain at least one non-natural mutation in the endogenous homologous domain-leucine zipper (HD-Zip) transcription factor and thus do not exhibit a reduced shade avoidance response, plants containing a DNA-binding mutation in the endogenous HD-Zip transcription factor, when planted adjacent to one or more other plants, may exhibit increased yield, reduced height, reduced stem-to-root ratio, reduced leaf length; increased stem mechanical strength; reduced lodging rate; delayed senescence; improved photosynthetic efficiency and grain filling; and / or enhanced defense responses against pathogens and herbivores. In some embodiments, the plant may be a maize plant, optionally containing a dwarf / semi-dwarf phenotype.
[0214] In some embodiments, the mutation that introduces the endogenous HD-Zip transcription factor and results in reduced DNA binding may be a non-naturally occurring mutation. In some embodiments, the mutation that introduces the endogenous HD-Zip transcription factor and results in reduced DNA binding may be a substitution, insertion, and / or deletion of one or more amino acid residues. In some embodiments, the mutation introduced into the endogenous HD-Zip transcription factor gene that results in reduced DNA binding may be a deletion, optionally a deletion of all or part of the amino acid sequence of SEQ ID NO:9 (e.g., deletion of 1, 2, 3, 4, 5, 6, 7, or 8 amino acid residues of SEQ ID NO:9).
[0215] In some implementations, the HD-Zip transcription factor may be an HD-Zip type II (HD-Zip II) transcription factor, optionally wherein the HD-Zip II transcription factor may be homeobox protein 53 (HB53) or homeobox protein 78 (HB78).
[0216] In some embodiments, the sequence-specific nucleic acid binding domain (DNA binding domain) that can be used in the editing system of the present invention may be derived from, for example, polynucleotide-guided endonucleases, CRISPR-Cas endonucleases (e.g., CRISPR-Cas effector proteins), zinc finger nucleases, transcription activator-like effector nucleases (TALEN), and / or Argonaute proteins.
[0217] In some embodiments, the sequence-specific DNA-binding domain may be a CRISPR-Cas effector protein, optionally wherein the CRISPR-Cas effector protein may be derived from a type I, type II, type III, type IV, type V, or type VI CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein of the present invention may be derived from a type II or type V CRISPR-Cas system. In some embodiments, the CRISPR-Cas effector protein may be a type II CRISPR-Cas effector protein, such as the Cas9 effector protein. In some embodiments, the CRISPR-Cas effector protein may be a type V CRISPR-Cas effector protein, such as the Cas12 effector protein.
[0218] In some implementations, CRISPR-Cas effector proteins may include, but are not limited to, Cas9, C2c1, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl, etc. Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG) and / or Csf5 nucleases, optionally wherein the CRISPR-Cas effector protein can be Cas9, Cas12a(Cpf1), Cas12b, Cas12c(C2c3), Cas12d(CasY), Cas12e(CasX), Cas12g, Cas12h, Cas12i, C2c4, C2c5, C2c8, C2c9, C2c10, Cas14a, Cas14b and / or Cas14c effector proteins.
[0219] In some embodiments, the CRISPR-Cas effector protein used in this invention may contain mutations at its nuclease active site (e.g., RuvC, HNH, for example, the RuvC site of the Cas12a nuclease domain, the RuvC site of the Cas9 nuclease domain, and / or the HNH site). CRISPR-Cas effector proteins with mutations at their nuclease active sites, and therefore no longer containing nuclease activity, are generally referred to as "dead," such as dCas. In some embodiments, the CRISPR-Cas effector protein domain or polypeptide with mutations in its nuclease active site may have impaired or reduced activity compared to the same unmutated CRISPR-Cas effector protein (e.g., nickases, such as Cas9 nickases, Cas12a nickases).
[0220] The CRISPR Cas9 effector protein or CRISPR Cas9 effector domain that can be used in this invention can be any known or later identified Cas9 nuclease. In some embodiments, the CRISPR Cas9 polypeptide can be a Cas9 polypeptide derived from, for example, certain species of the genus *Streptococcus* (e.g., *S. pyogenes*, *S. thermophilus*), certain species of the genus *Lactobacillus*, certain species of the genus *Bifidobacterium*, certain species of *Kandleria*, certain species of the genus *Leuconostoc*, certain species of the genus *Oenococcus*, certain species of the genus *Pediococcus*, certain species of the genus *Weissella*, and / or certain species of the genus *Olsenella*.
[0221] In some embodiments, the CRISPR-Cas effector protein may be a Cas9 polypeptide derived from Streptococcus pyogenes and recognizes the PAM sequence motifs NGG, NAG, and NGA (Mali et al., Science 2013; 339(6121):823-826). In some embodiments, the CRISPR-Cas effector protein may be a Cas9 polypeptide derived from Streptococcus thermophilus and recognizes the PAM sequence motifs NGNG and / or NNAGAAW (W=A or T) (see, for example, Horvath et al., Science, 2010; 327(5962):167-170, and Deveau et al., J Bacteriol 2008; 190(4):1390-1400). In some embodiments, the CRISPR-Cas effector protein may be a Cas9 polypeptide derived from *Streptococcus mutans* and recognizes the PAM sequence motifs NGG and / or NAAR (R = A or G) (see, for example, Deveau et al., J BACTERIOL 2008; 190(4):1390-1400). In some embodiments, the CRISPR-Cas effector protein may be a Cas9 polypeptide derived from *Streptococcus aureus* and recognizes the PAM sequence motif NNGRR (R = A or G). In some embodiments, the CRISPR-Cas effector protein may be a Cas9 protein derived from *S. aureus* that recognizes the PAM sequence motif NGRRT (R = A or G). In some embodiments, the CRISPR-Cas effector protein may be a Cas9 polypeptide derived from *S. aureus* that recognizes the PAM sequence motif NGRRV (R = A or G). In some embodiments, the CRISPR-Cas effector protein may be a Cas9 polypeptide derived from Neisseria meningitidis and recognizing a PAM sequence motif N GATT or N GCTT (R = A or G, V = A, G, or C) (see, for example, Hou et al., 2013, 1-6). In the above embodiments, N may be any nucleotide residue, such as any one of A, G, C, or T. In some embodiments, the CRISPR-Cas effector protein may be a Cas13a protein derived from Leptocrisium shahii that recognizes a single 3' A, U, or C protospacer flanking sequence (PFS) (or RNA PAM (rPAM)) sequence motif, which may be located within the target nucleic acid.
[0222] In some implementations, the CRISPR-Cas effector protein may be derived from Cas12a, a type V clustered, regularly spaced short palindromic repeat (CRISPR)-Cas nuclease. Cas12a differs from the more well-known type II CRISPR Cas9 nuclease in several ways. For example, Cas9 recognizes a G-rich protospacer adjacent motif (PAM) (3'-NGG) at the 3' of its guide RNA (gRNA, sgRNA, crRNA, crDNA, CRISPR array) binding site (protospacer, target nucleic acid, target DNA), while Cas12a recognizes a T-rich PAM (5'-TTN, 5'-TTTN) at the 5' of the target nucleic acid. In fact, the direction in which Cas9 and Cas12a bind their guide RNA is almost opposite to their N and C ends. Furthermore, the Cas12a enzyme uses a single guide RNA (gRNA, CRISPR array, crRNA) instead of the dual guide RNA (sgRNA (e.g., crRNA and tracrRNA)) found in the native Cas9 system, and Cas12a processes its own gRNA. In addition, Cas12a nuclease activity produces staggered DNA double-strand breaks, rather than blunt ends produced by Cas9 nuclease activity, and Cas12a relies on a single RuvC domain to cut both DNA strands, while Cas9 uses both HNH and RuvC domains for cutting.
[0223] The CRISPR Cas12a effector protein / domain that can be used in this invention can be any known or subsequently identified Cas12a polypeptide (formerly known as Cpf1) (see, for example, U.S. Patent No. 9,790,490, the disclosure of which, with respect to its Cpf1 (Cas12a) sequence, is incorporated herein by reference). The terms “Cas12a,” “Cas12a polypeptide,” or “Cas12a domain” refer to an RNA-guided nuclease containing a Cas12a polypeptide or a fragment thereof, comprising a Cas12a guide nucleic acid-binding domain and / or an active, inactive, or partially active DNA-cutting domain of Cas12a. In some embodiments, the Cas12a that can be used in this invention may contain a mutation at the nuclease's active site (e.g., the RuvC site of the Cas12a domain). A Cas12a domain or Cas12a polypeptide that has a mutation at its active site and therefore no longer contains nuclease activity is generally referred to as deadCas12a (e.g., dCas12a). In some implementations, a mutated Cas12a domain or Cas12a polypeptide in its nuclease active site may have impaired activity, for example, it may have nicking enzyme activity.
[0224] Any deaminase domain / peptide used for base editing can be used in this invention. In some embodiments, the deaminase domain may be a cytosine deaminase domain or an adenine deaminase domain. The cytosine deaminase (or cytidine deaminase) that can be used in this invention can be any known or subsequently identified cytosine deaminase from any organism (see, for example, U.S. Patent No. 0,167,457 and Thuronyi et al., Nat. Biotechnol. 37:1070–1079 (2019), each of which is incorporated herein by reference with respect to its disclosure of a cytosine deaminase). Cytosine deaminases can catalyze the hydrolysis and deamination of cytidine or deoxycytidine to uridine or deoxyuridine, respectively. Therefore, in some embodiments, the deaminase or deaminase domain that can be used in this invention may be a cytidine deaminase domain that catalyzes the hydrolysis and deamination of cytosine to uracil. In some embodiments, the cytosine deaminase may be a variant of a naturally occurring cytosine deaminase (including, but not limited to, variants of, primate (e.g., human, monkey, chimpanzee, gorilla), dog, cattle, rat, or mouse). Therefore, in some embodiments, the cytosine deaminase used in this invention may have about 70% to about 100% identity with the wild-type cytosine deaminase (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with naturally occurring cytosine deaminase, and any range or value thereof).
[0225] In some embodiments, the cytosine deaminase used in this invention may be an apolipoprotein B mRNA editing complex (APOBEC) family deaminase. In some embodiments, the cytosine deaminase may be APOBEC1 deaminase, APOBEC2 deaminase, APOBEC3A deaminase, APOBEC3B deaminase, APOBEC3C deaminase, APOBEC3D deaminase, APOBEC3F deaminase, APOBEC3G deaminase, APOBEC3H deaminase, APOBEC4 deaminase, human activation-inducible deaminase (hAID), rAPOBEC1, FERNY and / or CDA1, optionally pmCDA1, atCDA1 (e.g., At2g 19570), and their evolved forms. In some embodiments, the cytosine deaminase may be APOBEC1 deaminase having the amino acid sequence SEQ ID NO:188. In some embodiments, the cytosine deaminase may be APOBEC3A deaminase having the amino acid sequence SEQ ID NO:189. In some embodiments, the cytosine deaminase may be CDA1 deaminase, optionally having the amino acid sequence SEQ ID NO:190. In some embodiments, the cytosine deaminase may be FERNY deaminase, optionally having the amino acid sequence SEQ ID NO:191. In some embodiments, the cytosine deaminase may be an evolved deaminase, for example, SEQ ID NO:192, SEQ ID NO:193, or SEQ ID NO:194. In some embodiments, the cytosine deaminases used in this invention may have about 70% to about 100% identity with the amino acid sequence of naturally occurring cytosine deaminases (e.g., evolved deaminases) (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% identity).In some embodiments, the cytosine deaminase used in this invention may have about 70% to about 99.5% identity with the amino acid sequences of SEQ ID NO:188, SEQ ID NO:189, SEQ ID NO:190, SEQ ID NO:191, SEQ ID NO:192, SEQ ID NO:193, or SEQ ID NO:194 (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.5% identity) (e.g., with SEQ ID NO:188, SEQ ID NO:189, SEQ ID NO:190, SEQ ID NO:191, SEQ ID NO:192, SEQ ID NO:193, or SEQ ID NO:194). The amino acid sequences of SEQ ID NO:191, SEQ ID NO:192, SEQ ID NO:193, or SEQ ID NO:194 have at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity. In some embodiments, the polynucleotide encoding cytosine deaminase may be codon-optimized for expression in plants, and the codon-optimized polypeptide may have about 70% to 99.5% identity with the reference polynucleotide.
[0226] In some embodiments, the nucleic acid constructs of the present invention may also encode a uracil glycosylation inhibitor (UGI) polypeptide / domain (e.g., a uracil-DNA glycosylation inhibitor). Therefore, in some embodiments, a nucleic acid construct encoding a CRISPR-Cas effector protein and a cytosine deaminase domain (e.g., a fusion protein encoding a CRISPR-Cas effector protein domain fused with a cytosine deaminase domain, and / or a CRISPR-Cas effector protein domain fused with a peptide tag or an affinity polypeptide capable of binding to a peptide tag, and / or a deaminase protein domain fused with a peptide tag or an affinity polypeptide capable of binding to a peptide tag) may also encode a uracil-DNA glycosylation inhibitor (UGI), optionally wherein the UGI may be codon-optimized for expression in plants. In some embodiments, the present invention provides a fusion protein comprising a CRISPR-Cas effector polypeptide, a deaminase domain, and a UGI and / or one or more polynucleotides encoding therethe, optionally wherein said one or more polynucleotides may be codon-optimized for expression in plants. In some embodiments, the present invention provides fusion proteins in which a CRISPR-Cas effector peptide, a deaminase domain, and a UGI can be fused with any combination of peptide tags and affinity peptides described herein, thereby recruiting the deaminase domain and UGI to the CRISPR-Cas effector peptide and the target nucleic acid. In some embodiments, a guide nucleic acid can be linked to a recruitment RNA motif, and one or more deaminase domains and / or UGIs can be fused with an affinity peptide capable of interacting with the recruitment RNA motif, thereby recruiting the deaminase domain and UGI to the target nucleic acid.
[0227] The "uracil glycosylation enzyme inhibitor" that can be used in this invention can be any protein capable of inhibiting uracil-DNA glycosylation enzyme base-excision repair enzyme. In some embodiments, the UGI domain comprises wild-type UGI or a fragment thereof. In some embodiments, the UGI domain that can be used in this invention can have about 70% to about 100% identity with the amino acid sequence of a naturally occurring UGI domain (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% identity, and any range or value thereof). In some embodiments, the UGI domain may comprise the amino acid sequence of SEQ ID NO:195 or a polypeptide having about 70% to about 99.5% sequence identity with the amino acid sequence of SEQ ID NO:195 (e.g., having at least 80%, at least 85%, at least 90%, at least 92%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or at least 99.5% identity with the amino acid sequence of SEQ ID NO:195). For example, in some embodiments, the UGI domain may comprise a fragment of the amino acid sequence of SEQ ID NO:195, said fragment being 100% identical to a portion of a continuous nucleotide sequence of the amino acid sequence of SEQ ID NO:195 (e.g., 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80 consecutive nucleotides; e.g., about 10, 15, 20, 25, 30, 35, 40, 45 to about 50, 55, 60, 65, 70, 75, 80 consecutive nucleotides). In some embodiments, the UGI domain may be a variant of a known UGI (e.g., SEQ ID NO: 195) having about 70% to about 99.5% sequence identity with the known UGI (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5% sequence identity, and any range or value thereof). In some embodiments, the polynucleotide encoding the UGI may be codon-optimized for expression in a plant (e.g., a plant), and the codon-optimized polypeptide may have about 70% to about 99.5% identity with the reference polynucleotide.
[0228] The adenine deaminase (or adenosine deaminase) used in this invention can be any known or later-identified adenine deaminase from any organism (see, for example, U.S. Patent No. 10,113,163, the disclosure of which is incorporated herein by reference with respect to its adenine deaminase). Adenine deaminases can catalyze the hydrolytic deamination of adenine or adenosine. In some embodiments, adenine deaminases can catalyze the hydrolytic deamination of adenosine or deoxyadenosine to inosine or deoxyinosine, respectively. In some embodiments, adenosine deaminases can catalyze the hydrolytic deamination of adenine or adenosine in DNA. In some embodiments, the adenine deaminase encoded by the nucleic acid construct of this invention can generate an A→G transition in the meaningful (e.g., "+"; template) strand of the target nucleic acid or a T→C transition in the antisense (e.g., "-", complementary) strand of the target nucleic acid.
[0229] In some embodiments, adenosine deaminase may be a variant of naturally occurring adenine deaminase. Thus, in some embodiments, adenosine deaminase may have approximately 70% to 100% identity with wild-type adenine deaminase (e.g., approximately 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity with naturally occurring adenine deaminase, and any range or value therein). In some embodiments, one or more deaminases are not naturally occurring and may be referred to as engineered, mutant, or evolved adenosine deaminases. Therefore, for example, engineered, mutated, or evolved adenine deaminase peptides or adenine deaminase domains can have approximately 70% to 99.9% identity with naturally occurring adenine deaminase peptides / domains (e.g., approximately 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80% identity with naturally occurring adenine deaminase peptides or adenine deaminase domains). 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or 99.9% identity, and any range or value therein). In some embodiments, the adenosine deaminase may be derived from bacteria (e.g., *Escherichia coli*, *Staphylococcus aureus*, *Haemophilus influenzae*, *Caulobacter crescentus*, etc.). In some embodiments, the polynucleotide encoding the adenosine deaminase polypeptide / domain may be codon-optimized for expression in plants.
[0230] In some embodiments, the adenine deaminase domain may be a wild-type tRNA-specific adenine deaminase domain, such as tRNA-specific adenine deaminase (TadA), and / or a mutant / evolved adenine deaminase domain, such as a mutant / evolved tRNA-specific adenine deaminase domain (TadA*). In some embodiments, the TadA domain may be derived from *Escherichia coli*. In some embodiments, TadA may be modified, for example, truncated, by deleting one or more N-terminal and / or C-terminal amino acids relative to full-length TadA (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 6, 17, 18, 19, or 20 N-terminal and / or C-terminal amino acid residues relative to full-length TadA). In some embodiments, the TadA polypeptide or TadA domain does not contain an N-terminal methionine. In some embodiments, wild-type *E. coli* TadA comprises the amino acid sequence of SEQ ID NO:183. In some embodiments, mutant / evolved *E. coli* TadA* comprises the amino acid sequences of SEQ ID NO:184-187 (e.g., SEQ ID NO:184, 185, 186, or 187). In some embodiments, the polynucleotide encoding TadA / TadA* may be codon-optimized for expression in plants.
[0231] Cytosine deaminases catalyze the deamination of cytosine to produce thymidine (via a uracil intermediate), causing a C-to-T or G-to-A conversion in the complementary strand of the genome. Therefore, in some embodiments, the cytosine deaminase encoded by the polynucleotide of the present invention generates a C→T conversion in the meaningful (e.g., "+"; template) strand of the target nucleic acid or a G→A conversion in the antisense (e.g., "-", complementary) strand of the target nucleic acid.
[0232] In some embodiments, the adenine deaminase encoded by the nucleic acid construct of the present invention generates an A→G transition in the meaningful (e.g., "+"; template) strand of the target nucleic acid or a T→C transition in the antisense (e.g., "-", complementary) strand of the target nucleic acid.
[0233] The present invention relates to nucleic acid constructs encoding a base editor comprising a sequence-specific DNA-binding protein and a cytosine deaminase polypeptide, and nucleic acid constructs / expression cassettes / vectors encoding said base editor, which can be used in combination with guide nucleic acids to modify target nucleic acids, including but not limited to generating C→T or G→A mutations in target nucleic acids, said target nucleic acids including but not limited to plasmid sequences; generating C→T or G→A mutations in coding sequences to alter amino acid identity; generating C→T or G→A mutations in coding sequences to generate stop codons; generating C→T or G→A mutations in coding sequences to disrupt start codons; generating point mutations in genomic DNA to disrupt transcription factor binding; and / or generating point mutations in genomic DNA to disrupt splice sites.
[0234] The nucleic acid constructs of the present invention encoding a base editor comprising a sequence-specific DNA-binding protein and an adenine deaminase polypeptide, and expression cassettes and / or vectors encoding said base editor, can be used in combination with guide nucleic acids to modify target nucleic acids, including but not limited to generating A→G or T→C mutations in the target nucleic acids, said target nucleic acids including but not limited to plasmid sequences; generating A→G or T→C mutations in coding sequences to alter amino acid identity; generating A→G or T→C mutations in coding sequences to generate stop codons; generating A→G or T→C mutations in coding sequences to disrupt start codons; generating point mutations in genomic DNA to disrupt transcription factor binding; and / or generating point mutations in genomic DNA to disrupt splice sites.
[0235] The nucleic acid constructs of the present invention, comprising a CRISPR-Cas effector protein or a fusion protein thereof, can be used in combination with guide RNA (gRNA, CRISPR array, CRISPR RNA, crRNA) to modify target nucleic acids, said guide RNA being programmed to function in conjunction with the encoded CRISPR-Cas effector protein or domain. The guide nucleic acid used in the present invention comprises at least one spacer region sequence and at least one repeat sequence. The guide nucleic acid is capable of forming a complex with a CRISPR-Cas nuclease domain encoded and expressed by the nucleic acid constructs of the present invention, and the spacer region sequence is capable of hybridizing with the target nucleic acid to guide the complex (e.g., a CRISPR-Cas effector fusion protein (e.g., a CRISPR-Cas effector domain fused to a deaminase domain and optionally a UGI and / or a CRISPR-Cas effector domain fused to a peptide tag or affinity polypeptide) to the target nucleic acid, wherein the target nucleic acid can be modified (e.g., cleaved or edited) or regulated (e.g., regulated transcription) by the deaminase domain.
[0236] For example, a nucleic acid construct encoding a Cas9 domain linked to a cytosine deaminase domain (e.g., a fusion protein) can be combined with a Cas9 guide nucleic acid to modify a target nucleic acid, wherein the cytosine deaminase domain of the fusion protein deaminates a cytosine base in the target nucleic acid, thereby editing the target nucleic acid. In another example, a nucleic acid construct encoding a Cas9 domain linked to an adenine deaminase domain (e.g., a fusion protein) can be combined with a Cas9 guide nucleic acid to modify a target nucleic acid, wherein the adenine deaminase domain of the fusion protein deaminates an adenosine base in the target nucleic acid, thereby editing the target nucleic acid.
[0237] Similarly, the Cas12a domain (or other selected CRISPR-Cas nucleases, such as C2c1, C2c3, Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Cs) encoded by the cytosine deaminase domain or adenine deaminase domain is also encoded by the Cas12a domain linked to the cytosine deaminase domain or adenine deaminase domain. Nucleic acid constructs of m2, Csm3, Csm4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4(dinG) and / or Csf5 (e.g., fusion proteins) can be used in combination with Cas12a guide nucleic acid (or guide nucleic acid of other selected CRISPR-Cas nucleases) to modify target nucleic acids, wherein the cytosine deaminase domain or adenine deaminase domain of the fusion protein deaminates the cytosine bases in the target nucleic acid, thereby editing the target nucleic acid.
[0238] As used herein, “guide nucleic acid,” “guide RNA,” “gRNA,” “CRISPR RNA / DNA,” “crRNA,” or “crDNA” means a spacer sequence (e.g., a protospacer) that is complementary to (and hybridizes with) the target DNA and at least one repeat sequence (e.g., a repeat sequence, or a fragment or portion thereof, of a type V Cas12a CRISPR-Cas system; a repeat sequence, or a fragment thereof, of a type II Cas9 CRISPR-Cas system; a repeat sequence, or a fragment thereof, of a type V C2c1 CRISPR-Cas system). Repeated sequences or fragments thereof in the Cas system; for example, C2c3, Cas12a (also known as Cpf1), Cas12b, Cas12c, Cas12d, Cas12e, Cas13a, Cas13b, Cas13c, Cas13d, Cas1, Cas1B, Cas2, Cas3, Cas3', Cas3", Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csyl, Csy2, Csy3, Csel, Cse2, Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Cs Nucleic acids containing repeat sequences or fragments thereof from CRISPR-Cas systems of m4, Csm5, Csm6, Cmrl, Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, Csx10, Csx16, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4 (dinG) and / or Csf5, wherein the repeat sequences may be linked to the 5' and / or 3' ends of the spacer region sequence. The gRNAs of this invention can be designed based on type I, II, III, IV, V, or VI CRISPR-Cas systems.
[0239] In some implementations, the Cas12a gRNA may contain a repeat sequence (full length or a portion thereof (“handle”); e.g., a pseudoknot-like structure) and a spacer sequence from 5' to 3'.
[0240] In some embodiments, the guide nucleic acid may comprise not one repeat sequence-spacer sequence (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more repeat sequence-spacer sequences) (e.g., repeat sequence-spacer sequence-repeat sequence, e.g., repeat sequence-spacer sequence-repeat sequence-spacer sequence-repeat sequence-spacer sequence-repeat sequence-spacer sequence-repeat sequence-spacer sequence, etc.). The guide nucleic acid of the present invention is synthetic, artificial, and not found in nature. The gRNA can be quite long and can be used as an aptamer (as in the MS2 recruitment strategy) or other RNA structures suspended from the spacer region.
[0241] As used herein, "repetitive sequence" means, for example, any repetitive sequence of a wild-type CRISPR-Cas locus (e.g., Cas9, Cas12a, C2c1, etc.) or a repetitive sequence of a synthetic crRNA that functions together with the CRISPR-Cas effector protein encoded by the nucleic acid construct of this invention. The repetitive sequence that can be used in this invention can be any known or subsequently identified repetitive sequence of a CRISPR-Cas locus (e.g., type I, II, III, IV, V, or VI), or it can be a synthetic repetitive sequence designed to function in a type I, II, III, IV, V, or VI CRISPR-Cas system. The repetitive sequence may contain hairpin structures and / or stem-loop structures. In some embodiments, the repetitive sequence may form a pseudo-knot-like structure (i.e., a "stalk") at its 5' end. Therefore, in some embodiments, the repetitive sequence may be identical or substantially identical to repetitive sequences from wild-type CRISPR-Cas loci, including type I, type II, type III, type IV, type V, and / or type VI CRISPR-Cas loci. Repetitive sequences from wild-type CRISPR-Cas loci can be determined using established algorithms, such as CRISPRfinder provided via CRISPRdb (see Grissa et al., Nucleic Acids Res. 35 (WebServer issue): W52-7). In some embodiments, the repetitive sequence or a portion thereof is joined at its 3' end to the 5' end of a spacer sequence to form a repetitive sequence-spacer sequence (e.g., guide nucleic acid, guide RNA / DNA, crRNA, crDNA).
[0242] In some implementations, the repeat sequence comprises at least 10 nucleotides, or is substantially composed of at least 10 nucleotides, or is composed of at least 10 nucleotides, depending on whether the particular repeat sequence and the guide nucleic acid containing the repeat sequence are processed or unprocessed (e.g., about 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 to 100 or more nucleotides, or any range or value thereof).
[0243] In some embodiments, the repeating sequence comprises about 10 to about 20, about 10 to about 30, about 10 to about 45, about 10 to about 50, about 15 to about 30, about 15 to about 40, about 15 to about 45, about 15 to about 50, about 20 to about 30, about 20 to about 40, about 20 to about 50, about 30 to about 40, about 40 to about 80, about 50 to about 100 or more nucleotides, substantially composed of said nucleotides, or composed of said nucleotides.
[0244] The repeat sequence connected to the 5' end of the spacer sequence may contain a portion of the repeat sequence (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 or more consecutive nucleotides of a wild-type repeat sequence). In some embodiments, a portion of the repeat sequence linked to the 5' end of the spacer region sequence may be about 5 to about 10 consecutive nucleotides (e.g., about 5, 6, 7, 8, 9, or 10 nucleotides) of the wild-type CRISPR Cas repeat nucleotide sequence, and has at least 90% sequence identity with the same region (e.g., the 5' end) of the wild-type CRISPR Cas repeat nucleotide sequence (e.g., at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more). In some embodiments, a portion of the repeat sequence may include a pseudoknot-like structure (e.g., a "stalk") at its 5' end.
[0245] As used herein, a “spacer sequence” is a nucleotide sequence complementary to a target nucleic acid (e.g., target DNA) (e.g., the original spacer region) (e.g., any one of SEQ ID NO: 1-9; e.g., SEQ ID No: 175-182). The spacer sequence may be completely complementary or substantially complementary to the target nucleic acid (e.g., at least about 70% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more)). Therefore, in some embodiments, the spacer region sequence may have one, two, three, four, or five mismatches compared to the target nucleic acid, and these mismatches may be continuous or discontinuous. In some embodiments, the spacer region sequence may have 70% complementarity with the target nucleic acid. In other embodiments, the spacer region nucleotide sequence may have 80% complementarity with the target nucleic acid. In other embodiments, the spacer region nucleotide sequence may have 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 99.5% complementarity with the target nucleic acid (the original spacer region). In some embodiments, the spacer region sequence is 100% complementary to the target nucleic acid. The spacer region sequence may have a length of about 15 nucleotides to about 30 nucleotides (e.g., 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides, or any range or value thereof). Therefore, in some embodiments, the spacer sequence may be fully or substantially complementary over a region of at least about 15 to about 30 nucleotides in length on the target nucleic acid (e.g., the protospacer region). In some embodiments, the spacer length is about 20 nucleotides. In some embodiments, the spacer length is about 21, 22, or 23 nucleotides.
[0246] In some implementations, the 5' region of the spacer sequence of the guide nucleic acid may be identical to the target DNA, while the 3' region of the spacer may be substantially complementary to the target DNA (e.g., type V CRISPR-Cas). Alternatively, the 3' region of the spacer sequence of the guide nucleic acid may be identical to the target DNA, while the 5' region of the spacer may be substantially complementary to the target DNA (e.g., type II CRISPR-Cas). Thus, the overall complementarity of the spacer sequence to the target DNA may be less than 100%. Therefore, for example, in the guide nucleic acid of a type V CRISPR-Cas system, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 nucleotides in the 5' region (i.e., the seed region) of a 20-nucleotide spacer sequence may be 100% complementary to the target DNA, while the remaining nucleotides in the 3' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 70% complementary). In some embodiments, the first 1 to 8 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8 nucleotides and any range thereof) of the 5' end of the spacer sequence may be 100% complementary to the target DNA, while the remaining nucleotides of the 3' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 50% (e.g., 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) complementary).
[0247] As another example, in the guide nucleic acid of a type II CRISPR-Cas system, for example, the first 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10 nucleotides in the 3' region (i.e., the seed region) of a 20-nucleotide spacer sequence can be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 70% complementary). In some embodiments, the first 1 to 10 nucleotides (e.g., the first 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides and any range thereof) of the 3' end of the spacer sequence may be 100% complementary to the target DNA, while the remaining nucleotides in the 5' region of the spacer sequence are substantially complementary to the target DNA (e.g., at least about 50% (e.g., at least about 50%, 55%, 60%, 65%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more or any range or value thereof) complementary).
[0248] In some implementations, the seed region of the spacer may be about 8 to about 10 nucleotides long, about 5 to about 6 nucleotides long, or about 6 nucleotides long.
[0249] As used herein, “target nucleic acid,” “target DNA,” “target nucleotide sequence,” “target region,” or “target region in the genome” means that it is completely complementary (100% complementary) or substantially complementary (e.g., at least 70% (e.g., 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) to the spacer region sequence in the guide nucleic acid of the present invention. Target regions useful to CRISPR-Cas systems may be adjacent to the 3' (e.g., type V CRISPR-Cas system) or 5' (e.g., type II CRISPR-Cas system) of the PAM sequence in the genome of an organism (e.g., plant genome). The target region can be selected from any region of at least 15 consecutive nucleotides (e.g., 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides, etc.) located immediately adjacent to the PAM sequence.
[0250] "Protospacer sequence" refers to the target double-stranded DNA, specifically a portion of the target DNA (e.g., or a target region in the genome) that is completely or substantially complementary (and hybridizes) to the spacer sequence of a CRISPR repeat-spacer sequence (e.g., guide nucleic acid, CRISPR array, crRNA).
[0251] In the case of type V CRISPR-Cas (e.g., Cas12a) and type II CRISPR-Cas (Cas9) systems, the protospacer sequence is flanked by (e.g., adjacent) protospacer neighbor motifs (PAMs). For type IV CRISPR-Cas systems, the PAMs are located at the 5' end of the non-target strand and the 3' end of the target strand (e.g., see below).
[0252]
[0253] 3'AAANNNNNNNNNNNNNNNNNNNN-5' target strand (SEQ ID NO: 199)
[0254] ||||
[0255] 5'TTTNNNNNNNNNNNNNNNNNNN-3' non-target strand (SEQ ID NO: 199)
[0256] In the case of type II CRISPR-Cas systems (e.g., Cas9), the PAM is located immediately at the 3' of the target region. In type I CRISPR-Cas systems, the PAM is located at the 5' of the target strand. There is no known PAM for type III CRISPR-Cas systems. Makarova et al. described the nomenclature for all classes, types, and subtypes of CRISPR systems (Nature Reviews Microbiology 13: 722-736 (2015)). The guide nucleic acid structure and PAM are described by R. Barrangou (Genome Biol. 16: 247 (2015)).
[0257] Canonical Cas12a PAM is T-rich. In some embodiments, the canonical Cas12a PAM sequence can be 5'-TTN, 5'-TTTN, or 5'-TTTV. In some embodiments, a typical Cas9 (e.g., Streptococcus pyogenes) PAM can be 5'-NGG-3'. In some embodiments, non-canonical PAMs can be used, but may be less efficient.
[0258] Those skilled in the art can identify other PAM sequences using established experimental and computational methods. Thus, experimental methods, for example, include targeting sequences flanked by all possible nucleotide sequences and identifying sequence members that do not undergo targeting, such as through transformation of the target plasmid DNA (Esvelt et al. 2013. Nat. Methods 10: 1116-1121; Jiang et al. 2013. Nat. Biotechnol. 31: 233-239). In some aspects, computational methods may include performing a BLAST search on the natural spacer region to identify the original target DNA sequence in the phage or plasmid and comparing these sequences to determine conserved sequences of neighboring target sequences (Briner and Barrangou 2014. Appl. Environ. Microbiol. 80: 994-1001; Mojica et al. 2009. Microbiology 155: 733-740).
[0259] In some embodiments, the present invention provides an expression cassette and / or vector (e.g., one or more components of the editing system of the present invention) comprising the nucleic acid construct of the present invention. In some embodiments, an expression cassette and / or vector comprising the nucleic acid construct of the present invention and / or one or more guide nucleic acids may be provided. In some embodiments, the nucleic acid construct of the present invention encoding a base editor (e.g., a construct comprising a CRISPR-Cas effector protein and a deaminase domain (e.g., a fusion protein)) or components for base editing (e.g., a CRISPR-Cas effector protein fused with a peptide tag or affinity peptide, a deaminase domain fused with a peptide tag or affinity peptide, and / or a UGI fused with a peptide tag or affinity peptide) may be contained on an expression cassette or vector containing the same expression cassette or vector or on a separate expression cassette or vector containing the one or more guide nucleic acids. When a nucleic acid construct encoding a base editor or a component for base editing is contained on one or more expression cassettes or vectors separate from one or more expression cassettes or vectors containing a guide nucleic acid, the target nucleic acid may be contacted (e.g., provided together with) the one or more expression cassettes or vectors encoding the base editor or the component for base editing and the guide nucleic acid in any order between them, for example, before, simultaneously with, or after the expression cassette containing the guide nucleic acid is provided (e.g., contacted with the target nucleic acid).
[0260] As is known in the art, the fusion protein of the present invention may comprise a sequence-specific DNA-binding domain, a CRISPR-Cas peptide, and / or a deaminase domain fused to a peptide tag or an affinity peptide that interacts with the peptide tag, for recruiting the deaminase to a target nucleic acid. The recruitment method may further include a guide nucleic acid linked to an RNA recruitment motif and a deaminase (which is fused to an affinity peptide capable of interacting with the RNA recruitment motif) to recruit the deaminase to the target nucleic acid. Alternatively, chemical interactions may be used to recruit peptides (e.g., deaminases) to target nucleic acids.
[0261] Peptide tags (e.g., epitopes) that can be used in this invention may include, but are not limited to, GCN4 peptide tags (e.g., Sun tags), c-Myc affinity tags, HA affinity tags, His affinity tags, S affinity tags, methionine-His affinity tags, RGD-His affinity tags, FLAG octapeptides, strep tags or strep tag II, V5 tags, and / or VSV-G epitopes. Any epitope that can be linked to a polypeptide, and any epitope that can be linked to a corresponding affinity polypeptide for another polypeptide, can be used as a peptide tag in this invention. In some embodiments, the peptide tag may comprise one or two or more copies of the peptide tag (e.g., repeating units, polymerized epitopes (e.g., tandem repeats)) (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25 or more repeating units). In some embodiments, the affinity polypeptide interacting with / binding to the peptide tag may be an antibody. In some embodiments, the antibody may be an scFv antibody. In some embodiments, the affinity peptide that binds to the peptide tag may be synthetic (e.g., evolved for affinity interactions) and includes, but is not limited to, affinity peptides, anticalins, monomers, and / or DARPin (see, for example, Sha et al., Protein Sci. 26(5):910-924(2017)); Gilbreth (Curr Opin Struc Biol 22(4):413-420(2013)), U.S. Patent No. 9,982,053, each of which is incorporated herein by reference in its entirety with respect to teachings relating to affinity peptides, anticalins, monomers, and / or DARPin.
[0262] In some embodiments, the guide nucleic acid may be linked to an RNA recruitment motif, and the polypeptide to be recruited (e.g., a deaminase) may be fused to an affinity polypeptide that binds to the RNA recruitment motif. The guide nucleic acid binds to the target nucleic acid, and the RNA recruitment motif binds to the affinity polypeptide, thereby recruiting the polypeptide to the guide nucleic acid and contacting the target nucleic acid with the polypeptide (e.g., a deaminase). In some embodiments, two or more polypeptides may be recruited to the guide nucleic acid, thereby contacting the target nucleic acid with two or more polypeptides (e.g., deaminases).
[0263] In some embodiments, the peptide fused with the affinity peptide can be a reverse transcriptase, and the guide nucleic acid can be an extended guide nucleic acid linked to an RNA recruitment motif. In some embodiments, the RNA recruitment motif can be located at the 3' end of the extended portion of the extended guide nucleic acid (e.g., 5'-3', repeat sequence-spacer region-extended portion (RT template-primer binding site)-RNA recruitment motif). In some embodiments, the RNA recruitment motif can be embedded in the extended portion.
[0264] In some embodiments of the invention, the extended guide RNA and / or guide RNA may be linked to one or two or more RNA recruitment motifs (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more motifs, e.g., at least 10 to about 25 motifs), optionally wherein the two or more RNA recruitment motifs may be the same RNA recruitment motif or different RNA recruitment motifs. In some embodiments, the RNA recruitment motif and corresponding affinity peptide may include, but are not limited to, a telomerase Ku-binding motif (e.g., Ku-binding hairpin) and a corresponding affinity peptide Ku (e.g., Ku heterodimer), a telomerase Sm7-binding motif and a corresponding affinity peptide Sm7, an MS2 phage operon stem loop and a corresponding affinity peptide MS2 capsid protein (MCP), a PP7 phage operon stem loop and a corresponding affinity peptide PP7 capsid protein (PCP), an SfMu phage Com stem loop and a corresponding affinity peptide Com, RNA-binding proteins, PUF binding sites (PBS) and affinity peptide Pumilio / fem-3 mRNA binding factor (PUF), and / or synthetic RNA aptamers and aptamer ligands as the corresponding affinity peptides. In some embodiments, the RNA recruitment motif and corresponding affinity peptide may be an MS2 phage operon stem loop and an affinity peptide MS2 capsid protein (MCP). In some implementations, the RNA recruitment motif and the corresponding affinity peptide can be a PUF binding site (PBS) and the affinity peptide Pumilio / fem-3 mRNA binding factor (PUF).
[0265] In some implementations, the components used to recruit peptides and nucleic acids may be those that function through chemical interactions, including but not limited to, rapamycin-induced dimerization of FRB–FKBP; biotin-streptoacidin; SNAP tag; Halo tag; CLIP tag; compound-induced DmrA-DmrC heterodimer; bifunctional ligands (e.g., two protein-binding chemicals fused together, such as dihydrofolate reductase (DHFR)).
[0266] In some embodiments, the nucleic acid constructs, expression cassettes, or vectors of the present invention optimized for expression in plants may have about 70% to 100% (e.g., about 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100%) identity with nucleic acid constructs, expression cassettes, or vectors containing one or more of the same polynucleotides (but which have not been codon-optimized for expression in plants).
[0267] In some embodiments, the present invention provides cells comprising one or more polynucleotides, guide nucleic acids, nucleic acid constructs, expression cassettes, or vectors of the present invention.
[0268] In some embodiments, a method is provided for editing an endogenous HD-Zip transcription factor gene in a plant or plant part, the method comprising contacting a target site in the HD-Zip transcription factor gene in the plant or plant part with a cytosine base editing system, the cytosine base editing system comprising a cytosine deaminase and a nucleic acid binding domain that binds to the target site in the HD-Zip transcription factor, the HD-Zip transcription factor gene encoding (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:5). (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 8), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9) and / or (f) a polypeptide containing the amino acid sequence VWFQNRRA (SEQ ID NO: 9), thereby editing an endogenous HD-Zip transcription factor gene in a plant or a part thereof and producing a plant or a part thereof containing at least one cell with a mutation in the endogenous HD-Zip transcription factor gene.
[0269] In some embodiments, a method is provided for editing an endogenous HD-Zip transcription factor gene in a plant or plant part, the method comprising contacting a target site in the HD-Zip transcription factor gene in the plant or plant part with an adenosine base editing system, the adenosine base editing system comprising an adenosine aminoase and a nucleic acid binding domain that binds to the target site in the HD-Zip transcription factor, the HD-Zip transcription factor gene encoding (a) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:1 or SEQ ID NO:2; (b) a polypeptide comprising a sequence having at least 80% sequence identity with the amino acid sequence of SEQ ID NO:3; (c) a polypeptide comprising a sequence having at least 95% sequence identity with the nucleotide sequence of SEQ ID NO:4; (d) a polypeptide comprising a sequence having the amino acid sequence shown in SEQ ID NO:5, wherein X is L or S; (e) a polypeptide comprising: (i) having the amino acid sequence RKKLRLX1KX2QX3 (SEQ ID NO:5). (ii) a sequence having the amino acid sequence PX1X2X2LTX3CPX4CER (SEQ ID NO: 6), wherein X1 is S or T, X2 is D or E, and X3 is S or A; (iii) a sequence having the amino acid sequence ENRRLX1X2EX3 (SEQ ID NO: 7), wherein X1 is Q or H, X2 is R or K, and X3 is V or L; and a sequence having the amino acid sequence VWFQNRRA (SEQ ID NO: 9) and / or (f) a polypeptide containing the amino acid sequence VWFQNRRA (SEQ ID NO: 9), thereby editing an endogenous HD-Zip transcription factor gene in a plant or a part thereof and producing a plant or a part thereof containing at least one cell with a mutation in the endogenous HD-Zip transcription factor gene.
[0270] In some embodiments, a method is provided for detecting mutant HD-Zip (mutations in endogenous HD-Zip transcription factor genes), the method comprising detecting deletions in nucleic acids encoding any one of the amino acid sequences in SEQ ID NO:1-98 in a plant genome, wherein any one of the amino acid sequences in SEQ ID NO:1-98 contains the amino acid sequence of SEQ ID NO:9, and the deletion is in the nucleotide sequence encoding the amino acid sequence of SEQ ID NO:9.
[0271] In some embodiments, the present invention provides a method for detecting mutations in endogenous HD-Zip genes, comprising detecting in the plant genome a nucleotide sequence encoding at least one of the polypeptide sequences in SEQ ID NO:1-98.
[0272] In some embodiments, the present invention provides a method for generating plants containing a mutation in an endogenous HD-Zip transcription factor gene and at least one target polynucleotide, the method comprising crossing a plant of the present invention (a first plant) containing at least one mutation in an endogenous HD-Zip transcription factor gene with a second plant containing at least one target polynucleotide to generate offspring plants; and selecting offspring plants containing at least one mutation in an HD-Zip transcription factor gene and at least one target polynucleotide, thereby generating plants containing a mutation in an endogenous HD-Zip transcription factor gene and at least one target polynucleotide.
[0273] The present invention also provides a method for producing a plant containing a mutation in an endogenous HD-Zip transcription factor gene and at least one target polynucleotide, the method comprising introducing at least one target polynucleotide into a plant of the present invention containing at least one mutation in an HD-Zip transcription factor gene, thereby producing a plant containing at least one mutation in an HD-Zip transcription factor gene and at least one target polynucleotide.
[0274] In some embodiments, the present invention provides a method for producing a plant containing a mutation in an endogenous HD-Zip transcription factor gene and at least one target polynucleotide, the method comprising introducing at least one target polynucleotide into the present invention plant containing at least one mutation in an endogenous HD-Zip transcription factor gene, thereby producing a plant containing at least one mutation in an HD-Zip transcription factor gene and at least one target polynucleotide.
[0275] The target polynucleotide can be any polynucleotide capable of conferring a desired phenotype to a plant or otherwise altering the phenotype or genotype of a plant. In some embodiments, the target polynucleotide can be a polynucleotide that confers herbicide tolerance, insect resistance, disease resistance, increased yield, increased nutrient use efficiency, or resistance to abiotic stresses.
[0276] In some embodiments, a method is provided for generating plants containing mutations in the endogenous HD-Zip transcription factor gene and having a highly dwarf or extremely short phenotype, the method comprising crossing a plant of the present invention (a first plant) having at least one mutation in the endogenous HD-Zip transcription factor gene with a second plant containing a highly dwarf or extremely short phenotype to generate offspring plants; and selecting offspring plants containing at least one mutation in the HD-Zip transcription factor gene and a highly dwarf or extremely short phenotype to generate plants having a highly dwarf or extremely short phenotype and containing at least one mutation in the endogenous HD-Zip transcription factor gene.
[0277] The present invention also provides a method for controlling weeds in containers (e.g., cans or seed trays), growth chambers, greenhouses, fields (e.g., cultivated land), recreational areas, lawns, and / or roadsides containing one or more of the plants of the present invention, comprising applying a herbicide to one or more of the plants of the present invention growing in the containers, growth chambers, fields, or greenhouses, thereby controlling weeds in the containers, growth chambers, greenhouses, fields, recreational areas, lawns, and / or roadsides where one or more of the plants are growing.
[0278] In some embodiments, a method for reducing insect predation on plants (or plants) is provided, comprising applying an insecticide to one or more plants of the invention, thereby reducing insect predation on the plants. In some embodiments, the one or more plants may be grown in containers, growing rooms, fields, recreational areas (e.g., sports fields, golf courses), lawns, roadsides, or greenhouses.
[0279] In some embodiments, the present invention provides a method for reducing fungal diseases on plants, comprising applying a fungicide to one or more plants of the present invention, thereby reducing fungal diseases on one or more plants. In some embodiments, one or more plants may be grown in containers, growing rooms, fields, recreational areas (e.g., sports fields, golf courses), lawns, roadsides, or greenhouses.
[0280] The nucleic acid constructs of the present invention (e.g., constructs containing sequence-specific DNA-binding domains, CRISPR-Cas effector domains, deaminase domains, reverse transcriptase (RT), RT templates and / or guide nucleic acids, etc.) and expression cassettes / vectors containing them can be used as the editing system of the present invention for modifying target nucleic acids and / or their expression.
[0281] The target nucleic acids of any plant or plant part (or grouping of plants, e.g., genus or higher classification) can be modified (e.g., mutated, such as through base editing, cutting, nicking, etc.) using the peptides, polynucleotides, RNPs, nucleic acid constructs, expression cassettes, and / or vectors of the present invention. The plants include angiosperms, gymnosperms, monocots, dicots, C3 plants, C4 plants, CAM plants, bryophytes, ferns and / or fernoids, microalgae, and / or macroalgae. The plants and / or plant parts that can be modified as described herein can be plants and / or plant parts of any plant species / variety / cultivar. In some embodiments, the plants that can be modified as described herein are monocots. In some embodiments, the plants that can be modified as described herein are dicots.
[0282] As used herein, the term "plant part" includes, but is not limited to, reproductive tissues (e.g., petals, sepals, stamens, pistils, receptacles, anthers, pollen, flowers, fruits, flower buds, ovules, seeds, embryos, nuts, kernels, spikes, rachis, and shells); vegetative tissues (e.g., petioles, stems, roots, root hairs, root tips, pith, coleoptiles, stalks, buds, branches, bark, apical meristems, axillary buds, cotyledons, hypocotyls, and leaves); vascular tissues (e.g., phloem and xylem); specialized cells, such as epidermal cells, parenchyma cells, collenchyma cells, sclerenchyma cells, stomata, guard cells, cuticles, mesophyll cells; callus; and cuttings. The term "plant part" also includes plant cells, including intact plant cells, plant protoplasts, plant tissues, plant organs, plant cell tissue cultures, plant callus, plant masses, etc., in plants and / or plant parts. As used herein, "stem" refers to the above-ground part, including leaves and stems. As used in this article, the term "tissue culture" includes cultures of tissues, cells, protoplasts, and callus.
[0283] As used herein, "plant cell" refers to the structural and physiological unit of a plant, which typically includes a cell wall but also includes protoplasts. The plant cells of this invention may be in the form of isolated single cells, or may be cultured cells, or may be part of a higher tissue unit, such as plant tissue (including callus) or plant organ. In some embodiments, the plant cell may be an algal cell. A "protoplast" is an isolated plant cell without a cell wall or with only a partial cell wall. Therefore, in some embodiments of this invention, the transgenic cell containing the nucleic acid molecules and / or nucleotide sequences of this invention is a cell of any plant or plant part, including but not limited to root cells, leaf cells, tissue culture cells, seed cells, flower cells, fruit cells, pollen cells, etc. In some aspects of this invention, the plant part may be plant germplasm. In some aspects, the plant cell may be a non-reproductive plant cell that cannot regenerate into a plant.
[0284] "Plant cell culture" refers to a culture of plant units such as protoplasts, cultured cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at different developmental stages.
[0285] As used in this article, a “plant organ” is a unique and distinctly structured and differentiated part of a plant, such as a root, stem, leaf, flower bud, or embryo.
[0286] As used herein, “plant tissue” means a group of plant cells organized into structural and functional units. This includes any plant tissue in or cultured. The term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue cultures, and any group of plant cells organized into structural and / or functional units. The use of this term in conjunction with any particular type of plant tissue listed above or originally included in this definition, or in the absence of any such particular type of plant tissue, does not exclude any other type of plant tissue.
[0287] In some embodiments of the invention, transgenic tissue cultures or transgenic plant cell cultures are provided, wherein the transgenic tissue or cell culture contains the nucleic acid molecule / nucleotide sequence of the invention. In some embodiments, the transgene can be eliminated from plants developed from transgenic tissues or cells by mating transgenic plants with non-transgenic plants and selecting plants in the offspring that contain the desired gene edit but not the transgene used to produce said edit.
[0288] Any plant containing an endogenous HD-Zip transcription factor gene capable of regulating the shade avoidance response (SAR) can be modified as described herein to reduce / attenuate or eliminate SAR in the plant. In some embodiments, the plant may be a monocotyledonous plant. In some embodiments, the plant may be a dicotyledonous plant.
[0289] Non-limiting examples of plants that can be modified as described herein include, but are not limited to, turfgrass (e.g., bluegrass, creeping bentgrass, ryegrass, fescue), and feathery reeds. Grasses, including tufted grasses, miscanthus, reeds, switchgrass, and vegetable crops such as artichokes, turnips, arugula, leeks, asparagus, lettuce (e.g., head, leaf, and long-leaf lettuce), dark green taro (malanga), cucurbits (e.g., melons, watermelons, crenshaws, cantaloupes, and Roman melons), rapeseed crops (e.g., Brussels sprouts, cabbage, cauliflower, broccoli, kale, headless cabbage, Chinese cabbage, and bok choy), artichokes, carrots, napa cabbage, okra, onions, celery, parsley, chickpeas, parsley, European turnips, chicory, peppers, potatoes, cucurbitaceae plants (e.g., zucchini, cucumbers, dense zucchini, pumpkins, melons, watermelons, and Roman melons), radishes, dried onions, turnips, eggplants, ginseng, broadleaf endive, scallions, and chicory. Garlic, spinach, scallions, pumpkin, leafy greens, beets (sugar beets and feed beets), sweet potatoes, Swiss chard, wasabi, tomatoes, turnips, and spices; fruit crops such as apples, apricots, cherries, nectarines, peaches, pears, plums, apricots, cherries, quince, figs, nuts (e.g., chestnuts, pecans, pistachios, hazelnuts, peanuts, walnuts, macadamia nuts, almonds, etc.), citrus fruits (e.g., Clementine, kumquats, oranges, grapefruits, tangerines, lemons, limes, etc.), blueberries, black raspberries, boysonberries, cranberries, currants, currants, raspberries, strawberries, blackberries, grapes (for wine and table), avocados, bananas, kiwis, persimmons, pomegranates, pineapples, tropical fruits, pears, melons, mangoes, papayas, and lychees; field crops such as clover, alfalfa, timothy, evening primrose, and meadow foam. Foam), corn / maize (field corn, sweet corn, popcorn), hops, jojoba, buckwheat, safflower, quinoa, wheat, rice, barley, rye, millet, sorghum, oats, black wheat, tobacco, kapok, legumes (legumes (e.g., green and dried), lentils, peas, soybeans), oilseeds (rape, canola, mustard, poppy, olive, sunflower, coconut, castor oil plants, cocoa beans, peanuts, oil palm), duckweed, Arabidopsis, fiber plants (cotton, flax, hemp, jute), hemp (e.g., cannabis (Cannabis)). (sativa), Indian hemp and herbaceous hemp), Lauraceae (cinnamon, camphor) or plants such as coffee, sugarcane, tea and natural rubber plants; and / or flower bed plants and / or ornamental plants such as flowering plants, cacti, succulents and / or flowering plants, cacti, succulents (e.g. roses, tulips, violets), and forest trees such as broad-leaved trees and evergreen plants such as conifers;For example, trees such as elm, ash, oak, maple, fir, spruce, cedar, pine, birch, cypress, eucalyptus, and willow, as well as shrubs and other seedlings. In some embodiments, the nucleic acid constructs and / or expression cassettes and / or vectors encoding them of the present invention can be used to modify corn, soybean, wheat, brassica, rice, tomato, pepper, or sunflower.
[0290] In some implementations, plants that can be modified as described herein may include, but are not limited to, corn, soybean, rapeseed, wheat, rice, cotton, sugarcane, sugar beet, barley, oats, alfalfa, sunflower, safflower, oil palm, sesame, coconut, tobacco, potato, sweet potato, cassava, coffee, apple, plum, apricot, peach, cherry, pear, fig, banana, citrus, cocoa, avocado, olive, almond, walnut, strawberry, watermelon, pepper, grape, tomato, cucumber, or certain species of Brassica (e.g., European rapeseed, kale, turnip, mustard-type rapeseed, and / or black mustard).
[0291] In some implementations, the plant that can be modified as described herein is maize (i.e., corn, maize), optionally wherein the maize plant contains a dwarf / semi-dwarf phenotype.
[0292] In some embodiments, the plant that can be modified as described herein is wheat (e.g., common wheat, durum wheat, and / or dense-spike wheat (T. compactum)). In some embodiments, the wheat plant may contain at least one non-natural mutation in the endogenous HD-Zip transcription factor in its A, B, and / or D genomes.
[0293] Therefore, the plants or cultivars preferred according to the present invention include all plants for which genetic material has been acquired through genetic modification, said genetic material conferring particularly advantageous useful characteristics (“traits”) upon these plants. Examples of such traits are better plant growth, vigor, stress tolerance, standing ability, lodging resistance, nutrient uptake, plant nutrition and / or yield, particularly improved growth, enhanced tolerance to high or low temperatures, enhanced tolerance to drought or water or soil salinity levels, enhanced flowering performance, easier harvesting, accelerated ripening, higher yield, higher quality and / or higher nutritional value of harvested products, better shelf life and / or processability of harvested products.
[0294] A further and particularly emphasized example of this property is enhanced resistance to animal and microbial pests, such as insects, arachnids, nematodes, mites, slugs, and snails, attributable to toxins formed, for example, in plants. Among the DNA sequences encoding proteins conferring resistance to such animal and microbial pests (especially insects), the genetic material encoding Bt proteins from *Bacillus thuringiensis*, which is widely described in the literature and well-known to those skilled in the art, will be specifically mentioned. Proteins extracted from bacteria such as *Photorhabdus* (WO97 / 17432 and WO98 / 08932) will also be mentioned. Specifically, Bt Cry or VIP proteins, including Cry1A, CryIAb, CryIAc, CryIIA, CryIIIA, CryIIIB2, and Cry9c, will be mentioned. Cry2Ab, Cry3Bb, and CryIF proteins or their toxic fragments, as well as hybrids or combinations thereof, especially Cry1F proteins or hybrids derived from Cry1F proteins (e.g., hybrid Cry1A-Cry1F proteins or their toxic fragments), Cry1A type proteins or their toxic fragments, preferably Cry1Ac proteins or hybrids derived from Cry1Ac proteins (e.g., hybrid Cry1Ab-Cry1Ac proteins) or Cry1Ab or Bt2 proteins or their toxic fragments, Cry2Ae, Cry2Af, or Cry2Ag proteins or their toxic fragments, Cry1A.105 proteins or their toxic fragments, VIP3Aa19 proteins, VIP3Aa20 proteins, VIP3A proteins, VIP3Aa proteins or their toxic fragments generated in the COT202 or COT203 cotton events (e.g., Estruch et al. (1996), Proc Natl Acad Sci US). The Cry protein described in A.28;93(11):5389-94, the Cry protein described in WO2001 / 47952, insecticidal proteins from strains of the genus Xenorhabdus (as described in WO98 / 50427), the genus Serratia (especially from S. entomophila) or the genus Photorhabdus, and the Tc protein from the genus Photorhabdus as described in WO98 / 08932. Furthermore, any variants or mutants of these proteins that differ from any of the above sequences (especially the sequences of their toxic fragments) in certain amino acids (1-10, preferably 1-5) or that are fused to a transport peptide (such as a plasmid transport peptide) or another protein or peptide are also included herein.
[0295] Another particularly emphasized example of this property is the conferral of tolerance to one or more herbicides (e.g., imidazolinones, sulfonylureas, glyphosate, or glufosinate). In the DNA sequences (i.e., target polynucleotides) encoding proteins that confer tolerance to certain herbicides in transformed plant cells and plants, specific references will be made to the bar or PAT genes described in WO2009 / 152359 or the *Streptomyces coelicolor* gene (which confers tolerance to glufosinate-ammonium herbicide), genes encoding suitable EPSPS (5-enolpyruvylshikimate-3-phosphate synthase) (which confers tolerance to herbicides targeting EPSPS, particularly herbicides such as glyphosate and its salts), genes encoding glyphosate-n-acetyltransferase, or genes encoding glyphosate oxidoreductase. Other suitable herbicide tolerance traits include at least one ALS (acetyllactate synthase) inhibitor (e.g., WO2007 / 024782), a mutant Arabidopsis ALS / AHAS gene (e.g., U.S. Patent 6,855,533), a gene encoding a 2,4-D-monooxygenase that confers tolerance to 2,4-D (2,4-dichlorophenoxyacetic acid), and a gene encoding a dicamba monooxygenase that confers tolerance to dicamba (3,6-dichloro-2-methoxybenzoic acid).
[0296] A further, particularly emphasized example of this property is enhanced resistance to plant pathogenic fungi, bacteria, and / or viruses, which is attributed to, for example, systemically acquired resistance (SAR), systemins, phytoalexins, elicitors, and also resistance genes and the corresponding expressed proteins and toxins.
[0297] Transgenic events that are particularly useful in transgenic plants or plant cultivars that can be preferentially treated according to the present invention include event 531 / PV-GHBK04 (cotton, insect control, described in WO2002 / 040677), event 1143-14A (cotton, insect control, not preserved, described in WO2006 / 128569); event 1143-51B (cotton, insect control, not preserved, described in WO2006 / 128570); and event 1445 (cotton, herbicide tolerance, not preserved, described in US-A). Event 17053 (rice, herbicide tolerance, deposited with PTA-9843, described in WO2010 / 117737); Event 17314 (rice, herbicide tolerance, deposited with PTA-9844, described in WO2010 / 117735); Event 281-24-236 (cotton, insect control - herbicide tolerance, deposited with PTA-6233, described in WO2005 / 103266 or US-A 2005-216969); Event 3006-210-23 (cotton, insect control - herbicide tolerance, deposited with PTA-6233, described in US-A 2005 / 103266 or US-A 2005-216969). Event 3272 (maize, quality traits, preserved as PTA-9972, described in WO2006 / 098952 or US-A 2006-230473); Event 33391 (wheat, herbicide tolerance, preserved as PTA-2347, described in WO2002 / 027004); Event 40416 (maize, insect control - herbicide tolerance, preserved as ATCC PTA-11508, described in WO 11 / 075593); Event 43A47 (maize, insect control - herbicide tolerance, preserved as ATCC PTA-11509, described in WO2011 / 075595); Event 5307 (maize, insect control, preserved as ATCC PTA-11509, described in WO2011 / 075595); Event ASR-368 (Gnaphalium affine, herbicide tolerance, deposited with ATCC PTA-4816, described in US-A 2006-162007 or WO2004 / 053062); Event B16 (Maize, herbicide tolerance, not deposited, described in US-A 2003-126634); Event BPS-CV127-9 (Soybean, herbicide tolerance, deposited with NCIMB 41603, described in WO2010 / 080829); Event BLR1 (European rapeseed, restoration of male sterility, deposited with NCIMB 41193, described in WO2005 / 074671).Event CE43-67B (cotton, insect control, deposited under DSM ACC2724, described in US-A 2009-217423 or WO2006 / 128573); Event CE44-69D (cotton, insect control, not deposited, described in US-A 2010-0024077); Event CE44-69D (cotton, insect control, not deposited, described in WO2006 / 128571); Event CE46-02A (cotton, insect control, not deposited, described in WO2006 / 128572); Event COT102 (cotton, insect control, not deposited, described in US-A 2006-130175 or WO2004 / 039986); Event COT202 (cotton, insect control, not deposited, described in US-A 2009-217423 or WO2006 / 128573); Event COT203 (cotton, insect control, not deposited, described in WO2005 / 054480); Event DAS21606-3 / 1606 (soybean, herbicide tolerance, deposited as PTA-11028, described in WO2012 / 033794); Event DAS40278 (maize, herbicide tolerance, deposited as ATCCPTA-10244, described in WO2011 / 022469). (Described); Event DAS-44406-6 / pDAB8264.44.06.l (Soybean, Herbicide Tolerance, deposited with PTA-11336, described in WO2012 / 075426), Event DAS-14536-7 / pDAB8291.45.36.2 (Soybean, Herbicide Tolerance, deposited with PTA-11335, described in WO2012 / 075429), Event DAS-59122-7 (Maize, Insect Control - Herbicide Tolerance, deposited with ATCC) Events include: PTA 11384 (deposited in US-A2006-070139), DAS-59132 (maize, insect control - herbicide tolerance, not deposited, described in WO2009 / 100188); DAS68416 (soybean, herbicide tolerance, deposited in ATCC PTA-10442, described in WO2011 / 066384 or WO2011 / 066360); DP-098140-6 (maize, herbicide tolerance, deposited in ATCC PTA-8296, described in US-A 2009-137395 or WO08 / 112019); and DP-305423-1 (soybean, quality traits, not deposited, described in US-A 2009-070139). (As described in 2008-312082 or WO2008 / 054747); Event DP-32138-1 (maize, hybrid system, preserved with ATCC PTA-9158,(described in US-A 2009-0210970 or WO2009 / 103049); Event DP-356043-5 (soybean, herbicide tolerance, deposited with ATCC PTA-8287, described in US-A 2010-0184079 or WO2008 / 002872); Event EE-I (brinjal, insect control, not deposited, described in WO07 / 091277); Event Fil 17 (maize, herbicide tolerance, deposited with ATCC 209031, described in US-A 2009-0210970 or WO2009 / 103049); Event FG72 (soybean, herbicide tolerance, deposited as PTA-11041, described in WO2011 / 063413); Event GA21 (maize, herbicide tolerance, deposited as ATCC 209033, described in US-A 2005-086719 or WO98 / 044140); Event GG25 (maize, herbicide tolerance, deposited as ATCC 209032, described in US-A 2005-188434 or WO98 / 044140); Event GHB119 (cotton, insect control-herbicide tolerance, deposited as ATCC 2006-059581 or WO98 / 044140). PTA-8398 (extracted from WO2008 / 151780); Event GHB614 (cotton, herbicide tolerance, deposited with ATCC PTA-6878, described in US-A 2010-050282 or WO2007 / 017186); Event GJ11 (maize, herbicide tolerance, deposited with ATCC209030, described in US-A 2005-188434 or WO98 / 044140); Event GM RZ13 (sugar beet, virus resistance, deposited with NCIMB-41601, described in WO2010 / 076212); Event H7-l (sugar beet, herbicide tolerance, deposited with NCIMB 41158 or NCIMB 41159, described in US-A 2008 / 151780); Event JOPLIN1 (wheat, disease tolerance, not deposited, described in US-A 2008-064032); Event LL27 (soybean, herbicide tolerance, deposited with NCIMB 41658, described in WO2006 / 108674 or US-A 2008-320616); Event LL55 (soybean, herbicide tolerance, deposited with NCIMB 41660, described in WO2006 / 108675 or US-A 2008-196127); Event LLcotton25 (cotton, herbicide tolerance, deposited with ATCC PTA-3343).The following events are described in WO2003 / 013224 or USA 2003-097687: Event LLRICE06 (rice, herbicide tolerance, deposited with ATCC 203353, described in US Patent No. 6,468,747 or WO2000 / 026345); Event LLRICE62 (rice, herbicide tolerance, deposited with ATCC 20335, described in WO2000 / 026345); Event LLRICE601 (rice, herbicide tolerance, deposited with ATCC PTA-2600, described in US-A 2008-2289060 or WO2000 / 026356); Event LY038 (maize, quality traits, deposited with ATCC PTA-5623, described in US-A 2008-2289060 or WO2000 / 026356). Event MIR162 (maize, insect control, deposited as PTA-8166, described in US-A 2009-300784 or WO2007 / 142840); Event MIR604 (maize, insect control, not deposited, described in US-A 2008-167456 or WO2005 / 103301); Event MON15985 (cotton, insect control, deposited as ATCC PTA-2516, described in US-A 2004-250317 or WO2002 / 100163); Event MON810 (maize, insect control, not deposited, described in US-A 2002-102582); Event MON863 (maize, insect control, deposited as ATCC PTA-2516, described in US-A 2002-102582); PTA-2605 (extracted from WO2004 / 011601 or US-A2006-095986); Event MON87427 (maize, pollination control, deposited with ATCC PTA-7899, described in WO2011 / 062904); Event MON87460 (maize, stress tolerance, deposited with ATCC PTA-8910, described in WO2009 / 111263 or US-A 2011-0138504); Event MON87701 (soybean, insect control, deposited with ATCC PTA-8194, described in US-A 2009-130071 or WO2009 / 064652); Event MON87705 (soybean, quality trait - herbicide tolerance, deposited with ATCC PTA-9241, described in US-A... Event MON87708 (Soybean, herbicide tolerance, deposited with ATCC PTA-9670, described in WO2011 / 034704); Event MON87712 (Soybean, yield, deposited with PTA-10296).The following events are described in WO2012 / 051199: Event MON87754 (soybean, quality traits, deposited with ATCC PTA-9385, described in WO2010 / 024976); Event MON87769 (soybean, quality traits, deposited with ATCC PTA-8911, described in US-A 2011-0067141 or WO2009 / 102873); Event MON88017 (maize, insect control - herbicide tolerance, deposited with ATCC PTA-5582, described in US-A 2008-028482 or WO2005 / 059103); Event MON88913 (cotton, herbicide tolerance, deposited with ATCC PTA-4854, described in WO2004 / 072235 or US-A). Event MON88302 (European rapeseed, herbicide tolerance, deposited with PTA-10955, described in WO2011 / 153186), Event MON88701 (cotton, herbicide tolerance, deposited with PTA-11754, described in WO2012 / 134808), Event MON89034 (maize, insect control, deposited with ATCC PTA-7455, described in WO 07 / 140256 or US-A 2008-260932), Event MON89788 (soybean, herbicide tolerance, deposited with ATCC PTA-6708, described in US-A 2006-282915 or WO2006 / 130436), Event MS1 1 (European rapeseed, pollination control-herbicide tolerance, described in WO2001 / 031042 with ATCC PTA-850 or PTA-2485 accession); Event MS8 (European rapeseed, pollination control-herbicide tolerance, described in WO2001 / 041558 or US-A 2003-188347 with ATCC PTA-730 accession); Event NK603 (maize, herbicide tolerance, described in US-A 2007-292854 with ATCC PTA-2478 accession); Event PE-7 (rice, insect control, not accessed, described in WO2008 / 114282); Event RF3 (European rapeseed, pollination control-herbicide tolerance, described in WO2001 / 041558 or US-A 2003-188347 with ATCC PTA-730 accession); (described in 2003-188347); Event RT73 (European rapeseed, herbicide tolerance, not deposited, described in WO2002 / 036831 or US-A 2008-070260); Event SYHT0H2 / SYN-000H2-5 (soybean, herbicide tolerance, deposited with PTA-11226,Event T227-1 (sugar beet, herbicide tolerance, not deposited, described in WO2012 / 082548); Event T25 (maize, herbicide tolerance, not deposited, described in US-A 2001-029014 or WO2001 / 051654); Event T304-40 (cotton, insect control - herbicide tolerance, deposited with ATCC PTA-8171, described in US-A Event T342-142 (cotton, insect control, not deposited, described in WO2006 / 128568); Event TC1507 (maize, insect control - herbicide tolerance, not deposited, described in US-A 2005-039226 or WO2004 / 099447); Event VIP1034 (maize, insect control - herbicide tolerance, described in ATCC). PTA-3925 (extracted from WO2003 / 052073), event 32316 (maize, insect control - herbicide tolerance, extracted from PTA-11507, described in WO2011 / 084632), event 4114 (maize, insect control - herbicide tolerance, extracted from PTA-11506, described in WO2011 / 084621), optionally with event EE-GM1 / LL2 7 or event EE-GM2 / LL55 (WO2011 / 063413A2) superimposed with event EE-GM3 / FG72 (soybean, herbicide tolerance, ATCC accession number PTA-11041), event DAS-68416-4 (soybean, herbicide tolerance, ATCC accession number PTA-10442, WO2011 / 066360A1), event DAS-68416-4 (soybean, herbicide tolerance, A TCC accession number PTA-10442, WO2011 / 066384A1), event DP-040416-8 (maize, insect control, ATCC accession number PTA-11508, WO2011 / 075593A1), event DP-043A47-3 (maize, insect control, ATCC accession number PTA-11509, WO2011 / 075595A1), event DP-004114-3 (Maize, Insect Control, ATCC Registry No. PTA-11506, WO2011 / 084621A1), Event DP-032316-8 (Maize, Insect Control, ATCC Registry No. PTA-11507, WO2011 / 084632A1), Event MON-88302-9 (European rapeseed, Herbicide Tolerance, ATCC Registry No. PTA-10955, WO2011 / 153186A1)Event DAS-21606-3 (soybean, herbicide tolerance, ATCC Registry No. PTA-11028, WO2012 / 033794A2), Event MON-87712-4 (soybean, quality traits, ATCC Registry No. PTA-10296, WO2012 / 051199A2), Event DAS-44406-6 (soybean, superimposed herbicide tolerance, ATCC Registry No. PTA-11336, WO2012 / 075426A1), Event DAS-14536-7 (soybean, herbicide tolerance), Soybean, superimposed herbicide tolerance, ATCC accession number PTA-11335, WO2012 / 075429A1), event SYN-000H2-5 (soybean, herbicide tolerance, ATCC accession number PTA-11226, WO2012 / 082548A2), event DP-061061-7 (European rapeseed, herbicide tolerance, no accession number available, WO2012071039A1), event DP-073496-4 (European rapeseed, herbicide tolerance, no accession number available, US201 Event 2131692), Event 8264.44.06.1 (Soybean, superimposed herbicide tolerance, Registry No. PTA-11336, WO2012075426A2), Event 8291.45.36.2 (Soybean, superimposed herbicide tolerance, Registry No. PTA-11335, WO2012075429A2), Event SYHT0H2 (Soybean, ATCC Registry No. PTA-11226, WO2012 / 082548A2), Event MON88701 (Cotton, ATC Event C (ATCC accession number PTA-11754, WO2012 / 134808Al), Event KK179-2 (alfalfa, ATCC accession number PTA-11833, WO2013 / 003558Al), Event pDAB8264.42.32.1 (soybean, superimposed herbicide tolerance, ATCC accession number PTA-11993, WO2013 / 010094Al), Event MZDT09Y (maize, ATCC accession number PTA-13025, WO2013 / 012775Al).
[0298] Genes / events conferring the desired trait (e.g., target polynucleotides) can also be combined with each other in transgenic plants. Examples of transgenic plants that can be mentioned are important crop plants such as cereals (wheat, rice, triticale, barley, rye, oats), corn, soybeans, potatoes, sugar beets, sugarcane, tomatoes, peas and other types of vegetables, cotton, tobacco, canola, and fruit plants (such as apples, pears, citrus fruits and grapes), with particular emphasis on corn, soybeans, wheat, rice, potatoes, cotton, sugarcane, tobacco, and canola. The traits particularly emphasized are enhanced resistance of plants to insects, arachnids, nematodes, slugs, and snails, and enhanced resistance of plants to one or more herbicides.
[0299] Commercially available examples of such plants, plant parts or plant seeds that can be preferentially processed according to the present invention include those that are... RIB ROUNDUP VT DOUBLE VT TRIPLE BOLLGARD ROUNDUP READY 2 ROUNDUP 2 XTEN DTM INTACTA RR2 VISTIVE and / or XTENDFLEX TM Product name refers to goods sold or distributed, such as plant seeds.
[0300] The invention will now be described with reference to the following embodiments. It should be understood that these embodiments are not intended to limit the scope of the claims of the invention, but are intended to be examples of certain implementations. Any variations of the exemplary methods that may occur to those skilled in the art fall within the scope of the invention.
[0301] Example
[0302] Example 1. Design of genome editing constructs for HB53 and HB78
[0303] The genome sequences of Zm00001d002754(HB53) and Zm00001d029934(HB78) (maize) were identified from a proprietary maize line. From this reference line, the spacer region sequence SEQ ID NO:175-182 was identified. The edited constructs (pWISE444 targeting HB53 and pWISE451 targeting HB53 / HB78) contain CRISPR-Cas effector proteins (…). Figure 13 pWISE451 includes all guide nucleic acids from the pWISE shown. CRISPR-Cas effector proteins associate with spacer region nucleic acids specific to the DNA sequence of the maize HB53 / HB78 transcription factor. The DNA-binding domain of the gene target is uniquely disrupted using the Cpf1 cleavage enzyme.
[0304] Example 2. Transformation and selection of edited E0 plants
[0305] Dry, isolated maize embryos were transformed using *Agrobacterium* to deliver the edited construct. Healthy, non-chimeric plants (E0) were selected and inserted into growth trays. Genotyping of the E0 plants was performed to assess transgene copy number and editing efficacy. Plants identified as (1) healthy, non-chimeric, and fertile, with (2) low transgene copy number and (3) disrupted DNA-binding domains were advanced to the next generation. E0 plants meeting all the above criteria were self-pollinated to produce the E1 generation.
[0306] For pWISE444, 110 E0 plants were derived from a single transformation experiment. From this E0 plant library, we identified plants with out-of-frame deletions that resulted in disruption of the HB53 DNA-binding domain. Table 1 provides the plant identifiers, identified deletions, and the final progressive E1 alleles. Figure 10 The genotypes of the E1 offspring from the progressive E0 parents are shown.
[0307] Table 1. Parental plants of HB53 E0 and identified alleles
[0308] E0 parent E0 allele E1 allele (progressive) CE-9775 Hybrid (5bp and 7bp missing) homozygous 5bp deletion CE-9787 Heterogeneous (11bp and 14bp deletions) homozygous 11bp deletion CE-9778 Het-WT (23bp missing) homozygous 23bp deletion
[0309] For pWISE451, a single transformation experiment yielded 44 E0 events. Among them, single E0 plants with the desired out-of-frame deletions were improved. Figure 11 and Figure 12 The editing properties of the HB53 and HB78 E1 descendants are shown respectively.
[0310] Apart from pWISE444 and pWISE451, other transformation experiments produced the required edits on HB78. Figure 9 The genome sequences of representative edited plants are shown. This indicates that a 10 bp out-of-frame deletion upstream of the target DNA-binding domain resulted in an early stop codon, thus causing the loss of the DNA-binding domain.
[0311] Example 3. Phenotypic assessment of trait activity
[0312] For each E1 family, 100 seeds were planted and screened. Plants identified as (1) healthy, non-chimeric, and fertile, (2) non-GMO, and (3) possessing disrupted DNA-binding domains were advanced to the next generation. Trait activity was typically assessed at the E2 stage. To assess the activity of the desired progressively edited traits, 10-day-old seedlings were exposed to a simulated shaded environment. A custom-designed lighting array consisted of 12 CREE 6500k 36V COBs supplemented with over 330 2.25VCREE far-infrared LEDs (720-740) with Bluefish controllers. Other similar lighting was used to simulate the shaded environment.
[0313] Edited maize seedlings and controls (wild type and GUS control) maize seedlings were grown in a growing greenhouse with a simulated shade environment, wherein the temperature was 400 μmol / m³. -2 s -1 Under PAR (photosynthetically active radiation), they experienced a red-to-far-red wavelength ratio of 0.15. Edited maize seedlings and control maize seedlings were also grown in separate growing sheds, where they experienced a red-to-far-red wavelength ratio of 1.3 and 400 μmol m... -2 s -1 PAR (No Shading). Plant height was measured at three height markers (coleoptile, V1 sheath, V2 sheath), and control plants were compared with edited plants in simulated shade and unshading environments. An editing event was considered trait-active when the edited plant grown in simulated shade was 5%, 10%, 15% (or more) shorter than the control plant (wild type).
[0314] Example 4. Modification of endogenous maize homologous domain-leucine zipper (HD-Zip) transcription factor inhibits shade avoidance response.
[0315] Editing of HD-Zip transcription factors suppressed excessive stem elongation under simulated shading conditions. Homozygous editing of HD-Zip transcription factors HB53 and HB53 / HB78 showed a statistically significant suppressed shade-avoidance response, as demonstrated in stage V2. Figure 14 ), but no significant heterogeneity ( Figure 15 Conversely, wild-type plants (01DKD2 control) and plants transformed with vectors without guide nucleic acids (GUS control) showed statistically significant shade avoidance responses. In addition to the edited alleles generated from pWISE444 (HB53) and pWISE451 (HB53 / HB78), homozygous HB78 edits were also tested. Figure 8 and Figure 9Our results indicate that trait activity (i.e., no shade avoidance response) was present at E1 stage, but there was no significant difference in plant height between unedited and edited HB78 at E2 stage. Further testing is needed because the two tests (E1 stage vs. E2 stage) differed in growth chamber location.
[0316] In summary, our results demonstrate that HD-Zip transcription factors are highly efficient regulators of the shade avoidance response. Furthermore, we have shown that editing these transcription factors leads to a repression of the evolutionarily conserved over-response to shading.
[0317] The foregoing is a description of the present invention and should not be construed as limiting it. The present invention is defined by the following claims, wherein equivalents of the claims are included. sequence list <110> Paired Plant Services Co., Ltd. <120> Inhibition of plant shade avoidance response <130> 1499.17.WO <150> US 62 / 968,596 <151> 2020-01-31 <160> 329 <170> PatentIn version 3.5 <210> 1 <211> 116 <212> PRT <213> Corn (Zea mays) <400> 1 Arg Lys Lys Leu Arg Leu Ser Lys Asp Gln Ser Ala Val Leu Glu Asp 1 5 10 15 Ser Phe Arg Glu His Pro Thr Leu Asn Pro Arg Gln Lys Ala Ala Leu 20 25 30 Ala Gln Gln Leu Gly Leu Arg Pro Arg Gln Val Glu Val Trp Phe Gln 35 40 45 Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys Glu 50 55 60 Tyr Leu Lys Arg Cys Cys Glu Thr Leu Thr Glu Glu Asn Arg Arg Leu 65 70 75 80 Gln Lys Glu Val Gln Glu Leu Arg Ala Leu Lys Leu Val Ser Pro His 85 90 95 Leu Tyr Met His Met Ser Pro Pro Thr Thr Leu Thr Met Cys Pro Ser 100 105 110 Cys Glu Arg Val 115 <210> 2 <211> 116 <212> PRT <213> Maize <400> 2 Arg Lys Lys Leu Arg Leu Ser Lys Asp Gln Ala Ala Val Leu Glu Glu 1 5 10 15 Ser Phe Lys Glu His Asn Thr Leu Asn Pro Lys Gln Lys Ala Ala Leu 20 25 30 Ala Lys Gln Leu Asn Leu Lys Pro Arg Gln Val Glu Val Trp Phe Gln 35 40 45 Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys Glu 50 55 60 Phe Leu Lys Arg Cys Cys Glu Thr Leu Thr Glu Glu Asn Arg Arg Leu 65 70 75 80 Gln Arg Glu Val Ala Glu Leu Arg Val Leu Lys Leu Val Ala Pro His 85 90 95 His Tyr Ala Arg Met Pro Pro Pro Thr Thr Leu Thr Met Cys Pro Ser 100 105 110 Cys Glu Arg Leu 115 <210> 3 <211> 53 <212> PRT <213> Artificial sequence <220> <223> Synthetic HD-zip transcription factor peptide <400> 3 Leu Ala Lys Gln Leu Asn Leu Lys Pro Arg Gln Val Glu Val Trp Phe 1 5 10 15 Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys 20 25 30 Glu Phe Leu Lys Arg Cys Cys Glu Thr Leu Thr Glu Glu Asn Arg Arg 35 40 45 Leu Gln Arg Glu Val 50 <210> 4 <211> twenty four <212> PRT <213> Artificial sequence <220> <223> Synthetic HD-zip transcription factor peptide <400> 4 Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Leu 1 5 10 15 Lys Gln Thr Glu Val Asp Cys Glu 20 <210> 5 <211> twenty four <212> PRT <213> Artificial sequence <220> <223> Synthetic HD-zip transcription factor peptide <220> <221> MISC_FEATURE <222> (16)..(16) <223> Xaa represents Leu or Ser <400> 5 Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Xaa 1 5 10 15 Lys Gln Thr Glu Val Asp Cys Glu 20 <210> 6 <211> 11 <212> PRT <213> Artificial sequence <220> <223> Synthetic HD-zip transcription factor peptide <220> <221> MISC_FEATURE <222> (7)..(7) <223> Xaa represents Ser or Thr <220> <221> MISC_FEATURE <222> (9)..(9) <223> Xaa represents Asp or Glu. <220> <221> MISC_FEATURE <222> (11)..(11) <223> Xaa represents Ser or Ala <400> 6 Arg Lys Lys Leu Arg Leu Xaa Lys Xaa Gln Xaa 1 5 10 <210> 7 <211> 9 <212> PRT <213> Artificial sequence <220> <223> Synthetic HD-zip transcription factor peptide <220> <221> MISC_FEATURE <222> (6)..(6) <223> Xaa represents Gln or His <220> <221> MISC_FEATURE <222> (7)..(7) <223> Xaa represents Arg or Lys <220> <221> MISC_FEATURE <222> (9)..(9) <223> Xaa represents Val or Leu <400> 7 Glu Asn Arg Arg Leu Xaa Xaa Glu Xaa 1 5 <210> 8 <211> 13 <212> PRT <213> Artificial sequence <220> <223> Synthetic HD-zip transcription factor peptide <220> <221> MISC_FEATURE <222> (2)..(2) <223> Xaa represents Pro or Ala <220> <221> MISC_FEATURE <222> (3)..(4) <223> Xaa represents Thr or Ala <220> <221> MISC_FEATURE <222> (7)..(7) <223> Xaa represents Val or Met <220> <221> MISC_FEATURE <222> (10)..(10) <223> Xaa represents Gln, Ser, or Asn. <400> 8 Pro Xaa Xaa Xaa Leu Thr Xaa Cys Pro Xaa Cys Glu Arg 1 5 10 <210> 9 <211> 8 <212> PRT <213> Artificial sequence <220> <223> Synthetic DNA-binding domain peptides <400> 9 Val Trp Phe Gln Asn Arg Arg Ala 1 5 <210> 10 <211> 287 <212> PRT <213> Sweet orange (Citrus sinensis) <400> 10 Met Gly Glu Lys Asp Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly Cys 1 5 10 15 Ala Ala Arg Asn Glu Pro Ser Leu Arg Leu Asn His Met Pro Leu Ser 20 25 30 Ser Ser Gln Ser Met Gln Asn His His Lys Arg Ser Pro Trp Thr Glu 35 40 45 Leu Phe His Ser Ser Asp Arg Asn Ser Asp Thr Arg Ser Phe Leu Arg 50 55 60 Gly Ile Asp Val Asn Gln Ala Pro Thr Val Ala Asp Cys Glu Glu Glu 65 70 75 80 Asn Gly Val Ser Ser Pro Asn Ser Thr Val Ser Ser Ile Ser Gly Lys 85 90 95 Arg Ser Glu Arg Glu Pro Ile Gly Asp Glu Thr Glu Ala Glu Arg Ala 100 105 110 Ser Cys Ser Arg Gly Ser Asp Asp Glu Asp Gly Gly Ala Gly Asp Ala 115 120 125 Ser Arg Lys Lys Leu Arg Leu Ser Lys Glu Gln Ser Leu Leu Leu Glu 130 135 140 Glu Thr Phe Lys Glu His Ser Thr Leu Asn Pro Lys Gln Lys Leu Ala 145 150 155 160 Leu Ala Lys Gln Leu Asn Leu Arg Pro Arg Gln Val Glu Val Trp Phe 165 170 175 Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys 180 185 190 Glu Tyr Leu Lys Arg Cys Cys Glu Asn Leu Thr Glu Glu Asn Arg Arg 195 200 205 Leu Gln Lys Glu Val Gln Glu Leu Arg Ser Leu Lys Leu Ser Pro Gln 210 215 220 Leu Tyr Met Asn Met Asn Pro Pro Thr Thr Leu Thr Met Cys Pro Ser 225 230 235 240 Cys Glu Arg Val Ala Val Ser Ser Ser Ser Ser Ser Ser Ser Ala Ala 245 250 255 Ala Asn Gly Thr Thr Arg Leu Pro Ile Gly Pro Asn His Gln Arg Leu 260 265 270 Thr Pro Val Ser Pro Trp Ala Ala Leu Pro Ile His His Arg Ser 275 280 285 <210> 11 <211> 294 <212> PRT <213> Theobroma cacao <400> 11 Met Gly Ala Glu Lys Asp Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly 1 5 10 15 Cys Ala Gln Asn His Pro Ser Leu Lys Leu Asn Leu Met Pro Leu Ala 20 25 30 Ser Pro Arg Met Gln Asn Leu Gln Gln Lys Asn Thr Trp Asn Glu Leu 35 40 45 Phe Gln Ser Ser Asp Arg Asn Leu Asp Thr Arg Ser Phe Leu Arg Gly 50 55 60 Ile Asp Val Asn Arg Ala Pro Ala Thr Val Asp Cys Glu Glu Glu Gly 65 70 75 80 Gly Val Ser Ser Pro Asn Ser Thr Ile Ser Ser Ile Ser Gly Lys Arg 85 90 95 Asn Glu Arg Asp Pro Val Gly Asp Glu Thr Glu Ala Glu Arg Ala Ser 100 105 110 Cys Ser Arg Ala Ser Asp Asp Glu Asp Gly Gly Ala Gly Gly Asp Ala 115 120 125 Ser Arg Lys Lys Leu Arg Leu Ser Lys Glu Gln Ser Leu Leu Leu Glu 130 135 140 Glu Thr Phe Lys Glu His Ser Thr Leu Asn Pro Lys Gln Lys Leu Ala 145 150 155 160 Leu Ala Lys Gln Leu Asn Leu Arg Pro Arg Gln Val Glu Val Trp Phe 165 170 175 Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys 180 185 190 Glu Tyr Leu Lys Arg Cys Cys Glu Asn Leu Thr Glu Glu Asn Arg Arg 195 200 205 Leu Gln Lys Glu Val Gln Glu Leu Arg Ala Leu Lys Leu Ser Pro Gln 210 215 220 Leu Tyr Met His Met Asn Pro Pro Thr Thr Leu Thr Met Cys Pro Ser 225 230 235 240 Cys Glu Arg Val Ala Val Ser Ser Ser Ser Ser Ser Ala Ala Ala Thr 245 250 255 Ala Ser Ser Thr Pro Thr Ser Thr Val Pro Asn Arg His His Arg Thr 260 265 270 Ser Ser Val Ser Pro Trp Ala Ala Met Pro Ile Gly His Arg Pro Phe 275 280 285 His Ala Pro Ala Ser Arg 290 <210> 12 <211> 292 <212> PRT <213> Muskmelon (Cucumis melo) <400> 12 Met Gly Gly Arg Asp Asp Asp Leu Gly Leu Thr Leu Ser Leu Gly Phe 1 5 10 15 Gly Val Thr Thr Gln Pro Thr His Met Gln Arg Pro Ser Met His Asn 20 25 30 His Leu Arg Lys Thr Ser Trp Asn Glu Leu Phe Gln Phe Ser Asp Arg 35 40 45 Asn Ala Asp Ser Arg Ser Phe Leu Arg Gly Ile Asp Val Asn Arg Leu 50 55 60 Pro Thr Gly Val Asp Gly Glu Glu Glu Asn Gly Val Ser Ser Pro Asn 65 70 75 80 Ser Thr Ile Ser Ser Ile Ser Gly Lys Arg Ser Glu Arg Glu Ala Ala 85 90 95 Gly Asp Glu Ala Glu Ala Glu Ala Glu Ala Glu Ala Glu Ala Glu Ala 100 105 110 Glu Ala Glu Ala Glu Arg Ala Ser Cys Ser Arg Gly Ser Asp Asp Glu 115 120 125 Asp Gly Gly Gly Gly Asp Gly Asp Ala Ser Arg Lys Lys Leu Arg Leu 130 135 140 Ser Lys Glu Gln Ser Met Val Leu Glu Glu Thr Phe Lys Glu His Asn 145 150 155 160 Thr Leu Asn Pro Lys Gln Lys Leu Ala Leu Ala Lys Gln Leu Asn Leu 165 170 175 Thr Pro Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr 180 185 190 Lys Leu Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg Cys Cys 195 200 205 Glu Asn Leu Thr Glu Glu Asn Arg Arg Leu Gln Lys Glu Val Gln Glu 210 215 220 Leu Arg Ala Leu Lys Leu Ser Pro Gln Leu Tyr Met His Met Asn Pro 225 230 235 240 Pro Thr Thr Leu Thr Met Cys Pro Gln Cys Glu Arg Val Ala Val Ser 245 250 255 Ser Ser Ser Ser Thr Ser Ala Ala Thr Thr Thr Arg His Pro Ala Ala 260 265 270 Ala Gly Val Gln Arg Thr Ser Met Ala Ile Asn Pro Trp Ala Val Leu 275 280 285 Pro Ile Gln Arg 290 <210> 13 <211> 294 <212> PRT <213> Cucumber (Cucumis sativus) <400> 13 Met Gly Gly Arg Asp Asp Asp Val Gly Leu Thr Leu Ser Leu Gly Phe 1 5 10 15 Gly Val Thr Thr Gln Ser Thr His Met Gln Arg Pro Ser Ser Met His 20 25 30 Asn His His Leu Arg Lys Thr His Trp Asn Glu Leu Phe Gln Phe Ser 35 40 45 Asp Arg Asn Ala Asp Ser Arg Ser Phe Leu Arg Gly Ile Asp Val Asn 50 55 60 Arg Leu Pro Thr Gly Val Asp Gly Glu Glu Glu Asn Gly Val Ser Ser 65 70 75 80 Pro Asn Ser Thr Ile Ser Ser Ile Ser Gly Lys Arg Ser Glu Arg Glu 85 90 95 Ala Ala Gly Asp Glu Ala Glu Ala Glu Ala Glu Ala Glu Ala Glu Ala 100 105 110 Glu Ala Glu Ala Glu Ala Glu Arg Ala Ser Cys Ser Arg Gly Ser Asp 115 120 125 Asp Glu Asp Gly Gly Gly Gly Asp Gly Asp Ala Ser Arg Lys Lys Leu 130 135 140 Arg Leu Ser Lys Glu Gln Ser Met Val Leu Glu Glu Thr Phe Lys Glu 145 150 155 160 His Asn Thr Leu Asn Pro Lys Gln Lys Leu Ala Leu Ala Lys Gln Leu 165 170 175 Asn Leu Thr Pro Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala 180 185 190 Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg 195 200 205 Cys Cys Glu Asn Leu Thr Glu Glu Asn Arg Arg Leu Gln Lys Glu Val 210 215 220 Gln Glu Leu Arg Ala Leu Lys Leu Ser Pro Gln Leu Tyr Met His Met 225 230 235 240 Asn Pro Pro Thr Thr Leu Thr Met Cys Pro Gln Cys Glu Arg Val Ala 245 250 255 Val Ser Ser Ser Ser Ser Thr Ser Ala Ala Thr Thr Thr Arg His Gln 260 265 270 Ala Ala Ala Gly Val Gln Arg Pro Ser Met Ala Ile Asn Pro Trp Ala 275 280 285 Val Leu Pro Ile Gln Arg 290 <210> 14 <211> 298 <212> PRT <213> Populus trichocarpa <400> 14 Met Gly Asp Lys Asn Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly Phe 1 5 10 15 Asp Ala Thr Gln Gln Asn His Gln Gln Gln Pro Ser Leu Lys Leu Asn 20 25 30 Leu Met Pro Val Pro Ser Gln Asn Asn His Arg Lys Thr Ser Leu Thr 35 40 45 Asp Leu Phe Gln Ser Ser Asp Arg Ala Cys Gly Thr Arg Phe Phe Gln 50 55 60 Arg Gly Ile Asp Met Asn Arg Val Pro Ala Ala Val Thr Asp Cys Asp 65 70 75 80 Asp Glu Thr Gly Val Ser Ser Pro Asn Ser Thr Leu Ser Ser Leu Ser 85 90 95 Gly Lys Arg Ser Glu Arg Glu Gln Ile Gly Glu Glu Thr Glu Ala Glu 100 105 110 Arg Ala Ser Cys Ser Arg Asp Ser Asp Asp Glu Asp Gly Ala Gly Gly 115 120 125 Asp Ala Ser Arg Lys Lys Leu Arg Leu Ser Lys Glu Gln Ser Leu Val 130 135 140 Leu Glu Glu Thr Phe Lys Glu His Asn Thr Leu Asn Pro Lys Glu Lys 145 150 155 160 Leu Ala Leu Ala Lys Gln Leu Asn Leu Arg Pro Arg Gln Val Glu Val 165 170 175 Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val 180 185 190 Asp Cys Glu Tyr Leu Lys Arg Cys Cys Glu Asn Leu Thr Glu Glu Asn 195 200 205 Arg Arg Leu Gln Lys Glu Val Gln Glu Leu Arg Ala Leu Lys Leu Ser 210 215 220 Pro Gln Leu Tyr Met His Met Asn Pro Pro Thr Thr Leu Thr Met Cys 225 230 235 240 Pro Ser Cys Glu Arg Val Ala Val Ser Ser Ala Ser Ser Ser Ser Ala 245 250 255 Ala Ala Ala Ser Ser Ala Leu Ala Pro Thr Ala Ser Thr Arg Gln Pro 260 265 270 Gln Arg Pro Val Pro Ile Asn Pro Trp Ala Thr Met Pro Val His Gln 275 280 285 Arg Thr Phe Asp Ala Pro Ala Ser Arg Ser 290 295 <210> 15 <211> 318 <212> PRT <213> Arabidopsis thaliana <400> 15 Met Gly Glu Arg Asp Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly Asn 1 5 10 15 Ser Gln Gln Lys Glu Pro Ser Leu Arg Leu Asn Leu Met Pro Leu Thr 20 25 30 Thr Ser Ser Ser Ser Ser Ser Phe Gln His Met His Asn Gln Asn Asn 35 40 45 Asn Ser His Pro Gln Lys Ile His Asn Ile Ser Trp Thr His Leu Phe 50 55 60 Gln Ser Ser Gly Ile Lys Arg Thr Thr Ala Glu Arg Asn Ser Asp Ala 65 70 75 80 Gly Ser Phe Leu Arg Gly Phe Asn Val Asn Arg Ala Gln Ser Ser Val 85 90 95 Ala Val Val Asp Leu Glu Glu Glu Ala Ala Val Val Ser Ser Pro Asn 100 105 110 Ser Ala Val Ser Ser Leu Ser Gly Asn Lys Arg Asp Leu Ala Val Ala 115 120 125 Arg Gly Gly Asp Glu Asn Glu Ala Glu Arg Ala Ser Cys Ser Arg Gly 130 135 140 Gly Gly Ser Gly Gly Ser Asp Asp Glu Asp Gly Gly Asn Gly Asp Gly 145 150 155 160 Ser Arg Lys Lys Leu Arg Leu Ser Lys Asp Gln Ala Leu Val Leu Glu 165 170 175 Glu Thr Phe Lys Glu His Ser Thr Leu Asn Pro Lys Gln Lys Leu Ala 180 185 190 Leu Ala Lys Gln Leu Asn Leu Arg Ala Arg Gln Val Glu Val Trp Phe 195 200 205 Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys 210 215 220 Glu Tyr Leu Lys Arg Cys Cys Asp Asn Leu Thr Glu Glu Asn Arg Arg 225 230 235 240 Leu Gln Lys Glu Val Ser Glu Leu Arg Ala Leu Lys Leu Ser Pro His 245 250 255 Leu Tyr Met His Met Thr Pro Pro Thr Thr Leu Thr Met Cys Pro Ser 260 265 270 Cys Glu Arg Val Ser Ser Ser Ala Ala Thr Val Thr Ala Ala Pro Ser 275 280 285 Thr Thr Thr Thr Pro Thr Val Val Gly Arg Pro Ser Pro Gln Arg Leu 290 295 300 Thr Pro Trp Thr Ala Ile Ser Leu Gln Gln Lys Ser Gly Arg 305 310 315 <210> 16 <211> 321 <212> PRT <213> Brassica campestris <400> 16 Met Gly Glu Ser Glu Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly Leu 1 5 10 15 Ser Gln Leu Lys Glu Pro Ser Leu Gly Leu Gly Leu Asn Leu Leu Pro 20 25 30 Leu Arg Thr Ser Ser Ser Ser Phe Ser His Met His Asn His Asn Asn 35 40 45 Asn His Leu Gln Lys Lys Ile Asn His Asn Ser Trp Pro His Gln Phe 50 55 60 His Ser Ser Glu Arg Asn Ser Asp Val Gly Ser Leu Leu Arg Gly Leu 65 70 75 80 Glu Val Asn Arg Thr Pro Ser Ala Thr Val Val Ile Asn Leu Glu Glu 85 90 95 Asp Leu Ala Gly Val Ser Ser Pro Asn Ser Asn Ile Ser Ser Val Ser 100 105 110 Gly Asn Lys Arg Asp Leu Ala Ala Ala Arg Gly Asp Gly Gly Gly Asp 115 120 125 Glu Asn Glu Ala Glu Arg Ala Ser Cys Ser His Gly Gly Gly Ser Asp 130 135 140 Glu Glu Glu Gly Gly Asn Cys Glu Gly Thr Arg Lys Lys Leu Arg Leu 145 150 155 160 Ser Lys Glu Gln Ala Leu Val Leu Glu Asp Thr Phe Lys Glu His Ser 165 170 175 Thr Leu Asn Pro Lys Gln Lys Leu Ala Leu Ala Lys Gln Leu Asn Leu 180 185 190 Arg Thr Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr 195 200 205 Lys Leu Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg Cys Cys 210 215 220 Asp Thr Leu Thr Glu Glu Asn Arg Arg Leu His Lys Glu Val Ala Glu 225 230 235 240 Leu Arg Ala Leu Lys Leu Ser Pro His Leu Tyr Met His Met Thr Pro 245 250 255 Pro Thr Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val Ser Ala Ser 260 265 270 Ser Ser Ser Ser Ala Met Ala Ala Ala Ala Pro Pro Ser Ser Ile Thr 275 280 285 Ser Gly Gly Gly Gly Arg Ile Pro Thr Val Val Gly Arg Pro Ser Pro 290 295 300 Gln Arg Pro Thr Pro Cys Ala Ala Ile Ser Leu Gln Ser Arg Leu Ala 305 310 315 320 His <210> 17 <211> 318 <212> PRT <213> Brassica oleracea <400> 17 Met Gly Glu Ser Glu Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly Leu 1 5 10 15 Ser Gln Leu Lys Glu Pro Ser Leu Gly Leu Gly Leu Asn Leu Leu Pro 20 25 30 Leu Arg Thr Ser Ser Phe Ser His Met His Asn His Asn Asn Asn His 35 40 45 Leu Gln Lys Lys Ile Tyr His Asn Ser Trp Pro His Gln Phe Gln Ser 50 55 60 Ser Glu Arg Asn Ser Asp Val Gly Ser Leu Leu Arg Gly Leu Glu Val 65 70 75 80 Asn Arg Thr Pro Ser Ala Thr Val Val Ile Asn Leu Glu Glu Asp Ala 85 90 95 Ala Gly Val Ser Ser Pro Asn Ser Asn Val Ser Ser Val Ser Gly Asn 100 105 110 Lys Arg Asp Leu Ala Ala Ala Arg Gly Asp Gly Gly Gly Asp Glu Asn 115 120 125 Glu Ala Glu Arg Ala Ser Cys Ser His Arg Gly Gly Ser Asp Glu Glu 130 135 140 Glu Gly Gly Asn Cys Glu Gly Thr Arg Lys Lys Leu Arg Leu Ser Lys 145 150 155 160 Glu Gln Ala Leu Val Leu Glu Glu Thr Phe Lys Glu His Ser Thr Leu 165 170 175 Asn Pro Lys Gln Lys Leu Ala Leu Ala Lys Gln Leu Asn Leu Trp Thr 180 185 190 Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Leu 195 200 205 Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg Cys Cys Asp Thr 210 215 220 Leu Thr Glu Glu Asn Arg Arg Leu His Lys Glu Val Ser Glu Leu Arg 225 230 235 240 Ala Leu Lys Leu Ser Pro His Leu Tyr Met His Met Thr Pro Pro Thr 245 250 255 Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val Ser Ala Ser Ser Ser 260 265 270 Ser Ser Ala Met Ala Ala Ala Ala Pro Pro Ser Ser Thr Ala Ser Gly 275 280 285 Gly Gly Arg Ile Pro Thr Val Val Gly Arg Pro Ser Pro Gln Arg Pro 290 295 300 Thr Pro Cys Ala Ala Ile Ser Leu Gln Ser Arg Leu Ala His 305 310 315 <210> 18 <211> 319 <212> PRT <213> Peach (Prunus persica) <400> 18 Met Met Val Glu Arg Asp Gln Asp Leu Gly Leu Ser Leu Ser Leu Ser 1 5 10 15 Phe Pro Gln Thr His Asn His His Asn Asn Asn Asn Asn Asn Ser Ser 20 25 30 Ser Thr Thr Ser Thr Leu Gln Leu Asn Leu Met Pro Ser Leu Ala Pro 35 40 45 Thr Ser Ala Ser Ser Pro Ser Gly Phe Leu Pro Gln Lys Pro Ser Trp 50 55 60 Asn Glu Ala Leu Ile Ser Ser Asp Arg Asn Ser Asn Ser Glu Thr Phe 65 70 75 80 Arg Val Gly Pro Arg Ser Phe Leu Arg Gly Ile Asp Val Asn Arg Leu 85 90 95 Pro Ser Thr Gly Asp Cys Glu Asp Glu Ala Gly Val Ser Ser Pro Asn 100 105 110 Ser Thr Val Ser Ser Val Ser Gly Lys Arg Ser Glu Arg Glu Ala Asn 115 120 125 Gly Glu Asp Leu Asp Ile Glu Thr Arg Gly Ile Ser Asp Glu Glu Asp 130 135 140 Gly Glu Thr Ser Arg Lys Lys Leu Arg Leu Ser Lys Asp Gln Ser Ala 145 150 155 160 Ile Leu Glu Glu Ser Phe Lys Glu His Asn Thr Leu Asn Pro Lys Gln 165 170 175 Lys Leu Ala Leu Ala Lys Gln Leu Gly Leu Arg Pro Arg Gln Val Glu 180 185 190 Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu 195 200 205 Val Asp Cys Glu Phe Leu Lys Arg Cys Cys Glu Asn Leu Thr Glu Glu 210 215 220 Asn Arg Arg Leu Gln Lys Glu Val Gln Glu Leu Arg Ala Leu Lys Leu 225 230 235 240 Ser Pro Gln Phe Tyr Met Gln Met Thr Pro Pro Thr Thr Leu Thr Met 245 250 255 Cys Pro Ser Cys Glu Arg Val Ala Val Pro Pro Asn Ser Ser Ser Ser 260 265 270 Thr Val Glu Pro Arg Pro His Pro His Pro His Pro Gln Met Gly Ser 275 280 285 Val Gln Thr Arg Pro Val Pro Ile Asn Pro Trp Ala Ser Ala Thr Pro 290 295 300 Ile Pro His Arg Pro Leu Pro Phe Glu Ala Phe His Thr Arg Thr 305 310 315 <210> 19 <211> 299 <212> PRT <213> Chickpea (Cicer arietinum) <400> 19 Met Gly Asp Lys Glu Asp Glu Leu Gly Leu Gly Leu Ser Leu Ser Leu 1 5 10 15 Ser Leu Gly Tyr Gly Ala Asn Ala Asn Asn Ala Pro Leu Lys Val Thr 20 25 30 His Met His Lys Pro Pro Gln Ser Val Pro Asn Gln Arg Val Ser Phe 35 40 45 Asn Asn Phe Phe His Phe His Asp Leu Ser Ser Glu Thr Arg Ser Phe 50 55 60 Ile Gly Gly Ile Asp Val Asn Ser Pro Ala Thr Ala Ala Cys Asp Asp 65 70 75 80 Glu Asn Gly Gly Ser Ser Pro Asn Ser Thr Val Ser Ser Ile Ser Gly 85 90 95 Lys Arg Ser Glu Arg Glu Gly Asn Gly Glu Glu Asn Asp Ala Val Glu 100 105 110 Arg Ala Ser Cys Ser Arg Gly Gly Ser Asp Asp Asp Asp Gly Gly Gly 115 120 125 Cys Gly Gly Asp Gly Asp Ser Ser Arg Lys Lys Leu Arg Leu Ser Lys 130 135 140 Glu Gln Ser Val Leu Leu Glu Glu Thr Phe Lys Glu His Asn Thr Leu 145 150 155 160 Asn Pro Lys Gln Lys Gln Ala Leu Ala Lys Gln Leu Asn Leu Met Pro 165 170 175 Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Leu 180 185 190 Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg Cys Cys Glu Thr 195 200 205 Leu Thr Glu Glu Asn Arg Arg Leu Gln Lys Glu Val Gln Glu Leu Arg 210 215 220 Ala Leu Lys Leu Ser Pro Gln Leu Tyr Met His Met Asn Pro Pro Thr 225 230 235 240 Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val Ala Val Ser Ser Ala 245 250 255 Ser Ser Ser Ser Ala Asn Val Pro Ser Ala Leo Ala Pro Ala Asn Arg 260 265 270 Asn Ser Ile Gly Pro Ser Val Gln Arg Pro Val Pro Leu Asn Pro Trp 275 280 285 Ala Ala Met Ser Ile Gln Asn Arg Ser Arg Pro 290 295 <210> 20 <211> 317 <212> PRT <213> Soybean (Glycine max) <400> 20 Met Gly Glu Lys Asp Asp Gly Leu Gly Leu Gly Leu Ser Leu Lys Leu 1 5 10 15 Gly Trp Gly Glu Asn Asn Asp Asn Asn Asn Asn Gln Gln Gln His Pro 20 25 30 Phe Asn Val His Lys Pro Pro Gln Ser Val Pro Asn Gln Arg Val Ser 35 40 45 Val Asn Ser Leu Phe His Phe His Asp Glu Asn His Ala Met Arg Asn 50 55 60 Thr Asp Arg Ser Ser Glu Met Arg Ser Phe Phe Arg Gly Ile Asp Val 65 70 75 80 Asn Leu Pro Pro Pro Pro Pro Ser Ala Ala Leu Ala Ala Phe Asp Asp 85 90 95 Glu Asn Gly Val Ser Ser Pro Asn Ser Thr Ile Ser Ser Ile Ser Gly 100 105 110 Lys Arg Ser Glu Arg Glu Gly Asn Gly Glu Glu Asn Glu Arg Thr Ser 115 120 125 Ser Ser Arg Gly Gly Gly Gly Ser Asp Asp Asp Glu Gly Gly Ala Cys 130 135 140 Gly Gly Asp Ala Asp Ala Asp Ala Ser Arg Lys Lys Leu Arg Leu Ser 145 150 155 160 Lys Glu Gln Ala Leu Val Leu Glu Glu Thr Phe Lys Glu His Asn Thr 165 170 175 Leu Asn Pro Lys Gln Lys Gln Ala Leu Ala Lys Gln Leu Asn Leu Met 180 185 190 Pro Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys 195 200 205 Leu Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg Cys Cys Glu 210 215 220 Asn Leu Thr Glu Glu Asn Arg Arg Leu Gln Lys Glu Val Gln Glu Leu 225 230 235 240 Arg Ala Leu Lys Leu Ser Pro His Leu Tyr Met Gln Met Asn Pro Pro 245 250 255 Thr Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val Ala Val Ser Ser 260 265 270 Ala Ser Ser Ser Ser Ser Ala Thr Met Pro Ser Ala Leu Pro Pro Ala 275 280 285 Asn Leu Asn Pro Val Gly Pro Thr Ile Gln Arg Pro Met Pro Val Asn 290 295 300 Pro Trp Ala Ala Met Leu Asn Gln His Arg Gly Arg Pro 305 310 315 <210> 21 <211> 296 <212> PRT <213> Medicago truncatula <400> 21 Met Ser Ile Glu Lys Glu Asp Phe Gly Leu Ser Leu Ser Leu Ser Phe 1 5 10 15 Pro Gln Asn Pro Pro Asn Pro Gln Tyr Leu Asn Leu Met Ser Ser Ser 20 25 30 Thr His Ser Tyr Ser Pro Ser Thr Phe Asn Pro Gln Lys Pro Ser Trp 35 40 45 Asn Asp Val Phe Thr Ser Ser Asp Arg Asp Ser Glu Thr Cys Arg Ile 50 55 60 Glu Glu Arg Pro Leu Ile Leu Arg Gly Ile Asp Val Asn Arg Leu Pro 65 70 75 80 Ser Gly Ala Asp Cys Glu Glu Glu Ala Gly Val Ser Ser Pro Asn Ser 85 90 95 Thr Val Ser Ser Val Ser Gly Lys Arg Ser Glu Arg Glu Val Thr Gly 100 105 110 Glu Asp Leu Asp Met Glu Arg Asp Cys Ser Arg Gly Ile Ser Asp Glu 115 120 125 Glu Asp Ala Glu Thr Ser Arg Lys Lys Leu Arg Leu Thr Lys Asp Gln 130 135 140 Ser Ile Ile Leu Glu Glu Ser Phe Lys Glu His Asn Thr Leu Asn Pro 145 150 155 160 Lys Gln Lys Leu Ala Leu Ala Lys Gln Leu Gly Leu Arg Ala Arg Gln 165 170 175 Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Leu Lys Gln 180 185 190 Thr Glu Val Asp Cys Glu Phe Leu Lys Arg Cys Cys Glu Asn Leu Thr 195 200 205 Asp Glu Asn Arg Arg Leu Gln Lys Glu Val Gln Glu Leu Arg Ala Leu 210 215 220 Lys Leu Ser Pro Gln Phe Tyr Met Gln Met Thr Pro Pro Thr Thr Leu 225 230 235 240 Thr Met Cys Pro Ser Cys Glu Arg Val Ala Val Pro Ser Ser Ala Val 245 250 255 Asp Ala Ala Thr Arg Arg His Pro Met Ala Ser Asn His Pro Arg Thr 260 265 270 Phe Ser Val Gly Pro Trp Ala Thr Ala Ala Pro Ile Gln His Arg Thr 275 280 285 Phe Asp Thr Leu Arg Pro Arg Ser 290 295 <210> 22 <211> 305 <212> PRT <213> Phaseolus vulgaris <400> 22 Met Gly Glu Lys Asp Asp Gly Leu Gly Leu Arg Leu Ser Leu Arg Trp 1 5 10 15 Gly Glu Asn Asp Asp Asn Asn Met Asn Gln Gln His Pro Phe Asn Met 20 25 30 His Lys Pro Pro Gln Pro Val Pro Asn Gln Arg Thr Ser Phe Asn Asn 35 40 45 Leu Phe His Phe His Gly Ala Ser His Val Thr Asn Arg Asn Ser Glu 50 55 60 Pro Pro Pro Phe Phe Phe Gly Ile Asp Val Asn Leu Pro Pro Pro Pro 65 70 75 80 Thr Pro Thr Pro Ser Val Val Pro Cys Glu Glu Asp Asn Leu Val Ser 85 90 95 Ser Gln Asn Ser Ala Val Ser Ser Ile Ser Gly Lys Arg Ser Glu Arg 100 105 110 Glu Glu Asn Glu Arg Gly Ser Cys Ser His Gly Ser Glu Asp Glu Asp 115 120 125 Gly Gly Gly Phe Gly Gly Glu Gly Asp Gly Asp Met Ser Arg Lys Lys 130 135 140 Leu Arg Leu Ser Lys Glu Gln Ala Leu Val Leu Glu Glu Thr Phe Lys 145 150 155 160 Glu His Asn Thr Leu Asn Pro Lys Gln Lys Gln Ala Leu Ala Lys Gln 165 170 175 Leu Asn Leu Ser Pro Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg 180 185 190 Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys 195 200 205 Arg Cys Cys Glu Asn Leu Thr Glu Glu Asn Arg Arg Leu Gln Lys Glu 210 215 220 Val Gln Glu Leu Arg Ala Leu Lys Phe Ser Pro Gln Leu Tyr Met His 225 230 235 240 Met Asn Pro Pro Thr Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val 245 250 255 Ala Val Ser Ser Ala Ser Ser Ser Ser Ser Ala Ala Met Pro Ser Val 260 265 270 Pro Pro Pro Ala Asn His Asn Pro Leu Gly Pro Thr Ile Gln Arg Pro 275 280 285 Val Pro Val Asn Pro Trp Ala Ala Met Ser Ile Gln Arg Pro Cys Arg 290 295 300 Asn 305 <210> 23 <211> 294 <212> PRT <213> Ricinus communis <400> 23 Met Glu Asp Lys Asp Asp Gly Leu Gly Leu Gly Leu Ser Leu Ser Leu 1 5 10 15 Gly Gly Gln Glu Lys His Gln Asn Gln Pro Ser Leu Lys Leu Asn Leu 20 25 30 Met Pro Phe Pro Ser Leu Phe Met Gln Asn Thr His His Ser Thr Ser 35 40 45 Leu Asn Asp Leu Phe Gln Ser Ser Asp Arg Asn Ala Asp Thr Arg Ser 50 55 60 Phe Gln Arg Gly Ile Asp Met Asn Arg Met Pro Leu Phe Ala Asp Cys 65 70 75 80 Asp Asp Glu Asn Gly Val Ser Ser Pro Asn Ser Thr Ile Ser Ser Leu 85 90 95 Ser Gly Lys Arg Ser Glu Arg Glu Gln Ile Gly Gly Glu Glu Met Glu 100 105 110 Ala Glu Arg Ala Ser Cys Ser Arg Gly Gly Ser Asp Asp Glu Asp Gly 115 120 125 Gly Ala Gly Gly Asp Asp Gly Ser Arg Lys Lys Leu Arg Leu Ser Lys 130 135 140 Glu Gln Ser Leu Leu Leu Glu Glu Thr Phe Lys Glu His Asn Thr Leu 145 150 155 160 Asn Pro Lys Gln Lys Leu Ala Leu Ala Lys Gln Leu Asn Leu Lys Pro 165 170 175 Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala Arg Thr Lys Ser 180 185 190 Lys Gln Thr Glu Val Asp Cys Glu Tyr Leu Lys Arg Cys Cys Glu Asn 195 200 205 Leu Thr Gln Glu Asn Arg Arg Leu Gln Lys Glu Val Gln Glu Leu Arg 210 215 220 Ala Leu Lys Leu Ser Pro Gln Leu Tyr Met His Met Asn Pro Pro Thr 225 230 235 240 Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val Ala Val Ser Ser Ser 245 250 255 Ala Ala Pro Ser Arg Gln Pro Pro Asn Ser Gln Pro Gln Arg Pro Val 260 265 270 Pro Val Lys Pro Trp Ala Ala Leu Pro Ile Gln His Arg Pro Phe Asp 275 280 285 Thr Pro Ala Ser Arg Ser 290 <210> 24 <211> 293 <212> PRT <213> Nicotiana sylvestris <400> 24 Met Gly Gly Glu Lys Glu Asp Gly Leu Gly Leu Ser Leu Ser Leu Gly 1 5 10 15 Met Ser Cys Pro Gln Asn Asn Leu Lys Asn Asn Asp Phe Ile Ser Pro 20 25 30 Ser Thr Pro Pro Phe Leu Leu Pro Phe Met His Asn His Gln Ile Ser 35 40 45 Ala Glu Arg Asn Glu Glu Ala Arg Glu Phe Ile Arg Gly Glu Ile Asp 50 55 60 Met Asn Arg Pro Val Arg Met Ile Glu Ala Cys Asp Glu Glu Leu Glu 65 70 75 80 Asp Glu Ala Val Ile Met Val Ser Ser Pro Asn Asn Ser Thr Val Ser 85 90 95 Ser Val Ser Gly Lys Arg Ser His Asp Arg Glu Asp Asn Glu Gly Glu 100 105 110 Arg Pro Thr Ser Ser Leu Glu Asp Asp Gly Gly Asp Ala Ala Ala Arg 115 120 125 Lys Lys Leu Arg Leu Ser Lys Glu Gln Ala Ala Val Leu Glu Glu Thr 130 135 140 Phe Lys Glu His Asn Thr Leu Asn Pro Lys Gln Lys Leu Ala Leu Ser 145 150 155 160 Lys Gln Leu Asn Leu Arg Pro Arg Gln Val Glu Val Trp Phe Gln Asn 165 170 175 Arg Arg Ala Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys Glu Tyr 180 185 190 Leu Arg Arg Cys Cys Glu Asn Leu Thr Glu Glu Asn Arg Arg Leu Gln 195 200 205 Lys Glu Val Thr Glu Leu Arg Ala Leu Lys Leu Ser Pro Gln Met Tyr 210 215 220 Met Asn Met Thr Pro Pro Thr Thr Leu Thr Met Cys Pro Gln Cys Glu 225 230 235 240 Arg Val Ala Val Ser Ser Ser Ser Ser Ser Val Thr Ser Ala Gly Val 245 250 255 Ser Arg Ser Asn His Pro Val Gly Ala Leu His Gln Pro Pro Val Pro 260 265 270 Leu Asn Lys Pro Trp Ala Ala Ile Phe Ser Pro Lys Thr Leu Asp Asp 275 280 285 Gln Arg Thr Gln Leu 290 <210> 25 <211> 275 <212> PRT <213> Solanum tuberosum <400> 25 Met Val Glu Lys Glu Asp Leu Gly Leu Ser Leu Ser Leu Ser Phe Pro 1 5 10 15 Asp Asn Asn Asn Lys Lys Asn Thr Gln Leu Asn Leu Ser Pro Phe Asn 20 25 30 Leu Ile Gln Lys Thr Ser Trp Thr Asp Ser Leu Phe Pro Ser Ser Asp 35 40 45 Arg Asn Ile Glu Thr Cys Arg Val Glu Thr Arg Thr Phe Leu Lys Gly 50 55 60 Ile Asp Val Asn Arg Leu Pro Ala Thr Gly Glu Ala Asp Glu Glu Ala 65 70 75 80 Gly Val Ser Ser Pro Asn Ser Thr Ile Ser Ser Val Ser Gly Asn Lys 85 90 95 Arg Thr Glu Arg Glu Ala Asn Asn Cys Asp Gln Glu Glu His Glu Met 100 105 110 Glu Arg Gly Ser Asp Glu Glu Asp Gly Glu Thr Ser Arg Lys Lys Leu 115 120 125 Arg Leu Ser Lys Asp Gln Ser Ala Ile Leu Glu Glu Ser Phe Lys Glu 130 135 140 His Asn Thr Leu Asn Pro Lys Gln Lys Leu Ala Leu Ala Lys Arg Leu 145 150 155 160 Gly Leu Arg Pro Arg Gln Val Glu Val Trp Phe Gln Asn Arg Arg Ala 165 170 175 Arg Thr Lys Leu Lys Gln Thr Glu Val Asp Cys Glu Phe Leu Lys Arg 180 185 190 Cys Cys Glu Asn Leu Thr Glu Glu Asn Arg Arg Leu Gln Lys Glu Val 195 200 205 Gln Glu Leu Arg Ala Leu Lys Leu Ser Pro Gln Phe Tyr Met Gln Met 210 215 220 Thr Pro Pro Thr Thr Leu Thr Met Cys Pro Ser Cys Glu Arg Val Ala 225 230 235 240 Gly Pro Pro Ser Ser Ser Ser Gly Pro Thr Ser Thr Pro Met Gly Gln 245 250 255 Ala Gln Pro Arg Pro Arg Pro Phe Asn Leu Trp Ala Asn Ala Leu His 260 265 270 Pro Arg Serum 275 <210> 26 <211> 323 <212> PRT <213> Erythranthe guttata <400> 26 Met Met Thr Ala Gly Lys Glu Asp Leu Gly Leu Ser Leu Ser Leu Thr 1 5 10 15 Phe Pro Pro Glu Lys Lys Ala Val Ser Asn Ile Pro Ser Ser Leu His 20 25 30 Leu Asn Leu Met Pro Ser Ser Pro Ser Pro Asn Phe Thr Ile Phe Asn 35 40 45 Asn Asn Phe Asn Trp Thr Ser Thr Gln Gln Ala Ala Ala Ala Phe Pro 50 55 60 Phe Ser Asp Arg Ser Ser Glu Thr Cys Arg Val Glu Thr Thr Arg Ser 65 70 75 80 Phe Leu Lys Gly Ile Asp Val Asn Cys Leu Pro Ser Ala Ala ...
Claims
1. A method for producing maize plants, comprising: Regenerating maize plants from maize plant parts or maize plant cells, wherein the maize plant parts or maize plant cells contain at least one non-natural mutation in an endogenous homologous domain-leucine zipper (HD-Zip) transcription factor gene, wherein the mutation disrupts the binding of the HD-Zip transcription factor to DNA, wherein the mutant homologous domain-leucine zipper (HD-Zip) transcription factor protein encoded by the mutant homologous domain-leucine zipper (HD-Zip) transcription factor gene contains the amino acid sequence of SEQ ID NO: 312 or 316.
2. A method for producing corn plant parts, comprising: Regenerating maize plant parts from maize plant cells, wherein the maize plant cells contain at least one non-natural mutation in an endogenous homologous domain-leucine zipper (HD-Zip) transcription factor gene, wherein the mutation disrupts the binding of the HD-Zip transcription factor to DNA, wherein the mutant homologous domain-leucine zipper (HD-Zip) transcription factor protein encoded by the mutant homologous domain-leucine zipper (HD-Zip) transcription factor gene contains the amino acid sequence of SEQ ID NO: 312 or 316.
3. The method according to claim 1 or 2, wherein the mutant HD-Zip transcription factor gene comprises a nucleotide sequence selected from SEQ ID NO: 313-315 and 317-319.
4. The method according to claim 1 or 2, wherein the maize plant comprises a dwarf / semi-dwarf phenotype.
5. A method for producing a corn plant or a corn plant part, comprising regenerating the corn plant or corn plant part from corn plant cells: (1) A maize plant cell including an editing system, said editing system comprising: (a) CRISPR-related effector proteins; and (b) A guide nucleic acid having a spacer sequence complementary to an endogenous target gene encoding a wild-type HD-Zip transcription factor, wherein the guide nucleic acid comprises any of the nucleotide sequences shown in SEQ ID NO: 175-182.
6. The method of claim 5, wherein the wild-type HD-Zip transcription factor is the HB53 transcription factor.
7. The method according to claim 5, wherein the maize plant cell contains a mutant HD-Zip transcription factor gene encoding an HD-Zip transcription factor, wherein the HD-Zip transcription factor contains the amino acid sequence of SEQ ID NO: 312 or 316.
8. The method of claim 5, wherein the nucleic acid binding domain of the editing system is derived from a polynucleotide-guided endonuclease, a CRISPR-Cas endonuclease, a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), and / or an Argonaute protein.
9. The method according to claim 8, wherein the CRISPR-Cas endonuclease is a CRISPR-Cas effector protein.
10. The method according to any one of claims 1-9, wherein the maize plant contains a reduced shade avoidance response.
11. The method of claim 5, wherein the maize plant comprises a dwarf / semi-dwarf phenotype.
12. A method for producing / breeding non-GMO genome-edited maize plants, the method comprising: (a) Crossing maize plants produced by the method of any one of claims 1-11 with non-GMO maize plants, thereby introducing the mutation into the non-GMO maize plants; and (b) Select offspring maize plants containing the mutation but without transgenes to produce transgene-free genome-edited maize plants.
13. The method of claim 12, wherein the maize plant is base-edited.
14. A method for editing specific sites in the genome of a maize plant cell, the method comprising: The target site within the endogenous HD-Zip transcription factor gene in the maize plant cell is cut in a site-specific manner, wherein the target site is a DNA-binding domain, resulting in editing of the mutant endogenous HD-Zip transcription factor gene encoding a mutant HD-Zip transcription factor, wherein the mutant HD-Zip transcription factor contains the amino acid sequence of SEQ ID NO: 312 or 316.
15. The method of claim 14, further comprising regenerating maize plants from maize plant cells containing the edit in the endogenous HD-Zip transcription factor gene to produce maize plants containing the edit in their endogenous HD-Zip transcription factor gene.
16. The method of claim 15, wherein the maize plant containing the edit in its endogenous HD-Zip transcription factor gene has a weakened shade avoidance response compared to a control plant without the edit.
17. A method for reducing shade avoidance response in maize plants, comprising: (a) A maize plant cell containing a wild-type endogenous gene encoding an HD-Zip transcription factor is contacted with a nuclease targeting the wild-type endogenous gene, wherein the nuclease is linked to a DNA-binding domain that binds to a target site in the wild-type endogenous gene, the nuclease cleaves the endogenous HD-Zip transcription factor gene and introduces a mutation within the DNA-binding site of the endogenous HD-Zip transcription factor encoded by the endogenous HD-Zip transcription factor gene, wherein the mutated endogenous HD-Zip transcription factor gene encodes a mutated HD-Zip transcription factor protein containing the amino acid sequence SEQ ID NO: 312 or 316, thereby producing maize plant cells containing a mutation in the wild-type endogenous gene encoding the HD-Zip transcription factor; as well as (b) The corn plant cells are allowed to grow into corn plants, thereby reducing the shade avoidance response in the corn plants.
18. A method for producing a maize plant or a maize plant part comprising at least one maize plant cell having a mutation in an endogenous HD-Zip transcription factor gene, the method comprising: The target site of the endogenous HD-Zip transcription factor gene in the plant or plant part is contacted with a nuclease containing a cleavage domain and a DNA-binding domain, wherein the DNA-binding domain binds to the target site of the endogenous HD-Zip transcription factor gene. The nuclease cleaves the endogenous HD-Zip transcription factor gene and introduces a deletion within the DNA-binding site of the endogenous HD-Zip transcription factor encoded by the endogenous HD-Zip transcription factor gene. The endogenous HD-Zip transcription factor gene encodes the amino acid sequence of SEQ ID NO: 1, thereby producing a maize plant or a maize plant portion containing at least one cell with a mutation in the endogenous HD-Zip transcription factor gene, wherein the mutant HD-Zip transcription factor protein encoded by the mutant HD-Zip transcription factor gene contains the amino acid sequence of SEQ ID NO: 312 or 316.
19. The method of claim 18, wherein the maize plant exhibits a reduced shade avoidance response compared to the control plant.
20. The method of claim 18, wherein the maize plant comprising a reduced shade avoidance response comprises at least one of the following phenotypes: compared to plants without a reduced shade avoidance response planted adjacent to one or more plants with a reduced shade avoidance response, the maize plant, when planted adjacent to one or more plants with a reduced shade avoidance response, has increased yield, reduced height, reduced stem-to-root ratio, shortened leaf length; enhanced stem mechanical strength; reduced lodging rate; delayed senescence; improved photosynthetic efficiency and grain filling; and / or enhanced defense response against pathogens and herbivores.
21. The method according to any one of claims 18-20, wherein the nuclease is a zinc finger nuclease, a transcription activator-like effector nuclease (TALEN), a nuclease, or a CRISPR-Cas effector protein.
22. The method according to claim 21, wherein the endonuclease is Fok1.
23. The method of claim 18, wherein the HD-Zip transcription factor is capable of regulating the plant's response to light.
24. The method of claim 23, wherein the HD-Zip transcription factor is capable of regulating the shade avoidance response (SAR) in the plant.
25. A guide nucleic acid that binds to a target site in an HD-Zip transcription factor gene, the target site comprising a nucleotide sequence encoding a polypeptide having the amino acid sequence VWFQNRRA (SEQ ID NO:9), wherein the guide nucleic acid comprises a spacer having a nucleotide sequence shown in any of SEQ ID No:175-182.
26. The guide nucleic acid of claim 25, wherein the HD-Zip transcription factor is an HD-Zip type II (HD-Zip II) transcription factor, optionally wherein the HD-Zip II transcription factor is HB53.
27. An editing system comprising a guide nucleic acid of claim 25 or 26 and a CRISPR-Cas effector protein associated with said guide nucleic acid, wherein said guide nucleic acid comprises a spacer having a nucleotide sequence shown in any of SEQ ID No: 175-182.
28. The editing system of claim 27, further comprising a tracr nucleic acid associated with the guide nucleic acid and the CRISPR-Cas effector protein, optionally wherein the tracr nucleic acid and the guide nucleic acid are covalently linked.
29. An expression cassette comprising (a) a polynucleotide encoding a CRISPR-Cas effector protein containing a cleavage domain, and (b) a guide nucleic acid binding to a target site in an HD-Zip transcription factor gene, the target site comprising a nucleotide sequence encoding a polypeptide having the amino acid sequence VWFQNRRA (SEQ ID NO: 9), wherein the guide nucleic acid comprises a spacer sequence having a nucleotide sequence shown in any of SEQ ID No: 175-182.
30. The expression cassette of claim 29, wherein the HD-Zip transcription factor is an HD-Zip type II (HD-ZipII) transcription factor, optionally wherein the HD-Zip II transcription factor is HB53.