Expression of zygote preference
By using a zygotic-selected promoter to express gene editing system components in zygotic cells during haploid induction, the problem of low gene editing efficiency in existing technologies is solved, and efficient gene editing delivery is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SYNGENTA CROP PROTECITON AG
- Filing Date
- 2024-08-02
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies struggle to efficiently deliver gene-editing mechanisms to the plant genome during haploid induction, resulting in low gene-editing efficiency.
Zygote-preferred promoters, such as vacuole sorting protein promoters (VSP) or cyclic zinc finger domain protein promoters (RZDP), are used to preferentially express gene editing system components in zygote cells and deliver the editing mechanism to the recipient plant genome using haploid induction technology.
This technology enables the efficient delivery of gene editing system components to the plant genome during haploid induction, improving the efficiency and precision of gene editing.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_4
Abstract
Description
Expression of zygotic optimization Technical Field
[0001] This disclosure relates to the field of plant biotechnology, specifically agricultural biotechnology and gene editing, as well as plant breeding. The subject matter of this disclosure concerns the transformation of haploid inducible lines to contain DNA encoding a cellular mechanism capable of gene editing, which is preferentially expressed in zygotic cells. Priority is claimed.
[0002] This application claims priority to application serial number PCT / CN2023 / 110941, filed on August 3, 2023, which is incorporated herein by reference in its entirety. Sequence List
[0003] This application is accompanied by a sequence list entitled 82223PCT12mo.xml, created on July 26, 2024, which is approximately 555 kilobytes in size. This sequence list is incorporated herein by reference in its entirety. The sequence list is hereby submitted and conforms to 37 C.FR § 1.831-1.835. Background Technology
[0004] Haploid-induced genome editing (HI-editing) employs a modified haploid-inducing plant expressing a gene-editing mechanism to deliver that mechanism to the recipient plant's genome to be edited. In this procedure, a first plant is crossed with a second plant to obtain haploid progeny in which the chromosomes of the haploid-inducing line are eliminated and the recipient plant's haploid chromosomes have the desired editing. The HI-editing method is detailed in PCT Publication WO 2018 / 102816. See also Kelliher, T et al. (2019) One-step genome editing of elite crop germplasm during haploid induction. Nature Biotech 37: 287–292. Summary of the Invention
[0005] This disclosure is characterized by nucleic acid constructs comprising zygotic-preferred promoters, such as vacuole sorting protein promoters (VSPs), cyclic zinc finger domain protein promoters (RZDPs), or SUMO conjugates (SCE1), which preferentially drive the expression of transcripts from nucleic acids operatively linked to the promoter in early zygotes and male gametes. In some embodiments, zygotic-preferred promoters are employed in the nucleic acid constructs to drive the expression of gene editing system components used in gene editing procedures, such as HI-editing, which utilize haploid induction to deliver desired edits to plants.
[0006] In one aspect, this disclosure provides a synthetic DNA construct comprising a zygotic preferred promoter operatively linked to a first target nucleotide sequence (“NSOI”). In some embodiments, the zygotic preferred promoter is a vacuole sorting protein promoter (“VSP promoter”). In some cases, the VSP promoter comprises: a) a sequence selected from the group consisting of SEQ ID NO: 1-27 or a functional fragment thereof; or b) an orthologous promoter of SEQ ID NO: 1. In some embodiments, the VSP promoter comprises SEQ ID NO: 27. In alternative embodiments, the zygotic preferred promoter is a cyclic zinc finger domain protein promoter (“RZDP promoter”). In some cases, the RZDP promoter comprises: (a) a sequence selected from the group consisting of SEQ ID NO: 28-43 or a functional fragment thereof; or (b) an orthologous promoter of SEQ ID NO: 28. In some embodiments, the RZDP promoter comprises SEQ ID NO: 28. In some embodiments, the preferred promoter for the zygote is the SUMO conjugate 1 promoter (“SCE1 promoter”). In some cases, the SCE1 promoter comprises: a) a sequence selected from the group consisting of SEQ ID NO: 83-85 or a functional fragment thereof; or b) an orthologous promoter of SEQ ID NO: 83. In some embodiments, the SCE1 promoter comprises SEQ ID NO: 84. In some cases, the SCE1 promoter comprises SEQ ID NO: 85. In some embodiments, the synthetic DNA construct further comprises a U3 promoter operatively linked to a second NSOI. In some cases, the U3 promoter comprises: (a) a sequence selected from the group consisting of SEQ ID NO: 44-49 or a functional fragment thereof; or (b) an orthologous promoter of SEQ ID NO: 44. In some embodiments, the U3 promoter comprises SEQ ID NO: 45. In some embodiments, the synthetic DNA construct further comprises a terminator operatively linked to an NSOI. In some cases, the terminator is a ubiquitin terminator. In some embodiments, the ubiquitin terminator comprises SEQ ID NO: 88, 89, 90, 91, 92, or 93. In some embodiments of the synthetic DNA construct, the first NSOI comprises a sequence encoding a nuclease. In some cases, the nuclease is a zinc finger nuclease (“ZFN”), a large-scale nuclease (“MN”), a transcription activator-like effector nuclease (TALEN), or a CRISPR nuclease.In some embodiments, the CRISPR nuclease is selected from the group consisting of: Cas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12i, Cas12j, Cas12l, Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas10, Cas11, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, any other CRISPR-Cas nuclease, any mutants thereof, nicking enzyme variants, and inactivating variants. In some cases, the CRISPR nuclease further comprises a fusion domain, such as a fusion domain selected from the group consisting of deaminases, uracil DNA glycosylases, reverse transcriptases, and exonucleases. In some embodiments, the second NSOI of the synthetic DNA construct comprises a sequence encoding at least one guide RNA. In some embodiments, the at least one guide RNA is encoded by a sequence selected from the group consisting of SEQ ID NO: 50-61. In some embodiments, the synthetic DNA construct comprises a sequence selected from the group consisting of SEQ ID NO: 62-72.
[0007] In another aspect, this disclosure provides a plant cell comprising a synthetic DNA construct as described herein (e.g., in the preceding paragraphs). In some embodiments, the plant cell is a pollen cell or an egg cell. In another aspect, this disclosure provides a plant comprising the plant cell, such as a maize plant.
[0008] In another aspect, this disclosure provides a method for obtaining edited progeny plants, the method comprising: (a) providing a first plant, wherein the first plant is transformed to include the synthetic constructs as described in claims 1-18; (b) pollinating a second plant; and (c) selecting at least one progeny produced by the pollination in step (b), wherein the progeny has edits; thereby obtaining edited progeny plants. In some embodiments, the first plant is a haploid inducing line of said plant. In some cases, the haploid inducing line is a paternal haploid inducing line, such as a paternal haploid inducing line containing a mutation in the CENH3 gene. In other embodiments, the haploid inducing line is a maternal haploid inducing line, such as a maternal haploid inducing line containing a mutation in the MATL gene. In some embodiments, the second plant contains plant genomic DNA to be edited. In some embodiments, the edited progeny plant is a haploid progeny plant. In some embodiments, the haploid progeny plant contains the genome of the second plant but not the genome of the first plant. In some embodiments, the haploid progeny plant is a maize haploid progeny plant.
[0009] The above summary is provided to introduce the concept selection further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to help limit the scope of the claimed subject matter. Detailed Implementation
[0010] The following description illustrates various aspects and embodiments of the compositions and methods of the present invention. Specific embodiments are not intended to limit the scope of the compositions and methods. Rather, the embodiments provide only non-limiting examples of various compositions and methods that are at least within the scope of the disclosed compositions and methods. The description is intended to be read by one of ordinary skill in the art; therefore, it does not necessarily include information well known to those skilled in the art.
[0011] the term
[0012] Unless otherwise defined below, all technical and scientific terms used herein are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques used herein are intended to refer to techniques commonly understood in the art, including variations and / or equivalents of those techniques that are readily apparent to one of ordinary skill in the art. While it is believed that the following terms will be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate the interpretation of the subject matter disclosed herein.
[0013] Unless the context clearly specifies otherwise, as used herein, the singular forms “a / an” and “the” include plural indicators. For example, references to “cell” include one or more cells and may include tissues or organs.
[0014] As used herein, the term "about," when referring to a measurable value such as mass, weight, time, volume, concentration, or percentage, means to encompass variations suitable for performance of, in some embodiments, +20% of a specified amount, in some embodiments, +10% of a specified amount, in some embodiments, +5% of a specified amount, in some embodiments, +1% of a specified amount, in some embodiments, +0.5% of a specified amount, and in some embodiments, +0.1% of a specified amount, because such variations are suitable for performing the disclosed methods and / or using the disclosed compositions, nucleic acids, peptides, etc. Therefore, unless indicated to the contrary, the numerical parameters listed in this specification and the appended claims are approximate values, which may vary depending on the desired characteristics sought to be obtained through the subject matter of this disclosure.
[0015] Entities may exist individually or in combination. Thus, for example, the phrase “A, B, C and / or D” includes A, B, C, and D individually, but also any and all combinations and subcombinations of A, B, C, and D (e.g., AB, AC, AD, BC, BD, CD, ABC, ABD, and BCD). In some embodiments, one or more elements referred to by “and / or” may also exist individually in one or more occurrences in one or more combinations and / or one or more subcombinations.
[0016] As used herein, the term "plant" can refer to the whole plant at any stage of development or any part or component of a plant; and includes references to cell or tissue cultures derived from a plant. Thus, for example, "plant" can refer to components or organs such as leaves, stems, roots, plant tissues, seeds, and / or plant cells.
[0017] As used herein, the term "plant cell" refers to the structural and physiological unit of a plant, including the protoplast and cell wall. Plant cells can be in the form of isolated single cells or cultured cells, or as part of higher tissue units such as plant tissues, plant organs, or the whole plant. Plant cells can originate from angiosperms or gymnosperms, or be part of them. Plant cells can be monocotyledonous cells (e.g., corn cells, rice cells, sorghum cells, sugarcane cells, barley cells, wheat cells, oat cells, turfgrass cells, or ornamental grass cells) or dicotyledonous cells (e.g., tobacco cells, pepper cells, eggplant cells, sunflower cells, cruciferous plant cells, flax cells, potato cells, cotton cells, soybean cells, sugar beet cells, or oilseed rape cells). As used herein, the term "plant cell culture" refers to a culture of plant units such as protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at different developmental stages. As used herein, the term "plant tissue" refers to a group of plant cells organized into structural and functional units. This includes any plant tissue in or cultured. This term includes, but is not limited to, the whole plant, plant organs, plant seeds, tissue cultures, and any group of plant cells organized into structural and / or functional units. The use of this term in conjunction with (or where absent) any particular type of plant tissue listed above, or in other ways covered by this definition, is not intended to exclude any other type of plant tissue. As used herein, the term "plant part" refers to a portion of a plant, including single-celled and cellular tissues (such as intact plant cells in a plant), cell masses, and tissue cultures that can regenerate the plant. Examples of plant parts include, but are not limited to, single cells and tissues derived from: pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral parts, fruits, stems, buds, cuttings, and seeds; as well as pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, floral parts, fruits, stems, buds, cuttings, scions, rhizomes, seeds, protoplasts, callus, etc.
[0018] As used herein, the terms "offspring" and "offspring plant" refer to a plant produced from one or more parent plants through vegetative or sexual reproduction. The term "offspring" can refer to any descendant of a particular hybrid or parent plant. Typically, offspring plants are produced by breeding two individuals, but some species (particularly some plants and hermaphroditic animals) can self-fertilize (i.e., the same plant acts as a donor for both male and female gametes). This one or more descendants can be, for example, F1, F2, or any subsequent generations. In some embodiments, "offspring" plants are produced using Hi-editing methods.
[0019] Haploid induction (“HI”) is a plant phenomenon characterized by the deletion of a set of chromosomes from the embryo at some point during or after fertilization (often during early embryonic development) (the chromosomes of the haploid induction parent). Haploid induction is also known as monogyny if the inducing line is used as the male in the hybridization, or monoandry if the inducing line is used as the female. Haploid induction has been studied in numerous plant species, such as sorghum, barley, wheat, maize, Arabidopsis, and many others. Generally, during haploid induction, both parental lines used in the induced hybridization are diploid, and therefore their gametes (egg cells and sperm cells) are haploid. Haploid induction is usually a mediator of reducing the penetrance of the inducing line, so the resulting offspring, depending on the species or circumstances, can be diploid (if no genome deletion has occurred) or haploid (if a genome deletion has actually occurred). If the parental line crossed with a haploid inducer is not diploid, but rather tetraploid, hexaploid, or other plants of higher ploidy, the resulting "haploid" offspring will have gamete chromosome numbers such as diploid (if the parent is tetraploid) or triploid (if the parent is hexaploid). Therefore, as used herein, a "haploid" has half the chromosome number of its parent; thus, a haploid in a diploid organism (e.g., maize) exhibits haploidy; a haploid in a tetraploid organism (e.g., ryegrass) exhibits diploidy; and a haploid in a hexaploid organism (e.g., wheat) exhibits triploidy.
[0020] The term "HI-edit" refers to "haploid-induced editing" of the genome, which uses haploid-induced plants modified to express gene-editing mechanisms to deliver these mechanisms to the genome of the recipient plant to be edited. In this procedure, a first plant is crossed with a second plant to obtain haploid progeny, wherein the chromosomes of the haploid-induced line are eliminated and the recipient plant's haploid chromosomes have the desired editing. The HI-editing method is detailed in PCT Publication WO 2018 / 102816. See also Kelliher, T et al. (2019) One-step genome editing of elite crop germplasm during haploid induction. Nature Biotech 37: 287–292.
[0021] Plants referred to here as "double haploids" are produced by doubling the set of haploid chromosomes. Plants or seeds obtained from double haploid plants through self-pollination to any generation can still be identified as double haploid plants. Double haploid plants are considered homozygous plants. If a plant is fertile, it is considered double haploid even if the entire vegetative part of the plant is not composed of cells with doubled chromosome sets; that is, if the plant contains living gametes, it will be considered double haploid even if the plant is chimeric in its vegetative tissue.
[0022] A “promoter” is a polynucleotide sequence contained upstream of a 5' transcription start site that controls the transcription of a gene or sequence operatively linked to it. A promoter includes signals for RNA polymerase binding and transcription initiation. As used herein, a promoter may also contain regulatory elements such as enhancers and repressor binding sites. In some embodiments, the nucleotide sequence of a promoter may include a region encoding a 5' untranslated sequence of the transcript of the native gene derived from that promoter. For example, the VSP promoter sequence or a functional fragment thereof of any one of SEQ ID NO: 1-27 may include a region encoding a 5' untranslated sequence of the VSP transcript; or the RZDP promoter sequence or a functional fragment thereof of any one of SEQ ID NO: 28-43 may include a region encoding a 5' untranslated sequence of the RZDP transcript. In some embodiments, the promoter does not include a region encoding a 5' untranslated sequence. As another example, the promoter may be the U3 promoter sequence or a functional fragment thereof of any one of SEQ ID NO: 44-49. In some embodiments, such a promoter does not contain an untranscribed sequence.
[0023] As used herein, “target nucleotide sequence” refers to a polynucleotide to be expressed by the synthetic construct, including polynucleotides operatively linked to a promoter used to control polynucleotide transcription. The transcript produced by the synthetic construct can be protein-coding RNA, such as RNA encoding a nuclease, or a non-protein-coding transcript, such as guide RNA.
[0024] As used in this article, "endogenous" or "natural" nucleic acid sequences refer to nucleic acid sequences that are naturally present in the genome of an organism.
[0025] A "gene" is a defined region within the genome that, in addition to encoding a nucleic acid sequence, includes other sequences, primarily regulatory sequences responsible for controlling the expression (i.e., transcription) of the coding portion (where proteins are produced by the gene) and translation. A gene can include both coding and non-coding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). Genes typically express mRNA, functional RNA, or specific proteins, including regulatory sequences. Genes may or may not be used to produce functional proteins. In some embodiments, a gene refers only to the coding region. The term "natural gene" refers to a gene as found in nature. The term "chimeric gene" refers to any gene that includes: 1) a DNA sequence, including regulatory and coding sequences not found together in nature; or 2) a sequence encoding a portion of a protein not naturally adjacent to it; or 3) a portion of a promoter not naturally adjacent to it. Thus, a chimeric gene can contain regulatory and coding sequences derived from different sources, or it can contain regulatory and coding sequences derived from the same source but arranged in a different manner than those found in nature. Genes can be “isolated,” meaning that nucleic acid molecules do not substantially contain, or are substantially free of, the components typically found associated with nucleic acid molecules in their native state. Such components include other cellular material, culture media from recombinant production, and / or various chemicals used in the chemical synthesis of nucleic acid molecules.
[0026] The terms “nucleic acid” and “polynucleotide” are used interchangeably and, as used herein, refer to deoxyribonucleic acid (DNA) or ribonucleic acid (RNA) in single-stranded or double-stranded form and polymers thereof, and to both sense and antisense strands of RNA, cDNA, genomic DNA, and mitochondrial DNA, as well as synthetic forms and mixed polymers thereof. In higher plants, DNA is the genetic material, while RNA is involved in the transfer of information contained within DNA to proteins. A “genome” is the total genetic material contained in every cell of an organism. It should be understood that when RNA is described, its corresponding cDNA is also described, where uridine is represented as thymidine. In certain embodiments, a nucleotide refers to a ribonucleotide, deoxynucleotide, or a modified form of any type of nucleotide or a combination thereof. Additionally, the polynucleotides disclosed herein may include any or both of naturally occurring nucleotides and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide bonds. Nucleic acid molecules may be chemically or biochemically modified, or may contain non-natural or derivatized nucleotide bases, as will be readily understood by those skilled in the art. Such modifications include, for example, tagging, methylation, substitution of one or more naturally occurring nucleotides with analogs, internucleotide modifications such as uncharged linkages (e.g., methyl phosphonate, triphosphate, aminophosphate, carbamate, etc.), charged linkages (e.g., thiophosphate, dithiophosphate, etc.), side moieties (e.g., polypeptides), intercalating agents (e.g., acridine, psoralen, etc.), chelating agents, alkylating agents, and modified linkages (e.g., α-anomeric nucleic acids, etc.). The above terms are also intended to include any topological conformation, including single-stranded, double-stranded, partially double-stranded, triple-stranded, hairpin-shaped, circular, and padlock-shaped conformations. Unless otherwise stated, references to nucleic acid sequences cover their complements. Therefore, references to nucleic acid molecules having a specific sequence should be understood to cover their complementary strands having their complementary sequences. These nucleotide sequences are "complementary" (e.g., according to the Watson-Crick base pairing principle) when they specifically hybridize in solution. The term also includes codon-optimized nucleic acids encoding the same polypeptide sequence. It should also be understood that nucleic acids can be unpurified, purified, or attached to, for example, synthetic materials (such as bead or column matrices).
[0027] As used herein, in the context of nucleic acid sequences, the term "reference sequence" refers to a defined nucleotide sequence that serves as the basis for nucleotide sequence comparison. In some embodiments, the reference sequence may be a promoter sequence of any one of SEQ ID NO: 1-27, SEQ ID NO: 28-43, or SEQ ID NO: 44-49; a sequence encoding a guide RNA of any one of SEQ ID NO: 50-61; or a nucleic acid sequence encoding a construct of any one of SEQ ID NO: 62-72.
[0028] As used in this disclosure, in the context of nucleic acid sequences, the term "corresponds" means that, when two sequences are best aligned, certain positions or regions of the target nucleotide sequence align with those positions or regions of the reference sequence, but the position numbers of the two sequences are not necessarily precisely aligned. While best alignment and scoring can be performed manually, the process can be facilitated by computer-implemented alignment algorithms. Easily available sequence comparison and multiple sequence alignment algorithms include the Basic Local Alignment Search Tool (BLAST) and the ClustalW / ClustalW2 / Clustal Omega programs, which are available on the Internet (e.g., the EMBL-EBI website). Other suitable programs include, but are not limited to, GAP, BestFit, PlotSimilarity, and FASTA, which are part of the Accelrys GCG software package available from Accelrys Inc. (San Diego, California, USA). See also Smith and Waterman, 1981; Needleman and Wunsch, 1970; Pearson and Lipman, 1988; Ausubel et al., 1988; and Sambrook and Russell, 2001.
[0029] One example of an algorithm suitable for determining percentage sequence identity and sequence similarity is the BLAST algorithm, described in Altschul et al., 1990. In some embodiments, the percentage of sequence identity refers to sequence identity across the full length of a nucleic acid or polypeptide sequence.
[0030] Unless otherwise specified, a particular nucleic acid sequence also implicitly includes variants of its conserved modifications (e.g., degenerate codon substitutions in sequences encoding proteins), alleles, SNPs and complementary sequences, as well as explicitly specified sequences.
[0031] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, these terms cover amino acid chains of any length, including full-length proteins, where amino acid residues are linked by covalent peptide bonds.
[0032] As used in the context of polynucleotide or polypeptide sequences described herein, the term "identity" or "substantial identity" refers to a sequence having at least 60% sequence identity with a reference sequence. Alternatively, the identity percentage can be any integer from 60% to 100%. When using the procedure described herein (preferably BLAST with standard parameters as described below), the exemplary embodiments include at least the following identity percentages compared to the reference sequence: 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%. Those skilled in the art will recognize that these values can be appropriately adjusted by taking into account codon degeneracy, amino acid similarity, reading frame localization, etc., to determine the corresponding identity of proteins encoded by two nucleotide sequences.
[0033] For sequence comparisons, typically one sequence serves as a reference sequence to be compared with the test sequence. When using a sequence comparison algorithm, the test and reference sequences are input into the computer, with subsequence coordinates specified if necessary, and the sequence algorithm program parameters specified. Default program parameters can be used, or alternative parameters can be specified. The sequence comparison algorithm then calculates the percentage of sequence identity between the test sequence and the reference sequence based on the program parameters.
[0034] As used herein, a “comparison window” includes a segment that refers to any of the following groups of consecutive positions, ranging from 20 to 600, typically from about 50 to about 200, and more typically from about 100 to about 150, wherein, after optimal alignment of two sequences, the sequence can be compared with a reference sequence having the same number of consecutive positions. Sequence alignment methods used for comparison are well known in the art. The optimal alignment of sequences for comparison can be performed in the following ways: local homology algorithm by Smith and Waterman Add. APL. Math. [Advances in Applied Mathematics] 2:482 (1981); homology alignment algorithm by Needleman and Wunsch J. Mol. Biol. [Journal of Molecular Biology] 48:443 (1970); similarity method search by Pearson and Lipman Proc. Natl. Acad. Sci. [Proceedings of the National Academy of Sciences] (USA) 85: 2444 (1988); computer implementations of these algorithms (e.g., BLAST); or manual alignment and visual inspection.
[0035] Unless otherwise stated, identity and similarity will be calculated using the Needleman-Wunsch global alignment and scoring algorithm (Needleman and Wunsch (1970) J. Mol. Biol. [Molecular Biology Journal] 48(3):443-453), which is implemented as part of the "needle" program (Rice, P., Longden, I. and Bleasby, A., EMBOSS: The European Molecular Biology Open Software Suite, 2000, Trends in Genetics 16, (6) pp. 276-277, version 6.3.1, available from EMBnet (embnet.org / resource / emboss) and emboss.sourceforge.net, etc.), using the default vacancy penalty and scoring matrix (EBLOSUM62 for proteins and EDNAFULL for DNA). An equivalent program may also be used. "Equivalent procedure" refers to any sequence comparison procedure that, for any two sequences in question, generates an alignment with the same nucleotide residue match and the same percentage of sequence identity when compared to the corresponding alignment generated by needle in EMBOSS version 6.3.1.
[0036] Other mathematical algorithms are known in the art and can be used to compare two sequences. See, for example, the algorithm of Karlin and Altschul (1990) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 87:2264, which is an improvement on that of Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 90:5873-5877. Such algorithms are integrated into the BLAST program of Altschul et al. (1990) J. Mol. Biol. [Journal of Molecular Biology] 215:403. BLAST nucleotide searches can be performed using the following procedures: using the BLASTN program (a nucleotide query for searching nucleotide sequences) to obtain nucleotide sequences homologous to the nucleic acid molecules of the present invention, or using the BLASTX program (a nucleotide query for searching protein sequences for translation) to obtain protein sequences homologous to the nucleic acid molecules of the present invention. BLAST protein searches can be performed using the following procedures: the BLASTP procedure (searching for protein queries based on protein sequences) to obtain amino acid sequences homologous to the protein molecules of the present invention, or the TBLASTN procedure (searching for protein queries based on translated nucleotide sequences) to obtain nucleotide sequences homologous to the protein molecules of the present invention. For vacancy-free alignments for comparative purposes, Gapped BLAST (in BLAST 2.0) as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389 can be used. Alternatively, iterative searches can be performed using PSI-Blast, which detects distant relationships between molecules. See Altschul et al. (1997), ibid. When using BLAST, Gapped BLAST, and PSI-Blast procedures, the default parameters of the corresponding procedures (e.g., BLASTX and BLASTN) can be used. Alignments can also be performed manually by checking.
[0037] Therefore, an "isolated" nucleic acid molecule is a nucleic acid molecule or nucleotide sequence that is not adjacent to a nearby nucleotide sequence (either at the 5' end or the 3' end) in the naturally occurring genome of the organism from which it originates. Thus, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences that are adjacent to the coding sequence. Therefore, the term includes, for example, recombinant nucleic acids that are integrated into a vector, a self-replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or that exist as a separate molecule independent of other sequences (e.g., a fragment of cDNA or genomic DNA produced by PCR or restriction endonuclease treatment). It also includes recombinant nucleic acids that are portions of hybrid nucleic acid molecules encoding additional RNA or polypeptide sequences. "Isolated" nucleic acid molecules can also include polynucleotides that are derived from the same native cell type and inserted therein, but exist in a non-native state, for example, in different copy numbers, and / or under the control of regulatory sequences different from those found in the native state of nucleic acid molecules. "Isolated" does not necessarily mean that the preparation is industrially pure (homogeneous), but it is pure enough to provide nucleic acids in a form that can be used for the intended purpose.
[0038] A brief description of sequences in a sequence list.
[0039] SEQ ID NO: 1 is the nucleotide sequence of the VSP promoter from Arabidopsis thaliana.
[0040] SEQ ID NO: 2 is the nucleotide sequence of the VSP promoter from Arabidopsis thaliana.
[0041] SEQ ID NO: 3 is the nucleotide sequence of the VSP promoter from Arabidopsis thaliana.
[0042] SEQ ID NO: 4 is the nucleotide sequence of the VSP promoter from Arabidopsis thaliana.
[0043] SEQ ID NO: 5 is a nucleotide sequence from the VSP promoter of Arabidopsis thaliana.
[0044] SEQ ID NO: 6 is the nucleotide sequence of the VSP promoter from Arabidopsis thaliana.
[0045] SEQ ID NO: 7 is the nucleotide sequence of the VSP promoter from sunflower.
[0046] SEQ ID NO: 8 is the nucleotide sequence of the VSP promoter from sunflower.
[0047] SEQ ID NO: 9 is the nucleotide sequence of the VSP promoter from sunflower.
[0048] SEQ ID NO: 10 is the nucleotide sequence of the VSP promoter from sunflower.
[0049] SEQ ID NO: 11 is the nucleotide sequence of the VSP promoter from rice.
[0050] SEQ ID NO: 12 is a nucleotide sequence of the VSP promoter from rice.
[0051] SEQ ID NO: 13 is a nucleotide sequence of the VSP promoter from rice.
[0052] SEQ ID NO: 14 is a nucleotide sequence of the VSP promoter from rice.
[0053] SEQ ID NO: 15 is the nucleotide sequence of the VSP promoter from tomato.
[0054] SEQ ID NO: 16 is the nucleotide sequence of the VSP promoter from tomato.
[0055] SEQ ID NO: 17 is the nucleotide sequence of the VSP promoter from tomato.
[0056] SEQ ID NO: 18 is the nucleotide sequence of the VSP promoter from soybean.
[0057] SEQ ID NO: 19 is the nucleotide sequence of the VSP promoter from soybean.
[0058] SEQ ID NO: 20 is the nucleotide sequence of the VSP promoter from soybean.
[0059] SEQ ID NO: 21 is the nucleotide sequence of the VSP promoter from soybean.
[0060] SEQ ID NO: 22 is the nucleotide sequence of the VSP promoter from soybean.
[0061] SEQ ID NO: 23 is the nucleotide sequence of the VSP promoter from soybean.
[0062] SEQ ID NO: 24 is the nucleotide sequence of the VSP promoter from soybean.
[0063] SEQ ID NO: 25 is the nucleotide sequence of the VSP promoter from soybean.
[0064] SEQ ID NO: 26 is the nucleotide sequence of the VSP promoter from maize.
[0065] SEQ ID NO: 27 is the nucleotide sequence of the VSP promoter from maize.
[0066] SEQ ID NO: 28 is the nucleotide sequence of the RZDP promoter from maize.
[0067] SEQ ID NO: 29 is the nucleotide sequence of the RZDP promoter from Arabidopsis thaliana.
[0068] SEQ ID NO: 30 is the nucleotide sequence of the RZDP promoter from Arabidopsis thaliana.
[0069] SEQ ID NO: 31 is the nucleotide sequence of the RZDP promoter from rice.
[0070] SEQ ID NO: 32 is a nucleotide sequence from the RZDP promoter of rice.
[0071] SEQ ID NO: 33 is the nucleotide sequence of the RZDP promoter from tomato.
[0072] SEQ ID NO: 34 is the nucleotide sequence of the RZDP promoter from tomato.
[0073] SEQ ID NO: 35 is the nucleotide sequence of the RZDP promoter from tomato.
[0074] SEQ ID NO: 36 is the nucleotide sequence of the RZDP promoter from tomato.
[0075] SEQ ID NO: 37 is the nucleotide sequence of the RZDP promoter from soybean.
[0076] SEQ ID NO: 38 is the nucleotide sequence of the RZDP promoter from soybean.
[0077] SEQ ID NO: 39 is the nucleotide sequence of the RZDP promoter from soybean.
[0078] SEQ ID NO: 40 is the nucleotide sequence of the RZDP promoter from soybean.
[0079] SEQ ID NO: 41 is the nucleotide sequence of the RZDP promoter from sunflower.
[0080] SEQ ID NO: 42 is the nucleotide sequence of the RZDP promoter from sunflower.
[0081] SEQ ID NO: 43 is the nucleotide sequence of the RZDP promoter from sunflower.
[0082] SEQ ID NO: 44 is a nucleotide sequence from the U3 promoter of Arabidopsis thaliana.
[0083] SEQ ID NO: 45 is a nucleotide sequence from the U3 promoter (prOsU3-01) of rice.
[0084] SEQ ID NO: 46 is the nucleotide sequence of the U3 promoter (prOsU3-02) from rice.
[0085] SEQ ID NO: 47 is the nucleotide sequence of the U3 promoter (prOsU3-03) from rice.
[0086] SEQ ID NO: 48 is a nucleotide sequence from the U3 promoter of wheat.
[0087] SEQ ID NO: 49 is a nucleotide sequence from the U3 promoter of maize.
[0088] SEQ ID NO: 50 is a nucleotide sequence encoding a guide RNA targeting VLHP1-1 and VLHP1-2.
[0089] SEQ ID NO: 51 is a nucleotide sequence encoding a guide RNA targeting VLHP1-1 and VLHP1-2.
[0090] SEQ ID NO: 52 is a nucleotide sequence encoding a guide RNA targeting GW2-1 and GW2-2.
[0091] SEQ ID NO: 53 is the nucleotide sequence encoding the guide RNA targeting SBEIIb.
[0092] SEQ ID NO: 54 is a nucleotide sequence encoding a guide RNA targeting GL2.
[0093] SEQ ID NO: 55 is a nucleotide sequence encoding a guide RNA targeting Waxy1.
[0094] SEQ ID NO: 56 is a nucleotide sequence encoding a guide RNA targeting the first exon of O2.
[0095] SEQ ID NO: 57 is a nucleotide sequence encoding a guide RNA that targets the second exon of O2.
[0096] SEQ ID NO: 58 is a nucleotide sequence encoding a guide RNA that targets the third exon of O2.
[0097] SEQ ID NO: 59 is a nucleotide sequence encoding a guide RNA that targets YellowEndosperm.
[0098] SEQ ID NO: 60 is a nucleotide sequence encoding a guide RNA that targets UBL.
[0099] SEQ ID NO: 61 is a nucleotide sequence encoding a guide RNA targeting UPL3.
[0100] SEQ ID NO: 62 is the nucleotide sequence encoding construct 23396.
[0101] SEQ ID NO: 63 is the nucleotide sequence encoding construct 23397.
[0102] SEQ ID NO: 64 is the nucleotide sequence encoding construct 23399.
[0103] SEQ ID NO: 65 is the nucleotide sequence encoding construct 24520.
[0104] SEQ ID NO: 66 is the nucleotide sequence encoding construct 26258.
[0105] SEQ ID NO: 67 is the nucleotide sequence encoding construct 26296.
[0106] SEQ ID NO: 68 is the nucleotide sequence encoding construct 27145.
[0107] SEQ ID NO: 69 is the nucleotide sequence encoding construct 27146.
[0108] SEQ ID NO: 70 is the nucleotide sequence encoding construct 27226.
[0109] SEQ ID NO: 71 is the nucleotide sequence encoding construct 27234.
[0110] SEQ ID NO: 72 is the nucleotide sequence encoding construct 27241.
[0111] SEQ ID NO: 73 is the nucleotide sequence encoding construct 27680.
[0112] SEQ ID NO: 74 is the nucleotide sequence encoding construct 28255.
[0113] SEQ ID NO: 75 is the nucleotide sequence encoding construct 28291.
[0114] SEQ ID NO: 76 is the nucleotide sequence encoding construct 28292.
[0115] SEQ ID NO: 77 is the nucleotide sequence encoding construct 28293.
[0116] SEQ ID NO: 78 is the nucleotide sequence encoding construct 28510.
[0117] SEQ ID NO: 79 is the nucleotide sequence encoding construct 28520.
[0118] SEQ ID NO: 80 is the nucleotide sequence encoding construct 28560.
[0119] SEQ ID NO: 81 is the nucleotide sequence encoding construct 28825.
[0120] SEQ ID NO: 82 is the nucleotide sequence encoding construct 28834.
[0121] SEQ ID NO: 83 is the nucleotide sequence of the SCE1 promoter from maize.
[0122] SEQ ID NO: 84 is the nucleotide sequence of the SCE1 promoter from maize.
[0123] SEQ ID NO: 85 is the nucleotide sequence of the SCE1 promoter from maize.
[0124] SEQ ID NO: 86 is the nucleotide sequence of the VSP-01 promoter from maize.
[0125] SEQ ID NO: 87 is the nucleotide sequence of the VSP-02 promoter from maize.
[0126] SEQ ID NO: 88 is the nucleotide sequence of a ubiquitin terminator (“tUbi1-04”) from maize.
[0127] SEQ ID NO: 89 is the nucleotide sequence of a ubiquitin terminator (“tUbi1-06”) from maize.
[0128] SEQ ID NO: 90 is the nucleotide sequence of the ubiquitin terminator (“tUbi1-09”) from the maize gene GRMZM2G409726.
[0129] SEQ ID NO: 91 is the nucleotide sequence of the ubiquitin terminator from sorghum bicolor.
[0130] SEQ ID NO: 92 is the nucleotide sequence of the ubiquitin terminator from Medicago truncatula.
[0131] SEQ ID NO: 93 is the nucleotide sequence of the ubiquitin terminator from soybean (Glycine max).
[0132] This disclosure is characterized by nucleic acid constructs comprising zygotic-preferred promoters, such as vacuole sorting protein promoters (VSPs), cyclic zinc finger domain protein promoters (RZDPs), or SUMO conjugates (SCE1), which preferentially drive the expression of transcripts from nucleic acids operatively linked to the promoter in early zygotes and male gametes. In some embodiments, zygotic-preferred promoters are employed in the nucleic acid constructs to drive the expression of gene editing system components used in gene editing procedures, such as HI-editing, which utilize haploid induction to deliver desired edits to plants.
[0133] The zygotic-preferred promoter causes its downstream sequences to be preferentially expressed in the zygotic. When used in conjunction with HI-editing, the zygotic-preferred promoter disclosed herein provides a zygotic editing rate of at least 10%. The zygotic editing rate is related to HI-editing efficiency; and therefore, in some embodiments, the zygotic-preferred promoter as described herein can be used in HI-editing methods.
[0134] The term "HI-editing efficiency" refers to a measurement of progeny plants produced by HI-editing hybridization that are both edited and haploid (or, if doubled, were haploid before doubling), usually expressed as a percentage. "HI-editing efficiency," "haploid editing rate," and "HI-editing rate" are used interchangeably throughout the text.
[0135] For the purposes of this disclosure, the zygotic-preferred promoters used in the HI-editing method have a zygotic editing rate of at least about 10%. The zygotic editing rate can be determined using feasible methods, such as those described in Part III of Example 1. For example, the relative chimerism of new edits in diploid F1 progeny is assessed after crossbreeding a transgenic parent with a non-transgenic parent line. The transgenic parent expresses DNA-modifying enzymes (such as Case proteins, e.g., Cas9 or Cas12a) and guide RNA to target genes in the non-transgenic parent for modification. Next-generation sequencing (NGS) is typically used to determine whether new edits have occurred in the target genes of the non-transgenic parent (those produced in the F1 hybrid). Because the non-transgenic parent does not contain the desired CRISPR-Cas editing mechanism, the editing in the non-transgenic target genes may only occur in the F1 progeny. This assay allows not only the measurement of the frequency of this editing but also the approximate developmental timing at which the editing occurs during F1 embryonic development. Editing occurring early (e.g., at the 1-cell zygote stage) should produce pure biallelic results, with nearly 50% of reads being novel edit types not found in the transgenic parent, and the other 50% matching the edit results from the E0 transgenic parent generation. Editing occurring later (in the early multicellular embryo) should not produce such a high percentage of reads, but is more likely to show multiple edit results mixed with lower read ratios (i.e., it will be chimera or mosaic).
[0136] Then, a zygotic preferred promoter exhibiting an editing rate of 10% or higher can be selected for use in any context where the zygotic preferred expression is intended to be expressed. In a preferred embodiment, such a promoter is used in the HI-editing program.
[0137] I.VSP promoter
[0138] In some embodiments, the promoters used in the HI-editing procedure include VSP promoters that exhibit a zygotic editing rate of 10% or higher. In some embodiments, the VSP promoters are derived from endogenous genes in rice or maize, or are orthologous promoters. In some embodiments, the VSP promoters are derived from endogenous genes in Arabidopsis thaliana, sunflower, rice, tomato, or maize. For example, in some embodiments, the VSP promoter is a rice VSP promoter derived from LOC_Os09g09480 / Os09g0267600, which is expressed in both parents in the zygotic (Anderson et al. (2017) The Zygotic Transition Is Initiated in Unicellular Plant Zygotes with Asymmetric Activation of Parental Genomes. Dev Cell 43(3):349-358). In some embodiments, VSP is a promoter derived from the maize gene Zm00001d011353 / Zm00001eb358220. Further examples of orthologous VSP genes are provided in Table 1.
[0139] Table 1. Example orthologs of OsVSP and ZmVSP in different crops.
[0140]
[0141] The promoters and terminators of the genes listed above or orthologous genes (e.g., those that provide at least 10% zygotic editing efficiency) can serve as highly efficient HI-editing regulatory elements to drive robust expression of CRISPR enzymes and / or guide RNA in sperm and zygotic cell types.
[0142] In some embodiments, the VSP promoter comprises a sequence of any one of SEQ ID NO: 1-27 or a functional fragment thereof. In some embodiments, the VSP promoter comprises a functional fragment of any one of SEQ ID NO: 1-27, 86, or 87 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the VSP promoter comprises a functional fragment of any one of SEQ ID NO: 1-27, 86, or 87 with a length of at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides, or a length of at least 2000, 2100, 2200, 2300, 2400, or 2500 nucleotides. In some embodiments, the VSP promoter comprises a variant of a functional fragment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides or at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides in length, having at least 80% or at least 85%, 90%, or at least 95% identity with the corresponding segment of any one of SEQ ID NO: 1-27. In some embodiments, such a functional fragment has at least 96% or at least 97%, 98%, 99%, or greater identity with the corresponding segment of any one of SEQ ID NO: 1-27.
[0143] In some embodiments, the functional fragment of the VSP promoter lacks the 5' untranslated region sequence of the mRNA. The 5' untranslated region can be readily identified / determined using known techniques such as 5' RACE analysis. Additional functional fragments (e.g., deletions or variants of any of SEQ ID NO: 1-27) can be determined by evaluating the mutagenic and / or truncated regions of the promoter to identify those fragments exhibiting at least about 10% zygotic editing activity.
[0144] In some embodiments, the VSP promoter comprises SEQ ID NO: 27 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 27 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 27 with a length of at least 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 27 with a length of at least 1100, 1200, 1300, 1400, or 1500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 27 with a length of at least 1600 nucleotides or at least 1700, 1800, or 1900 nucleotides. In some embodiments, the VSP promoter includes a region having at least 90% or at least 95% identity with a segment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length as specified in SEQ ID NO: 27; or a region having at least 90% or at least 95% identity with a segment of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides in length as specified in SEQ ID NO: 27. In some embodiments, the VSP promoter does not include a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 27, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 27 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0145] In some embodiments, the VSP promoter comprises SEQ ID NO: 12 or a functional fragment thereof. In some embodiments, the fragment comprises a region of SEQ ID NO: 12 with a length of at least 100, 150, 200, 250, 300, 350, 400, or 450 nucleotides. In some embodiments, the functional VSP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 12 or a segment of SEQ ID NO: 12 with a length of at least 100, 150, 200, 250, 300, 350, 400, or 450 nucleotides. In some embodiments, the VSP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 12, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 12 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0146] In some embodiments, the VSP promoter comprises SEQ ID NO: 14 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 14 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 14 with a length of at least 600, 700, 800, 900, or 1000 nucleotides, or a length of at least 1100, 1200, 1300, 1400, or 1500 nucleotides. In some embodiments, the functional VSP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 14 or a segment of SEQ ID NO: 14 with a length of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional VSP promoter comprises a region having at least 90% or at least 95% identity with a segment of at least 1100, 1200, 1300, 1400, or 1500 nucleotides in length as SEQ ID NO: 14. In some embodiments, the VSP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 14, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 14 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0147] In some embodiments, the VSP promoter comprises SEQ ID NO: 11 or SEQ ID NO: 13, or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 11 or SEQ ID NO: 13 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 11 or SEQ ID NO: 13 with a length of at least 500, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 11 or SEQ ID NO: 13 with a length of at least 1100, 1200, 1300, 1400, or 1500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 11 or SEQ ID NO: 13 with a length of at least 1600 nucleotides or at least 1700, 1800, or 1900 nucleotides. In some embodiments, the VSP promoter includes a region having at least 90% or at least 95% identity with a segment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length, which is identical to or identical to the segment of SEQ ID NO: 11 or SEQ ID NO: 13; or a segment of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides in length, which is identical to or identical to the segment of SEQ ID NO: 11 or SEQ ID NO: 13. In some embodiments, the VSP promoter does not include a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 11 or SEQ ID NO: 13, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 11 or SEQ ID NO: 13 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0148] In some embodiments, the VSP promoter comprises SEQ ID NO: 26 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 26 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 26 with a length of at least 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 26 with a length of at least 1100, 1200, 1300, 1400, or 1500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 26 with a length of at least 1600 nucleotides or at least 1700, 1800, or 1900 nucleotides. In some embodiments, the VSP promoter includes a region having at least 90% or at least 95% identity with a segment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length as specified in SEQ ID NO: 26; or a region having at least 90% or at least 95% identity with a segment of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides in length as specified in SEQ ID NO: 26. In some embodiments, the VSP promoter does not include a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 26, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 26 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0149] In some embodiments, the VSP promoter comprises any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24, or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of at least 100, 200, 300, 400, or 500 nucleotides in length from any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24. In some embodiments, the functional fragment comprises a region of at least 500, 600, 700, 800, 900, or 1000 nucleotides in length from any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24. In some embodiments, the functional fragment comprises a region of at least 1100, 1200, 1300, 1400, or 1500 nucleotides in length, any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24. In some embodiments, the functional fragment comprises a region of at least 1600 nucleotides in length or at least 1700, 1800, or 1900 nucleotides in length, any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24. In some embodiments, the VSP promoter includes a region having at least 90% or at least 95% identity with any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24, or a region having a length of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides; or a region having a length of at least 95% identity with any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24; or a region having a length of at least 90% or at least 95% identity with any one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24. The segments of length of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides from any one of 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 1800, or 24 have at least 90% or at least 95% identity. In some embodiments, the VSP promoter does not contain a 5' untranslated region sequence.In some embodiments, the VSP promoter is a functional fragment of one of SEQ ID NO: 1, 3, 5, 6, 7, 8, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, or 24 that lacks one or more nucleotides from the 5' and / or 3' ends of the full-length sequence, wherein the functional fragment constitutes at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment constitutes at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0150] In some embodiments, the VSP promoter comprises SEQ ID NO: 2 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 2 with a length of at least 100, 200, 300, 400, or 450 nucleotides. In some embodiments, the functional VSP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 2 or a segment of SEQ ID NO: 2 with a length of at least 100, 200, 300, 400, or 450 nucleotides. In some embodiments, the VSP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 2, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 2 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0151] In some embodiments, the VSP promoter comprises SEQ ID NO: 4 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 4 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 4 with a length of at least 600, 700, or 800 nucleotides. In some embodiments, the VSP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 4 or a segment of SEQ ID NO: 4 with a length of at least 100, 200, 300, 400, or 500 nucleotides or at least 600, 700, or 800 nucleotides. In some embodiments, the VSP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 4, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 4 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0152] In some embodiments, the VSP promoter comprises SEQ ID NO: 9 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 9 with a length of at least 100, 200, 300, 400, 500, or 600 nucleotides. In some embodiments, the functional VSP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 9 or a segment of SEQ ID NO: 9 with a length of at least 100, 200, 300, 400, 500, or 600 nucleotides. In some embodiments, the VSP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 9, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 9 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0153] In some embodiments, the VSP promoter comprises SEQ ID NO: 25 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 25 with a length of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 25 with a length of at least 1100, 1200, 1300, 1400, 1500, 1600, or 1700 nucleotides. In some embodiments, the VSP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 25 or a segment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length, or a segment of at least 1100, 1200, 1300, 1400, 1500, 1600, or 1700 nucleotides in length. In some embodiments, the VSP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 25, wherein one or more nucleotides are deleted from the 5' end and / or 3' end of the full-length SEQ ID NO: 25 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0154] In some embodiments, the VSP promoter comprises SEQ ID NO: 86 or 87 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 86 or 87 with a length of at least 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 86 or 87 with a length of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides. In some embodiments, the functional fragment comprises a region with a length of at least 2000, 2100, 2200, 2300, 2400, or 2500 nucleotides. In some embodiments, the VSP promoter comprises a segment of at least 500, 600, 700, 800, 900, or 1000 nucleotides in length, corresponding to SEQ ID NO: 86 or 87; or a segment of at least 1100, 1200, 1300, 1400, 1500, 1600, 170, 1800, or 1900 nucleotides in length, corresponding to SEQ ID NO: 25; or a segment of at least 2000, 2100, 2200, 2300, 2400, or 2500 nucleotides in length, having at least 90% or at least 95% identity. In some embodiments, the VSP promoter is a functional fragment of SEQ ID NO: 86 or 87, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 86 or 87 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0155] II. RZDP promoter
[0156] In some embodiments, the zygotic preferred promoters disclosed herein (e.g., for HI-editing procedures) include RZDP promoters exhibiting about 10% or higher zygotic editing. In some embodiments, the RZDP promoter is derived from an endogenous maize gene or from an orthologous promoter. In some embodiments, the RZDP promoter is derived from an endogenous gene in rice, maize, Arabidopsis thaliana, soybean, tomato, or sunflower. For example, in some embodiments, the RZDP promoter is derived from the Zm00001d050090 gene. Further examples of orthologous RZDP genes are provided in Table 2.
[0157] Table 2. Example orthologs of ZmRZDP in different crops.
[0158]
[0159] The promoters and terminators of the genes listed above or orthologous genes (e.g., those with at least 10% zygotic editing efficiency) can serve as highly efficient HI-editing regulatory elements to drive high expression of CRISPR enzymes and / or guide RNA in sperm and zygotic cell types.
[0160] In some embodiments, the RZD promoter comprises the sequence of any one of SEQ ID NO: 28-43 or a functional fragment thereof. In some embodiments, the RZD promoter comprises a functional fragment of any one of SEQ ID NO: 28-43 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the RZD promoter comprises a functional fragment of any one of SEQ ID NO: 28-43 with a length of at least 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, or 1500 nucleotides. In some embodiments, the RZD promoter comprises a functional fragment of any one of SEQ ID NO: 28-43 with a length of at least 1600, 1700, 1800, or 1900 nucleotides. In some embodiments, the RZD promoter comprises a variant of a functional fragment of any one of SEQ ID NO: 28-43 with a length of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides, or a length of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides, which has at least 80% identity, or at least 85%, at least 90%, or at least 95% identity, with the corresponding segment of any one of SEQ ID NO: 28-43. In some embodiments, such a functional fragment has at least 96% identity, or at least 97%, 98%, 99%, or greater identity, with the corresponding segment of any one of SEQ ID NO: 28-43.
[0161] In some embodiments, the functional fragment of the RZDP promoter lacks the 5' untranslated region sequence of the mRNA. The 5' untranslated region can be readily identified / determined using known techniques such as 5' RACE analysis. Additional functional fragments (e.g., deletions or variants of any of SEQ ID NO: 28-43) can be determined by evaluating the mutagenic and / or truncated regions of the promoter to identify those fragments exhibiting at least about 10% zygotic editing activity.
[0162] In some embodiments, the RZD promoter comprises SEQ ID NO: 28 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 28 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 28 with a length of at least 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 28 with a length of at least 1100, 1200, 1300, 1400, or 1500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 28 with a length of at least 1600 nucleotides or at least 1700, 1800, or 1900 nucleotides. In some embodiments, the RZD promoter includes a region having at least 90% or at least 95% identity with a segment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length as specified in SEQ ID NO: 28; or a region having at least 90% or at least 95% identity with a segment of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides in length as specified in SEQ ID NO: 28. In some embodiments, the RZD promoter does not include a 5' untranslated region sequence. In some embodiments, the RZD promoter is a functional fragment of SEQ ID NO: 28, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 28 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0163] In some embodiments, the RZDP promoter comprises any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41, or 42, or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of at least 100, 200, 300, 400, or 500 nucleotides in length from any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41, or 42. In some embodiments, the functional fragment comprises a region of at least 600, 700, 800, 900, or 1000 nucleotides in length from any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41, or 42. In some embodiments, the functional fragment comprises a region of at least 1100, 1200, 1300, 1400, 1500 nucleotides or at least 1600, 1700, 1800 or 1900 nucleotides in length, any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41 or 42. In some embodiments, the RZDP promoter includes a region having at least 90% or at least 95% identity with any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41 or 42 or any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41 or 42. In some embodiments, the RZDP promoter includes a region having at least 90% or at least 95% identity with any one of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41, or 42, or with a length of at least 11, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides. In some embodiments, the RZDP promoter does not include a 5' untranslated region sequence. In some embodiments, the RZD promoter is a functional segment of SEQ ID NO: 29, 31, 32, 33, 35, 36, 37, 39, 40, 41, or 42 that lacks one or more nucleotides from the 5' and / or 3' ends of the full-length sequence, wherein the segment constitutes at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional segment constitutes at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0164] In some embodiments, the RZDP promoter comprises SEQ ID NO: 30 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 30 with a length of at least 100, 200, 300, 400, or 500 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 30 with a length of at least 600, 700, or 800 nucleotides. In some embodiments, the RZDP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 30 or a segment of SEQ ID NO: 30 with a length of at least 100, 200, 300, 400, 500, 600, 700, or 800 nucleotides. In some embodiments, the RZDP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the RZD promoter is a functional fragment of SEQ ID NO: 30, wherein one or more nucleotides are deleted from the 5' and / or 3' end of the full-length SEQ ID NO: 30 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0165] In some embodiments, the RZDP promoter comprises SEQ ID NO: 34 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of at least 100, 200, or 300 nucleotides in length of SEQ ID NO: 34. In some embodiments, the RZDP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 34 or a segment of at least 100, 200, or 300 nucleotides in length of SEQ ID NO: 34. In some embodiments, the RZDP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the RZD promoter is a functional fragment of SEQ ID NO: 34, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 34 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0166] In some embodiments, the RZDP promoter comprises SEQ ID NO: 38 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 38 with a length of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the fragment comprises a region of SEQ ID NO: 38 with a length of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, or 1800 nucleotides. In some embodiments, the RZD promoter comprises a region of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 nucleotides in length, or a region of at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, or 1800 nucleotides in length, having at least 90% or at least 95% identity with the region of SEQ ID NO: 38. In some embodiments, the RZDP promoter does not comprise a 5' untranslated region sequence. In some embodiments, the RZD promoter is a functional fragment of SEQ ID NO: 38, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 38 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0167] In some embodiments, the RZDP promoter comprises SEQ ID NO: 43 or a functional fragment thereof. In some embodiments, the fragment comprises a region of at least 1100 nucleotides in length as SEQ ID NO: 43. In some embodiments, the functional fragment comprises a region of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 1100 nucleotides in length as SEQ ID NO: 43. In some embodiments, the RZDP promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 43 or a segment of at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or 1100 nucleotides in length as SEQ ID NO: 43. In some embodiments, the RZDP promoter does not comprise a 5' untranslated region. In some embodiments, the RZD promoter is a functional fragment of SEQ ID NO: 43, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 43 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0168] III. SCE1 promoter
[0169] In some embodiments, the zygotic preferred promoters disclosed herein (e.g., for HI-editing procedures) include the SCE1 promoter, which exhibits about 10% or higher zygotic editing. In some embodiments, the SCE1 promoter is derived from an endogenous maize gene or from an orthologous promoter. In some embodiments, the SCE1 promoter is derived from an endogenous gene in rice, maize, Arabidopsis thaliana, soybean, tomato, or sunflower. For example, in some embodiments, the SCE1 promoter is derived from the Zm00001d002570 gene.
[0170] In some embodiments, the SCE1 promoter comprises the sequence of any one of SEQ ID NO: 83-85 or a functional fragment thereof. In some embodiments, the SCE1 promoter comprises a functional fragment of any one of SEQ ID NO: 83-85 with a length of at least 300, 400, or 500 nucleotides. In some embodiments, the SCE1 promoter comprises a functional fragment of any one of SEQ ID NO: 83-85 with a length of at least 600, 700, 800, 900, 1000, 1100, 1200, or 1300 nucleotides. In some embodiments, the SCE1 promoter comprises a functional fragment of SEQ ID NO: 84 or 85 with a length of 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides. In some embodiments, the SCE1 promoter comprises a functional fragment of SEQ ID NO: 84 or 85 with a length of 2000, 2100, 2200, 2300, or 2400 nucleotides; or (in some embodiments) a functional fragment of at least 2500, 2600, 2700, 2800, 2900, or 3000 nucleotides. In some embodiments, the SCE1 promoter comprises a sequence having at least 90%, 92%, 92%, 93%, or 94% identity with any one of SEQ ID NO: 83-85. In some embodiments, the SCE1 promoter comprises a sequence having at least 95%, 96%, 97%, 98%, or 99% identity with any one of SEQ ID NO: 83-85. In some embodiments, the SCE1 promoter comprises SEQ ID NO: 83, 84, or 85.
[0171] In some embodiments, the functional fragment of the SCE1 promoter lacks the 5' untranslated region sequence of the mRNA. The 5' untranslated region can be readily identified / determined using known techniques such as 5' RACE analysis. Additional functional fragments (e.g., deletions or variants of any of SEQ ID NO: 83-) can be determined by evaluating the mutagenic and / or truncated regions of the promoter to identify those fragments exhibiting at least about 10% zygotic editing activity.
[0172] In some embodiments, the SCE1 promoter comprises SEQ ID NO: 83 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 83 with a length of at least 500, 600, 700, 800, 900, or 1000 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 83 with a length of at least 1100 nucleotides or at least 1200 or 1300 nucleotides. In some embodiments, the SCE1 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 83 or a segment of SEQ ID NO: 83 with a length of at least 500, 600, 700, 800, 900, or 1000 nucleotides; or having at least 90% or at least 95% identity with a segment of SEQ ID NO: 83 with a length of at least 1100, 1200, or 1300 nucleotides. In some embodiments, the SCE1 promoter is a functional fragment of SEQ ID NO: 83, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 83 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0173] In some embodiments, the SCE1 promoter comprises SEQ ID NO: 84 or SEQ ID NO: 85, or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of at least 500, 750, or 1000 nucleotides in length, as defined in SEQ ID NO: 84 or SEQ ID NO: 85. In some embodiments, the functional fragment comprises a region of at least 1100, 1200, 1300, 1400, or 1500 nucleotides in length, as defined in SEQ ID NO: 84 or SEQ ID NO: 85; or a region of at least 1600, 1700, 1800, 1900, or 2000 nucleotides in length, as defined in SEQ ID NO: 84 or SEQ ID NO: 85. In some embodiments, the functional fragment comprises a region of at least 2100, 2200, 2300, 2400, or 2500 nucleotides in length, as defined in SEQ ID NO: 84 or SEQ ID NO: 85; or a region of at least 2600, 2700, 2800, 2900, 3000, or 3100 nucleotides in length, as defined in SEQ ID NO: 84 or SEQ ID NO: 85. In some embodiments, the SCE1 promoter includes a region having at least 90% or at least 95% identity with the functional fragment of SEQ ID NO: 84 or 85. In some embodiments, the SCE1 promoter is a functional fragment of SEQ ID NO: 84 or 85, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 84 or 85 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0174] IV. Expression constructs containing promoters
[0175] In another aspect, this disclosure provides expression constructs and transgenic plant cells comprising a zygotic preferred promoter (e.g., the VSP promoter or RZD promoter as described herein) to drive the expression of a target nucleic acid sequence in plants. In some embodiments, the target nucleic acid sequence is a nuclease, such as a nuclease.
[0176] In some embodiments, the expression construct may be a single vector encoding two or more target expression products. In some embodiments, the expression of one target product is driven by a zygotic preferred promoter as described herein, and the expression of a second target product is driven by a different promoter. In some embodiments, the expression system comprises a binary vector encoding two target expression products, for example, wherein the expression of the desired gene product (e.g., a nucleomodifying enzyme, such as a DNA editing nuclease) is driven by a zygotic preferred promoter, and the expression of the target RNA product (e.g., guide RNA) is driven by an RNA polymerase III promoter.
[0177] In some embodiments, the expression construct includes a zygotic-preferred promoter as described herein, which drives the expression of a zinc finger nuclease (ZFN). A ZFN is a fusion between a cleavage domain of FokI and a DNA recognition domain containing three or more zinc finger motifs. Examples of ZFNs include, but are not limited to, those described in the following literature: Urnov et al., Nature Reviews Genetics, 2010, 11:636-646; Gaj et al., Nat Methods, 2012. 9(8):805-7; U.S. Patent Nos. 6,534,261, 6,607,882, 6,746,838, 6,794,136, 6,824,978, 6,866,997, 6,933,113, 6,979,539, 7,013,219, 7,030,215, 7,220,719, 7,241,573, 7,241,574, 7,585,849, 7,595,376, 6,903,185, 6,479,626, and U.S. Application Publications Nos. 2003 / 0232410 and 2009 / 0203140.
[0178] In some embodiments, a zygotic-preferred promoter can be used to drive the expression of a TAL effector nuclease (TALEN). TALEN is an engineered transcription activator-like effector nuclease containing a central domain of a DNA-binding tandem repeat sequence, a nuclear localization signal, and a C-terminal transcription activation domain. TALEN can be generated by fusing the TAL effector DNA-binding domain with a DNA cleavage domain. For example, the TALE protein can be fused with a nuclease such as wild-type or mutant FokI endonuclease or the catalytic domain of FokI. Detailed descriptions of TALEN and its use in gene editing can be found in, for example, the following publications: U.S. Patent Nos. 8,440,431, 8,440,432, 8,450,471, 8,586,363, and 8,697,853; Scharenberg et al., Curr Gene Ther, 2013, 13(4):291-303; Gaj et al., Nat Methods, 2012, 9(8):805-7; Beurdeley et al., Nat Commun, 2013, 4:1762; and Joung and Sander, Nat Rev Mol Cell Biol, 2013, 14(1):49-55.
[0179] In some embodiments, a zygotic preferred promoter can be used to drive the expression of a wide range of nucleases. The wide range of nucleases can be modular DNA-binding nucleases, such as any fusion protein comprising at least one catalytic domain of a nuclease and at least one DNA-binding domain, or a protein specifying a nucleic acid target sequence. The DNA-binding domain may contain at least one motif recognizing single-stranded or double-stranded DNA. The wide range of nucleases can be monomeric or dimer. Detailed descriptions of useful, wide-ranging nucleases and their applications in gene editing can be found in, for example, the following publications: Silva et al., Curr Gene Ther [Current Gene Therapy], 2011, 11(1):11-27; Zaslavoskiy et al., BMC Bioinformatics [BMC Bioinformatics], 2014, 15:191; Takeuchi et al., Proc Natl Acad SciUSA [Proceedings of the National Academy of Sciences of the United States of America], 2014, 111(11):4061-4066; and US Patent Nos. 7,842,489; 7,897,372; 8,021,867; 8,163,514; 8,133,697; 8,021,867; 8,119,361; 8,119,381; 8,124,36; and 8,129,134.
[0180] In some embodiments, the expression construct may include a zygotic preferred promoter as described herein, which drives the expression of CRISPR nucleases used in the CRISPR editing system. In some embodiments, the Cas protein expressed under the control of the zygotic preferred promoter is Cas9, Cas12a (formerly known as Cpf1), Cas12b (formerly known as C2c1), Cas13a (formerly known as C2c2), C2c3, Cas13b, or a Cas protein or orthologous protein derived from prokaryotes. In some embodiments, the Cas protein is (modified) Cas9, such as (modified) Staphylococcus aureus Cas9 (SaCas9) or (modified) Streptococcus pyogenes Cas9 (SpCas9). In some embodiments, the Cas protein is Cas12a, optionally derived from an Acidaminococcus sp. species, such as Acidaminococcus sp. BV3L6 Cpf1 (AsCas12a), or from Lachnospiraceae bacteria Cas12a, such as Lachnospiraceae bacteria MA2020 or Lachnospiraceae bacteria MD2006 (LBCas12a). See U.S. Patent No. 10,669,540, which is incorporated herein by reference in its entirety. Alternatively, the Cas12a protein may be derived from Moraxella bovoculi AAX08_00205 [Mb2Cas12a] or Moraxella bovoculi AAX11_00205 [Mb3Cas12a]. See WO 2017 / 189308, which is incorporated herein by reference in its entirety. In some embodiments, the Cas protein is a (modified) C2c2, such as Leptotrichiawadei C2c2 (LwC2c2) or Listeria newyorkensis FSL M6-0635 C2c2 (LbFSLC2c2). In some embodiments, the (modified) Cas protein is C2c1. In some embodiments, the (modified) Cas protein is C2c3. In some embodiments, the (modified) Cas protein is Cas13b. Other Cas enzymes are available to those skilled in the art.
[0181] In some embodiments, the CRISPR nuclease further comprises a fusion domain, such as a deaminase, uracil DNA glycosylase, reverse transcriptase, or exonuclease. Examples of modified CRISPR nucleases include chimeric Cas proteins, such as dCas9-FokI, dCpf1-FokI, chimeric Cas9-cytidine deaminase, chimeric Cas9-adenine deaminase, nickase Cas9 (nCas9), chimeric dCas9 non-FokI nuclease, and dCpf1 non-FokI nuclease.
[0182] In some embodiments, the expression construct (such as a binary vector) comprises a first polynucleotide and a second polynucleotide, the first polynucleotide comprising a zygotic-preferred promoter operatively linked to a sequence encoding a DNA-modifying enzyme, and the second polynucleotide comprising an RNA polymerase III promoter operatively linked to a nucleic acid sequence encoding at least one guide RNA. In some embodiments, the zygotic-preferred promoter is a VSP or RZD promoter as described herein, for example, a VSP promoter or a functional fragment or variant thereof of any one of SEQ ID NO: 1-27, 86, or 87; or an RZD promoter or a functional fragment or variant thereof of any one of SEQ ID NO: 28-43. In some embodiments, the zygotic-preferred promoter is an SCE1 promoter as described herein, for example, an SCE1 promoter or a functional fragment or variant thereof of any one of SEQ ID NO: 83-85. In some embodiments, the second polynucleotide comprises a U3 promoter operatively linked to a sequence encoding one or more guide RNAs.
[0183] A. U3 bootloader
[0184] In some embodiments, the U3 promoter is derived from rice, wheat, maize, or Arabidopsis thaliana. In some embodiments, the U3 promoter driving the expression of one or more guide RNAs comprises a sequence having at least 70% or at least 75%, 80%, or 85% identity with the U3 promoter sequence of any one of SEQ ID NO: 44-49. In some embodiments, the U3 promoter sequence has at least 90% or at least 95% identity with any one of SEQ ID NO: 44-49. In some embodiments, the U3 promoter sequence has at least 96%, 97%, 98%, or 99% identity with any one of SEQ ID NO: 44-49, or comprises any one of SEQ ID NO: 44-49.
[0185] Functional fragments (e.g., deletions or variants of any of SEQ ID NO: 44-49) can be determined by evaluating the mutagenic and / or truncated regions of the promoter to identify those fragments that exhibit at least about 50% or at least about 70% promoter activity compared to the full-length promoter sequence.
[0186] In some embodiments, the U3 promoter comprises SEQ ID NO: 44 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of at least 100, 150, 200, 250, or 300 nucleotides in length of SEQ ID NO: 44. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 44 or a segment of at least 100, 150, 200, 250, or 300 nucleotides in length of SEQ ID NO: 44. In some embodiments, the U3 promoter is a functional fragment of SEQ ID NO: 44, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 44 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0187] In some embodiments, the U3 promoter comprises SEQ ID NO: 45 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of at least 100, 150, 200, 250, 300, or 350 nucleotides in length of SEQ ID NO: 45. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 45 or a segment of at least 100, 150, 200, 250, 300, or 350 nucleotides in length of SEQ ID NO: 45. In some embodiments, the U3 promoter is a functional fragment of SEQ ID NO: 45, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 45 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0188] In some embodiments, the U3 promoter comprises SEQ ID NO: 46 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 46 with a length of at least 100, 150, 200, 250, 300, or 350 nucleotides. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 46 or a segment of SEQ ID NO: 46 with a length of at least 100, 150, 200, 250, 300, or 350 nucleotides. In some embodiments, the U3 promoter is a functional fragment of SEQ ID NO: 46, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 46 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0189] In some embodiments, the U3 promoter comprises SEQ ID NO: 47 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 47 with a length of at least 100, 150, 200, 250, 300, or 350 nucleotides. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 47 or a segment of SEQ ID NO: 47 with a length of at least 100, 150, 200, 250, 300, or 350 nucleotides. In some embodiments, the U3 promoter is a functional fragment of SEQ ID NO: 47, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 47 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0190] In some embodiments, the U3 promoter comprises SEQ ID NO: 48 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 48 with a length of at least 200, 250, 300, 350, 400, or 450 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 48 with a length of at least 500, 550, 600, 650, 700, 750, 800, or 850 nucleotides. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 48 or a segment of SEQ ID NO: 48 with a length of at least 200, 250, 300, 350, 400, or 450 nucleotides. In some embodiments, the U3 promoter includes a region having at least 90% or at least 95% identity with SEQ ID NO: 48 or a segment of at least 500, 550, 600, 650, 700, 750, 800, or 850 nucleotides in length. In some embodiments, the U3 promoter is a functional fragment of SEQ ID NO: 48, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 48 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0191] In some embodiments, the U3 promoter comprises SEQ ID NO: 49 or a functional fragment thereof. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 49 with a length of at least 200, 250, 300, 350, 400, or 450 nucleotides. In some embodiments, the functional fragment comprises a region of SEQ ID NO: 49 with a length of at least 500, 550, 600, 650, 700, or 750 nucleotides. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 49 or a segment of SEQ ID NO: 49 with a length of at least 200, 250, 300, 350, 400, or 450 nucleotides. In some embodiments, the U3 promoter comprises a region having at least 90% or at least 95% identity with SEQ ID NO: 49 or a segment of SEQ ID NO: 49 with a length of at least 500, 550, 600, 650, 700, or 750 nucleotides. In some embodiments, the U3 promoter is a functional fragment of SEQ ID NO: 49, wherein one or more nucleotides are deleted from the 5' and / or 3' ends of the full-length SEQ ID NO: 49 sequence, wherein the functional fragment is at least 90%, at least 91%, at least 92%, at least 93%, or at least 94% of the length of the full-length sequence. In some embodiments, the functional fragment is at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% of the length of the full-length sequence.
[0192] B. Other expression construct regulatory elements
[0193] Expression constructs as described herein may contain additional regulatory elements. As used herein, a “regulatory element” includes any sequence that affects the construction or function of the expression construct, such as affecting transcription and / or translation in cells in which the expression construct is expressed. Such sequences include transcriptional or translational enhancers, additional promoters or promoter elements that drive the expression of other gene products encoded by the construct, such as genes encoding selectable markers. Additional regulatory elements include introns and transcription terminators. Such regulatory elements may be endogenous or heterologous to the host cell, or these regulatory elements may be endogenous or heterologous to each other.
[0194] A variety of transcription terminators are available for use in expression constructs. These terminators are responsible for transcription termination outside of transgenes and proper mRNA polyadenylation. Termination regions can be naturally associated with transcription initiation regions, naturally associated with operatively linked target DNA sequences, naturally associated with plant hosts, or possibly derived from another source (i.e., foreign or heterologous to the promoter, target DNA sequence, plant host, or any combination thereof). Suitable transcription terminators are those known to function in plants and include the CAMV pSOY1 terminator, the tml terminator, the carmine synthase terminator, and the pea rbcs E9 terminator. These terminators can be used in both monocotyledonous and dicotyledonous plants. Additionally, natural transcription terminators of the gene can be used. Termination regions used in expression cassettes can be obtained from Ti-plasmids of, for example, Agrobacterium tumefaciens, such as octopaline synthase and carmine synthase terminators. See also Guerineau et al. (1991) Mol. Gen. Genet. [Molecular Genetics and General Genetics] 262: 141-144; Proudfoot (1991) Cell [Cell] 64:671-674; Sanfacon et al. (1991) Genes Dev. [Genes and Development] 5: 141-149; Mogen et al. (990) Plant Cell [Plant Cell] 2: 1261-1272; Munroe et al. (1990) Gene [Genes] 91: 151-158; Ballas et al. (1989) Nucleic Acids Res. [Nucleic Acid Research] 17:7891-7903; and Joshi et al. (1987) Nucleic Acid Res. [Nucleic Acid Research] 15:9627-9639.
[0195] Table 3 provides illustrative expression constructs. Table 3 describes additional vector configurations and regulatory elements provided in constructs 23396, 23397, 23399, 24520, 26258, 26296, 27145, 27146, 27226, 272234, 27241, 27680, 28255, 28291, 28292, 28293, 28510, 28520, 28560, 28825, and 28834. In some embodiments, the expression construct used in the HI-editing program comprises the binary vector shown in Table 3, which has a CRISPR / Cas driver promoter, CRISPR / Cas, and terminator; and a driver promoter, gRNA target gene, and gRNA sequence for the construct. In some embodiments, the construct includes the carrier component of construct 27145 or 27146, as listed in Table 3. In some embodiments, the expression construct used in the HI-editing program includes SEQ ID NO: 68 or SEQ ID NO: 69 or a variant thereof, the variant having at least 75%, at least 80%, or at least 85% identity with SEQ ID NO: 68 or SEQ ID NO: 69. In some embodiments, the expression construct sequence has at least 90% or at least 95% identity with SEQ ID NO: 68 or SEQ ID NO: 69.
[0196] V. Gene Editing Programs
[0197] In another aspect, this disclosure provides a method for performing gene editing to obtain progeny plants with a desired genotype. In some embodiments, the method is HI-editing. As indicated above, the HI-editing method is detailed in WO 2018 / 102816. Briefly, in this disclosure, pollen from a first plant is pollinated to a second plant, which is then transformed with any synthetic expression construct of this disclosure to express a DNA-modifying enzyme (e.g., a gene-editing nuclease) under the control of a zygotic preferred promoter. In some embodiments, the promoter is a VSP promoter comprising a sequence or a functional fragment thereof of any one of SEQ ID NO: 1-27, 86, or 87. In some embodiments, the promoter has at least 70%, 75%, 80%, or 85% identity or at least 90% or at least 95% identity with any one of SEQ ID NO: 1-27, 86, or 87. In some embodiments, the promoter comprises a variant of a functional fragment of at least 100, 200, 300, or 500 nucleotides, or at least 600, 700, 800, 900, or 1000 nucleotides, or at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides, which has at least 70%, 75%, 80%, or 85% identity with the corresponding segment of any one of SEQ ID NO: 1-27, 86, or 87, or at least 90% or at least 95% identity with the corresponding segment of any one of SEQ ID NO: 1-27, 86, or 87. In some embodiments, the VSP promoter comprises SEQ ID NO: 1, or an orthologous promoter of SEQ ID NO: 1. In some embodiments, the VSP promoter comprises SEQ ID NO: 27 or a functional fragment thereof. In some embodiments, the VSP promoter comprises SEQ ID NO: 86 or 87 or a functional fragment thereof (e.g., at least 2100, 2200, 2300, 2400, or 2500 nucleotides in length). In some embodiments, the promoter is an RZD promoter comprising the sequence of any one of SEQ ID NO: 28-43 or a functional fragment thereof. In some embodiments, the promoter has at least 70%, 75%, 80%, or 85% identity with any one of SEQ ID NO: 28-43, or at least 90% or at least 95% identity.In some embodiments, the RZD promoter comprises a variant of a functional fragment of at least 100, 200, 300, or 500 nucleotides, or at least 600, 700, 800, 900, or 1000 nucleotides, or at least 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, or 1900 nucleotides, which has at least 70%, 75%, 80%, or 85% identity with the corresponding segment of any one of SEQ ID NO: 28-43, or at least 90% or at least 95% identity with the corresponding segment of any one of SEQ ID NO: 28-43. In some embodiments, the RZD promoter comprises SEQ ID NO: 28 or a functional fragment thereof, or an orthologous promoter of SEQ ID NO: 28. In a typical embodiment, the first plant is also genetically modified to express one or more gRNA sequences of the target gene to be edited in the second plant. In some embodiments, the promoter is an SCE1 promoter comprising the sequence of SEQ ID NO: 83, 84, or 85 or a functional fragment thereof. In some embodiments, the promoter has at least 70%, 75%, 80%, or 85% identity or at least 90% or at least 95% identity with SEQ ID NO: 83, 84, or 85 or a functional fragment thereof (e.g., a fragment of SEQ ID NO: 83 with a length of at least 1100, 1200, or 1300 nucleotides); or a fragment of SEQ ID NO: 84 with a length of at least 2700, 2800, 2900, 300, or 3100 nucleotides; or a fragment of SEQ ID NO: 85 with a length of at least 2800, 2900, 300, 3100, 3200, or 3300 nucleotides.
[0198] After pollinating the second plant, the edit can be detected using well-known techniques, and progeny with the desired edit can be selected. In the HI-editing method, the second plant may contain the genomic DNA to be edited. In some embodiments, the edited progeny plants are haploid progeny plants. Haploid progeny contain the genome of the second plant but not the genome of the first plant.
[0199] In some embodiments, the first plant is a haploid inducible line. The haploid inducible line can be a paternal haploid inducible line, for example, having a tail exchange mutation in the CENH3 gene (see, for example, Ravi and Chan, Nature 464:615-618, 2010) or another CENH3 mutation (see, for example, Maheshwari et al., Genome Research 27(3), 471-478, 2017). In other embodiments, the haploid inducible line can be a maternal haploid inducible line, for example, a line containing a mutation in the gene encoding the Patatin-like phospholipase A2α (also referred to herein as the MATRILINEAL (MATL) gene). In some embodiments, the MATL gene contains a loss-of-function mutation. In some embodiments, the haploid inducible line can be an ig-type haploid inducible line caused by a mutation in the INDETERMINATE GAMETOPHYTE 1 gene.
[0200] The zygotic preferred promoter of the present invention can be used to modify any monocot or dicot to express the target nucleic acid. In some embodiments, the plant is corn, wheat, rice, barley, oats, turfgrass, brassica, tomato, pepper, lettuce, eggplant, soybean, sunflower, beet, cotton, alfalfa, tobacco and many other plants.
[0201] Example I. Promoter Mining and Selection
[0202] To identify promoters with sperm- and / or zygotic-specific expression patterns, promoters were searched and selected from maize and rice transcriptome databases for this study. The putative sperm-specific ZmDUO1A promoter and two constitutively active promoters, prSoUBI4 (sugarcane polyubiquitin 4) and prOsACT1 (rice actin 1, i.e., LOC_Os03g50885), were selected. The prSoUBI4 promoter was used as a positive control because it induced at least 3% haploid editing (“HI-editing efficiency”) when operatively linked to the CRISPR-Cas9 enzyme in five or six maize lines tested [Kelliher, T et al. (2019) One-step genome editing of elite crop germplasm during haploid induction. Nature Biotech 37: 287–292; see also U.S. Patent No. 10,285,348], using a series of guide RNA (gRNA) targets driven by the rice U3 promoter prOsU3. In this disclosure, the combination of prSoUBI4:cCas9 and prOsU3:gRNA also induced editing of haploid wheat embryos generated via distant hybridization with maize pollen—indicating strong expression of CRISPR-Cas system components (gRNAs and proteins) in sperm cells and / or zygotes before, during, and / or after fertilization.
[0203] Although the HI-editing performance of prZmDUO1A and prOsAct1 during HI-editing has not been tested, AtDUO1A is specifically expressed in sperm cells and activates genes expressed in Arabidopsis sperm [Borg, M et al. (2011) The R2R3 MYB Transcription Factor DUO1 Activates a Male Germline-Specific Regulon Essential for Sperm Cell Differentiation in Arabidopsis Plant Cell. 23(2):534–549], and OsAct1 expression in 2.5-hour-old rice zygotes is at the 98th percentile of the genes studied [Anderson SN et al. (2017) The Zygotic Transition Is Initiated in Unicellular Plant Zygotes with Asymmetric Activation of Parental Genomes]. [Initiation of zygotic transition in zygotes of single-celled plants is accompanied by asymmetric activation of the parental genome]. Dev Cell [Developmental Cell] 43(3):349-358. Natural maize expression of prSoUBI4 is unavailable because it is derived from sugarcane, but the homolog ZmUbi1 (Zm00001d015327) is at the 99th percentile of genes expressed in maize sperm and 12-hour zygotes [Chen J et al. (2017) Zygotic Genome Activation Occurs Shortly after Fertilization in Maize]. Plant Cell [Plant Cell] 29(9):2106-2125.
[0204] In addition to these well-characterized promoters, we identified three novel promoters that had not previously been characterized in any transgenic context and continued to test whether their HI-editing efficiency could be higher than that of prSoUBI4. First, ZmVSP (vacuole sorting protein; Zm00001d011353 / Zm00001eb358220) was identified as a maize ortholog of a rice gene (LOC_Os09g09480 / Os09g0267600) expressed in zygotes by both parents [Anderson SN et al. (2017) The Zygotic Transition Is Initiated in Unicellular Plant Zygotes with Asymmetric Activation of Parental Genomes. Dev Cell 43(3):349-358]. Expression data from this study showed that the OsVsp transcript was highly expressed in rice zygote cells 2.5 hours after pollination, and that its transcripts were produced by chromosomes derived from both sperm and egg cells. This is noteworthy because most gene transcripts are maternally inherited 2.5 hours after pollination. Maize RNA-seq data [Chen J et al. (2017) Zygotic Genome Activation Occurs Shortly after Fertilization in Maize. Plant Cell 29(9):2106-2125] showed that the ZmVsp transcript accumulated to high levels of expression in maize egg cells and to extremely high levels in maize sperm cells and zygotes; Vsp expression was at the 75th, 99th, and 97th percentiles of all expressed genes in maize eggs, sperm cells, and zygotes, respectively. Therefore, the prVSP expression profiles in rice and maize conform to the expected pattern of male HI-editing promoters, i.e., high expression in male gametes and early zygotes, derived from paternally inherited chromosomes. Table 1 provides illustrative gene names of VSP homologs. Promoters and terminators from these genes can also serve as highly efficient HI-editing regulatory elements to drive high expression of CRISPR enzymes and / or guide RNA in sperm and zygotic cell types.
[0205] Besides VSP, the transcript of the maize cyclic zinc finger domain protein gene (“ZmRZDP”, Zm00001d050090) was at the 93rd and 98th percentiles of genes expressed in maize sperm and 12-hour zygotes, respectively [Chen J et al. (2017) Zygotic Genome Activation Occurs Shortly after Fertilization in Maize. Plant Cell 29(9):2106-2125]. According to the public maize gene map study [Hoopes GM et al. (2019) An updated geneatlas for maize reveals organ-specific and stress-induced genes. Plant J 97(6): 1154-1167], ZmRZDP expression was not detected in the shoot apical meristems of V1 and V3. Therefore, the expression profile of prZmRZDP closely matches the expected expression pattern of male HI-editing promoters. However, zygotic expression may not be derived from the paternal chromosome, as the OsRZDP ortholog is expressed as a maternal transcript [Anderson SN et al. (2017) The Zygotic Transition Is Initiated in Unicellular Plant Zygotes with Asymmetric Activation of Parental Genomes. DevCell [Developmental Cell] 43(3):349-358]. This has not yet been tested in maize. Table 2 shows illustrative orthologs of ZmRZDP. Promoters and terminators from these genes can also serve as highly efficient HI-editing regulatory elements to drive high expression of CRISPR enzymes and / or guide RNA in sperm and zygotic cell types.
[0206] Furthermore, by mining the mRNASeq dataset to search for genes moderately / highly expressed in sperm cells and early zygotes (preferably from paternal alleles, a characteristic previously unknown for this gene), we discovered SCE1, characterized as SUMO conjugate 1 (“ZmSCE1”, Zm00001d002570) in maize. The mined data showed that the SCE1 transcript was at the 99th percentile of genes expressed in both maize sperm and zygotes. We tested two versions of the promoter. The short version (SEQ ID NO: 83) was not expressed in pollen. The long version (SEQ ID NO: 85) was expressed in pollen, and, intriguingly, the ZER was higher when expressed from the female side of the zygote than from the male side.
[0207] II. Carrier Construction and Transformation
[0208] Binary vectors were constructed using coding sequences for CRISPR / Cas enzymes driven by constitutive, sperm-specific, or zygotic-specific promoters. The gRNA cassette was driven by the rice U6 or U3 promoter. This included phosphomannose isomerase (PMI) cassettes [Negrotto, D et al. (2000) The use of phosphomannose isomerase as a selectable marker to recover transgenic maize plants (Zea mays L.) via Agrobacterium transformation], Plant Cell Reports. 19, 798–803. In most vectors, the CRISPR enzyme used was LbCas12a, optimized for maize (vectors 26296, 27145, 27146, and 27234). All Cas12a vectors use an optimized long adapter (6x GGGGS) between the SV40 NLS and LbCas12a sequence at the N-terminus, and another adapter (GSPKK KRKVS GGSSG GSPKK KRKV) with two SV40 NLS sequences at the C-terminus. Vectors 26296, 27145, 27146, and 27234 also contain a GLOSSY2 (ZmGL2) DNA donor with 400 bp homologous arms on each side of the target site for inducing homologous recombination. The donor's flanks are gRNA cleavage sites, allowing it to be released from the T-DNA.
[0209] The Cas9 gRNA targets are ZmVLHP1 / 2-01 (5'-GCAGGAGGCGTCGAGCAGCG-3') (in vector 23396), ZmVLHP1 / 2-02 (5'-GCTGGAGCTGAGCTTCCGGG-3') (in vector 23397), ZmGW2-1 / 2 (5'-AAGCTCGCGCCCTGCTACCC-3') (in vector 23399), and ZmSBEIIb (5'-ATTGATAGAGCACATGAGCT-3') (in vector 24520). Cas12a The gRNA targets are ZmGL2 (5'-GTCACAGATCACAAACTTCAAATG-3') (in vectors 26296, 27226, 27145, 27146, and 27234), ZmWaxy1 (5'-GGGAAAGACCGAGGAGAAGATCT-3') (in vectors 26258, 27226, 27145, 27146, 27234, and 27241), ZmO2-01 (targeting the OPAQUE2 gene) (5'-CTGTATCTCGAGCGTCTGGCTGA-3') (in vector 26258), and ZmYellowEndosperm1 (5'-CTATCTT). ATCCTAAAGATGGTGG-3' (in vector 26258), ZmUBL (ubiquitin ligase) (5'-GGAAGGAAAAGGTATCTGAAGG-3') (in vectors 26258 and 27241), ZmUPL3 (ubiquitin ligase) (5'-GGAGGGAAAAGGTGTCTGAGGC-3') (in vectors 26258 and 27241), ZmO2-02 (5'-GGGCGCCTGAGCAACAAGAGTTC-3') (in vector 27241), and ZmO2-03 (5'-CTCACTCTTTCCTCGGTAG-3') (in vector 27241) (Table 3).
[0210]
[0211] To generate transgenic events, the inbred line NP2222 and a novel inbred haploid inducer derived from material 20BD917233 were grown in a greenhouse at a 16:8 photoperiod (light:dark) and 26°C / 16°C (day / night). Material 20BD917233 is an F7 stage inbred line derived from the parental breeding population SYN-INBB23 x RWKS / Z21S / / RWKS. Ears were harvested 10 or 11 days after pollination, husks and silks were removed, and the ears were sterilized with sodium hypochlorite solution. Immature embryos were isolated. Transformation was performed using the vectors listed in Table 3 as described in the published protocol [Zhong, H et al. (2018) Advances in Agrobacterium-mediated Maize Transformation. Methods Mol Biol 1676:41-59]. The zygosity and editing checks of the transgenic E0 events were characterized by TaqMann assay. Select single-copy and double-copy events (Table 4) and send them to the greenhouse for further growth.
[0212] Table 4. E0 events sent to GH for cross-pollination (nt = untested)
[0213] III. The zygotic editing and HI-editing tests using the prSoUbi4 and prOsU6 combination show the correlation between the zygotic editing test and the HI-editing rate.
[0214] The transformation event was acclimatized in a growth chamber for 1-2 weeks, after which the seedlings were transplanted into pots in a greenhouse equipped with high-pressure sodium (HPS) lamps as supplemental light for plant growth. Growth conditions were as described in Table 9. The tassels and ears of the E0 plants were bagged before pollen shedding and silking. The NP2222 plants and other inbred lines of the test species were also grown under the same conditions, with the ears and tassels bagged for intercrossing.
[0215] Table 5. Greenhouse conditions for E0 plant growth.
[0216]
[0217] In this study, to assess the potential HI-editing efficiency of different promoters, a rapid substitution assay was initially used, followed by the selection of promoters and events to be tested for HI-editing efficiency. The purpose of this substitution assay was to rapidly assess the potential of promoters to drive good HI-editing rates while avoiding the inherent problem of reduced sample size when assessing haploid editing results, since haploids are always a minority in the offspring produced by haploid-induced hybridization. By evaluating the relative chimerism of new edits in diploid F1 progeny after crossbreeding transgenic E0 parents with non-transgenic lines, the relative efficiency of promoters for early zygotic editing could be assessed in the absence of haploid induction. For this purpose, next-generation sequencing (NGS) was used to determine whether new edits occurred in the target genes of the non-transgenic parents (those produced in F1 hybrids). Because the non-transgenic parents do not contain the required CRISPR-Cas editing mechanism, edits in the non-transgenic target genes may only occur in F1 progeny. Furthermore, this substitution assay not only allows for the measurement of the frequency of this editing but also allows for the measurement of the approximate developmental time of the edit during F1 embryonic development. Editing occurring early (e.g., at the 1-cell zygote stage) should produce pure biallelic results, with nearly 50% of reads being novel edit types not found in the transgenic parent, and the other 50% matching the edit results from the E0 transgenic parent generation. Editing occurring later (in the early multicellular embryo) should not produce such a high percentage of reads, but is more likely to show multiple edit results mixed with lower read ratios (i.e., it will be chimera or mosaic).
[0218] To test zygotic editing rates, we performed crossbreeding and editing pattern analysis on five T1 events from vector 26258 (see Table 4, except for PLANTIE31) (females (4) or males (1) used as test species SYN-INBG78). F1 ears harvested 17 days post-pollination were dehulled, silks removed, and kernel crowns cut off. Individual embryos (approximately 3–7 mm) were isolated and placed in 96-well plates for genomic DNA extraction. T-DNA presence and target site mutations were analyzed using TaqMan probes. Editing rates were low in most events for O2 and YellowEndosperm1, but >50% for Waxy1, UBL, and UPL3; therefore, the edited genotypes of these three target sites were determined by NGS (next-generation sequencing). Parental edits at the target sites were found to be consistently present in F1 offspring with read abundances of approximately 50%. Zygotic editing was defined as the discovery of novel edits with read abundances >30% in addition to parental edits. According to this standard, the average zygotic editing rate of the three targets was 17% (18 / 105), with Waxy1 gRNA having the highest zygotic editing rate (40%) (Table 6).
[0219] Table 6. Editing rates of F1 zygotes from three targets of vector 26258.
[0220]
[0221] In contrast, using those events as male pollen donors, the average F1 zygotic editing rate for events PLANTHIE23, PLANTHIE24, and PLANTHIE25 (23396, 23397, and 23399, respectively) was 56% (55 / 99) (Table 7).
[0222] Table 7. F1 zygote editing rates from three target sites of vectors 23396, 23397 and 23399.
[0223]
[0224] The PLANTIE25 lineage has two CRISPR T-DNA insertions, which may increase the editing rate.
[0225] Variations in zygotic editing rates can be attributed to the efficiency of the promoter, Cas enzyme, or guide RNA used, germplasm, environmental factors, or a combination of these factors. To determine whether zygotic editing alternative assays run as reliable predictors of HI-editing efficiency, we attempted to compare the HI-editing efficiency of 26258 with three Cas9 constructs that expressed Cas9 using prSoUbi4-04 and their guide RNA using prOsU3. These three vectors were originally reported in HI-editing publications [Kelliher, T et al. (2019) One-step genome editing of elite cropgermplasm during haploid induction. Nature Biotech 37: 287–292], with average HI-editing efficiencies of approximately 3%–6% in five of the six maize lines tested. If the F1 zygotic editing assay is a good predictor of the HI-editing rate, we would expect the HI-editing rate of 26258 to be about one-third of that of the disclosed material, since the F1 zygotic editing rate of 26258 (17%) is about one-third of that of the three vectors in the disclosed material (56%).
[0226] To test the HI-editing efficiency of two events carrying vector 26258, CRISPR-transgenic homozygous T2 generation plants were crosscrossed as males (pollen donors) onto ears of test inbred lines from the sturdy stem line NP2222 and the non-sturdy stem line SYN-INBG78. Haploids were color-sorted, and the haploid induction rate was approximately 15.5%. Haploids were submitted for molecular analysis, and edited haploids were determined based on the results of Taqman assays targeting the editing target sites. This assay detected the WT sequence; mutations that could not amplify or bind to the Taqman probe resulted in a zero copy number for the haploid. Therefore, the edited haploids were those with a zero copy number for the target gene. In Table 8, the average haploid editing rate for the Waxy1 target site was approximately 0.8%. This low HI-editing rate is approximately one-third to one-quarter of the HI-editing rate in the original publication, which corresponds well to one-third to one-quarter of the editing rate in the F1 zygotes of 26258. Therefore, the zygotic editing rate may be a useful measure for predicting HI-editing efficiency.
[0227] Table 8. HI-editing rates of two events (PLANTHIE30 and PLANTHIE31) at the Waxy1 target site in the 20BD917233 line carrying vector 26258 (obtained from two test strains).
[0228]
[0229] Several potential variables could explain the lower HI-editing rate in 26258: different promoters used to express the guide RNA (26258 used prOsU6, while the original vector used prOsU3), different guide RNAs used, different Cas enzymes (26258 used Cas12a; the original vector used Cas9), different test species and inducible lines used in the experiment, differences in environmental conditions, or some or all of these variables. In any case, based on this comparison, F1 zygote assays appear to be a good surrogate indicator of HI-editing efficiency.
[0230] We suspected that the key factor driving F1 zygote and HI-editing rates was the choice of promoter used to express the CRISPR mechanism and guide RNA. To test the guide RNA expression factor, a new vector, 27241, was constructed, which, compared to vector 26258, has the same promoter driving the Cas12a enzyme and the same or similar gRNA, but in which we replaced the prOsU6 promoter with prOsU3. Homozygous T2 generation plants (males, pollen donors) were crosscrossed onto female ears of inbred lines from test species with different genetic backgrounds in maize, and haploid editing rates across test species were determined. The results indicated that prOsU3 and prOsU6 did not differ significantly in driving zygote editing and HI-editing.
[0231] IV. E0 cross-crossing and editing analysis of novel sperm, zygote, and constitutive promoters.
[0232] Select T-DNA positive F1 (hybrid embryos with editing mechanisms) and perform NGS to check the editing type (as shown in Table 10).
[0233] Table 9. Analysis of NGS data of F1 zygotes after E0 crossover.
[0234]
[0235] Here we can see the high F1 zygotic editing rate of the positive control vector 27680 and the high zygotic editing rate of vector 27146, with a zygotic editing rate of 90% when used as a male. This vector has prZmVSP (Zm00001d011353) driving Cas12a and prOsU3 driving gRNAs targeting GL2 and Waxy1. Similarly, vector 27145 (prZmRZDP) exhibits high F1 zygotic editing in both male and female cases compared to 27226 (rice actin 1) and 27234 (prZmDUO1A). This demonstrates that these two newly discovered promoters are more valuable than the previously thought to have high constitutive activity and be highly expressed in zygotes (rice actin 1) and highly expressed in sperm cells (prZmDUO1A).
[0236] V. E1 heterocrossing and Hi-editing rates of new sperm, zygotes, and constitutive promoters
[0237] E1 seeds were sown from haploid inducible lines transformed with constructs containing novel sperm, zygotes, and constitutive promoters. Transgenic homozygous E1 plants were selected and crossed with WT test strains, and haploids were identified by colorimetric observation during early embryonic development. Immature haploid embryos were then characterized by TaqMan and Sanger sequencing. Haploid editing rates are presented in Table 10.
[0238] Table 10. Haploid induction rate and haploid editing rate: This table summarizes the two inducers and different promoter combinations as well as gRNAs.
[0239]
[0240] Detailed results of the tests performed on prVSP (27146), prRZDP (27145), and prSoUbi4 (27680) are shown in Table 11, including F1 zygotic editing rates (testing diploids) and HI-editing rates (testing haploids). The results in Table 11 are broken down by event. This data indicates significant variability between events: one or two events for each vector showed much higher HI-editing rates than other events, indicating that T-DNA insertion sites can play a role in HI-editing efficiency. Importantly, for each of the following examples, F1 zygotic data predict HI-editing efficiency. For example, in construct 27146 (which corresponds to prVSP), according to Taqman and NGS data, event PLANTIE36- showed a significantly higher HI-editing rate in both test species, and the F1 zygotic editing rate was also the highest for any event in this construct (75% and 43% in test species 8 and 2, respectively). In 27145, event PLANTIE37 showed the highest HI-editing rate and the highest F1-zygosity rate in both the test species and the gRNA target site. In 27680, event PLANTIE38 had the highest HI-editing rate and F1-zygosity rate. This correlation suggests that F1 zygosity assay is a viable surrogate indicator for HI-editing. The 27146 VSP promoter results also showed the potential for this promoter to drive an exceptionally high HI-editing rate (compared to the prSoUbi4 standard control) when the T-DNA insertion site is correct.
[0241]
[0242]
[0243] VI. Efficient zygotic editing through maternal hybridization.
[0244] Although the zygotic-selected promoters were initially used for paternal expression in maize sperm, we also tested their applicability to maternal expression in maize oocytes. For this purpose, T2 seeds of event NP3003RS (homozygous for Cas12a and guide RNA) were sown and then used as females to cross with pollen from test species 1. The resulting F1 embryos were isolated and genotyped to detect zygotic editing. The table below shows the zygotic editing rate (“ZER”) results.
[0245] Table 12. Promoter zygotic editing rate (“ZER”) in hybridization.
[0246]
[0247] "F1#" refers to the number of F1 zygotes obtained from hybridization.
[0248] Clearly, the promoters prZmRZDP, prZmVSP, and prSCE1 can enable efficient zygotic editing from the female side.
[0249] VII. Construct Annotations
[0250] Table 13-33 provides the construct characteristics. The terms "minimum" and "maximum" in the table refer to the positions of the first and last nucleotides inserted in the construct, respectively.
[0251] Table 13. Construct 23396
[0252] Table 14. Construct 23397
[0253] Table 15. Construct 23399
[0254] Table 16. Construct 24520
[0255] Table 17. Construct 26258
[0256] Table 18. Construct 26296
[0257] Table 19. Construct 27145
[0258] Table 20. Construct 27146
[0259] Table 21. Construct 27226
[0260] Table 22. Construct 27234
[0261] Table 23. Construct 27241
[0262] Table 24. Construct 27680
[0263] Table 25. Construct 28255
[0264] Table 26. Construct 28291
[0265] Table 27. Construct 28292
[0266] Table 28. Construct 28293
[0267] Table 29. Construct 28510
[0268] Table 30. Construct 28520
[0269] Table 31. Construct 28560
[0270] Table 32. Construct 28825
[0271] Table 33. Construct 28834
[0272] Table 34. Illustrative promoter and terminator sequences. Lowercase nucleotide sequences in VSP promoter sequences SEQ ID NO: 1-27 and RZDP promoter sequences SEQ ID NO: 28-43: 5'UTR or intron; Uppercase nucleotide sequences in VSP promoter sequences SEQ ID NO: 1-27 and RZDP promoter sequences SEQ ID NO: 28-43: exons.
[0273] All patents, patent publications, patent applications, journal articles, books, technical references, etc., discussed in this disclosure are incorporated herein by reference in their entirety for all purposes.
[0274] It should be understood that the description in this disclosure has been simplified and is intended to illustrate only the elements relevant to a clear understanding of this disclosure. Omitted details and modifications or alternative embodiments are within the knowledge of those skilled in the art.
[0275] It is understood that, in certain aspects of this disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component, to provide an element or structure or to perform a given one or more functions. Such substitution is considered to be within the scope of this disclosure unless it would render certain embodiments of this disclosure inoperable.
[0276] The examples presented herein are intended to illustrate potential and specific implementations of this disclosure. It will be understood that these examples are primarily intended for the purpose of explaining this disclosure to those skilled in the art. Variations may be made to these figures or the operations described herein without departing from the spirit of this disclosure. For example, in some cases, method steps or operations may be performed or carried out in a different order, or operations may be added, deleted, or modified.
[0277] Where a numerical range is provided, it should be understood that each intermediate value between the upper and lower limits of the range (the smallest decimal place to the units digit of the lower limit, unless the context explicitly states otherwise) is also specifically disclosed. This covers any smaller range between any stated or non-statement intermediate value in the stated range and any other stated or intermediate value in the stated range. The upper and lower limits of these smaller ranges may be independently included or excluded from the range, and each range in which no one or both limit values are included is also covered by this technique, depending on any limit values specifically excluded from the stated range. Where the stated range includes one or both limit values, it also includes ranges that exclude one or both of those included limit values.
[0278] Numerous specific details have been set forth in the foregoing description to provide a more thorough understanding of the invention. However, it will be apparent to those skilled in the art that the invention described herein can be practiced without one or more of these specific details. In other instances, features and procedures well-known to those skilled in the art have not been described to avoid obscuring the invention. Embodiments of this disclosure have been described for illustrative and not restrictive purposes. While the invention has been described primarily with reference to specific embodiments, other embodiments are contemplated to become apparent to those skilled in the art upon reading this disclosure, and it is intended that such embodiments be included within the methods of the invention. Therefore, this disclosure is not limited to the embodiments depicted above or in the accompanying drawings, and various embodiments and modifications may be made without departing from the scope of the following claims.
Claims
1. A synthetic DNA construct comprising a zygotic preferred promoter operatively linked to a first target nucleotide sequence ("NSOI").
2. The synthetic DNA construct of claim 1, wherein the promoter of the zygote is preferably a vacuole sorting protein (VSP) promoter.
3. The synthetic DNA construct of claim 1, wherein the VSP promoter comprises the following sequences: a) a sequence selected from the group consisting of SEQ ID NO: 1-27, 86, 87 or a functional fragment thereof; or b) an orthologous promoter of SEQ ID NO:
1.
4. The synthetic DNA construct of claim 1, wherein the VSP promoter comprises SEQ ID NO:
86.
5. The synthetic DNA construct of claim 1, wherein the promoter of the zygote is preferably a cyclic zinc finger domain protein (RZDP) promoter.
6. The synthetic DNA construct of claim 5, wherein the RZDP promoter comprises the following sequences: (a) a sequence selected from the group consisting of SEQ ID NO: 28-43 or a functional fragment thereof; or (b) an orthologous promoter of SEQ ID NO:
28.
7. The synthetic DNA construct of claim 6, wherein the RZDP promoter comprises SEQ ID NO:
28.
8. The synthetic DNA construct of claim 1, further comprising a U3 promoter operatively linked to a second NSOI.
9. The synthetic DNA construct of claim 8, wherein the U3 promoter comprises the following sequences: (a) a sequence selected from the group consisting of SEQ ID NO: 44-49 or a functional fragment thereof; or (b) an orthologous promoter of SEQ ID NO:
44.
10. The synthetic DNA construct of claim 9, wherein the U3 promoter comprises SEQ ID NO:
45.
11. The synthetic DNA construct of claim 1, wherein the first NSOI comprises a sequence encoding a nuclease.
12. The synthetic DNA construct of claim 11, wherein the nuclease is selected from the group consisting of zinc finger nucleases ("ZFN"), large-scale nucleases ("MN"), transcription activator-like effector nucleases (TALEN), and CRISPR nucleases.
13. The synthetic DNA construct of claim 12, wherein the CRISPR nuclease is selected from the group consisting of: Cas9, Cas12a, Cas12b, Cas12c, Cas12d, Cas12e, Cas12f, Cas12i, Cas12j, Cas12l, Cas1, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas10, Cas11, Cas13a, Cas13b, Cas13c, Cas13d, Cas14, any other CRISPR-Cas nuclease, any mutants thereof, nickase variants, and inactivating variants.
14. The synthetic DNA construct of claim 13, wherein the CRISPR nuclease further comprises a fusion domain.
15. The synthetic DNA construct of claim 14, wherein the fusion domain is selected from the group consisting of: deaminases, uracil DNA glycosylases, reverse transcriptases, ubiquitin receptors, and exonucleases.
16. The synthetic DNA construct of claim 8, wherein the second NSOI comprises a sequence encoding at least one guide RNA.
17. The synthetic DNA construct of claim 16, wherein the at least one guide RNA is encoded by a sequence selected from the group consisting of SEQ ID NO: 50-61.
18. The synthetic DNA construct of claim 17, wherein the synthetic DNA construct comprises a sequence selected from the group consisting of SEQ ID NO:62-72.
19. A plant cell comprising the synthetic DNA construct as described in claims 1-18.
20. The plant cell of claim 19, wherein the plant cell is a pollen cell or an egg cell.
21. A plant comprising plant cells as described in claims 19-20.
22. The plant of claim 21, wherein the plant is a maize plant.
23. A method for obtaining edited progeny plants, the method comprising: (a) Providing a first plant, wherein the first plant is transformed to include the synthetic constructs as described in claims 1-18; (b) Pollinate the second plant; and (c) Select at least one offspring produced by the pollination in step (b), wherein the offspring has edited features; thereby obtaining edited offspring plants.
24. The method of claim 23, wherein the first plant is a haploid induction line of the plant.
25. The method of claim 24, wherein the haploid inducing line is a paternal haploid inducer.
26. The method of claim 25, wherein the paternal haploid inducible line contains a mutation in the CENH3 gene.
27. The method of claim 24, wherein the haploid inducing line is a maternal haploid inducer.
28. The method of claim 27, wherein the maternal haploid inducible line contains a mutation in the MATL gene.
29. The method of claim 23, wherein the second plant contains plant genomic DNA to be edited.
30. The method of claim 23, wherein the edited progeny plant is a haploid progeny plant.
31. The method of claim 23, wherein the haploid progeny plant contains the genome of the second plant but not the genome of the first plant.
32. The method of claim 30, wherein the haploid progeny plant is a maize haploid progeny plant.
33. The synthetic DNA construct of claim 1, wherein the promoter of the zygote is preferably the SUMO conjugate 1 (SCE1) promoter.
34. The synthetic DNA construct of claim 33, wherein the SCE1 promoter comprises the following sequences: a) a sequence selected from the group consisting of SEQ ID NO: 83-85 or a functional fragment thereof; or b) an orthologous promoter of SEQ ID NO:
83.
35. The synthetic DNA construct of claim 33, wherein the SCE1 promoter comprises SEQ ID NO:
84.
36. The synthetic DNA construct of claim 33, wherein the SCE1 promoter comprises SEQ ID NO:
85.
37. The synthetic DNA construct of claim 1, further comprising a terminator operatively linked to the NSOI.
38. The synthetic DNA construct of claim 37, wherein the terminator is a ubiquitin terminator.
39. The synthetic DNA construct of claim 37, wherein the ubiquitin terminator comprises SEQ ID NO: 88, 89, 90, 91, 92 or 93.
Citation Information
Patent Citations
Simultaneous gene editing and haploid induction
US10285348B2
CRISPR enzymes and systems
US10669540B2
Methods and compositions for using zinc finger endonucleases to enhance homologous recombination
US20030232410A1
Genomic editing in zebrafish using zinc finger nucleases
US20090203140A1
Poly zinc finger proteins with improved linkers
US6479626B1