Induction of plant units
Combining mutated CENH3 and ig genes in plants like maize and rapeseed significantly boosts haploid induction rates, addressing the inefficiencies of previous methods and enhancing breeding efficiency.
Patent Information
- Application Number
- JP2026092166
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-05-29
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-25
AI Technical Summary
Existing methods for haploid induction in crops like maize and rapeseed yield significantly lower rates compared to Arabidopsis, with paternal induction rates particularly low, limiting the efficiency of breeding and genetic manipulation.
Combining a mutated centromere or kinetochore gene, such as CENH3, with a mutated gametophyte (ig) gene to induce higher rates of paternal haploid induction in plants like maize, sorghum, and rapeseed.
Enhances haploid induction rates, particularly paternal induction, facilitating more efficient breeding and genetic manipulation by increasing the production of phantom embryos, which are crucial for genetic diversity and trait incorporation.
Smart Images

Figure 2026136325000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of plant breeding, and more particularly to the development of haploid inducing substances and their use in generating haploid plants and in doubled haploid technology.
[0002] Background Art The generation and use of haploids is one of the most powerful biotechnological approaches for improving cultivated plants. The advantage of haploids for breeders is that doubled haploid plants can be created without the need for several generations of backcrossing required to obtain a high degree of homozygosity, and homozygosity can already be achieved in the first generation after diploidization of the digenomic haploid. Furthermore, the value of haploids in plant research and breeding lies in the fact that the founder cells of the doubled haploids are the products of meiosis, and the resulting population constitutes a pool of diverse recombinants and at the same time genetically fixed individuals. Therefore, the generation of doubled haploids not only provides a completely useful genetic variability for selection in crop improvement, but is also a valuable means for generating mapping populations, recombinant inbred lines, and immediate homozygous mutants and transgenic lines.
[0003] Monomers can be obtained by in vitro or in vivo approaches. However, many species and genotypes are refractory to these processes. Instead, substantial alterations of the centromere-specific histone H3 variant (CENH3, also known as CENP-A) create monomer-inducible lineages in the model plant Arabidopsis thaliana by exchanging its N-terminal domain and fusing it with GFP ("GFP-tailswap" CENH3) (Ravi and Chan, Nature, 464 (20 10), 615 - 618; Comai, L, "Genome elimination: translating basic research into a future tool for plant breeding.", PLoS biology, 12. 6 (2014)). The CENH3 protein is a variant of the H3 histone protein that is a member of the active centromere kinetochore complex. Regarding these "GFP-tail exchange" singularity-inducing plant lines, when these singularity-inducing plants were crossed with wild-type plants, singularity occurred in the offspring. The singularity-inducing plant lines were stable during self-pollination, suggesting that competition between the modified centromere and the wild-type centromere in the developing hybrid embryo leads to centromere inactivation of the inducing parent, resulting in uniparental chromosome removal. Consequently, the chromosome containing the modified CENH3 protein is lost during early embryonic development, producing singular offspring containing only the wild-type parent's chromosomes. Therefore, singular plants can be obtained by crossing "GFP-tail exchange" plants as singularity-inducing plants with wild-type plants.
[0004] International Publication Nos. 2016 / 030019 and 2016 / 102665 describe alternative non-transgenic methods for modifying the endogenous CENH3 gene in plants for the creation of singularity-inducing lineages. The authors show that when mutant plants are crossed with wild-type plants, one or more single amino acid substitutions, particularly in diverse domains of the CENH3 protein, result in singularity induction.
[0005] CENH3 mutants function as haploid inducers in Arabidopsis species, either as transgenic "tail-swap" inducers or as non-transgenic inducers possessing the mutated endogenous CENH3 gene, reaching rates of up to 10%. However, these data could not be transferred to crops. In both maize and rapeseed, the haploid induction rates were significantly lower than in Arabidopsis species, reaching up to 3.6% for the transgenic "tail-swap" inducer (Kelliher et al. (2016) "Maternal haploids are preferentially induced by CENH3-tailswap transgenic complementation in maize", Frontiers in plant science, 7, 414), and up to 2% for the non-transgenic inducer, with haploid induction primarily observed on the maternal side.
[0006] Another possibility for singular induction in maize is the indeterminate gametophyte (ig) system. The so-called mutated ig gene induces singularities of both male (androgenic) and female (gynogenic) origins. The ig gene was first described by Kermicle (1969, "Androgenesis conditioned by a mutation in maize", Science, 166(3911), 1422-1424) as occurring spontaneously in the highly inbred Wisconsin-23 (W23) strain. The ig gene is essential for the normal growth and development of the gametophyte, and loss of ig gene function results in too many or too few nuclei to be produced. In ig lines, the developing female gametophyte is released from its normal three mitotic divisions. Lin (1981, Rev. Brasil. Biol. 41(3): 557-63) observed that the presence of a mutated ig allele causes a variable number of mitosis and partial degeneration of the nucleus. Following fertilization of the female gametophyte, the sperm nucleus develops into a patrilineal phantom embryo, occasionally through androgenesis. Embryonic development of the sperm nucleus in the maternal cytoplasm leads to the formation of a patrilineal phantom. Kermicle et al. (1980, Maize Genet. Coop. Newsl. 54: 84-85) determined that the ig allele is located on the long arm of chromosome 3, 90 cM from the most distal locus on the short arm designed by g2 (EP0831689). The presence of the ig allele increases the occurrence of patrilineal phantoms from a spontaneous occurrence rate of approximately 1 per 80,000 to a frequency of 1-3% in observed maize plants. This is significantly lower than the maternal induction rate, which is typically around 10%.
[0007] Therefore, an object of the present invention is to address one or more of the drawbacks of the prior art.
[0008] Summary of the Invention Surprisingly, the inventors found that a combination of a mutated centromere or kinetochore gene, such as CENH3, and a mutated undetermined gametophyte (ig) gene is particularly suitable for generating singular-inducing plants, especially paternal singular-inducing plants, such as maize (e.g., Zea mays), sorghum (e.g., Sorghum bicolor), and rapeseed (e.g., Brassica napus). The singular induction rate was found to be much higher than that resulting from either mutation alone, and even higher than realistically expected with such combinations.
[0009] Accordingly, in one embodiment, the present invention relates to a plant or plant part comprising a polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and a polynucleic acid encoding a mutated centromere or kinetochore protein, wherein the mutated centromere or kinetochore protein is preferably CENH3. Both the mutated ig and the centromere or kinetochore protein result in singular induction activity, for example, paternal singular induction activity.
[0010] In one embodiment, the present invention relates to a method for producing a plant or plant part, particularly a singular plant or plant part, comprising crossing a first plant containing a polynucleic acid encoding a mutated uncertain gametophyte (ig) protein and a polynucleic acid encoding a mutated centromere or kinetochore protein with a second plant, and selecting singular offspring, wherein the mutated centromere or kinetochore protein is preferably CENH3. Optionally, the singular offspring may be converted into a double singular plant or plant part.
[0011] In one embodiment, the present invention relates to a plant or plant part, particularly a plant or plant part obtained or obtainable by a method for producing a singular plant or plant part, comprising crossing a first plant containing a polynucleic acid encoding a mutated uncertain gametophyte (ig) protein and a polynucleic acid encoding a mutated centromere or kinetochore protein with a second plant and selecting singular offspring, wherein the mutated centromere or kinetochore protein is preferably CENH3. Optionally, the singular offspring may be converted into a doubled singular plant or plant part.
[0012] In one embodiment, the present invention relates to the use of a plant or part of a plant comprising a polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and a polynucleic acid encoding a mutated centromere or kinetochore protein as a singular inducer, preferably a paternal singular inducer, wherein the mutated centromere or kinetochore protein is preferably CENH3.
[0013] In one embodiment, the present invention relates to zea mays seeds named igEIN, whose representative sample is deposited under NCIMB accession number NCIMB43772, or to plants or plant parts grown therefrom. In one embodiment, the present invention relates to zea mays seeds deposited under NCIMB accession number NCIMB43772, or to plants or plant parts grown therefrom.
[0014] In one embodiment, the present invention relates to a method for identifying suitable centromere or kinetochore protein variants or mutations, preferably of CENH3, for combination with ig variants or mutations described elsewhere herein to increase singular-inducing activity or capacity, by analyzing the singular-inducing activity or capacity obtained by combining such mutations.
[0015] The inventors have surprisingly found that the plants and methods described herein increase the phantom induction rate, particularly the paternal phantom induction rate. This makes it possible to increase the efficiency of cytoplasmic male sterility (CMS) conversion based on paternal phantom induction. Furthermore, the provision of paternal phantom inducers is of particular importance. When it is necessary to produce many phantoms from a single isolate, the use of the maternal line is limited to only one cross, yielding an average of 1-2 haploid plants. The paternal system allows for several crosses using plant pollen, providing the possibility of pollinating with the paternal inducer. With the use of a high-performance inducer, more phantoms can be obtained per single isolate. Such systems offer an opportunity to optimize breeding schemes by more efficiently using whole-genome prediction or trait incorporation. Furthermore, paternal induction systems are preferred for crops where castration systems are difficult. They can be applied using sterile inducers based on nuclear sterility that can be pollinated by any fertile line. Furthermore, after the introduction of singular selection markers such as red roots in maize, the present invention can be used in special cases in novel breeding or trait gene transfer programs for the generation of doubled singular (DH) from a single isolated plant. Ultimately, efficient paternality inducers with high induction rates can be used in genome editing, particularly when the paternity inducer simultaneously includes a genome editing mechanism.
[0016] The present invention is particularly captured by one or more of the following numbered claims 1 to 125 and any one or more combinations of any other claims and / or embodiments presented herein.
[0017] 1. A plant or plant part containing a polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and a polynucleic acid encoding a mutated centromere or kinetochore protein.
[0018] 2. The plant or plant part according to Claim 1, wherein the polynucleic acid encoding the mutated ig protein includes one or more nucleic acid insertions (compared to the polynucleic acid encoding the wild-type undetermined gametophyte (ig) protein).
[0019] 3. The plant or plant part according to claim 1 or 2, wherein the polynucleic acid encoding the mutated ig protein contains a frameshift mutation or a nonsense mutation (compared to the polynucleic acid encoding the wild-type undetermined gametophyte (ig) protein).
[0020] 4. The plant or plant part according to any one of claims 1 to 3, wherein the polynucleic acid encoding the mutated ig protein includes a knockout mutation or a knockdown mutation.
[0021] 5. The plant or plant part according to any one of claims 1 to 4, wherein the polynucleic acid encoding the mutated ig protein includes one or more nucleic acid insertions in the ig-coding sequence (compared to the polynucleic acid encoding the wild-type undetermined gametophyte (ig) protein).
[0022] 6. The plant or plant part according to any one of claims 1 to 5, wherein the polynucleic acid encoding the mutated ig protein includes one or more nucleic acid insertions in the sequence encoding the LOB domain (compared to the polynucleic acid encoding the wild-type undetermined gametophyte (ig) protein).
[0023] 7. The plant or plant part according to any one of claims 1 to 6, wherein the polynucleic acid encoding the mutated ig protein comprises one or more nucleic acid insertions in a first protein encoding an exon, for example, a first protein encoding an exon in the range of nucleotide positions 431 to 841 of the reference maize (Zea mays) sequence shown in Sequence ID No. 6.
[0024] 8. The plant or plant part according to any one of claims 1 to 7, wherein the polynucleic acid encoding the mutated Ig protein contains an insertion of one or more nucleic acids in an intron preceding a first protein encoding an exon.
[0025] 9. The plant or plant part according to any one of claims 1 to 8, wherein the polynucleic acid encoding the mutated Ig protein contains an Ig-O allele.
[0026] 10. The plant or plant part according to any one of claims 1 to 9, wherein the polynucleic acid encoding the mutated Ig protein contains an Ig-mum allele.
[0027] 11. The plant or plant part according to any one of claims 1 to 10, wherein the polynucleic acid encoding the mutated Ig protein contains an insertion of one or more nucleic acids in an Ig codon corresponding to a codon selected from codons 118, 119 or 120 of the wild-type maize (Zea mays) Ig protein as shown in SEQ ID NO: 7 or 8, or corresponding to a codon selected from codons 191, 192 or 193 of the wild-type sorghum (Sorghum bicolor) Ig protein as shown in SEQ ID NO: 22, or corresponding to a codon selected from codons 143, 144 or 145 of the wild-type sorghum (Sorghum bicolor) Ig protein as shown in SEQ ID NO: 25, or corresponding to a codon selected from codons 94, 95 or 96 of the wild-type rapeseed (Brassica napus) Ig protein as shown in SEQ ID NO: 28 or 31.
[0028] 12. The plant or plant part according to any one of claims 1 to 11, wherein the polynucleic acid encoding the mutated Ig protein contains an insertion of at least 100 nucleotides, preferably at least 200 nucleotides (compared to the polynucleic acid encoding the wild-type indeterminate gametophyte (Ig) protein).
[0029] 13. The plant or plant part according to any one of claims 1 to 12, wherein the polynucleic acid encoding the mutated Ig protein comprises one or more amino acid insertions and / or one or more amino acid substitutions (compared to the wild-type Ig protein).
[0030] 14. The plant or plant part according to any one of claims 1 to 13, wherein the mutated Ig protein comprises one or more amino acid insertions and / or one or more amino acid substitutions in a region corresponding to amino acid residues 110 - 130 of the wild-type Zea mays Ig protein as shown in SEQ ID NO: 9 or 10, or corresponding to amino acid residues 183 - 203 of the wild-type Sorghum bicolor Ig protein as shown in SEQ ID NO: 23, or corresponding to amino acid residues 135 - 155 of the wild-type Sorghum bicolor Ig protein as shown in SEQ ID NO: 26, or corresponding to amino acid residues 86 - 106 of the wild-type Brassica napus Ig protein as shown in SEQ ID NO: 29 or 32.
[0031] 15. The plant or plant part according to any one of claims 1 to 14, wherein the mutated Ig protein comprises one or more amino acid insertions and / or one or more amino acid substitutions in a region corresponding to amino acid residues 116 - 120, preferably 117 - 119 of the wild-type Zea mays Ig protein as shown in SEQ ID NO: 9 or 10, or corresponding to amino acid residues 189 - 193, preferably 190 - 192 of the wild-type Sorghum bicolor Ig protein as shown in SEQ ID NO: 23, or corresponding to amino acid residues 141 - 145, preferably 142 - 144 of the wild-type Sorghum bicolor Ig protein as shown in SEQ ID NO: 26, or corresponding to amino acid residues 92 - 96, preferably 93 - 95 of the wild-type Brassica napus Ig protein as shown in SEQ ID NO: 29 or 32.
[0032] 16. The plant or plant part described in any of claims 1 to 15, wherein the mutated Ig protein is a cleaved Ig protein.
[0033] 17. The plant or plant part described in any of claims 1 to 16, wherein the aforementioned ig is ig1.
[0034] 18. The plant or plant part according to any of claims 1 to 16, wherein the aforementioned ig is ig2.
[0035] 19. The plant is of the genus Zea, preferably Zea mays, and the wild-type undetermined gametophyte (ig) protein is a) Encoded by a polynucleic acid comprising the nucleotide sequence of SEQ ID NO: 6, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 6, b) Derived from the nucleotide sequence of SEQ ID NO: 7 or 8, or a coding sequence that is at least 90% identical, preferably at least 95%, more preferably at least 98% identical to SEQ ID NO: 7 or 8, or c) Having the amino acid sequence of SEQ ID NO: 9 or 10, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 9 or 10, The plant or plant part described in any of claims 1 through 18.
[0036] 20. The plant is of the genus Sorghum, preferably Sorghum bicolor, and the wild-type undetermined gametophyte (ig) protein is a) Encoded by a polynucleic acid comprising the nucleotide sequence of SEQ ID NO: 21 or 24, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 21 or 24, b) Derived from a coding sequence that includes the nucleotide sequence of SEQ ID NO: 22 or 25, or a sequence that is at least 90% identical, preferably at least 95% identical, more preferably at least 98% identical to SEQ ID NO: 22 or 25, or c) Having the amino acid sequence of SEQ ID NO: 23 or 26, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 23 or 26. The plant or plant part described in any of claims 1 through 18.
[0037] 21. The plant is of the genus Brassica, preferably Brassica napus, and the wild-type undetermined gametophyte (ig) protein is a) Encoded by a polynucleic acid comprising the nucleotide sequence of SEQ ID NO: 27 or 30, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 27 or 30, b) Derived from a coding sequence that includes the nucleotide sequence of SEQ ID NO: 28 or 31, or a sequence that is at least 90% identical, preferably at least 95%, more preferably at least 98% identical to SEQ ID NO: 28 or 31, or c) Having the amino acid sequence of SEQ ID NO: 29 or 32, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 29 or 32 The plant or plant part described in any of claims 1 through 18.
[0038] 22. The plant is of the genus Zea, preferably Zea mays, and the mutated undetermined gametophyte (ig) protein is a) Encoded by a polynucleic acid comprising the nucleotide sequence of SEQ ID NO: 1, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 1, b) Derived from the nucleotide sequence of SEQ ID NO: 2 or 3, or a coding sequence that is at least 90% identical, preferably at least 95%, more preferably at least 98% identical to SEQ ID NO: 2 or 3, or c) Having an amino acid sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to the amino acid sequence of SEQ ID NO: 4 or 5. The plant or plant part described in any of claims 1 through 18.
[0039] 23. The plant or plant part according to any one of claims 1 to 18, wherein the plant is of the genus Sorghum, preferably Sorghum bicolor, and the mutated undetermined gametophyte (ig) protein is at least 90% identical, preferably at least 95% identical, more preferably at least 98% identical to SEQ ID NO: 23 or 26, and has an amino acid sequence that is not 100% identical to the amino acid sequence of SEQ ID NO: 23 or 26, respectively.
[0040] 24. The plant or plant part according to any one of claims 1 to 18, wherein the plant is of the genus Brassica, preferably Brassica napus, and the mutated undetermined gametophyte (ig) protein is at least 90% identical, preferably at least 95%, more preferably at least 98% identical to SEQ ID NO: 29 or 32, and has an amino acid sequence that is not 100% identical to the amino acid sequence of SEQ ID NO: 29 or 32, respectively.
[0041] 25. The plant or plant part according to any of claims 1 to 24, wherein the mutated centromere protein is a mutated histone protein.
[0042] 26. The plant or plant part according to any one of claims 1 to 25, wherein the protein of the mutated centromere or kinetochore is selected from the group including CENH3 or proteins that interact with CENH3.
[0043] 27. The plant or plant part according to any one of claims 1 to 26, wherein the protein of the mutated centromere or kinetochore is selected from the group comprising CENH3, CENP-C, KNL2, SCM3, SAD2, and SIM3.
[0044] 28. The plant or plant part according to any of claims 1 to 27, wherein the mutated centromere protein is a mutated CENH3 protein.
[0045] 29. The plant or plant part according to any one of claims 1 to 28, wherein the mutated CENH3 protein contains one or more mutated amino acids in one or more of the N-terminal domain, αN helix, α1 helix, loop 1 domain, α2 helix, loop 2 domain, α3 helix, or C-terminal domain of CENH3.
[0046] 30. The mutated CENH3 protein has an N-terminal domain corresponding to amino acids 1-82 of Arabidopsis thaliana CENH3, an αN helix corresponding to amino acids 83-97 of Arabidopsis thaliana CENH3, an α1 helix corresponding to amino acids 103-113 of Arabidopsis thaliana CENH3, a loop 1 domain corresponding to amino acids 114-126 of Arabidopsis thaliana CENH3, an α2 helix corresponding to amino acids 127-155 of Arabidopsis thaliana CENH3, and a loop 2 domain corresponding to amino acids 156-162 of Arabidopsis thaliana CENH3. The plant or plant part according to any one of claims 1 to 29, comprising one or more mutated amino acids in one or more of the α3 helix corresponding to amino acids 163-172 of Arabidopsis thaliana CENH3 and one or more of the C-terminal domain corresponding to amino acids 173-178 of Arabidopsis thaliana CENH3, wherein the Arabidopsis thaliana CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in Sequence ID No. 12.
[0047] 31. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids in one or more of the following: the N-terminal domain corresponding to amino acids 1 to 62 of corn (Zea mays) CENH3, the αN-helix corresponding to amino acids 63 to 77 of corn CENH3, the α1-helix corresponding to amino acids 83 to 93 of corn CENH3, the loop 1 domain corresponding to amino acids 94 to 106 of corn CENH3, the α2-helix corresponding to amino acids 107 to 135 of corn CENH3, the loop 2 domain corresponding to amino acids 136 to 142 of corn CENH3, the α3-helix corresponding to amino acids 143 to 152 of corn CENH3, and the C-terminal domain corresponding to amino acids 153 to 157 of corn CENH3, and the corn CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in Sequence ID No. 14.
[0048] 32. The mutated CENH3 protein is found in sorghum. The plant or plant part according to any one of claims 1 to 29, wherein one or more of the following contain mutated amino acids: the N-terminal domain corresponding to amino acids 1 to 62 of bicolor)CENH3, the αN helix corresponding to amino acids 63 to 77 of sorghum CENH3, the α1 helix corresponding to amino acids 83 to 93 of sorghum CENH3, the loop 1 domain corresponding to amino acids 94 to 106 of sorghum CENH3, the α2 helix corresponding to amino acids 107 to 135 of sorghum CENH3, the loop 2 domain corresponding to amino acids 136 to 142 of sorghum CENH3, the α3 helix corresponding to amino acids 143 to 152 of sorghum CENH3, and the C-terminal domain corresponding to amino acids 153 to 157 of sorghum CENH3, and the sorghum CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in Sequence ID No. 18.
[0049] 33. The mutated CENH3 protein is found in rapeseed (Brassica napus). A plant or plant part according to any one of claims 1 to 29, wherein one or more of the following contain mutated amino acids: an N-terminal domain corresponding to amino acids 1 to 84 of napus)CENH3, an αN helix corresponding to amino acids 85 to 99 of rapeseed CENH3, an α1 helix corresponding to amino acids 105 to 115 of rapeseed CENH3, a loop 1 domain corresponding to amino acids 116 to 128 of rapeseed CENH3, an α2 helix corresponding to amino acids 129 to 157 of rapeseed CENH3, a loop 2 domain corresponding to amino acids 158 to 164 of rapeseed CENH3, an α3 helix corresponding to amino acids 165 to 174 of rapeseed CENH3, and a C-terminal domain corresponding to amino acids 175 to 180 of rapeseed CENH3, and the rapeseed CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in Sequence ID No. 16.
[0050] 34. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids in the N-terminal domain of CENH3.
[0051] 35. The plant or plant part according to claim 34, wherein the N-terminal domain of CENH3 corresponds to amino acids 1-82 of the reference Arabidopsis thaliana CENH3 protein, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98% identical to the sequence shown in SEQ ID NO: 12.
[0052] 36. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to positions 3, 17, 32, 35, 9, 24, 29, 40, 42, 50, 55, 57, 61, 74, or 82 of the reference Arabidopsis thaliana CENH3 protein, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0053] 37. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to positions 3, 17, 32, or 35 of the Arabidopsis thaliana CENH3 protein when the plant or plant part is from the genus Zea, preferably Zea mays, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0054] 38. The plant or plant part according to any one of claims 1 to 37, wherein the mutated CENH3 protein contains one or more mutated amino acids at position 3, 16, 32, or 35 of the CENH3 protein of a plant or plant part from the genus Zea, preferably Zea mays, and preferably the maize CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 14.
[0055] 39. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to positions 9, 24, 29, 32, 40, 42, 50, 55, 57, or 61 of the reference Arabidopsis thaliana CENH3 protein when the plant or plant part is from the genus Brassica, preferably Brassica napus, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0056] 40. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids at positions 9, 24, 29, 30, 33, 41, 43, 50, 55, 57, or 61 of the CENH3 protein of a plant or plant part from the genus Brassica, preferably Brassica napus, and preferably the maize CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 16.
[0057] 41. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to position 42 or 74 of the Arabidopsis thaliana CENH3 protein when the plant or plant part is from the genus Sorghum, preferably Sorghum bicolor, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in Sequence ID No. 12.
[0058] 42. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids at position 42 or 55 of the CENH3 protein of a plant or plant part from the genus Sorghum, preferably Sorghum bicolor, and preferably the Sorghum CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98% identical to the sequence shown in SEQ ID NO: 18.
[0059] 43. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids corresponding to positions 104, 109, 120, 148, 175, 130, 151, 157, 158, 164, 166, 83, 86, 124, 127, 132, 136, 152, 155, or 172 of the reference Arabidopsis thaliana CENH3 protein, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0060] 44. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to positions 104, 109, 120, 148, or 175 of the reference Arabidopsis thaliana CENH3 protein when the plant or plant part is from the genus Zea, preferably Zea mays, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0061] 45. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids at positions 84, 89, 100, 128, or 155 of the CENH3 protein of a plant or plant part from the genus Zea, preferably Zea mays, and preferably the maize CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 14.
[0062] 46. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids corresponding to position 130 of the Arabidopsis thaliana CENH3 protein when the plant or plant part is from the genus Sorghum, preferably Sorghum bicolor, and preferably the Arabidopsis thaliana CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98% identical to the sequence shown in Sequence ID No. 12.
[0063] 47. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein contains one or more mutated amino acids at position 110 or 157 of the CENH3 protein of a plant or plant part from the genus Sorghum, preferably Sorghum bicolor, and preferably the Sorghum CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 18.
[0064] 48. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to positions 130, 151, 157, 158, 164, or 166 of the reference Arabidopsis thaliana CENH3 protein when the plant or plant part is from the genus Brassica, preferably Brassica napus, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0065] 49. The plant or plant part according to any one of claims 1 to 29, wherein the mutated CENH3 protein comprises one or more mutated amino acids corresponding to positions 132, 153, 159, 160, 166, or 168 of the CENH3 protein of a plant or plant part from the genus Brassica, preferably Brassica napus, and preferably the maize CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 16.
[0066] 50. The plant or plant part according to any one of claims 25 to 49, wherein the mutated protein contains one or more amino acid substitutions, or the one or more mutated amino acids contain one or more amino acid substitutions.
[0067] 51. A plant or plant part as described in any of claims 25 to 49, comprising 1 to 7 mutations, e.g., 1 to 7 amino acid substitutions.
[0068] 52. A plant or plant part according to any of claims 25 to 49, comprising one mutation, for example, one amino acid substitution.
[0069] 53. The plant or plant part according to any one of claims 1 to 29, wherein the plant is zea mays, and the protein of the mutated centromere or kinetochore is a mutated CENH3 protein having an amino acid substitution corresponding to position 35 of maize CENH3, preferably corresponding to position 35 of SEQ ID NO: 14 or having an amino acid substitution at position 35 of SEQ ID NO: 14, and preferably the amino acid substitution is 35K, for example, E35K.
[0070] 54. The plant or plant part according to any one of claims 1 to 53, wherein the polynucleic acid encoding the mutated undetermined gametophyte (ig) protein and the polynucleic acid encoding the mutated centromere or kinetochore protein are operationally linked to one or more control sequences.
[0071] 55. The plant or plant part according to any one of claims 1 to 54, wherein the mutated undetermined gametophyte (ig) protein and the mutated centromere or kinetochore protein can be expressed in the plant or plant part.
[0072] 56. The plant or plant part according to any one of claims 1 to 55, wherein the mutated undetermined gametophyte (ig) protein confers singular induction activity or enhances singular induction ability.
[0073] 57. The plant or plant part according to any one of claims 1 to 56, wherein the mutated centromere or kinetochore protein confers singular induction activity or enhances singular induction ability.
[0074] 58. A plant or plant part according to any one of claims 1 to 57, wherein the polynucleic acid encoding the mutated undetermined gametophyte (ig) protein encodes the mutated endogenous undetermined gametophyte (ig) protein.
[0075] 59. A plant or plant part according to any one of claims 1 to 58, wherein the polynucleic acid encoding the mutated undetermined gametophyte (ig) protein encodes the mutated endogenous undetermined gametophyte (ig) protein at its native locus.
[0076] 60. The plant or plant part according to any one of claims 1 to 59, wherein the polynucleic acid encoding the mutated kinetochore or centromere protein encodes the mutated endogenous kinetochore or centromere protein.
[0077] 61. A plant or plant part according to any one of claims 1 to 60, wherein the polynucleic acid encoding the mutated kinetochore or centromere protein encodes the mutated endogenous kinetochore or centromere protein at its native locus.
[0078] 62. The plant or plant part according to any one of claims 1 to 61, wherein the polynucleic acid encoding the mutated undetermined gametophyte (ig) protein and / or the polynucleic acid encoding the mutated centromere or kinetochore protein is homozygous.
[0079] 63. The plant or plant part according to any one of claims 1 to 62, wherein the polynucleic acid encoding the mutated undetermined gametophyte (ig) protein and / or the polynucleic acid encoding the mutated centromere or kinetochore protein is heterozygous.
[0080] 64. The plant or plant part described in any of claims 1 to 63, wherein the plant or plant part is a crop or plant part.
[0081] 65. The plant or plant part described in any of claims 1 to 64, wherein the plant or plant part is selected from the group including the genera Zea, Sorghum, and Brassica.
[0082] 66. The plant or plant part according to claim 65, wherein the plant or plant part is selected from the group including the genera Zea and Sorghum.
[0083] 67. The plant or plant part described in claim 66, wherein the plant or plant part is from the genus Zea.
[0084] 68. The plant or plant part according to claim 65, wherein the plant or plant part is selected from the group including the species of maize (Zea mays), sorghum bicolor, and rapeseed (Brassica napus).
[0085] 69. The plant or plant part according to claim 66, wherein the plant or plant part is selected from the group including the species of maize (Zea mays) and sorghum (Sorghum bicolor).
[0086] 70. The plant or plant part described in Claim 67, wherein the plant or plant part is derived from corn (Zea mays).
[0087] 71. The plant or plant part according to any of claims 1 to 70, wherein the plant part is a plant cell, tissue, organ, or seed.
[0088] 72. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is diploid.
[0089] 73. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is singular.
[0090] 74. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is a digenous singular.
[0091] 75. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is a trigeminal singular.
[0092] 76. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is a doubled singular.
[0093] 77. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is a doubling digenomic singular.
[0094] 78. The plant or plant part described in any of claims 1 to 71, wherein the plant or plant part is a doubled trigeminal singular.
[0095] 79. A plant according to any one of claims 1 to 78, further comprising polynucleic acids encoding site-specific DNA or RNA-binding proteins.
[0096] 80. A plant according to any one of claims 1 to 79, further comprising polynucleic acids encoding site-directed (mutant) DNA or RNA nucleases.
[0097] 81. The plant according to claim 80, wherein the site-specific (mutant) nuclease is selected from the group comprising meganuclease (MN), zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), (mutant) Cas nuclease / effector proteins, such as Cas9 nuclease, Cfp1 nuclease, MAD7 nuclease, dCas9-FokI, dCpf1-FokI, dMAD7 nuclease-FokI, chimeric Cas9-cytidine deaminase, chimeric Cas9-adenine deaminase, chimeric FENI-FokI, and mega-TAL, nickase Cas9 (nCas9), chimeric dCas9 non-FokI nuclease, dCpf1 non-FokI nuclease, and dMAD7 non-FokI nuclease.
[0098] 82. The plant according to claim 80 or 81, wherein the site-specific (mutant) nuclease is a (mutant) Cas effector protein, the plant further comprises a polynucleic acid encoding gRNA and optionally a polynucleic acid encoding tracrRNA.
[0099] 83. A plant or plant part obtained by crossing a first plant, which is a plant described in any of claims 1 to 82, with a second plant.
[0100] 84. A method for producing a plant or plant part, comprising: preparing a singular, digenomic singular, or trigenomic singular plant obtained from a hybrid of a first plant which is a plant described in any of claims 1 to 72 or 79 to 82 and a second plant; and converting a singular, digenomic singular, or trigenomic singular plant or plant part into a doubled singular, doubled digenomic singular, or doubled trigenomic singular plant or plant part.
[0101] 85. A method for producing a plant or plant part, comprising crossing a first plant, which is a plant described in any of claims 1 to 72 or 76 to 82, with a second plant.
[0102] 86. A method for producing a plant or plant part, comprising crossing a first plant or plant part which is a plant as described in any of claims 1 to 72 or 76 to 82 with a second plant, and selecting a phantom, digenous phantom or trigenous phantom offspring plant or plant part.
[0103] 87. A method for producing a plant or plant part, comprising: crossing a first plant or plant part which is a plant as described in any of claims 1 to 72 or 76 to 82 with a second plant; selecting a singular, digenomic singular or trigenomic singular offspring plant or plant part; and converting a singular, digenomic singular or trigenomic singular plant or plant part into a doubled singular, doubled digenomic singular or doubled trigenomic singular plant or plant part.
[0104] 88. A method for modifying plant genomic DNA, comprising: a) preparing a first plant which is a plant described in any of claims 76 to 82; b) preparing a second plant (containing the plant genomic DNA to be modified); c) pollinating the second maize plant with pollen from the first plant; and d) selecting at least one phantom, digenomic phantom, or trigenomic phantom offspring produced by the pollination in step (c) (wherein the phantom, digenomic phantom, or trigenomic phantom offspring contains the genome of the second plant rather than the genome of the first plant, and the genome of the phantom, digenomic phantom, or trigenomic phantom offspring is modified by a site-specific DNA or RNA-binding protein delivered by the first plant).
[0105] 89. The method of claim 88, wherein the offspring of a modified singular organism are treated with a chromosome doubling agent to produce offspring of a modified doubling singular organism.
[0106] 90. The method according to claim 89, wherein the chromosome doubling agent is colchicine, pronamide, diticill, trifluralin, or other known microtubule inhibitor.
[0107] 91. The method according to any one of claims 84 to 90, wherein the second plant is of the same species as the first plant.
[0108] 92. The method according to any one of claims 84 to 91, wherein the second plant has a different haplotype from the first plant.
[0109] 93. The method according to any one of claims 84 to 92, wherein the second plant is diploid, tetraploid, or hexaploid.
[0110] 94. The plant or plant part according to any one of claims 84 to 93, wherein the second plant does not contain polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and / or polynucleic acid encoding a mutated centromere or kinetochore protein.
[0111] 95. The method according to any of claims 84 to 94, wherein the second plant is not a singular inducer.
[0112] 96. A plant or plant part obtained by any method described in Claims 84 to 95.
[0113] 97. Use of any of the plants or plant parts described in paragraphs 1 to 83 as a claim for singular inducer.
[0114] 98. Use of any of the plants or plant parts described in paragraphs 1 to 83 as a paternal singular inducer.
[0115] 99. The plant or plant part described in claim 71, wherein the plant part is pollen.
[0116] 100. A plant or plant part described in any of claims 1 to 82, which is not necessarily obtained by biological means.
[0117] 101. A method for identifying a plant or plant part, comprising detecting mutated unspecified gametophyte proteins and mutated centromere or kinetochore proteins (in a sample from a plant or plant part, for example, a sample containing (genomic) DNA from a plant or plant part), or detecting polynucleic acids encoding mutated unspecified gametophyte proteins and polynucleic acids encoding mutated centromere or kinetochore proteins.
[0118] 102. The method according to Claim 101, comprising detecting mutated unspecified gametophyte proteins and mutated centromere or kinetochore proteins as defined in any of Claims 1 to 63, or detecting polynucleic acids encoding mutated unspecified gametophyte proteins and polynucleic acids encoding mutated centromere or kinetochore proteins.
[0119] 103. The method according to any one of claims 101 to 102, wherein the plant or plant part is a plant or plant part described in any one of claims 1 to 83, 96, or 100.
[0120] 104. A method for detecting plants or plant parts having singular induction activity or enhanced singular induction activity, according to any of claims 101 to 103.
[0121] 105. A method for detecting a plant or plant part having paternal singularity-inducing activity or enhanced paternal singularity-inducing activity, according to any of claims 101 to 104.
[0122] 106. A method according to any of claims 101 to 105, including marker-assisted selection.
[0123] 107. The method according to any one of claims 101 to 106, comprising detecting a (molecular or genetic) marker supported by or linked to a polynucleic acid encoding an unspecified gametophyte protein containing a mutation, and detecting a (molecular or genetic) marker supported by or linked to a polynucleic acid encoding a centromere or kinetochore protein containing a mutation.
[0124] 108. The method according to claim 107, wherein the (molecular or genetic) marker comprises or encodes a polynucleic acid containing the mutation, its complement, or its reverse complement.
[0125] 109. The method according to claim 107 or 108, wherein the (molecular or genetic) marker comprises a primer or probe.
[0126] 110. The method according to any of claims 101 to 109, wherein the detection includes splicing, hybridization-based methods (e.g., (dynamic) allele-specific hybridization, molecular beacons, SNP microarrays), enzyme-based methods (e.g., PCR, KASP (Kompetitive AlleleSpecific PCR), RFLP, ALFP, RAPD, flap endonuclease, primer extension, 5'-nuclease, oligonucleotide ligation assay), and post-amplification methods based on the physical properties of DNA (e.g., single nucleotide polymorphism, temperature gradient gel electrophoresis, denaturing high-performance liquid chromatography, high-resolution melting of whole amplicons, use of DNA mismatch-binding proteins, SNPlex, surveyor nuclease assay).
[0127] 111. A method for producing a plant or plant part, A) (i) A step of preparing a plant or plant part, and (ii) Muting one or more (endogenous) ig alleles, genes, or protein-coding polynucleic acids, and mutating one or more (endogenous) centromere or kinetochore protein alleles, genes, or protein-coding polynucleic acids, and / or introducing one or more mutated ig alleles, genes, or protein-coding polynucleic acids, and one or more mutated centromere or kinetochore protein alleles, genes, or protein-coding polynucleic acids (genome), or B) (i) A step of preparing one or more (endogenous) mutated ig alleles, genes, or proteins encoding polynucleic acids, and / or one or more (genomically) introduced mutated ig alleles, genes, or proteins encoding polynucleic acids, and (ii) the step of mutating one or more (endogenous) centromere or kinetochore protein alleles, genes or polynucleic acids encoding a protein, and / or the step of introducing one or more mutated centromere or kinetochore protein alleles, genes or polynucleic acids encoding a protein into the (genome), or C) (i) A step of preparing polynucleic acids encoding one or more (endogenous) mutant centromere or kinetochore protein alleles, genes or proteins, and / or polynucleic acids encoding one or more (genomically) introduced mutant centromere or kinetochore protein alleles, genes or proteins, and (ii) The step of mutating one or more (endogenous) ig alleles, genes, or polynucleic acids encoding a protein, and / or introducing one or more mutated ig alleles, genes, or polynucleic acids encoding a protein into a (genome). A method for producing plants or plant parts, including the above.
[0128] 112. A method for producing the plant or plant part described in claim 111, wherein the plant or plant part is the plant or plant part described in any of claims 1 to 82.
[0129] 113. A method for producing a plant or plant part according to any of claims 11 to 112, wherein the mutation is as defined in any of claims 1 to 63.
[0130] 114. A method for producing a plant or plant part, preferably a plant or plant part as described in any of claims 1 to 82, a) A step of mutagenerating a plant or parts thereof and identifying a plant that contains nucleic acids encoding a mutated undetermined gametophyte (ig) protein, preferably as defined in any of claims 2 to 24, 54, 55, 56, 58, 59, 62, or 63, and b) Mutagenesis of the plants identified in step a) or their parts or their offspring, which include polynucleic acids encoding mutated undetermined gametophyte (ig) proteins, and further identification of plants which preferably include polynucleic acids encoding mutated centromere or kinetochore proteins as defined in any of claims 25-53, 54, 55, 57, 60, 61, 62, or 63. or A) A step of inducing mutagenesis in plants or parts thereof and identifying plants containing polynucleic acids encoding mutated centromere or kinetochore proteins as defined in any of claims 25-53, 54, 55, 57, 60, 61, 62, or 63, and B) Mutagenesis of the plants identified in step a) or their parts or their offspring, which contain polynucleic acids encoding mutated centromere or kinetochore proteins, and further identification of plants containing nucleic acids encoding mutated undetermined gametophyte (ig) proteins as defined in any of claims 2 to 24, 54, 55, 56, 58, 59, 62, or 63. or Steps to induce mutagenesis in a plant or plant part and identify a plant or plant part containing a polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and a nucleic acid encoding a mutated centromere or kinetochore protein, preferably a plant or plant part as described in any of claims 1 to 82. A method for producing plants or plant parts, including the above.
[0131] 115. The method according to any one of claims 111 to 114, wherein the mutation or mutagenesis includes random mutagenesis or site-directed mutagenesis.
[0132] 116. The method according to any one of claims 111 to 115, wherein the mutation or mutagenesis includes irradiation, e.g., UV, X-ray or gamma-ray radiation, or chemical mutagenesis, e.g., ethyl methanesulfonate (EMS), ethyl nitrosourea (ENU), or dimethyl sulfate (DMS).
[0133] 117. The method according to any of claims 111 to 116, wherein the mutation or mutagenesis includes tilling.
[0134] 118. The method according to any one of claims 111 to 115, wherein the mutation or mutagenesis involves the use of site-specific (mutant) DNA or RNA nuclease.
[0135] 119. The method according to claim 118, wherein the site-directed (mutant) DNA or RNA nuclease is selected from the group comprising meganucleases (MN), zinc finger nucleases (ZFN), transcriptional activator-like effector nucleases (TALEN), (mutant) Cas nuclease / effector proteins, such as Cas9 nuclease, Cfp1 nuclease, MAD7 nuclease, dCas9-FokI, dCpf1-FokI, dMAD7 nuclease-FokI, chimeric Cas9-cytidine deaminase, chimeric Cas9-adenine deaminase, chimeric FENI-FokI, and mega-TAL, nickase Cas9 (nCas9), chimeric dCas9 non-FokI nuclease, dCpf1 non-FokI nuclease, and dMAD7 non-FokI nuclease.
[0136] 120. The method according to any of claims 111 to 115, wherein the mutation or mutagenesis involves the use of a CRISPR / Cas system.
[0137] 121. The method according to claim 120, wherein the CRISPR / Cas system comprises a guide RNA and a Cas effector protein, and optionally tracrRNA.
[0138] 122. The method according to claim 121, wherein the Cas effector protein is Cas9 or Cas12(Cpf1).
[0139] 123. The method according to claim 121 or 122, wherein the Cas effector protein is a nickase or a catalytically inactive Cas-effective protein.
[0140] 124. The method according to any one of claims 121 to 123, wherein the Cas effector protein is a heterogeneous protein (domain), preferably a heterogeneous protein domain having enzymatic activity.
[0141] 125. The method according to any one of claims 121 to 124, wherein the Cas effector protein fuses to an adenine deaminase or cytidine deaminase (domain).
[0142] 126. Maize (Zea mays) seeds deposited under NCIMB deposit number NCIMB43772.
[0143] 127. (igEIN) Maize (Zea Mays) seeds, a representative sample deposited under NCIMB deposit number NCIMB43772.
[0144] 128. A zea mays plant grown from or obtained from the seeds described in claim 126 or 127.
[0145] 129. Plant parts of zea mays grown from or obtained from seeds described in claim 126 or 127, or obtained from the plant described in claim 128.
[0146] 130. A method for identifying or selecting a plant or plant part, for example, a plant or plant part having (enhanced) singular-inducing activity or the ability thereof, i) To prepare a plant or plant part having reduced expression, stability, and / or activity of genes, mRNA, or proteins of an undetermined gametophyte (ig). ii) Mutating a centromere or kinetochore protein, preferably a gene encoding CENH3, and iii) Analyze the singular induction activity or ability in the plant or plant part, or their offspring. This includes, and optionally further iv) Selecting a plant or plant part that has enhanced singular induction activity or the ability thereof. A method for identifying or selecting a plant or plant part, including the method described above.
[0147] 131. A method for identifying or selecting a plant or plant part, for example, a plant or plant part having (enhanced) singular-inducing activity or the ability thereof, i) Prepare a first plant or plant part having reduced expression, stability, and / or activity of genes, mRNA, or proteins of an undetermined gametophyte (ig). ii) Crossing the first plant with a second plant having a gene encoding a mutated centromere or kinetochore protein, preferably CENH3, and iii) Analyze the singular induction activity or ability in the offspring obtained, This includes, and optionally further iv) Selecting a plant or plant part that has enhanced singular induction activity or the ability thereof. A method for identifying or selecting a plant or plant part, including the method described above.
[0148] 132. Use of plants or plant parts having reduced expression, stability, and / or activity of undetermined gametophyte (ig) genes, mRNA, or proteins to screen or identify mutations in centromere or kinetochore proteins, preferably CENH3, that confer or enhance singular-inducing activity or its ability. [Brief explanation of the drawing]
[0149] [Figure 1] Protein alignments of various CENH3 orthologues. The amino acid sequences shown are the wild-type CENH3 protein sequences provided as SEQ ID NO: 12 for Arabidopsis thaliana, SEQ ID NO: 34 for beet (Beta vulgaris), SEQ ID NO: 16 for Brassica napus, SEQ ID NO: 14 for maize (Zea mays), and SEQ ID NO: 18 for sorghum bicolor.
[0150] Detailed description of the invention Before describing the systems and methods of the present invention, it should be understood that the present invention is not limited to the specific systems and methods or combinations described, for such systems and methods and combinations are, of course, subject to change. It should also be understood that the scope of the present invention is limited only by the appended claims, and therefore the terms used herein are not intended to be restrictive.
[0151] As used herein, the singular forms "a," "an," and "the" include both singular and plural referents unless otherwise clearly indicated herein.
[0152] As used herein, the terms “comprising,” “comprises,” and “comprised of” are synonymous with “including,” “includes,” or “containing,” and are inclusive or unenclosed, and do not exclude any additional uncited members, elements, or method steps. As used herein, the terms “comprising,” “comprises,” and “comprised of” are understood to include the terms “consisting of,” “consists,” and “consists of,” as well as the terms “consisting essentially of,” “consists essentially,” and “consists essentially of.”
[0153] The enumeration of numerical ranges by endpoints includes all numerical values and fractions contained within each range, as well as the enumerated endpoints.
[0154] The terms “about” or “approximately” as used herein to indicate measurable values, such as parameters, quantities, time durations, etc., mean to include variations of + / -20%, preferably + / -10%, more preferably + / -5%, and even more preferably + / -1% or less of the specified value, insofar as such variations are appropriate for carrying out the disclosed invention. It should be understood that the values referred to by the modifiers “about” or “approximately” are themselves also specifically and preferably disclosed.
[0155] The terms “one or more” or “at least one,” for example, one or more or at least one group element of a group of group elements, are clear in themselves, and by further examples, the terms encompass, among other things, any one of the group elements, or any two or more of the group elements, for example, any ≥3, ≥4, ≥5, ≥6, or ≥7 of the group elements, and references to all of the group elements.
[0156] All references cited herein are incorporated herein by reference in their entirety. In particular, all teachings of references specifically mentioned herein are incorporated herein by reference.
[0157] Unless otherwise defined, all terms used in disclosing this invention, including technical and scientific terms, have meanings generally understood by those skilled in the art. Further instructions may include definitions of terms to better understand the teachings of this invention.
[0158] Standard reference books explaining the general principles of recombinant DNA technology include the following molecular cloning: A Laboratory Manual, 4th ed., (Green and Sambrook et al., 2012, Cold Spring Harbor Laboratory Press); Current Protocols in Molecular Biology, ed. Ausubel et al., Greene Publishing and Wiley-Interscience, New York, 1992 (with periodic updates) ("Ausubel et al. 1992"); the series Methods in Enzymology (Academic Press, Inc.); Innis et al., PCR Protocols: A Guide to Methods and Applications, Academic Press: San Diego, 1990; PCR 2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995); Harlow and Lane, eds. (1988) Antibodies, a Laboratory Manual; and Animal Cell Culture (RI Freshney, ed. (1987). General principles of microbiology are outlined, for example, in Davis, BD et al., Microbiology, 3rd edition, Harper & Row, publishers, Philadelphia, Pa. (1980).
[0159] Different embodiments of the present invention are defined in more detail in the following sections. Each of the embodiments defined in this way may be combined with other embodiments unless otherwise explicitly indicated. In particular, any feature indicated as preferred or advantageous may be combined with any other or more features indicated as preferred or advantageous.
[0160] Throughout this specification, any reference to “one embodiment” or “embodiment” means that a particular feature, structure, or characteristic described in relation to an embodiment is included in at least one embodiment of the present invention. Therefore, the appearance of the phrase “in one embodiment” or “in an embodiment” in various places throughout this specification does not necessarily refer to the same embodiment, but may be so. Furthermore, certain features, structures, or characteristics may be combined in one or more embodiments in any suitable manner, as will be apparent to those skilled in the art from this disclosure. Moreover, while some embodiments described herein include some features but not others included in other embodiments, combinations of features from different embodiments mean that they form different embodiments, within the scope of the present invention and as will be understood by those skilled in the art. For example, in the appended claims, any of the claimed embodiments may be used in any combination.
[0161] In the following detailed description of the invention, references are made to the accompanying drawings, which form part of this specification and are shown solely for illustrative purposes of specific embodiments that can carry out the invention. It should be understood that other embodiments may be used, and structural or logical modifications may be made, without departing from the scope of the invention. Accordingly, the following detailed description should not be construed as restrictive, and the scope of the invention is defined by the appended claims.
[0162] Preferred claims (features) and embodiments of the present invention are listed below herein. Each of the claims and embodiments of the present invention as defined herein may be combined with other claims and / or embodiments unless otherwise expressly indicated. In particular, any feature shown as preferred or advantageous may be combined with any other feature or feature or claim shown as preferred or advantageous.
[0163] In one embodiment, the present invention relates to a plant or plant part that comprises or expresses a polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and a mutated centromere or kinetochore protein, preferably a polynucleic acid encoding a mutated CENH3.
[0164] In one embodiment, the present invention relates to a plant or plant part that contains or expresses a mutated undetermined gametophyte (ig) allele and a mutated centromere or kinetochore protein allele, preferably a mutated CENH3.
[0165] In one embodiment, the present invention relates to a plant or plant part that contains or expresses a mutated undetermined gametophyte (ig) gene and a mutated centromere or kinetochore gene, preferably a mutated CENH3.
[0166] In one embodiment, the present invention relates to a plant or plant part that contains or expresses a mutated undetermined gametophyte (ig) protein and a mutated centromere or kinetochore protein, preferably a mutated CENH3.
[0167] In one embodiment, the present invention relates to a plant or plant part that contains, or expresses, a polynucleic acid encoding an undetermined gametophyte (ig) protein that confers or enhances singular-inducing activity or ability, and a polynucleic acid encoding a centromere or kinetochore protein, preferably CENH3, that confers or enhances singular-inducing activity or ability.
[0168] In one embodiment, the present invention relates to a plant or plant part that contains or expresses an undetermined gametophyte (ig) allele that confers or enhances singular-inducing activity or the ability thereof, and a centromere or kinetochore protein allele, preferably CENH3, that confers or enhances singular-inducing activity or the ability thereof.
[0169] In one embodiment, the present invention relates to a plant or plant part that contains or expresses an undetermined gametophyte (ig) gene that confers or enhances singular induction activity or the ability thereof, and a centromere or kinetochore gene, preferably CENH3, that confers or enhances singular induction activity or the ability thereof.
[0170] In one embodiment, the present invention relates to a plant or plant part that contains or expresses an undetermined gametophyte (ig) protein that confers or enhances singular induction activity or the ability thereof, and a centromere or kinetochore protein, preferably CENH3, that confers or enhances singular induction activity or the ability thereof.
[0171] In one embodiment, the present invention relates to a plant or plant part comprising a polynucleic acid encoding a mutated centromere or kinetochore protein, preferably a mutated CENH3, which has reduced expression, stability, and / or activity of an undetermined gametophyte (ig) gene, mRNA, or protein.
[0172] In one embodiment, the present invention relates to a plant or plant part comprising a mutated centromere or kinetochore protein allele, preferably mutated CENH3, which has reduced expression, stability, and / or activity of an undetermined gametophyte (ig) gene, mRNA, or protein.
[0173] In one embodiment, the present invention relates to a plant or plant part comprising a mutated centromere or kinetochore gene, preferably a mutated CENH3, which has reduced expression, stability, and / or activity of an undetermined gametophyte (ig) gene, mRNA, or protein.
[0174] In one embodiment, the present invention relates to a plant or plant part having reduced expression, stability, and / or activity of an undetermined gametophyte (ig) gene, mRNA, or protein, and comprising a mutated centromere or kinetochore protein, preferably a mutated CENH3.
[0175] In one embodiment, the present invention relates to a plant or plant part comprising a centromere or kinetochore protein, preferably a polynucleic acid encoding CENH3, which has reduced expression, stability and / or activity of an undetermined gametophyte (ig) gene, mRNA or protein, and which confers or enhances singular induction activity or its ability.
[0176] In one embodiment, the present invention relates to a plant or plant part comprising a centromere or kinetochore protein allele, preferably CENH3, which has reduced expression, stability and / or activity of an undetermined gametophyte (ig) gene, mRNA or protein, and which confers or enhances singular induction activity or its ability.
[0177] In one embodiment, the present invention relates to a plant or plant part comprising a centromere or kinetochore gene, preferably CENH3, which has reduced expression, stability and / or activity of an undetermined gametophyte (ig) gene, mRNA or protein, and which confers or enhances singular induction activity or its ability.
[0178] In one embodiment, the present invention relates to a plant or plant part comprising a centromere or kinetochore protein, preferably CENH3, which has reduced expression, stability and / or activity of an undetermined gametophyte (ig) gene, mRNA or protein, and which confers or enhances singular induction activity or its ability.
[0179] In one embodiment, the present invention is a method for identifying or selecting a plant or plant part having (enhanced) singular induction activity or the ability thereof, i) To prepare a plant or plant part having reduced expression, stability and / or activity of an undetermined gametophyte (ig) gene, mRNA or protein, for example, the ig gene according to the present invention as described herein. ii) Mutating a centromere or kinetochore protein, preferably a gene encoding CENH3, and iii) Analyze the singular induction activity or ability in the plant or plant part, or their offspring. This includes, and optionally further iv) Selecting a plant or plant part that has enhanced singular induction activity or the ability thereof. The present invention relates to a method for identifying or selecting plants or plant parts, including the present invention.
[0180] Such methods enable the identification of centromere or kinetochore proteins, preferably CENH3, suitable for combination with mutated ig to generate a singular inducer or to enhance singular induction. Mutagenesis of centromere or kinetochore proteins can be carried out as described elsewhere herein, but is not limited to, random mutagenesis, e.g., tilling, or site-directed mutagenesis, e.g., genome editing (e.g., CRISPR / Cas-mediated).
[0181] In one embodiment, the present invention is a method for identifying or selecting a plant or plant part having (enhanced) singular induction activity or the ability thereof, i) To prepare a plant having reduced expression, stability, and / or activity of an undetermined gametophyte (ig) gene, mRNA, or protein, such as the ig gene according to the present invention as described herein. ii) Crossing the plant with a plant having a gene encoding a mutated centromere or kinetochore protein, preferably CENH3, and iii) Analyze the singular induction activity or ability in the offspring obtained, This includes, and optionally further iv) Selecting a plant or plant part that has enhanced singular induction activity or the ability thereof. The present invention relates to a method for identifying or selecting plants or plant parts, including the present invention.
[0182] Such a method allows for the identification of centromere or kinetochore proteins, preferably CENH3, that are suitable for combination with mutated ig to generate a singular inducer or to enhance singular induction.
[0183] In related embodiments, the present invention relates to the use of centromere or kinetochore proteins, preferably CENH3, to confer or enhance singular-inducible activity or ability to confer singular-inducible activity to plants or plant parts having reduced expression, stability and / or activity of undeterministic gametophyte (ig) genes, mRNA or proteins, such as the ig gene according to the present invention as described herein.
[0184] Those skilled in the art will understand that the analysis of (enhanced) singular-inducing activity or capacity may involve determining the amount or proportion of singular-inducing substance obtained from a population of seeds or other plant parts, such as reproductive plant parts. The enhanced singular-inducing activity or capacity may be identified by a (relative) increase in the amount of singular-inducing substance (offspring).
[0185] The term “plant” according to the present invention includes the entire plant or parts of such plant. The entire plant is preferably a seed plant or a crop. The term “plant” according to the present invention includes the entire plant or parts of such plant. The entire plant is preferably a seed plant or a crop. “Plant parts” include, for example, the vegetative organs / structures of shoots, e.g., leaves, stems and tubers; roots, flowers and floral organs / structures, e.g., bracts, sepals, petals, stamens, carpels, anthers and ovules; pollen including embryos, endosperm and seed coats, seeds; fruits and mature ovaries; plant tissues, e.g., vascular tissue, basal tissue, etc.; and cells, e.g., guard cells, egg cells, pollen, trichomes, etc.; and similar offspring. Plant parts may be attached to or separated from an intact plant. Such parts of a plant include, but are not limited to, the organs, tissues and cells of a plant, preferably pollen (or seeds). “Plant cells” are the structural and physiological units of a plant, including protoplasts and cell walls. Plant cells may be in the form of isolated single cells or cultured cells, or parts of more highly organized units, such as plant tissue, plant organs, or whole plants. “Plant cell culture” means plant units, such as protoplasts, cell cultures, cells in plant tissue, pollen, pollen tubes, ovaries, embryo sacs, zygotes, and cultures of embryos at various stages of development. “Plant material” means leaves, stems, roots, flowers or parts of flowers, fruits, pollen, egg cells, zygotes, pollen, seeds, cuttings, cell or tissue cultures, or other parts or products of a plant. It also includes callus or callus tissue, and extracts (e.g., extracts from the taproot) or samples. “Plant organ” is a clearly visually structured and differentiated part of a plant, such as a root, stem, leaf, flower bud, or embryo. As used herein, “plant tissue” means a group of plant cells organized into structural and functional units. It includes plant tissue in plants or cultures. This term includes, but is not limited to, the whole plant, plant organs, plant pollen, plant seeds, tissue cultures, and any group of plant cells organized into structural and / or functional units. The use of this term in combination with, or in the absence of, any particular type of plant tissue described above or included in this definition is not intended to exclude any other type of plant tissue.In certain embodiments, the plant part or derivative is not (functional) reproductive material, such as germplasm, seeds, or plant embryos, or other material capable of regenerating a plant. In certain embodiments, the plant part or derivative does not include (functional) male and female reproductive organs. In certain embodiments, the plant part or derivative is or includes reproductive material, but is reproductive material that is no longer used or can not be used to produce or generate new plants, such as reproductive material that has been rendered nonfunctional by other means, such as chemical, mechanical, or heat treatment, acid treatment, compression, crushing, shredding, etc. In certain embodiments, the plant part or derivative is (functional) reproductive material, such as germplasm, seeds, or plant embryos, or other material capable of regenerating a plant. In certain embodiments, the plant part or derivative includes (functional) male and female reproductive organs.
[0186] As used herein, the terms “offspring” and “offspring plant” refer to a plant produced from vegetative or sexual reproduction from one or more parent plants. In phantom induction via gynogenesis, the phantom embryo of the female parent contains female chromosomes but excludes male chromosomes, and therefore it is not an offspring of a male phantom induction line. Phantom maize seeds typically still have the normal triploid endosperm containing the male genome. Edited phantom offspring and subsequent edited doubled phantom plants and their seeds are not the only desired offspring. Often, seeds from the phantom induction material line itself, which often possesses the Cas9 transgene, and offspring of subsequent plants and seeds of phantom induction plants are also present. Both phantom seeds and phantom induction material (self-pollination derived) seeds may be offspring. Offspring plants are obtained by cloning or self-pollinating a single parent plant, or by crossing two or more parent plants. For example, offspring plants are obtained by cloning or self-pollination of parent plants, or by crossing two parent plants, and include self-pollination and F1 or F2 or yet another generation. F1 is the first generation offspring produced from a parent in which at least one of the parents is initially used as a donor of a genetic trait, while the offspring of the second generation (F2) or subsequent generations (F3, F4, etc.) are samples made from self-pollination, hybridization, backcrossing, and / or other crosses such as F1 and F2. Thus, F1 (and in some embodiments) may be a hybrid resulting from a cross between two purebred parents (i.e., the purebred parents are homozygous for the trait of interest or their alleles, respectively), while F2 (and in some embodiments) may be an offspring resulting from self-pollination of the F1 hybrid. The term “offspring” can be used indistinguishably from “offspring” in certain embodiments, particularly when the plant or plant material originates from a sexual cross of parent plants.
[0187] In certain embodiments, plants include crop plants, such as cash crops or subsistence plants, such as food or non-food crops, including agricultural, horticultural, floricultural, or industrial crops. The term crop plant has the common meaning known in the art. With further guidance, without limitation, crops are plants cultivated by humans for food and other resources, and may be cultivated and harvested extensively for profit or subsistence, typically in agricultural environments or situations.
[0188] In the description of this invention, unless otherwise specified, "plant" may be any species from dicotyledonous plants, monocotyledonous plants, and gymnosperms. Examples without restrictions include barley (Hordeum vulgare), sorghum (Sorghum bicolor), rye (Secale cereale), rye wheat (Triticale), sugarcane (Saccharum officinarium), maize (Zea mays), millet (Setaria italic), rice (Oryza sativa), Oryza minuta, Oryza australiensis, Oryza alta, bread wheat (Triticum aestivum), durum wheat (Triticum durum), Hordeum bulbosum, Brachypodiurn distachyon, seaweed (Hordeum marinum), and corn wheat (Aegilops). tauschii), sugar beet (Beta vulgaris), sunflower (Helianthus annuus), Australian bush louse (Daucus glochidiatus), Daucus pusillus, Daucus muricatus, wild carrot (Daucus carota), Eucalyptus grandis, Erythranthe guttata, Genlisea aurea, Gossypium species, Musa species, Avena species, Nicotiana sylvestris, tobacco (Nicotiana tabacum), Nicotiana tomentosiformis (Nicotiana (Tomentosiformis), tomato (Solanum lycopersicum), potato (Solanum tuberosum), robusta coffee tree (Coffea canephora), grape (Vitis vinifera), cucumber (Cucumis)sativus), Morus notabilis, Arabidopsis thaliana, Arabidopsis lyrata, Arabidopsis arenosa, Crucihimalaya himalaica, Crucihimalaya wallichii, Cardamine flexuosa, Lepidiurn virginicum, Capsella bursa-pastoris, Olmarabidopsis pumila, Arabis hirsuta, Brassica napus, Brassica oleracea This includes oleracea, Brassica rapa, Brassica juncacea, Brassica nigra, radish (Raphanus sativus), arugula (Eruca vesicaria sativa), orange (Citrus sinensis), tung tree (Jatropha curcas), soybean (Glycine max), and black cottonwood (Populus trichocarpa). Preferably, the plants used herein belong to the genera Zea, preferably maize (Zea mays), sorghum, preferably sorghum bicolor, and Brassica, preferably rapeseed (Brassica napus).
[0189] As used herein, “maize” means the Zea mays species, preferably the Zea mays ssp. mays plant.
[0190] As used herein, “sorghum” means plants of the genus Sorghum, including but not limited to Sorghum bicolor, Sorghum sudanense, Sorghum bicolor × Sorghum sudanense, Sorghum × almum (Sorghum bicolor × Sorghum halepense), Sorghum arundinaceum, Sorghum × drummondii, Sorghum halepense, and / or Sorghum propinquum.
[0191] As used herein, “rapeseed” means plants of the genus Brassica, including, but not limited to, Brassica napus, preferably Brassica napus ssp. napus. Rapeseed includes canola, Brassica oleracea, Brassica rapa, Brassica juncacea, and / or Brassica nigra.
[0192] As used herein, unless otherwise specified, the term “plant” is intended to mean a plant at any stage of development.
[0193] As used herein, the term “plant (part) group” may be used without distinction from a group of plants or plant parts. A plant (part) group preferably includes a large number of individual plants (or their plant parts), for example, preferably at least 10, for example 20, 30, 40, 50, 60, 70, 80, or 90, more preferably at least 100, for example 200, 300, 400, 500, 600, 700, 800, or 900, and even more preferably at least 1000, for example at least 10000 or at least 100000.
[0194] In certain embodiments, a plant population (or a part thereof) is a line, strain, or variety of a plant. In certain embodiments, a plant population (or a part thereof) is not a line, strain, or variety of a plant. In certain embodiments, a plant population (or a part thereof) is a line, strain, or variety of an inbred plant. In certain embodiments, a plant population (or a part thereof) is not a line, strain, or variety of an inbred plant. In certain embodiments, a plant population (or a part thereof) is a line, strain, or variety of a non-inbred plant. In certain embodiments, a plant population (or a part thereof) is not a line, strain, or variety of a non-inbred plant.
[0195] As used herein, the terms “phenotype,” “phenotypic trait,” or “trait” refer to one or more traits of a plant or plant cell. Phenotypes may be observable with the naked eye or by other means of evaluation known in the art, such as microscopy, biochemical analysis, or electromechanical assays. In some cases, the phenotype is directly controlled by a single gene or locus (i.e., corresponding to a “single gene trait”). In the case of singular induction, the use of color markers, such as R Navajo, and other markers, including transgenes visualized by the presence or absence of color in the seed, indicates that the seed is an induced singular seed. The use of R Navajo as a color marker and the use of transgenes are well known in the art as means of detecting the induction of singular seeds in female plants. In other cases, the phenotype is the result of interactions between several genes, and in some embodiments, also arises from interactions between the plant and / or plant cell and its environment.
[0196] As used herein, the term “sequence” refers to nucleotide sequences, polynucleotides, nucleic acid sequences, nucleic acids, nucleic acid molecules, peptides, polypeptides, and proteins, depending on how the term “sequence” is used in this description.
[0197] The terms “polynucleic acid,” “nucleotide sequence,” “polynucleotide,” “nucleic acid sequence,” “nucleic acid,” and “nucleic acid molecule” are used interchangeably herein and refer to any combination of nucleotides, ribonucleotides, or deoxyribonucleotides in an unbranched polymer form of any length. A nucleic acid sequence may include DNA, RNA, cDNA, genomic DNA, RNA, synthetic forms, and mixed polymers, both sense and antisense strands, or non-natural or derived nucleotide bases, as will be readily understood by those skilled in the art.
[0198] As used herein, the terms “polypeptide” or “protein” (both terms are used interchangeably herein) mean a peptide, protein, or polypeptide comprising an amino acid chain of a given length, in which amino acid residues are linked by covalent peptide bonds. However, peptide mimes of such proteins / polypeptides in which amino acids and / or peptide bonds are replaced by functional analogs are also included in the present invention and include, in addition to the 20 gene-coding amino acids, e.g., selenocysteine. Peptides, oligopeptides, and proteins are sometimes referred to as polypeptides. The term polypeptide is also not exclusive and may include modifications of polypeptides, e.g., glycosylation, acetylation, phosphorylation, etc. Such modifications are described in detail in the basic text and more detailed monographs, as well as in the research literature.
[0199] As used herein, the term “gene” means a polymer form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term includes double-stranded and single-stranded DNA and RNA. It also includes known types of modifications, such as methylation, “caps,” and substitution of one or more naturally occurring nucleotides in analogs. Preferably, a gene includes a coding sequence that encodes a polypeptide as defined herein. A “coding sequence” is a nucleotide sequence that, when under the control of or with appropriate regulatory sequences, is transcribed into mRNA and / or translated into a polypeptide. The boundaries of a coding sequence are determined by a translation start codon at the 5' end and a translation stop codon at the 3' end. A coding sequence may include, but is not limited to, mRNA, cDNA, recombinant nucleic acid sequences, or genomic DNA, while introns may be present under certain circumstances.
[0200] As used herein, the term “endogenous” refers to a gene or allele that is present in its natural genomic location. The term “endogenous” can be used interchangeably with “natural.” However, this does not preclude the existence of one or more nucleic acid differences from the wild-type allele due to naturally occurring polymorphisms. In certain embodiments, the difference from the wild-type allele may be limited to less than 9 nucleotides, preferably less than 6 nucleotides, and more particularly less than 3 nucleotides. More specifically, the difference from the wild-type sequence may be as little as 1 nucleotide. As used herein, the term “endogenous” may refer to a gene or allele that has not been introduced into the plant (or its ancestor) by genetic engineering techniques or (artificial) mutagenesis. Naturally occurring variations / variations may also be considered endogenous. The term “endogenous” can be used interchangeably with “natural” or “wild-type.” In contrast to artificially introduced mutations or polymorphisms, all naturally occurring polymorphisms may be considered endogenous, natural, and / or wild-type. Nevertheless, if a naturally occurring polymorphism (e.g., a naturally occurring ig mutation that confers singularity-inducing activity) has a specific phenotypic effect, such polymorphism may be considered a mutation in the context of this invention. Polymorphisms or mutations that do not occur naturally, such as those introduced by random mutagenesis, may be considered exogenous, non-natural, or genetically modified.
[0201] The term “locus” (plural of “loci”) refers to a specific location or site on a chromosome where a genomic region of interest, such as a QTL, gene, or genetic marker, is found. Haplotypes can be defined by the distinctive fingerprint of an allele at each marker within a specific window. As used herein, the terms “(one) allele” or “(multiple) alleles” refer to one or more alternative forms of a locus, i.e., different nucleotide sequences. Typically, an allele refers to an alternative form of a gene or any kind of identifiable genetic element that is an alternative in heredity because it is located at the same locus on homologous chromosomes. In diploid cells or organisms, two alleles of a given gene (or marker) typically occupy corresponding loci on a pair of homologous chromosomes.
[0202] A “marker” is a location (or means of finding it) on a genetic or physical map, or a link between a marker and a trait locus (a locus that affects a trait). The location detected by the marker may be known through the detection of polymorphic alleles and their genetic mapping, or by hybridization, sequence matching, or amplification of physically mapped sequences. Markers may be DNA markers (detecting DNA polymorphisms), proteins (detecting mutations in encoded polypeptides), or simply inherited phenotypes (e.g., the “waxy” phenotype). DNA markers may be developed from genomic nucleotide sequences or expressed nucleotide sequences (e.g., from spliced RNA or cDNA). Depending on the DNA marker technology, a marker may consist of complementary primers adjacent to the locus and / or complementary probes that hybridize to the polymorphic allele at the locus. The term marker locus is the locus (gene, sequence, or nucleotide) detected by the marker. A “marker,” “molecular marker,” or “marker locus” may be used to indicate a nucleic acid sequence or amino acid sequence that is sufficiently unique to characterize a particular locus on the genome. Detectable polymorphic traits can be used as markers insofar as they are inherited differently and exhibit a linkage imbalance with the desired phenotypic trait.
[0203] Markers for detecting genetic polymorphisms among members of a population are well-established in the art. Markers can be defined by the type of polymorphism to be detected and the marker technology used to detect it. Marker types include, but are not limited to, the detection of restriction fragment length polymorphisms (RFLPs), isozyme markers, randomly amplified polymorphic DNA (RAPDs), amplified fragment length polymorphisms (AFLPs), simple sequence repeats (SSRs), amplified variable sequences of the plant genome, auto-persistent sequence replication, or single nucleotide polymorphisms (SNPs). SNPs may be detected by, for example, DNA sequencing, PCR-based sequence-specific amplification, allele-specific hybridization (ASH) for detecting polynucleotide polymorphisms, dynamic allele-specific hybridization (DASH), molecular beacons, microarray hybridization, oligonucleotide ligase assays, flap endonucleases, 5' endonucleases, primer extension, single-strand conformational polymorphisms (SSCPs), or temperature gradient gel electrophoresis (TGGE). DNA sequencing, such as pyrosequencing, has the advantage of being able to detect a series of linked SNP alleles that constitute a haplotype. Haplotypes tend to be more useful than SNPs (they detect higher levels of polymorphism). A "marker allele," or "allele of a marker locus," may represent one of several polymorphic nucleotide sequences found at a marker locus in a population. With respect to SNP markers, the allele indicates the specific nucleotide base present at that SNP locus in that individual plant.
[0204] Marker-assisted selection (MAS) is a process of selecting individual plants based on marker genotypes. Marker-assisted counter-selection is a process used to identify plants that are not selected for a marker genotype, allowing them to be removed from breeding programs or planting. Marker-assisted selection uses the presence of molecular markers genetically linked to specific loci or chromosomal regions (e.g., introgression fragments, transgenes, polymorphisms, mutations, etc.) to select plants for the presence of a specific locus or region (e.g., introgression fragments, transgenes, polymorphisms, mutations, etc.). For example, plants containing a target genomic region can be detected and / or selected using molecular markers genetically linked to a target genomic region as defined herein. The closer the genetic linkage of the molecular marker to a locus (e.g., approximately 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.5 cM or less), the less likely the marker is to be separated from the locus by meiotic recombination. Similarly, the closer two markers are to each other (e.g., within the range of 7 or 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, or less), the less likely the two markers are to separate from each other (and the more likely they are to separate simultaneously as a single unit). A marker "within 7 cM or within 5 cM, 3 cM, 2 cM, or 1 cM" of another marker refers to a marker that is genetically located within a region of 7 cM or 5 cM, 3 cM, 2 cM, or 1 cM adjacent to (i.e., on both sides of) the marker. Similarly, markers within the ranges of 5Mb, 3Mb, 2.5Mb, 2Mb, 1Mb, 0.5Mb, 0.4Mb, 0.3Mb, 0.2Mb, 0.1Mb, 50kb, 20kb, 10kb, 5kb, 2kb, 1kb or less refer to markers physically located within the ranges of 5Mb, 3Mb, 2.5Mb, 2Mb, 1Mb, 0.5Mb, 0.4Mb, 0.3Mb, 0.2Mb, 0.1Mb, 50kb, 20kb, 10kb, 5kb, 2kb, 1kb or less in genomic DNA regions adjacent to the marker (i.e., on both sides of the marker).The "LOD score" (logarithm of odds, base 10) is a statistical test often used in linkage analysis of animal and plant populations. The LOD ("logarithm of odds") score compares the probability of obtaining test data when two loci (molecular marker loci and / or phenotypic trait loci) are actually linked to the probability of observing the same data purely by chance. A positive LOD score supports the existence of linkage, and an LOD score greater than 3.0 is considered evidence of linkage. An LOD score of +3 indicates that the probability of the observed linkage not occurring by chance is 1 in 1000.
[0205] Centi Morgan ("cM") is a unit of measurement for recombination frequency. 1 cM corresponds to a 1% probability that a marker at one gene locus will be separated from a marker at a second gene locus in a single generation of crosses.
[0206] The “physical distance” between loci on the same chromosome (e.g., between molecular markers and / or phenotypic markers) is the actual physical distance expressed in bases or base pairs (bp), kilobases or kilobase pairs (kb), or megabases or megabase pairs (Mb).
[0207] The "genetic distance" between loci on the same chromosome (e.g., between molecular markers and / or phenotypic markers) is measured by the crossover frequency, or recombination frequency (RF), expressed in centimorgan units (cM). 1 cM corresponds to a 1% recombination frequency. If no recombinants are found, the RF is zero, and the loci are physically very close or identical. The further apart the two loci are, the higher the RF.
[0208] A "marker haplotype" refers to the combination of alleles in a marker location.
[0209] A "marker locus" is a specific chromosomal location within a species' genome where a particular marker can be found. Marker loci can be used to track the presence of a second linked locus, for example, one that influences the expression of a phenotypic trait. For instance, marker loci can be used to monitor allele segregation at genetically or physically linked loci.
[0210] A "marker probe" is a nucleic acid sequence or nucleic acid molecule that can be used by nucleic acid hybridization to identify the presence of a nucleic acid probe at a marker locus, for example, a marker locus sequence that is complementary to it. A marker probe containing 30 or more consecutive nucleotides of a marker locus ("all or part" of the marker locus sequence) may be used in nucleic acid hybridization. Alternatively, in some embodiments, a marker probe refers to any type of probe (i.e., genotype) that can distinguish a particular allele present at a marker locus.
[0211] The term “molecular marker” may be used to refer to a genetic marker or its encoding product (e.g., a protein) used as a reference point when identifying linked loci. Markers may originate from genomic nucleotide sequences or expressed nucleotide sequences (e.g., spliced RNA, cDNA, etc.), or encoded polypeptides. The term refers to nucleic acid sequences that are complementary to or adjacent to the marker sequence, for example, nucleic acids used as probe or primer pairs that can amplify the marker sequence. A “molecular marker probe” is a nucleic acid sequence or diffusing molecule that can be used to identify the presence of a marker locus, for example, a nucleic acid probe that is complementary to the marker locus sequence. Alternatively, in some embodiments, a marker probe refers to any type of probe (i.e., genotype) that can distinguish a particular allele present at the marker locus. Nucleic acids are “complementary” if they hybridize specifically in solution, for example, according to the Watson-Crick base pairing rule. Some of the markers described herein are also called hybridization markers if they are located in indel regions, for example, non-collinear regions as described herein. This is because, by definition, the insertion region is a polymorphism compared to plants without the insertion. Therefore, the marker only needs to indicate whether or not the indel region is present. Any suitable marker detection technique can be used to identify such hybridization markers, for example, SNP technique is used in the examples provided herein.
[0212] A “genetic marker” is a nucleic acid that is polymorphic in a population and whose allele can be detected and identified by one or more analytical methods, such as RFLP, AFLP, isozymes, SNPs, SSRs, etc. The terms “molecular marker” and “genetic marker” are used interchangeably herein. This term also refers to nucleic acid sequences complementary to the genome sequence, such as nucleic acids used as probes. Markers corresponding to genetic polymorphisms among members of a population can be detected by methods well established in the art. These include, for example, sequence-specific amplification based on PCR, detection of restriction fragment length polymorphisms (RFLPs), detection of isozyme markers, detection of polynucleotide polymorphisms by allele-specific hybridization (ASH), detection of amplified variable sequences in plant genomes, detection of auto-persistent sequence replication, detection of simple sequence repeats (SSRs), detection of single nucleotide polymorphisms (SNPs), or detection of amplified fragment length polymorphisms (AFLPs). Well-established methods are also known for detecting expression sequence tags (ESTs) and SSR markers derived from EST sequences and randomly amplified polymorphic DNA (RAPDs). Without limitation, screening may include or encompass splicing, hybridization-based methods (e.g., (dynamic) allele-specific hybridization, molecular beacons, SNP microarrays), enzyme-based methods (e.g., PCR, KASP (Kompetitive AlleleSpecific PCR), RFLP, ALFP, RAPD, flap endonuclease, primer extension, 5'-nuclease, oligonucleotide ligation assays), and post-amplification methods based on the physical properties of DNA (e.g., single nucleotide polymorphism, temperature gradient gel electrophoresis, denaturing high-performance liquid chromatography, high-resolution melting of whole amplicons, use of DNA mismatch-binding proteins, SNPlex, surveyor nuclease assays).
[0213] In this application, the terms “linked” or “closely linked” mean that recombination between two linked loci occurs at a frequency of about 20% or less (i.e., they are separated by 20 cM or less on the gene map). In other words, closely linked loci co-separate with a probability of at least 80%. Marker loci are particularly useful with respect to the subject matter of this disclosure if they indicate a significant possibility of co-separation (linking) with a desired trait. Closely linked loci, such as a marker locus and a second locus, may exhibit locus recombination frequencies of 20% or less, for example, 10% or less, preferably about 9% or less, more preferably about 8% or less, more preferably about 7% or less, more preferably about 6% or less, more preferably about 5% or less, more preferably about 4% or less, more preferably about 3% or less, and more preferably about 2% or less. In a very preferred embodiment, the relevant loci exhibit recombination at a frequency of about 1% or less, for example, about 0.75% or less, more preferably about 0.5% or less, or more preferably about 0.25% or less. Two loci that are localized on the same chromosome and are located close enough that recombination between them occurs at a frequency of less than 20%, for example, less than 10% (e.g., approximately 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or less), are also said to be "proximally close" to each other. In some cases, two different markers may have the same gene map coordinates. In that case, the two markers are very close to each other, and recombination between them occurs at a frequency too low to be detected.
[0214] "Linkage" refers to the tendency for alleles to separate more frequently than would be expected by chance if allele transmission were independent. Typically, linkage refers to alleles on the same chromosome. Genetic recombination is assumed to occur at random frequencies throughout the genome. Gene maps are constructed by measuring the frequency of recombination between pairs of traits or markers. The closer traits or markers are to each other on a chromosome, the lower the frequency of recombination and the greater the degree of linkage. Traits or markers are considered linked if they generally segregate simultaneously. A 1 / 100 probability of recombination per generation is defined as a gene map distance of 1.0 centimorgan (1.0 cM). The term "linkage disequilibrium" refers to the non-random segregation of a locus or trait (or both). In either case, linkage disequilibrium means that related loci segregate together at a higher frequency (i.e., non-random) than random because they are in sufficient physical proximity along the length of the chromosome. Markers exhibiting linkage disequilibrium are considered linked. Linkage loci co-separate with a probability greater than 50%, for example, between approximately 51% and 100%. In other words, two markers that co-separate have a recombination frequency of less than 50% (by definition, they are in the same linkage group and less than 50 cM apart). As used herein, linkage may be between two markers, or between a marker and a phenotypic locus, such as a genomic region of interest as defined elsewhere herein. Marker loci can be “associated” (linked) with a trait. The degree of linkage between a marker locus and a phenotypic locus is measured, for example, as the statistical probability of co-separation of the molecular marker and the phenotype (e.g., F-statistic or LOD score).
[0215] Genetic elements or genes located on a single chromosome segment are physically linked. In some embodiments, the two loci are located close together so that recombination between homologous chromosome pairs does not occur frequently between the two loci during meiosis, for example, so that the linked loci are simultaneously separated with a probability of at least about 80%, preferably at least 90%, for example, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.75%, or higher. Genetic elements located within chromosome segments are also "genetically linked," typically below 50 cM, for example, around 49 cM, 48 cM, 47 cM, 46 cM, 45 cM, 44 cM, 43 cM, 42 cM, 41 cM, 40 cM, 39 cM, 38 cM, 37 cM, 36 cM, 35 cM, 34 cM, 33 cM, 32 cM, 31 cM, 30 cM, 29 cM, 28 cM, and 27 cM. , 26cM, 25cM, 24cM, 23cM, 22cM, 21cM, 20cM, 19cM, 18cM, 17cM, 16cM, 15cM, 14cM, 13cM, 12cM, 11cM, 10cM, 9cM, 8cM, 7cM, 6cM, 5cM, 4cM, 3cM, 2cM, 1cM, 0.75cM, 0.5cM, 0.25cM or less in terms of genetic recombination distance of several hundred meters. In other words, two genetic elements within a single chromosome segment undergo recombination with each other during meiosis at frequencies of approximately 50% or less, for example, approximately 49%, 48%, 47%, 46%, 45%, 44%, 43%, 42%, 41%, 40%, 39%, 38%, 37%, 36%, 35%, 34%, 33%, 32%, 31%, 30%, 29%, 28%, 27%, 26%, 25%, 24%, 23%, 22%, 21%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or less."Closely linked" markers exhibit crossover frequencies with a given marker of approximately 10% or less, e.g., 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or less (a given marker locus is within approximately 10 cM of a closely linked marker locus, e.g., 9 cM, 8 cM, 7 cM, 6 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.75 cM, 0.5 cM, 0.25 cM, or less of a closely linked marker locus). In other words, closely linked marker loci will separate simultaneously with a probability of at least approximately 80%, for example, at least 90%, for example, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.75%, or higher.
[0216] As used herein, the terms “introgression,” “introgressed,” and “intogressing” refer to both natural and artificial processes in which a chromosomal fragment or gene from one species, variety, or cultivar is transferred to the genome of another species, variety, or cultivar by crossing those species. This process may be optionally terminated by backcrossing with a repeating parent. For example, the transfer of a desired allele at a particular locus can be transmitted to at least one offspring through sexual crossing between two parents of the same species, such that at least one of the parents has the desired allele in its genome. Alternatively, for example, allele transfer may occur by recombination between two donor genomes in a fused protoplast, for example, in which at least one of the donor protoplasts has the desired allele in its genome. The desired allele may be detected by markers associated with, for example, phenotype, QTL, or transgene. In any case, offspring containing a desired allele can be repeatedly backcrossed with a line having a desired genetic background, and the desired allele can be selected to be fixed to the selected genetic background. The process of “transfer” is often referred to as “backcrossing” when this process is repeated two or more times. A “transferred fragment,” “transferred segment,” or “transferred region” is a chromosome fragment (or chromosome part or region) that has been introduced into another plant of the same or related species, artificially or naturally, for example, by hybridization or traditional breeding techniques, such as backcrossing; that is, a transferred fragment is the result of a breeding method (e.g., backcrossing) referred to by the verb “transfer.” It is understood that the term “transferred fragment” never includes an entire chromosome, but only a portion of a chromosome.The gene transfer fragment may be large, even three-quarters or half the size of a chromosome, but it is preferable that it be smaller than approximately 15 Mb or less, for example, approximately 10 Mb or less, approximately 9 Mb or less, approximately 8 Mb or less, approximately 7 Mb or less, approximately 6 Mb or less, approximately 5 Mb or less, approximately 4 Mb or less, approximately 3 Mb or less, approximately 2.5 Mb or 2 Mb or less, approximately 1 Mb (corresponding to 1,000,000 base pairs) or approximately 0.5 Mb (corresponding to 500,000 base pairs), for example, approximately 200,000 bp (corresponding to 200 kilobase pairs) or less, approximately 100,000 bp (100 kb) or less, approximately 50,000 bp (50 kb) or less, approximately 25,000 bp (25 kb) or less.
[0217] A genetic element, gene transfer fragment, or gene or allele conferring a trait described herein is said to be “obtained from,” or may be “available from,” or “inducible from,” or may be “derived from,” or “present in,” or “found in,” a plant or plant part described elsewhere herein, if, apart from the addition of a trait conferred by a genetic element, locus, gene transfer fragment, gene, or allele described herein, it can be transferred using traditional breeding techniques from a plant in which it exists to another plant (e.g., a lineage or variety) in which it does not exist, without causing a change in the phenotypic characteristics of the recipient plant. This term is used without distinction, and a genetic element, locus, gene transfer fragment, gene, marker, or allele can be transferred to another genetic background lacking the trait. Not only plants containing a genetic element, locus, gene transfer fragment, gene, or allele, but also progeny / offspring from such plants selected to retain a genetic element, locus, gene transfer fragment, gene, or allele can be used and are included herein. Whether a plant (or the genomic DNA, cells, or tissues of a plant) contains the same genetic elements, loci, gene transfer fragments, genes, or alleles as those obtained from such plant can be determined by a person skilled in the art using one or more techniques known in the art, such as phenotypic assays, whole-genome sequencing, molecular marker analysis, trait mapping, chromosome coloring, allele testing, or a combination of such techniques. It will be understood that transgenic plants may also be included.
[0218] As used herein, the terms “genetic engineering,” “transformation,” and “genetic modification” are all synonymous with the transfer of isolated and loaned genes into the DNA, usually chromosomal DNA or genome, of another organism.
[0219] As used herein, “transgenic” or “genetically modified” (GMO) refers to an organism whose genetic material has been altered using a technique commonly known as “recombinant DNA technology.” Recombinant DNA technology encompasses the ability to combine DNA molecules from different sources into a single molecule ex vivo (e.g., in a test tube). This term generally does not refer to organisms whose genetic composition has been altered by conventional crossbreeding or “mutagenic” breeding, as these methods predate the discovery of recombinant DNA technology. As used herein, “non-transgenic” refers to plants and plant-derived foods that are not “transgenic” or “genetically modified” as defined above.
[0220] An "introduced gene" or "chimeric gene" refers to a DNA sequence, such as a recombinant gene, that has been introduced into the plant genome through transformation, for example, Agrobacterium-mediated transformation. Plants containing such introduced genes that have been stably incorporated into their genome are called "transgenic plants."
[0221] As used herein, the term “homozygous” means an individual cell or plant having identical alleles at one or more or all loci. When this term is used in relation to a particular locus or gene, it means that at least that locus or gene has identical alleles. As used herein, the term “homozygous” means the genetic state that exists when identical alleles are located at corresponding loci on homologous chromosomes. Thus, for diploid organisms, two alleles are identical, and for tetraploid organisms, four alleles are identical. As used herein, the term “heterozygous” means an individual cell or plant having different alleles at one or more or all loci. When this term is used in relation to a particular locus or gene, it means that at least that locus or gene has different alleles. Thus, for diploid organisms, two alleles are not identical, and for tetraploid organisms, four alleles are not identical (i.e., at least one allele is different from the others). As used herein, the term “heterozygous” means the genetic state that exists when different alleles are located at corresponding loci on homologous chromosomes. In some embodiments, the proteins, genes, or coding sequences described herein are homozygous. In some embodiments, the proteins, genes, or coding sequences described herein are heterozygous. In some embodiments, the protein, gene, or coding sequence alleles described herein are homozygous. In some embodiments, the protein, gene, or coding sequence alleles described herein are heterozygous. It will be understood that homozygosity or heterozygosity preferably relates to a locus containing at least a gene, i.e., a gene (or a coding sequence derived therefrom, or a protein encoded therefrom). However, to elaborate further, homozygosity or heterozygosity may refer equally to a particular mutation, for example, the mutations described herein. Therefore, a particular mutation may be considered homozygous (i.e., all alleles harbor the mutation), but the remainder of the gene, coding sequence, or protein, for example, may include differences between alleles.
[0222] In some embodiments, the mutations defined herein are homozygous. Thus, in diploid plants, two alleles are identical (at least with respect to a particular mutation), in tetraploid plants, four alleles are identical, and in hexaploid plants, six alleles are identical with respect to the mutation or marker. In some embodiments, the mutations / markers defined herein are heterozygous. Thus, in diploid plants, two alleles are not identical, in tetraploid plants, four alleles are not identical (e.g., only one, two, or three alleles constitute a particular mutation / marker), and in hexaploid plants, six alleles are not identical with respect to the mutation or marker (e.g., only one, two, three, four, or five alleles constitute a particular mutation / marker). Similar considerations apply to pseudoploid plants.
[0223] The term "singular" refers to a state in which a plant or a plant cell, organ, or tissue has the number of chromosome sets normally found in its gametes, i.e., pollen or ovules. Typically, singular refers to half the number of chromosomes normally found in somatic cells. A singular cell (or plant) can have more than one set of chromosomes, especially in polyploid plants. For example, a plant whose somatic cells are tetraploid (four sets of chromosomes) produces gametes containing two sets of chromosomes through meiosis. These gametes are numerically diploid, but are still sometimes called singular. Therefore, a singular plant derived from a plant that is normally tetraploid contains two sets of chromosomes. Another name for such a plant is a digenomic singular. Similarly, a singular plant derived from a plant that is normally hexaploid contains three sets of chromosomes. Another name for such a plant is a trigenomic singular.
[0224] The terms “singularity inducer” and “singularity inducer” are used herein as synonyms and refer to plants that can produce fertilized seeds or embryos having a set of singular chromosomes through crosses with plants of the same genus, preferably plants of the same species that are not singularity inducers. Mechanistically, singularity induction results from the removal of one parental chromosome after fertilization. Since singularity induction is often a trait of moderate to low penetrance in derivative lines, the resulting offspring may be either diploid (when no genome loss occurs) or singular (when genome loss actually occurs), depending on the species or circumstances. Singularity can be selected by any suitable means known in the art (e.g., by markers, cytology, karyotype analysis, etc.). In one embodiment, the singularity inducer used herein can produce at least 0.1% singular offspring. In another embodiment, the singularity inducer used herein can produce at least 0.5% singular offspring. In one embodiment, the singularity inducer used herein can produce at least 1% singular offspring. In one embodiment, the singularity inducer used herein can produce at least 2% singular offspring. In one embodiment, the singularity inducer used herein can produce at least 3% singular offspring. In one embodiment, the singularity inducer used herein can produce at least 4% singular offspring. In one embodiment, the singularity inducer used herein can produce at least 5%, for example, at least 6%, or at least 7% singular offspring. It is understood that a particular gene or protein encoded by it, in particular the (mutated) gene described herein, is a singularity inducer or confers the ability thereof, or is an enhancer of a singularity inducer or its ability. Therefore, in one embodiment, each of the genes or protein products encoded thereby, individually or in combination, confers at least 0.1%, for example, at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or at least 7% of singular inducible substance / inducible activity or ability thereof.In one embodiment, the combined gene or protein product encoded thereby enhances the singular inducer / inducing activity or its ability by at least 0.1%, for example, at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or at least 7%, compared to the singular inducer rate of a plant containing only one of such genes or protein products encoded thereby.
[0225] As used herein, the term “singularity-inducing ability or activity enhancer” means a (mutated) gene of the protein it encodes, which may or may not confuse singularity-inducing activity on its own, but which, when combined with another (mutated) gene or the protein it encodes, increases singularity-inducing ability or activity compared to the presence of the other (mutated) gene or the protein it encodes alone. In some embodiments, the increase in singular offspring is at least 0.1%, e.g., at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or at least 7% (meaning the final (average) singularity induction rate of plants containing both (mutated) proteins). "Enhancing or increasing the haploid induction ability of a haploid inducer" or "mediating the enhancer properties of the singular induction ability of a singular inducer" means that by using polynucleic acids encoding mutated proteins as described herein, the haploid induction rate of a singular inducer can be increased, preferably by at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%, preferably by at least 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, or 5%, more preferably by at least 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, or 50% (an increase in induction rate compared to a single (mutated) protein). The number of fertilized seeds or embryos having a phantom chromosome set and resulting from the cross between a phantom inducer and a plant of the same genus (preferably the same species) that is not a phantom inducer is therefore at least 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, or 0.9%, preferably at least 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, or 5%, more preferably at least 6%, 7%, 8%, 9%, 10%, 15%, 20%, 30%, or 50%, which is greater than the number of phantom fertilized seeds or embryos achieved without using the nucleic acids described herein.
[0226] The term “singular induction rate” refers to the (average) percentage of singular offspring produced or that can be produced by a singular inducer. In some embodiments, each of such genes or protein products encoded by it confers or enhances singular inducer / inducing activity or its capacity by at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or at least 7%. The term “singular induction rate” refers to the (average) percentage of singular offspring produced or that can be produced by a singular inducer. In some embodiments, a combination of such genes or protein products encoded by it confers or enhances singular inducer / inducing activity or its capacity by at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, or at least 7%. The term "singular induction rate" refers to the (average) percentage of singular offspring produced or that can be produced by a singular inducer.
[0227] The term "patrilineal singularity inducer" or "patrilineal singularity inducer" refers to a male plant being a singularity inducer. Therefore, after fertilizing a female non-singularity inducer with a patrilineal (i.e., male) singularity inducer, chromosomes originating from the male / patrilineal singularity inducer are lost. Consequently, the resulting singular plant contains only chromosomes derived from the female. This process of singularity inducement is also called gynogenesis. The term "patrilineal singularity inducer rate" refers to the (average) percentage of singular offspring produced or capable of being produced by a patrilineal singularity inducer.
[0228] The terms "maternal singular inducer" or "maternal singular induction" refer to the female plant being a singular inducer. Therefore, after fertilizing a maternal (i.e., female) singular inducer with a non-singular inducer, the chromosomes originating from the female / maternal singular inducer are lost. Consequently, the resulting singular plant contains only male-derived chromosomes. This process of singular induction is also called androgenesis. The term "maternal singular induction rate" refers to the (average) percentage of singular offspring produced or capable of being produced by a maternal singular inducer.
[0229] As used herein, the terms “mutation” or “mutated” refer to a gene or protein product that has been altered or modified in such a way that the function normally associated with the gene or protein product is changed, or that the expression, stability, and / or activity normally associated with the gene or protein product is altered. Typically, mutations as referred herein result in phenotypic effects, such as singular induction, as described elsewhere herein. Mutations in a gene or protein product are understood to be referred to in comparison to a gene or protein product that does not have such mutation, such as a wild-type or endogenous gene or protein product. Typically, mutations refer to alterations at the DNA level and include genetic and / or epigenetic changes. Genetic changes may include insertions, deletions, introduction of stop codons, base changes (e.g., transitions or transversions), or changes at splice junctions. These changes may occur in coding or non-coding regions of an endogenous DNA sequence (e.g., promoter regions, exons, introns, or splice junctions). For example, a genetic change may be an exchange (including insertions and deletions) of at least one nucleotide in the endogenous DNA sequence or in a regulatory sequence of the endogenous DNA sequence. For example, if such a nucleotide exchange occurs in a promoter, it may lead to a change in promoter activity, for example, because the cis regulatory element has been modified such that the affinity of the transcription factor to the mutated cis regulatory element is altered compared to a wild-type promoter, and as a result, the activity of the promoter containing the mutated cis regulatory element increases or decreases depending on whether the transcription factor is a repressor or an inducer, or whether the affinity of the transcription factor to the mutated cis regulatory element is increased or decreased. If such a nucleotide exchange occurs, for example, in the coding region of the endogenous DNA sequence, it may lead to an amino acid exchange in the encoded protein, which may result in a change in the protein's activity or stability compared to a wild-type protein. An epigenetic change may occur via the altered DNA methylation pattern.In some embodiments, the mutations referred to herein relate to the insertion of one or more nucleotides in a gene. In some embodiments, the mutations referred to herein relate to the deletion of one or more nucleotides in a gene. In some embodiments, the mutations referred to herein relate to the deletion and insertion of one or more nucleotides. In some embodiments, a specific nucleotide elongation is deleted, for example, encoding a particular protein region. In some embodiments, a specific nucleotide elongation is deleted, for example, encoding a particular protein region, and replaced with a nucleotide sequence encoding a different protein region (see, for example, the “GFP-tailswap” CENH3 variant described elsewhere in this specification, for example, the whole of which is incorporated by reference, see Kelliher et al. (2016). “Maternal haploids are preferentially induced by CENH3-tailswap transgenic complementation in maize.” Frontiers in plant science, 7, 414). In some embodiments, the mutations referred to herein relate to the substitution of one or more nucleotides in a gene by different nucleotides. In some embodiments, the mutation is a nonsense mutation (i.e., the mutation results in the generation of a stop codon in a protein-coding sequence). In some embodiments, the mutation is a frameshift mutation (i.e., the insertion or deletion of one or more nucleotides (not 3 and / or its products) in a protein-coding sequence). In some embodiments, the mutation results in a cleaved protein product. In some embodiments, the mutation results in an N-terminal cleaved protein product. In some embodiments, the mutation results in a C-terminal cleaved protein product. In some embodiments, the mutation results in both N-terminal and C-terminal cleaved protein products. In some embodiments, the mutation results in an altered splice site (e.g., an altered splice donor and / or splice acceptor site). In some embodiments, the mutation is in an exon. In some embodiments, the mutation is in an intron.In some embodiments, the mutation is located in a regulatory sequence, such as a promoter. In some embodiments, the mutation results in a codon encoding a different amino acid. In some embodiments, the mutation results in the insertion or deletion of one or more codons (i.e., a nucleotide triplet). In some embodiments, the mutation is a knockout mutation. Both frameshift mutations and nonsense mutations may be considered knockout mutations in some embodiments, particularly when the mutation is located in an early exon. As used herein, a knockout mutation preferably means that a functional gene product, such as a functional protein, is no longer produced. In particular, frameshift mutations and nonsense mutations lead to premature termination of protein translation, resulting in a cleaved protein that often lacks the stability and / or activity necessary to perform its naturally occurring function. In some embodiments, the mutation is a knockdown mutation. In contrast to a knockout mutation, a knockdown mutation results in a decrease in the activity, stability, and / or expression rate of a naturally occurring functional gene product, such as a protein, thereby ultimately resulting in a decrease in function. For example, a mutation in a promoter region affecting activator binding (or other regulatory sequences), particularly a decrease in transcription rate, may be considered a knockdown mutation. Mutations that adversely affect protein stability (e.g., increased ubiquitination and subsequent proteolysis) are also considered knockdown mutations. Furthermore, mutations that adversely affect protein activity (e.g., binding strength or enzyme activity) are also considered knockdown mutations. Mutations described herein in accordance with the present invention are understood to confer singular inducer or inducible activity or ability thereof, or to enhance haploid inducer or inducible activity or ability thereof, as described elsewhere herein. Mutations described herein may not occur naturally, but this is not necessarily required. For example, as described elsewhere herein, several naturally occurring mutations that confer singular inducer activity are described for undeterministic gametophyte (ig) genes. In some embodiments, the term “mutated protein” may be used interchangeably with “singular inducer protein” or “singular donor protein,” etc.As used herein, mutated proteins, genes, alleles, or coding sequences (i.e., polynucleic acids encoding proteins) may be used without distinction from proteins, genes, alleles, or coding sequences that confer or enhance singular-inducing activity or ability as described elsewhere herein.
[0230] In some embodiments, the wild-type / endogenous allele is replaced by a mutated allele, preferably all wild-type / endogenous alleles are replaced by the mutated allele. The replacement may be carried out by any means known in the art, as described elsewhere in this specification. The replacements used herein include (direct) mutagenesis of the wild-type / endogenous allele at its natural genomic locus. Thus, in some embodiments, the wild-type / endogenous allele is mutated, preferably all wild-type / endogenous alleles are mutated, as described elsewhere in this specification. Those skilled in the art will understand that only one copy of the wild-type / endogenous allele may be mutated, and homozygosity (if desired) may be obtained by self-pollination and subsequent selection. In some embodiments, a reduced number of wild-type / endogenous alleles exist (i.e., the wild-type / endogenous alleles are heterozygous).
[0231] In one embodiment, the wild-type / endogenous allele is knocked out, preferably all wild-type / endogenous alleles are knocked out, and the mutated allele is introduced transgenically and transiently or genomically, preferably genomically. In another embodiment, the wild-type / endogenous allele is knocked out, preferably all wild-type / endogenous alleles are knocked out, and transgenically replaced by the mutated allele (at the natural genomic location of the wild-type allele). Those skilled in the art will understand that only one copy of the wild-type / endogenous allele may be knocked out, and homozygosity (if desired) may be obtained by self-pollination and subsequent selection.
[0232] In some embodiments, the mutations described herein, such as the ig mutation or the CENH3 mutation, are amino acid substitutions or result in amino acid substitutions (compared to the wild-type or unmutated protein, gene, or coding sequence). In some embodiments, the mutation is a point mutation. Preferably, the mutation is a missense mutation (i.e., the mutation results in a codon encoding a different amino acid). In some embodiments, there is one or more mutations. In some embodiments, there are 1 to 10 mutations. In some embodiments, there are 1 to 9 mutations. In some embodiments, there are 1 to 8 mutations. In some embodiments, there are 1 to 7 mutations. In some embodiments, there are 1 to 6 mutations. In some embodiments, there are 1 to 5 mutations. In some embodiments, there are 1 to 4 mutations. In some embodiments, there are 1 to 3 mutations. In some embodiments, there are 1 to 2 mutations. In some embodiments, there is 1 mutation. In some embodiments, 1 to 10 amino acid substitutions are present in the mutated protein. In some embodiments, 1 to 9 amino acid substitutions are present in the mutated protein. In some embodiments, 1 to 8 amino acid substitutions are present in the mutated protein. In some embodiments, 1 to 7 amino acid substitutions are present in the mutated protein. In one embodiment, 1 to 6 amino acid substitutions are present in the mutated protein. In one embodiment, 1 to 5 amino acid substitutions are present in the mutated protein. In one embodiment, 1 to 4 amino acid substitutions are present in the mutated protein. In one embodiment, 1 to 3 amino acid substitutions are present in the mutated protein. In one embodiment, 1 to 2 amino acid substitutions are present in the mutated protein. In one embodiment, 1 amino acid substitution is present in the mutated protein. In one embodiment, 1 to 10 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 9 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 8 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence.In one embodiment, 1 to 7 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 6 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 5 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 4 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 3 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 to 2 point mutations, preferably missense point mutations, are present in the mutated gene, allele, or coding sequence. In one embodiment, 1 point mutation, preferably missense point mutation, is present in the mutated gene, allele, or coding sequence.
[0233] The terms “indeterminate gametophyte” or “ig” refer to wild-type indeterminate gametophyte genes or the protein products encoded therein. In the literature, the term indeterminate gametophyte is recognized as potentially referring to mutant genes or their phenotypes, i.e., singular induction; however, as used herein, unless otherwise explicitly specified, the term refers to unmutated genes (or the proteins encoded therein), i.e., ig1 genes that do not or do not confer singular induction activity. In this description, it will be understood that ig1 genes that do not or do not confer singular induction activity refer to ig1 genes with a singular induction rate of less than 1%, preferably less than 0.5%, and more preferably less than 0.1%. In contrast, the term “mutant indeterminate gametophyte” refers to mutant genes, such as naturally occurring mutations, e.g., ig-O (ig1-O) or ig-mum (ig1-mum) that confer or enhance singular induction activity, and artificially induced mutations. At least three ig genes have been identified (see, for example, U.S. Patent Application Publication No. 2009 / 0151025, which is incorporated herein by reference in its entirety): ig1, ig2, and ig3. Preferably, according to the present invention, the ig gene is ig1. Ig1 promotes the transition from proliferation to differentiation in the embryo sac. It is a negative regulator of cell proliferation on the adaxial side of the leaf and regulates the formation of symmetrical leaflets and the establishment of the venous phase. Ig1 directly interacts with RS2 (rough sheath2) and represses several knox homeobox genes (see Evans (2007) “The indeterminate gametophyte1 Gene of Maize Encodes a LOB Domain Protein Required for Embryo Sac and Leaf Development”; The Plant Cell; 19:46-62 (which is incorporated herein by reference in its entirety)). The ig1 gene is also known as “LOB Domain Protein 6”.
[0234] In plants of the genus Zea, such as maize (Zea mays), the ig protein (i.e., wild-type ig) may have, contain, or consist of the protein sequence shown in SEQ ID NO: 9 or 10, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 9 or 10. In plants of the genus Zea, such as maize (Zea mays), the ig gene (i.e., wild-type ig) may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 6, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 6. In plants of the genus Zea, such as maize (Zea mays), the ig coding sequence (i.e., wild-type ig) may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 7 or 8, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 7 or 8. The maize (Zea mays) ig protein, gene, or coding sequence is preferably the ig1 protein, gene, or coding sequence. In plants of the genus Brassica, such as rapeseed (Brassica napus), the ig protein (i.e., wild-type ig) may have, contain, or consist of the protein sequence shown in SEQ ID NO: 29 or 32, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 29 or 32. In plants of the genus Brassica, such as Brassica napus, the ig gene (i.e., wild-type ig) may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 27 or 30, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 27 or 30.In plants of the genus Brassica, such as Brassica napus, the ig coding sequence (i.e., wild-type ig) may have, contain, or consist of a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to the nucleic acid sequence shown in SEQ ID NO: 28 or 31. The Brassica napus ig protein, gene, or coding sequence is preferably an ortholog of the Zea mays ig (preferably ig1) protein, gene, or coding sequence. In plants of the genus Sorghum, such as Sorghum bicolor, the ig protein (i.e., wild-type ig) may have, contain, or consist of a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to the protein sequence shown in SEQ ID NO: 23 or 26. In plants of the genus Sorghum, such as Sorghum bicolor, the ig gene (i.e., wild-type ig) may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 21 or 24, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 21 or 24. In plants of the genus Sorghum, such as Sorghum bicolor, the ig coding sequence (i.e., wild-type ig) may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 22 or 25, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 22 or 25. The protein, gene, or coding sequence of Sorghum bicolor ig is preferably an ortholog of the protein, gene, or coding sequence of Zea mays ig (preferably ig1).
[0235] In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is preferably at least 80% identical to the sequence shown in SEQ ID NOs: 9, 10, 29, 32, 23, or 26, over its entire length. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is preferably at least 85% identical to the sequence shown in SEQ ID NOs: 9, 10, 29, 32, 23, or 26, over its entire length. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is preferably at least 90% identical to the sequence shown in SEQ ID NOs: 9, 10, 29, 32, 23, or 26, over its entire length. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is preferably at least 95% identical to the sequence shown in SEQ ID NOs: 9, 10, 29, 32, 23, or 26, over its entire length. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is preferably at least 98% identical to the sequence shown in SEQ ID NOs: 9, 10, 29, 32, 23, or 26, over its entire length. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 99% identical to the sequence shown in SEQ ID NOs: 9, 10, 29, 32, 23, or 26, preferably over its entire length.
[0236] In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 80% identical to the sequence shown in sequence numbers 9, 10, 29, 32, 23, or 26, preferably, with respect to the sequence of the LOB domain of ig. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 85% identical to the sequence shown in sequence numbers 9, 10, 29, 32, 23, or 26, preferably, with respect to the sequence of the LOB domain of ig. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 90% identical to the sequence shown in sequence numbers 9, 10, 29, 32, 23, or 26, preferably, with respect to the sequence of the LOB domain of ig. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 95% identical to the sequence shown in sequence numbers 9, 10, 29, 32, 23, or 26, preferably, with respect to the sequence of the LOB domain of ig. In one embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 98% identical to the sequence shown in sequence numbers 9, 10, 29, 32, 23, or 26, preferably, with respect to the sequence of the LOB domain of ig. In another embodiment, the indeterminate gametophyte gene encodes a protein having a sequence that is at least 99% identical to the sequence shown in sequence numbers 9, 10, 29, 32, 23, or 26, preferably, with respect to the sequence of the LOB domain of ig.
[0237] In one embodiment, the indeterminate gametophyte gene encodes a protein containing a region having a sequence that is at least 80% identical to amino acids 30-145 of the sequence shown in SEQ ID NO: 9 or 10, or a corresponding region in SEQ ID NO: 23, 26, 29, or 31. In another embodiment, the indeterminate gametophyte gene encodes a protein containing a sequence that is at least 85% identical to amino acids 30-145 of the sequence shown in SEQ ID NO: 9 or 10, or a corresponding region in SEQ ID NO: 23, 26, 29, or 31. In yet another embodiment, the indeterminate gametophyte gene encodes a protein containing a sequence that is at least 90% identical to amino acids 30-145 of the sequence shown in SEQ ID NO: 9 or 10, or a corresponding region in SEQ ID NO: 23, 26, 29, or 31. In yet another embodiment, the indeterminate gametophyte gene encodes a protein containing a sequence that is at least 95% identical to amino acids 30-145 of the sequence shown in SEQ ID NO: 9 or 10, or a corresponding region in SEQ ID NO: 23, 26, 29, or 31. In one embodiment, the undetermined gametophyte gene encodes a protein containing a sequence that is at least 98% identical to amino acids 30-145 of the sequence shown in SEQ ID NO: 9 or 10, or a corresponding region in SEQ ID NO: 23, 26, 29, or 31. In another embodiment, the undetermined gametophyte gene encodes a protein containing a sequence that is at least 99% identical to amino acids 30-145 of the sequence shown in SEQ ID NO: 9 or 10, or a corresponding region in SEQ ID NO: 23, 26, 29, or 31. It will be understood that the sequence variant still maintains wild-type ig function. In one embodiment, ig is an ortholog of zea mays ig, sorghum bicolor ig, or Brassica napus ig. In another embodiment, ig1 is an ortholog of zea mays ig1, sorghum bicolor ig1, or Brassica napus ig1.
[0238] In some embodiments, a mutated Ig gene, or an Ig gene that confers or enhances singular-inducing activity or its ability, includes the insertion of one or more nucleotides. In some embodiments, a mutated Ig coding sequence, or an Ig coding sequence that confers or enhances singular-inducing activity or its ability, includes the insertion of one or more nucleotides. In some embodiments, a polynucleic acid encoding a mutated Ig protein, or a polynucleic acid encoding an Ig protein that confers or enhances singular-inducing activity or its ability, includes the insertion of one or more nucleotides. In some embodiments, the insertion is the insertion of 1 to 1000 nucleotides. In some embodiments, the insertion is the insertion of 1 to 500 nucleotides. In some embodiments, the insertion is the insertion of 1 to 300 nucleotides. In some embodiments, the insertion is the insertion of 1 to 200 nucleotides. In some embodiments, the insertion is the insertion of 10 to 1000 nucleotides. In some embodiments, the insertion is the insertion of 10 to 500 nucleotides. In some embodiments, the insertion is the insertion of 10 to 300 nucleotides. In one embodiment, the insertion is an insertion of 10 to 200 nucleotides. In one embodiment, the insertion is an insertion of 10 to 100 nucleotides. In one embodiment, the insertion is an insertion of 10 to 100 nucleotides. In one embodiment, the insertion is an insertion of 100 to 1000 nucleotides. In one embodiment, the insertion is an insertion of 100 to 500 nucleotides. In one embodiment, the insertion is an insertion of 100 to 300 nucleotides. In one embodiment, the insertion is an insertion of 100 to 200 nucleotides. In one embodiment, the insertion is an insertion of 200 to 1000 nucleotides. In one embodiment, the insertion is an insertion of 200 to 500 nucleotides. In one embodiment, the insertion is an insertion of 200 to 300 nucleotides. Preferably, the insertion is not a 3-nucleotide product. Those skilled in the art will understand that the presence of an insertion is comparable to that of a mutant or wild-type or ig that does not confuse or enhance singular-inducible activity or its ability.
[0239] In one embodiment, the insertion of one or more nucleotides is the insertion of one or more nucleotides in a region or sequence that encodes a LOB domain. In maize (Zea mays), the LOB domain corresponds to amino acids 32-133, for example, amino acids 32-133 of SEQ ID NO: 9 or 10. Those skilled in the art can determine the corresponding positions that depict the LOB domain in an orthologous ig gene or protein.
[0240] In one embodiment, the insertion of one or more nucleotides is the insertion of one or more nucleotides in the exon encoding the first protein. In zea mays, the exon encoding the first protein is exon 2 (exon 1 is the 5'UTR exon). In zea mays, the exon encoding the first protein corresponds to nucleotide positions 431-841 of the ig gene, for example, nucleotide positions 431-841 of sequence number 6. Those skilled in the art can determine the corresponding positions that describe the exon encoding the first protein in an orthologous ig gene or protein.
[0241] In one embodiment, the insertion of one or more nucleotides is the insertion of one or more nucleotides in an intron, for example, an intron that precedes the exon encoding the first protein. In zea mays, the intron that precedes the exon encoding the first protein is intron 1. The insertion of one or more nucleic acids in an intron preferably affects splicing and results in a decrease in (wild-type) Ig expression.
[0242] In one embodiment, the mutated ig gene (or coding sequence), or the ig gene (or coding sequence) that confers or enhances singular induction activity or its ability, corresponds to the ig1-O allele. In another embodiment, the mutated ig gene, or the ig gene that confers or enhances singular induction activity or its ability, corresponds to the ig1-mum allele.
[0243] In one embodiment, the mutated ig gene (or coding sequence) or the ig gene (or coding sequence) that confers or enhances singular induction activity or its ability includes the insertion of one or more nucleic acids in the ig codon corresponding to the codon shown in, for example, SEQ ID NO: 7 or 8, selected from codons 118, 119, or 120 of the wild-type maize (Zea mays) ig coding sequence.
[0244] In one embodiment, the mutated ig gene (or coding sequence) or the ig gene (or coding sequence) that confers or enhances singular induction activity or its ability includes the insertion of one or more nucleic acids in an ig codon corresponding to the codon shown in, for example, SEQ ID NO: 22, selected from codons 191, 192, or 193 of the wild-type sorghum (Sorghum bicolor) ig coding sequence.
[0245] In one embodiment, the mutated ig gene (or coding sequence) or the ig gene (or coding sequence) that confers or enhances singular induction activity or its ability includes the insertion of one or more nucleic acids in an ig codon corresponding to the codon shown in, for example, SEQ ID NO: 25, selected from codons 143, 144, or 145 of the wild-type sorghum (Sorghum bicolor) ig coding sequence.
[0246] In one embodiment, the mutated ig gene (or coding sequence) or the ig gene (or coding sequence) that confers or enhances singular induction activity or its ability comprises the insertion of one or more nucleic acids in an ig codon corresponding to a codon shown in, for example, SEQ ID NO: 28 or 31, selected from codons 94, 95, or 96 of the wild-type Brassica napus ig coding sequence.
[0247] In some embodiments, a mutated Ig gene, or an Ig gene that confers or enhances singular-inducing activity or its ability, includes a frameshift mutation. In some embodiments, a mutated Ig coding sequence, or an Ig coding sequence that confers or enhances singular-inducing activity or its ability, includes a frameshift mutation. In some embodiments, a polynucleic acid encoding a mutated Ig protein, or a polynucleic acid encoding an Ig protein that confers or enhances singular-inducing activity or its ability, includes a frameshift mutation. A frameshift mutation is an insertion or deletion of one or more nucleotides that is not a three-nucleotide product. Preferably, a frameshift mutation is an insertion or deletion of one or two nucleotides. Those skilled in the art will understand that the presence of a frameshift mutation is compared to a mutant or wild-type Ig, or Ig that does not confer or enhance singular-inducing activity or its ability.
[0248] In some embodiments, a mutated Ig gene, or an Ig gene that confers or enhances singular-inducing activity or its ability, includes a nonsense mutation. In some embodiments, a mutated Ig coding sequence, or an Ig coding sequence that confers or enhances singular-inducing activity or its ability, includes a nonsense mutation. In some embodiments, a polynucleic acid encoding a mutated Ig protein, or a polynucleic acid encoding an Ig protein that confers or enhances singular-inducing activity or its ability, includes a nonsense mutation. A nonsense mutation is a mutation in which the amino acid encoding a codon is mutated into a stop codon. Those skilled in the art will understand that the presence of a nonsense mutation is compared to a mutant or wild-type Ig, or Ig that does not confer or enhance singular-inducing activity or its ability.
[0249] In some embodiments, a mutated Ig gene, or an Ig gene that confers or enhances singular-inducible activity or ability, includes a point mutation. In some embodiments, a mutated Ig coding sequence, or an Ig coding sequence that confers or enhances singular-inducible activity or ability, includes a point mutation. In some embodiments, a polynucleic acid encoding a mutated Ig protein, or a polynucleic acid encoding an Ig protein that confers or enhances singular-inducible activity or ability, includes a point mutation. A point mutation is a single nucleotide substitution. Preferably, a point mutation is a missense mutation (i.e., a mutation in a codon resulting in a different codon encoding a different amino acid). Those skilled in the art will understand that the presence of a point mutation is compared to a mutant or wild-type Ig, or Ig that does not confer or enhance singular-inducible activity or ability.
[0250] In some embodiments, a mutated Ig gene, or an Ig gene that confers or enhances singular-inducing activity or its ability, includes a knockout mutation. In some embodiments, a mutated Ig coding sequence, or an Ig coding sequence that confers or enhances singular-inducing activity or its ability, includes a knockout mutation. In some embodiments, a polynucleic acid encoding a mutated Ig protein, or a polynucleic acid encoding an Ig protein that confers or enhances singular-inducing activity or its ability, includes a knockout mutation. Those skilled in the art will understand that the presence of a knockout mutation is compared to a mutant or wild-type Ig, or Ig that does not confer or enhance singular-inducing activity or its ability.
[0251] In some embodiments, a mutated ig gene, or an ig gene that confers or enhances singular-inducing activity or its ability, includes a knockdown mutation. In some embodiments, a mutated ig coding sequence, or an ig coding sequence that confers or enhances singular-inducing activity or its ability, includes a knockdown mutation. In some embodiments, a polynucleic acid encoding a mutated ig protein, or a polynucleic acid encoding an ig protein that confers or enhances singular-inducing activity or its ability, includes a knockdown mutation. Those skilled in the art will understand that the presence of a knockout mutation is compared to a mutant or wild-type ig, or an ig that does not confer or enhance singular-inducing activity or its ability. Those skilled in the art will understand that the same effect can be achieved instead of a knockdown mutation by, for example, RNAi (e.g., siRNA, shRNA) or by using site-specific nucleases as described elsewhere in this specification, such as RNA-specific CRISPR / Cas systems.
[0252] In one embodiment, the (wild-type) ig gene, mRNA, and / or protein have reduced expression or transcription (rate), reduced stability, and reduced activity.
[0253] As used herein, “reduce expression (rate),” “reduction in expression rate,” “suppression of expression,” “reduced expression (rate),” “suppression,” or equivalent phrases mean a reduction in the expression level or rate of a nucleotide or protein sequence by more than 10%, 15%, 20%, 25%, or 30%, preferably more than 40%, 45%, 50%, 55%, 60%, or 65%, more preferably more than 70%, 75%, 80%, 85%, 90%, 92%, 94%, 96%, or 98%, compared to a specific reference, such as a plant without the genetic modification or other modification according to the present invention as described elsewhere herein, or a reference plant (such as BL73 of maize). However, it may also mean a reduction in the expression rate of a nucleotide sequence or protein up to 100%. The reduction in expression rate preferably leads to a change in the phenotype of the plant with reduced expression rate. In the context of the present invention, the modified phenotype may be an enhanced inducible capacity of the singular inducer.
[0254] "Decrease in transcription rate" or "decreased transcription rate" or equivalent phrases mean a decrease in the transcription rate of a nucleotide sequence by more than 10%, 15%, 20%, 25%, or 30%, preferably more than 40%, 45%, 50%, 55%, 60%, or 65%, more preferably more than 70%, 75%, 80%, 85%, 90%, 92%, 94%, 96%, or 98%, compared to a plant or reference plant (such as BL73 of maize) that does not contain the genetic modification or other modification according to the present invention as described elsewhere in this specification. However, it may also mean a decrease in the transcription rate of a nucleotide sequence up to 100%. The decrease in transcription rate preferably leads to a change in the phenotype of the plant with the decreased transcription rate. In the context of the present invention, the altered phenotype may be an enhanced inducible capacity of the singular inducer.
[0255] As used herein, “reduced (protein) activity” means a reduction in activity of at least about 10%, preferably at least 30%, more preferably at least 50%, for example, at least 20%, 40%, 60%, 80%, or more, for example, at least 85%, at least 90%, at least 95%, or more. If the activity is reduced by at least 80%, preferably at least 90%, more preferably at least 95%, the activity is (substantially) absent or eliminated. In some embodiments, if no activity, particularly wild-type or native protein activity, is detected, the activity is (substantially) absent. The (protein) activity level may be determined by any means known in the art, depending on the type of protein, by standard detection methods including, for example, enzyme assays (for enzymes), transcription assays (for transcription factors), and assays for analyzing phenotypic output. The activity may be compared to the criteria defined above.
[0256] As used herein, “reduced stability” may refer to reduced protein stability or reduced RNA stability, such as mRNA stability. Protein or RNA stability may be determined by means known in the art, such as determining the protein / RNA half-life. Reduced protein or RNA stability in a given embodiment means a reduction in stability of at least about 10%, preferably at least 30%, more preferably at least 50%, for example, at least 20%, 40%, 60%, 80%, or more, for example, at least 85%, at least 90%, or at least 95%. Stability may be compared to the criteria defined above.
[0257] In some embodiments, a mutated Ig protein, or an Ig protein that confers or enhances singular-inducing activity or its ability, includes the insertion of one or more amino acids. In some embodiments, the insertion is of 1 to 350 amino acids. In some embodiments, the insertion is of 1 to 250 amino acids. In some embodiments, the insertion is of 1 to 150 amino acids. In some embodiments, the insertion is of 1 to 50 amino acids. In some embodiments, the insertion is of 10 to 350 amino acids. In some embodiments, the insertion is of 10 to 250 amino acids. In some embodiments, the insertion is of 10 to 150 amino acids. In some embodiments, the insertion is of 10 to 50 amino acids. In some embodiments, the insertion is of 50 to 350 amino acids. In some embodiments, the insertion is of 50 to 250 amino acids. In some embodiments, the insertion is of 50 to 150 amino acids. In some embodiments, the insertion is of 100 to 350 amino acids. In one embodiment, the insertion is the insertion of 100 to 250 amino acids. In another embodiment, the insertion is the insertion of 100 to 150 amino acids. Those skilled in the art will understand that the presence of the insertion is compared to the mutant or wild type, or to ig that does not confuse or enhance singular-inducible activity or its ability.
[0258] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular-inducing activity or ability thereof comprises the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 110-130 of the wild-type maize (Zea mays) ig protein, as shown in SEQ ID NO: 9 or 10.
[0259] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular induction activity or its ability comprises the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 183-203 of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 23.
[0260] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular-inducing activity or ability thereof includes the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 135-155 of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 26.
[0261] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular-inducing activity or ability thereof includes the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 86-106 of the wild-type Brassica napus ig protein, as shown in SEQ ID NO: 29 or 32.
[0262] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular-inducing activity or ability thereof comprises the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 116-120, preferably 117-119, of the wild-type maize (Zea mays) ig protein, as shown in SEQ ID NO: 9 or 10.
[0263] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular induction activity or its ability comprises the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 189-193, preferably 190-192, of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 23.
[0264] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular-inducing activity or ability thereof comprises the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 141-145, preferably 142-144, of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 26.
[0265] In one embodiment, a mutated ig protein or an ig protein that confers or enhances singular-inducing activity or ability thereof comprises the insertion of one or more amino acids and / or the substitution of one or more amino acids in the region corresponding to amino acid residues 92-96, preferably 93-95, of the wild-type Brassica napus ig protein, as shown in SEQ ID NO: 29 or 32.
[0266] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular-inducing activity or its ability, is a cleaved ig protein. In another embodiment, the mutated ig protein, or the ig protein that confers or enhances singular-inducing activity or its ability, is a C-terminal cleaved ig protein (i.e., the mutated protein contains only the N-terminal region, for example, only the LOB domain).
[0267] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular-inducing activity or its ability, consists of a protein sequence corresponding to amino acid residues 1-116, 1-117, 1-118, 1-119, or 1-120, preferably 1-117, 1-118, or 1-119, of the wild-type maize (Zea mays) ig protein, as shown in SEQ ID NO: 9 or 10.
[0268] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular induction activity or its ability, consists of a protein sequence corresponding to amino acid residues 1-189, 1-190, 1-191, 1-192, or 1-193, preferably 1-190, 1-191, or 1-192, of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 23.
[0269] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular induction activity or its ability, consists of a protein sequence corresponding to amino acid residues 1-141, 1-142, 1-143, 1-144, or 1-145, preferably 1-142, 1-143, or 1-144, of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 26.
[0270] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular induction activity or its ability, consists of a protein sequence corresponding to amino acid residues 1-92, 1-93, 1-94, 1-95, or 1-96, preferably 1-93, 1-94, or 1-95, of the wild-type Brassica napus ig protein, as shown in SEQ ID NO: 29 or 32.
[0271] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular-inducing activity or ability thereof, does not contain the protein sequence corresponding to amino acid residues 117-260, 118-260, 119-260, 120-260, or 121-260, preferably 118-260, 119-260, or 120-260, of the wild-type maize (Zea mays) ig protein, as shown in SEQ ID NO: 9 or 10.
[0272] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular induction activity or its ability, does not contain the protein sequence corresponding to amino acid residues 190-332, 1-191-332, 192-332, 193-332, or 194-332, preferably 191-332, 192-332, or 193-332, of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 23.
[0273] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular induction activity or its ability, does not contain the protein sequence corresponding to amino acid residues 142-308, 143-308, 144-308, 145-308, or 146-308, preferably 143-308, 144-308, or 145-308, of the wild-type sorghum bicolor ig protein, as shown in SEQ ID NO: 26.
[0274] In one embodiment, the mutated ig protein, or the ig protein that confers or enhances singular-inducing activity or its ability, does not contain the protein sequence corresponding to amino acid residues 93-202, 94-202, 95-202, 96-202, or 97-202, preferably 94-202, 95-202, or 96-202, of the wild-type Brassica napus ig protein, as shown in SEQ ID NO: 29 or 32.
[0275] In plants of the genus Zea, such as maize (Zea mays), a mutated ig protein, or an ig protein that confers or enhances singular induction activity or its ability, may have, contain, or consist of the protein sequence shown in SEQ ID NO: 4 or 5, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, most preferably at least 98% identical to SEQ ID NO: 4 or 5. In plants of the genus Zea, such as maize (Zea mays), a mutated ig gene, or an ig gene that confers or enhances singular induction activity or its ability, may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 1, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, most preferably at least 98% identical to SEQ ID NO: 1. In plants of the genus Zea, such as maize (Zea mays), a mutated ig coding sequence, or an ig coding sequence that confers or enhances singular induction activity or its ability, may have, contain, or consist of the nucleic acid sequence shown in SEQ ID NO: 2 or 3, or a sequence that is at least 80%, preferably at least 90%, more preferably at least 95%, and most preferably at least 98% identical to SEQ ID NO: 2 or 3. The mutated maize (Zea mays) ig protein, gene, or coding sequence is preferably the ig1 protein, gene, or coding sequence.
[0276] In one embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular inducing activity or its ability, encodes a protein having a sequence that is preferably at least 80% identical to the sequence shown in SEQ ID NO: 4 or 5, over its entire length. In one embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular inducing activity or its ability, encodes a protein having a sequence that is preferably at least 85% identical to the sequence shown in SEQ ID NO: 4 or 5, over its entire length. In one embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular inducing activity or its ability, encodes a protein having a sequence that is preferably at least 90% identical to the sequence shown in SEQ ID NO: 4 or 5, over its entire length. In one embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular inducing activity or its ability, encodes a protein having a sequence that is preferably at least 95% identical to the sequence shown in SEQ ID NO: 4 or 5, over its entire length. In one embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular induction activity or its ability, encodes a protein having a sequence that is preferably at least 98% identical to the sequence shown in SEQ ID NO: 4 or 5 over its entire length. In another embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular induction activity or its ability, encodes a protein having a sequence that is preferably at least 99% identical to the sequence shown in SEQ ID NO: 4 or 5 over its entire length. In yet another embodiment, the mutated Ig gene or allele, or the Ig gene or allele that confers or enhances singular induction activity or its ability, encodes a protein having a sequence that is identical to the sequence shown in SEQ ID NO: 4 or 5.
[0277] The term “centromere protein” refers to any protein associated with the centromere. These may be proteins associated with DNA in the centromere region, such as centromere histone proteins (e.g., CENH3). The term “kinetocore protein” refers to any protein associated with the kinetocore. These may be proteins present in the kinetocore, with the exception of microtubule proteins, such as tubulin. In some embodiments, the centromere or kinetocore protein is a histone protein. In some embodiments, the centromere or kinetocore protein is not a histone protein. In some embodiments, the centromere or kinetocore protein is CENP. In the course of this invention, it will be understood that mutated centromere or kinetocore proteins confer or enhance singular induction activity. In some embodiments, the centromere or kinetocore protein is selected from CENH3, or any centromere or kinetocore that directly or indirectly interacts with CENH3, preferably directly. In one embodiment, the centromere or kinetochore protein is selected from CENH3, CENP-C, KNL2, SCM3, SAD2, and SIM3.
[0278] As used herein, "CENP-C" or "CENPC" refers to centromere protein C. By way of example and not limitation, Zea mays CENP-C may have the amino acid sequence set forth in NCBI reference sequence XP_008656649.1 (SEQ ID NO: 36). Sorghum bicolor CENP-C may have the amino acid sequence set forth in GenBank accession number AAU04623.1 (SEQ ID NO: 38). One of ordinary skill in the art will be able to readily identify orthologs in different plant species. Variants of CENP-C that confer monoploid-inducing activity are described, for example, in Wang, N., & Dawe, R. K. (2018). “Centromere size and its relationship to haploid formation in plants.” Molecular plant, 11(3), 398-406, and International Publication No. WO 2017 / 058022, which is incorporated herein by reference in its entirety. Nucleic acid molecules encoding CENP-C proteins may be selected from the group consisting of the following i) to iv): i) a nucleic acid molecule having the coding sequence of SEQ ID NO: 35 or 37; ii) a nucleic acid molecule having a coding sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98% or 99% identical to the sequence of SEQ ID NO: 35 or 37; iii) a nucleic acid molecule encoding a protein having the amino acid sequence of SEQ ID NO: 36 or 38; or iv) a nucleic acid molecule encoding a protein having an amino acid sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98% or 99% identical to the sequence of SEQ ID NO: 36 or 38.
[0279] As used herein, "KNL2" refers to the kinetochore-related protein KNL-2 homolog, or kinetochore null2. By way of non-limiting example, Arabidopsis thaliana KNL2 can have the amino acid sequence shown in UniProtKB / Swiss-Prot accession number F4KCE9.1 (SEQ ID NO: 40). One of ordinary skill in the art will be able to readily identify orthologs in different plant species. Variants of KNL2 that confer monoploid-inducing activity are described, for example, in Sandmann et al. (2017) "Targeting of Arabidopsis KNL2 to Centromeres Depends on the Conserved CENPC-k Motif in Its C Terminus" Plant Cell, 29(1):144-155, and U.S. Patent Application Publication No. 2019 / 0075744 (incorporated herein by reference in its entirety). A nucleic acid molecule encoding the KNL2 protein may be selected from the group consisting of the following i) to iv): i) a nucleic acid molecule having the nucleotide sequence of SEQ ID NO: 41, 43, 45, or 47, or a nucleotide sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to the sequence of SEQ ID NO: 41, 43, 45, or 47; ii) a nucleic acid molecule having the coding sequence of SEQ ID NO: 39, or a coding sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to the sequence of SEQ ID NO: 39; iii) a nucleic acid molecule encoding a protein having the amino acid sequence of SEQ ID NO: 40, 42, 44, 46, or 48, or iv) a nucleic acid molecule encoding a protein having an amino acid sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to the sequence of SEQ ID NO: 40, 42, 44, 46, or 48.
[0280] As used herein, “Scm3” refers to the suppressor of chromosomal missegregation protein 3, which was first identified in Saccharomyces cerevisiae, see, for example, https: / / www.yeastgenome.org / locus / S000002298 (SEQ ID NO: 50). This is a homolog of HJURP. Scm3 is the chaperone protein of CENH3. The nucleic acid molecule encoding the Scm3 protein may be selected from the group consisting of i) and ii) below: i) A nucleic acid molecule having a coding sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to the coding sequence of sequence number 49. ii) A nucleic acid molecule encoding a protein having the amino acid sequence of SEQ ID NO: 50, or an amino acid sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to the sequence of SEQ ID NO: 50.
[0281] As used herein, “SAD2” refers to “Sensitive to ABA (abscisic acid) and Drought2” as described in Verslues et al. (2006). Mutation of SAD2, an importin β-domain protein in Arabidopsis, alters abscisic acid sensitivity. The Plant Journal, 47(5), 776-787. SAD2 encodes an importin β-domain family protein likely involved in nuclear transport. SAD2 was expressed at low levels in all tissues examined except flowers, but SAD2 expression was not induced by ABA or stress. The intracellular localization of GFP-tagged SAD2 showed predominantly nuclear localization, consistent with SAD2’s role in nuclear transport. SAD2 is on the same pathway as two transcription factors (GLABROUS1 (GL1) and GLABRA3 (GL3)). A recent publication showed that the mutant sad2 gene affects singular induction in plants (European Patent Application Publication No. 3794939). The nucleic acid molecule encoding the SAD2 protein may be selected from the group consisting of i) and ii) below: i) A nucleic acid molecule having a coding sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to the coding sequence of sequence number 51, ii) A nucleic acid molecule encoding a protein having an amino acid sequence that is 80%, 85%, 90%, 92%, 94%, 96%, 98%, or 99% identical to any one of the sequences in SEQ ID NOs. 52-71.
[0282] As used herein, “SIM3” refers to the NASP-associated protein sim3. SIM3 is a histone H3 and H3-like CENP-A specific chaperone. SIM3 facilitates the delivery and integration of CENP-A in centromere chromatin, presumably by escorting nascent CENP-A to the CENP-A chromatin assembly factor. This is necessary for central nucleus silencing and normal chromosome segregation.
[0283] As used herein, “CENH3” refers to centromere-specific histone H3. It is also known as CENPA or CENP-A (centromere protein A). CENH3 is a centromere protein containing a histone H3-related histone fold domain necessary for centromere targeting. Centromere protein A has been proposed as a modified nucleosome or nucleosome-like structural component that replaces one or both copies of conventional histone H3 in the (H3-H4)2 tetrameric core of a nucleosome particle. This protein is a replication-independent histone that is a member of the histone H3 family. In Arabidopsis thaliana, CENH3 may have the protein sequence shown in SEQ ID NO: 12. In Zea mays, CENH3 may have the protein sequence shown in SEQ ID NO: 14. In Brassica napus, CENH3 may have the protein sequence shown in SEQ ID NO: 16. In Sorghum bicolor, CENH3 may have the protein sequence shown in SEQ ID NO: 18. Therefore, in some embodiments, CENH3 encodes a protein having a sequence that is preferably at least 80% identical to the sequence shown in SEQ ID NO: 12, 14, 16, or 18, over its entire length. In some embodiments, CENH3 encodes a protein having a sequence that is preferably at least 85% identical to the sequence shown in SEQ ID NO: 12, 14, 16, or 18, over its entire length. In some embodiments, CENH3 encodes a protein having a sequence that is preferably at least 90% identical to the sequence shown in SEQ ID NO: 12, 14, 16, or 18, over its entire length. In some embodiments, CENH3 encodes a protein having a sequence that is preferably at least 95% identical to the sequence shown in SEQ ID NO: 12, 14, 16, or 18, over its entire length. In one embodiment, CENH3 encodes a protein having a sequence that is at least 98% identical to the sequence shown in SEQ ID NOs: 12, 14, 16, or 18, preferably over its entire length.In one embodiment, CENH3 encodes a protein having a sequence that is at least 99% identical to the sequence shown in SEQ ID NOs. 12, 14, 16, or 18, preferably over its entire length. In one embodiment, CENH3 is an ortholog of zea mays CENH3, sorghum bicolor CENH3, or brassica napus CENH3.
[0284] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in one or more of the N-terminal domain, αN helix, α1 helix, loop 1 domain, α2 helix, loop 2 domain, α3 helix, or C-terminal domain, as defined in Table 1.
[0285] [Table 1]
[0286] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the N-terminal domain corresponding to amino acids 1 to 82 of Arabidopsis thaliana CENH3.
[0287] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the αN helix corresponding to amino acids 83-97 of Arabidopsis thaliana CENH3.
[0288] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α1 helix corresponding to amino acids 103-113 of Arabidopsis thaliana CENH3.
[0289] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 1 domain corresponding to amino acids 114-126 of Arabidopsis thaliana CENH3.
[0290] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α2 helix corresponding to amino acids 127-155 of Arabidopsis thaliana CENH3.
[0291] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 2 domain corresponding to amino acids 156-162 of Arabidopsis thaliana CENH3.
[0292] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α3 helix corresponding to amino acids 163-172 of Arabidopsis thaliana CENH3.
[0293] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances monomer-inducing activity or its ability comprises one or more mutated amino acids, preferably one or more amino acid substitutions, in the C-terminal domain of CENH3 corresponding to amino acids 173-178 of Arabidopsis thaliana CENH3.
[0294] Preferably, the wild-type Arabidopsis thaliana CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98% identical to the sequence shown in SEQ ID NO: 12.
[0295] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances monomer-inducing activity or its ability comprises one or more mutated amino acids, preferably one or more amino acid substitutions, in the N-terminal domain corresponding to amino acids 1-62 of Zea mays CENH3.
[0296] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances monomer-inducing activity or its ability comprises one or more mutated amino acids, preferably one or more amino acid substitutions, in the αN helix corresponding to amino acids 63-77 of Zea mays CENH3.
[0297] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances monomer-inducing activity or its ability comprises one or more mutated amino acids, preferably one or more amino acid substitutions, in the α1 helix corresponding to amino acids 83-93 of Zea mays CENH3.
[0298] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 1 domain corresponding to amino acids 94-106 of maize (Zea mays) CENH3.
[0299] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α2 helix corresponding to amino acids 107-135 of maize (Zea mays) CENH3.
[0300] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 2 domain corresponding to amino acids 136-142 of maize (Zea mays) CENH3.
[0301] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α3 helix corresponding to amino acids 143-152 of maize (Zea mays) CENH3.
[0302] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the C-terminal domain of CENH3 at amino acids 153-157 of maize (Zea mays) CENH3.
[0303] Preferably, wild-type maize (Zea mays) CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, and more preferably at least 98% identical to the sequence shown in Sequence ID No. 14.
[0304] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the N-terminal domain corresponding to amino acids 1-62 of sorghum bicolor CENH3.
[0305] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the αN helix corresponding to amino acids 63-77 of sorghum bicolor CENH3.
[0306] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or ability thereof contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α1 helix corresponding to amino acids 83-93 of sorghum bicolor CENH3.
[0307] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 1 domain corresponding to amino acids 94-106 of sorghum bicolor CENH3.
[0308] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α2 helix corresponding to amino acids 107-135 of sorghum bicolor CENH3.
[0309] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular induction activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 2 domain corresponding to amino acids 136-142 of sorghum bicolor CENH3.
[0310] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular induction activity or its ability includes one or more mutated amino acids, preferably one or more amino acid substitutions, in the α3 helix corresponding to amino acids 143-152 of sorghum (Sorghum bicolor) CENH3 and in the C-terminal domain of CENH3 corresponding to amino acids 153-157 of sorghum (Sorghum bicolor) CENH3.
[0311] Preferably, wild-type sorghum (Sorghum bicolor) CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, and more preferably at least 98% identical to the sequence shown in Sequence ID No. 18.
[0312] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the N-terminal domain corresponding to amino acids 1-84 of Brassica napus CENH3.
[0313] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the αN helix corresponding to amino acids 85-99 of Brassica napus CENH3.
[0314] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α1 helix corresponding to amino acids 105-115 of Brassica napus CENH3.
[0315] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 1 domain corresponding to amino acids 116-128 of Brassica napus CENH3.
[0316] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α2 helix corresponding to amino acids 129-157 of Brassica napus CENH3.
[0317] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the loop 2 domain corresponding to amino acids 158-164 of Brassica napus CENH3.
[0318] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the α3 helix corresponding to amino acids 165-174 of Brassica napus CENH3.
[0319] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in the C-terminal domain of CENH3 relative to amino acids 175-180 of Brassica napus CENH3.
[0320] Preferably, wild-type Brassica napus CENH3 has an amino acid sequence that is at least 90%, preferably at least 95%, and more preferably at least 98% identical to the sequence shown in SEQ ID NO: 16.
[0321] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability contains one or more mutated amino acids, preferably one or more amino acid substitutions, in one or more of the N-terminal domain, αN helix, α1 helix, loop 1 domain, α2 helix, loop 2 domain, α3 helix, or C-terminal domain, as defined in Table 2.
[0322] [Table 2-1] [Table 2-2]
[0323] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or its ability comprises one or more mutated amino acids, preferably one or more amino acid substitutions, or a corresponding mutation in a CENH3 ortholog, as described in International Publication No. 2016 / 030019, International Publication No. 2016 / 102665, or International Publication No. 2016 / 138021 (each incorporated by reference in its entirety).
[0324] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or its ability comprises one or more mutated amino acids, preferably one or more amino acid substitutions, corresponding to positions 3, 17, 32, 35, 9, 24, 29, 40, 42, 50, 55, 57, 61, 74, 82, 104, 109, 120, 148, 175, 130, 151, 157, 158, 164, 166, 83, 86, 124, 127, 132, 136, 152, 155, or 172 of the reference Arabidopsis thaliana CENH3 protein, and preferably has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0325] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof contains one or more mutated amino acids, preferably one or more amino acid substitutions, corresponding to positions 3, 17, 32, 35, 104, 109, 120, 148, or 175 of the Arabidopsis thaliana CENH3 protein, if the plant or plant part containing such sequence is from the genus Zea, preferably from maize (Zea mays), and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0326] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof comprises one or more mutated amino acids, preferably one or more amino acid substitutions, at positions 3, 16, 32, 35, 84, 89, 100, 128, or 155 of a CENH3 protein from the genus Zea, preferably from maize (Zea mays), preferably the maize CENH3 protein having an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 14.
[0327] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof contains one or more mutated amino acids, preferably one or more amino acid substitutions, corresponding to positions 9, 24, 29, 32, 40, 42, 50, 55, 57, 61, 130, 151, 157, 158, 164, or 166 of the reference Arabidopsis thaliana CENH3 protein, if the plant or plant part containing such sequence is from the genus Brassica, preferably Brassica napus, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 12.
[0328] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof comprises one or more mutated amino acids, preferably one or more amino acid substitutions, at positions 9, 24, 29, 30, 33, 41, 43, 50, 55, 57, 61, 132, 153, 159, 160, 166, or 168 of a CENH3 protein from the genus Brassica, preferably Brassica napus, and preferably having an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO: 16.
[0329] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof contains one or more mutated amino acids, preferably one or more amino acid substitutions, corresponding to positions 42, 74, or 130 of the Arabidopsis thaliana CENH3 protein, if the plant or plant part containing such sequence is from the genus Sorghum, preferably Sorghum bicolor, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in Sequence ID No. 12.
[0330] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof comprises one or more mutated amino acids, preferably one or more amino acid substitutions, at positions 42, 55, 110, or 157 of a CENH3 protein from the genus Sorghum, preferably Sorghum bicolor, and preferably having an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98% identical to the sequence shown in SEQ ID NO: 18.
[0331] In one embodiment, a mutated CENH3 protein or a CENH3 protein that confers or enhances singular-inducing activity or ability thereof includes an amino acid substitution corresponding to position 35 of maize (Zea mays) CENH3, preferably corresponding to position 35 of SEQ ID NO: 14 or an amino acid substitution at position 35 of SEQ ID NO: 14, preferably the amino acid substitution being 35K, for example, E35K in maize (Zea mays). Such sequences are preferably found in plants of the genus Zea, preferably from maize (Zea mays).
[0332] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or ability thereof includes an amino acid substitution corresponding to position 35 of sorghum bicolor CENH3, preferably corresponding to position 35 of SEQ ID NO: 18 or an amino acid substitution at position 35 of SEQ ID NO: 18, preferably the amino acid substitution being 35K, for example, E35K in sorghum bicolor. Such sequences are preferably found in plants of the genus Sorghum, preferably from sorghum bicolor.
[0333] In one embodiment, the mutated CENH3 protein or the CENH3 protein that confers or enhances singular-inducing activity or ability thereof includes an amino acid substitution corresponding to position 36 of Brassica napus CENH3, preferably corresponding to position 36 of SEQ ID NO: 16 or an amino acid substitution at position 36 of SEQ ID NO: 16, preferably the amino acid substitution being 35K, for example, T35K in Brassica napus. Such sequences are preferably found in plants of the genus Brassica, preferably Brassica napus.
[0334] Those skilled in the art will understand how to determine the corresponding location in the CENH3 ortholog.
[0335] In a preferred embodiment, the mutant CENH3 protein, or the CENH3 protein that confers or enhances singular induction activity or its ability, includes the amino acid sequence shown in SEQ ID NO: 20. In a preferred embodiment, the mutant CENH3 protein, or the CENH3 protein that confers or enhances singular induction activity or its ability, includes the amino acid sequence described in SEQ ID NO: 20, the amino acid sequence corresponding to the amino acid sequence described in SEQ ID NO: 20, or the amino acid sequence that is at least 80%, for example at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence described in SEQ ID NO: 20, and includes the amino acid at position 35 or the corresponding amino acid position that is not E. In a preferred embodiment, the mutant CENH3 protein, or the CENH3 protein that confers or enhances singular induction activity or its ability, includes the amino acid sequence shown in SEQ ID NO: 20. In a preferred embodiment, the mutant CENH3 protein, or the CENH3 protein that confers or enhances singular-inducing activity or ability thereof, comprises the amino acid sequence described in SEQ ID NO: 20, the amino acid sequence corresponding to the amino acid sequence described in SEQ ID NO: 20, or the amino acid sequence that is at least 80%, for example, at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence described in SEQ ID NO: 20, and comprises the corresponding amino acid position which is the amino acid or K at position 35 (for example, amino acid position 36 in some species, including Brassica napus). Those skilled in the art can determine the corresponding amino acid position by a suitable alignment algorithm as described elsewhere in this specification.
[0336] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 1, 2, or 3, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 4 or 5, and further comprising a polynucleic acid encoding a CENH3 protein having the sequence shown in SEQ ID NO: 20.
[0337] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 1, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 4 or 5, and further comprising a polynucleic acid encoding a CENH3 protein having the sequence shown in SEQ ID NO: 20.
[0338] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 2, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 4, and also comprising a polynucleic acid encoding a CENH3 protein having the sequence shown in SEQ ID NO: 20.
[0339] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 3, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 5, and also comprising a polynucleic acid encoding a CENH3 protein having the sequence shown in SEQ ID NO: 20.
[0340] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 1, 2, or 3, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 4 or 5, and having an amino acid at position 35 that is different from E, preferably the amino acid being K, which encodes a CENH3 protein.
[0341] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 1, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 4 or 5, and having an amino acid at position 35 that is different from E, preferably the amino acid being K, which encodes a CENH3 protein.
[0342] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 2, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 4, and having an amino acid at position 35 that is different from E, preferably the amino acid being K, which encodes a CENH3 protein.
[0343] In one embodiment, the present invention relates to a corn (Zea mays) plant or plant part (e.g., pollen or seed) comprising a polynucleic acid encoding a mutated ig1 protein having the sequence shown in SEQ ID NO: 3, or a polynucleic acid encoding a sequence protein shown in SEQ ID NO: 5, and having an amino acid at position 35 that is different from E, preferably the amino acid being K, which encodes a CENH3 protein.
[0344] In some embodiments, the plant or plant part of the present invention described herein further comprises a polynucleic acid encoding a site-specific DNA or RNA-binding protein, or a site-specific DNA or RNA-binding protein, preferably a site-specific DNA or RNA-editing or modification protein. Thus, in some embodiments, the plant or plant part of the present invention described herein further comprises a polynucleic acid encoding a site-specific DNA or RNA-binding protein, or a site-specific DNA or RNA-editing or modification protein. Such plants and methods for producing such plants are described, for example, in U.S. Patent No. 1,0285,348, which is incorporated herein by reference in its entirety.
[0345] As used herein, the term “site-specific DNA or RNA-binding protein” means a protein that binds to DNA or RNA in a sequence-specific manner, or is recruited to DNA or RNA directly (e.g., in the case of TALENS or zinc finger nucleases) or indirectly (in the case of the CRISPR / Cas system, where a Cas effector protein binds to guide RNA (including guide sequences and direct repeat sequences) that hybridizes DNA or RNA, and optionally (as needed) to a tracr sequence) in a sequence-specific manner. Site-specific DNA or RNA-binding proteins may directly edit or modify DNA or RNA (i.e., a DNA or RNA-binding protein may inherently exhibit the ability to edit or modify DNA or RNA, like a Cas effector protein), or may be fused to other proteins or domains that have the ability to edit or modify DNA or RNA (e.g., TALEN or ZFN containing TALE or ZF fused to FokI, respectively). As used herein, the term “site-directed DNA or RNA editing or modification protein” conventionally refers to a protein that directly or indirectly binds to DNA or RNA in a sequence-specific manner and directly or indirectly edits or modifies DNA or RNA (e.g., via a fusion partner, i.e., a chimeric protein), and may instead be called an “editing mechanism.”
[0346] In one embodiment, the site-specific DNA or RNA-binding protein, or the DNA or RNA site-specific editing or modification protein, is a nuclease (i.e., a DNA or RNA nuclease). In another embodiment, the site-specific DNA or RNA-binding protein, or the DNA or RNA site-specific editing or modification protein, is an endonuclease (i.e., a DNA or RNA endonuclease).
[0347] In some embodiments, the site-specific DNA or RNA binding protein, or the DNA or RNA site-specific editing or modification protein, is a mutant nuclease (i.e., a DNA or RNA nuclease). In some embodiments, the site-specific DNA or RNA binding protein, or the DNA or RNA site-specific editing or modification protein, is a mutant endonuclease (i.e., a DNA or RNA endonuclease). Such mutant (endo)nucleases may include mutations that alter DNA or RNA binding specificity (e.g., to alter PAM specificity at the site of a Cas effector protein), stability (e.g., destabilizing mutants), and / or activity (e.g., mutants that increase or (partially) eliminate enzyme activity, e.g., catalytically inactive Cas effector proteins or nickase Cas effector proteins). The advantage of catalytically inactive mutants is that they can function as a medium for recruiting fusion partners in a sequence-specific manner. Such fusion partners may have different DNA or RNA editing or modification activities, or other activities, such as transcriptional activation or repression activity, or even chromatin remodeling activity.
[0348] In one embodiment, the site-specific DNA or RNA-binding protein, or the DNA or RNA site-specific editing or modification protein, is a meganuclease (MN), zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), (mutated) Cas nuclease / effector protein, Cas9, Cfp1 (Cas12a), MAD7, Cas13 (e.g., Cas13a or Cas13b), dCas9-FokI (a "dead" or catalytically inactive Cas9 fused to FokI), dCpf1-FokI (a "dead" or catalytically inactive Cpf1 fused to FokI), dMAD7-FokI (F The group includes "dead" or catalytically inactive MAD7 fused to okI, nickase Cas effector proteins (e.g., Cas9 or Cpf1), chimeric Cas effector (e.g., Cas9, Cpf1, Cas13)-cytidine deaminase (Cas effector protein is catalytically inactive), chimeric Cas effector (e.g., Cas9, Cpf1, Cas13)-adenine deaminase (Cas effector protein is catalytically inactive), chimeric FENI-FokI, and mega-TAL, chimeric dCas9 non-FokI nuclease, dCpf1 non-FokI nuclease, and dMAD7 non-FokI nuclease. For example, fusion proteins of Cas effectors (e.g., Cas9, Cas12, or Cas13) and deaminases, such as adenine or cytidine deaminase, enable base editing, particularly the introduction of point mutations.
[0349] As described elsewhere in this specification, when the site-specific DNA or RNA-binding protein is a (mutated) Cas effector protein, the sequence-specific DNA or RNA binding requires the presence of a guide RNA (gRNA) that hybridizes to a specific target sequence and recruits the Cas effector protein to this target sequence. The gRNA typically includes a guide sequence (which hybridizes to the target sequence) and a direct repeat sequence (or tracr-mate sequence) (which binds to and recruits the Cas effector protein). As is known in the art, depending on the type of Cas effector protein, a tracr sequence may or may not be required. The gRNA and tracr sequences may be provided on the same or different polynucleic acids. Chimeric gRNA (i.e., a fusion of gRNA and tracr) is also within the scope of this invention. Those skilled in the art will understand that the gRNA (and optionally the tracr) may also be included in or expressed in the singular-inducible plant according to this invention. However, it is not necessarily essential to be so. For example, only the Cas effector protein may be included in or expressed in the singular inducible plant according to the present invention, while the appropriate gRNA (and tracrRNA as needed) may be provided (e.g., inserted, converted) at a separate time.
[0350] The plant or plant part according to the present invention, in particular the singular-inducing plant described herein, for example, the paternal singular-inducing plant described herein, further comprising a site-specific DNA or RNA-binding, edited or modified protein, or a polynucleic acid encoding a site-specific DNA or RNA-binding, edited or modified protein, enables simultaneous singular induction and gene editing. The editing mechanism is delivered by an inducer line. The editing mechanism is encoded by and present in the inducer line, since it is stably inserted into the inducer, for example, via irradiation or Agrobacterium-mediated transformation. In other examples, the editing mechanism is transiently introduced (by exogenous application) or transiently expressed in the gametophyte before fertilization. After fertilization, editing is performed by the editing mechanism in the target gene of the non-inducing substance before or during the removal of the inducer chromosome. As a result, a singular embryo or plant or seed is obtained, which contains only the chromosome set from the non-inducing parent containing the edited DNA sequence. These edited singularities may be identified, propagated, and have their chromosomes doubled, preferably with colchicine, pronamide, diticyl, trifluralin, or other known microtubule inhibitors. This line can then be used directly in downstream breeding programs.
[0351] In one embodiment, the editing mechanism is any DNA-modifying enzyme, but preferably a site-specific nuclease. The site-specific nuclease is preferably CRISPR-based, but may be a meganuclease, a transcription activator-like effector nuclease (TALEN), or a zinc finger nuclease. The nuclease used in the present invention may be Cas9, Cfp1, dCas9-FokI, or a chimeric FEN1-FokI. In one embodiment, the DNA modifying enzyme is a site-directed base editing enzyme, such as a Cas9 (or Cpf1, etc.)-cytidine deaminase fusion protein or a Cas9 (or Cpf1, etc.)-adenine deaminase fusion protein, where Cas9 (or Cpf1, etc.) may have one or both of its inactivated nuclease activity, i.e., a chimeric Cas9 (or Cpf1, etc.) nickase (nCas9, nCpf1, etc.) fused to cytidine deaminase or adenine deaminase, or inactivated Cas9 (dCas9, dCpf1, etc.). Any guide RNA targets the genome at the specific site to be edited.
[0352] In one embodiment, the present invention relates to a plant or plant part obtained from a hybridization of a first plant, which is a plant according to the present invention as described herein, and a second plant. In one embodiment, the present invention relates to a plant or plant part obtained from a hybridization of a first female plant, which is a plant according to the present invention as described herein, and a second male plant. In one embodiment, the present invention relates to a plant or plant part obtained from the pollination of a second plant by pollen from a first plant, which is a plant according to the present invention as described herein.
[0353] In one embodiment, the present invention relates to a method for producing a plant or plant part, including the hybridization of a first plant and a second plant, which are plants according to the present invention as described herein. In one embodiment, the present invention relates to a method for producing a plant or plant part, including the hybridization of a first female plant and a second male plant, which are plants according to the present invention as described herein. In one embodiment, the present invention relates to a method for producing a plant or plant part, including the pollination of a second plant with pollen from a first plant, which are plants according to the present invention as described herein.
[0354] In one embodiment, the present invention relates to zea mays seeds named igEIN, whose representative sample was deposited on May 11, 2021, at NCIMB (National Collection of Industrial Food and Marine Bacteria; Ltd. Ferguson Building, Craibstone Estate, Bucksburn, Aberdeen, AB21 9YA Scotland) under accession number NCIMB43772, or to plants or plant parts grown therefrom. In one embodiment, the present invention relates to zea mays seeds deposited at NCIMB accession number NCIMB43772, or to plants or plant parts grown therefrom. Plants grown from or obtained from seeds deposited at NCIMB accession number NCIMB43772 exhibit a phenotype (on average) of (increased) singular inducer. Seeds deposited at NCIMB accession number NCIMB43772 contain a CENH3 mutation resulting in an E35K amino acid exchange (sequence number 20) and contain the ig nucleotide sequence shown in sequence number 1, as described in Example 1.
[0355] In one embodiment, the present invention relates to a method for producing a singular plant or plant part, comprising crossing a first plant, which is a plant according to the present invention as described herein, with a second plant, and selecting a singular offspring plant or plant part. In one embodiment, the present invention relates to a method for producing a singular plant or plant part, comprising crossing a first female plant, which is a plant according to the present invention as described herein, with a second male plant, and selecting a singular offspring plant or plant part. In one embodiment, the present invention relates to a method for producing a singular plant or plant part, comprising pollinating a second plant with pollen from a first plant, which is a plant according to the present invention as described herein, and selecting a singular offspring plant or plant part. It will be understood that singular offspring include diploid, triploid, and other offspring as described elsewhere in this specification. Optionally, the method further comprises producing a double singular plant or plant part from the singular plant or plant part, or converting the singular plant or plant part into a double singular plant or plant part.
[0356] In one embodiment, the present invention provides a singular plant or plant part obtained from a hybrid of a first plant and a second plant, which are plants according to the present invention as described herein, and a method for producing a plant or plant part, which includes converting a singular plant or plant part into a doubled singular plant or plant part. In one embodiment, the present invention provides a singular plant or plant part obtained from a hybrid of a first female plant and a second male plant, which are plants according to the present invention as described herein, and a method for producing a plant or plant part, which includes converting a singular plant or plant part into a doubled singular plant or plant part. In one embodiment, the present invention provides a singular plant or plant part obtained from the pollination of a second plant with pollen from a first plant, which are plants according to the present invention as described herein, and a method for producing a plant or plant part, which includes converting a singular plant or plant part into a doubled singular plant or plant part. It will be understood that singular plants or plant parts include diploid, triploid, and other plants or plant parts as described elsewhere in this specification.
[0357] In one embodiment, the present invention relates to a method for producing a (double singular) plant or plant part, comprising crossing a first plant, which is a plant according to the present invention as described herein, with a second plant, and converting singular offspring into a double singular plant or plant part. In one embodiment, the present invention relates to a method for producing a (double singular) plant or plant part, comprising crossing a first female plant, which is a plant according to the present invention as described herein, with a second male plant, and converting singular offspring into a double singular plant or plant part. In one embodiment, the present invention relates to a method for producing a (double singular) plant or plant part, comprising pollinating a second plant with pollen from a first plant, which is a plant according to the present invention as described herein, and converting singular offspring into a double singular plant or plant part.
[0358] In one embodiment, the present invention provides a method for editing the genomic DNA of a plant. This is done by obtaining a first plant, which is a singular inducer plant and has the necessary mechanisms (e.g., Cas9 enzyme and guide RNA) encoded in its DNA to achieve the editing, and then pollinating a second plant using the pollen of the first plant. The second plant is the plant to be edited. From the pollination event, offspring (e.g., embryos or seeds) are produced, at least one of which becomes a singular seed. This singular seed contains only the chromosomes of the second plant, with the chromosomes of the first plant absent (removed, lost, or degraded), but prior to this, the chromosomes of the first plant either allowed the expression of a gene editing mechanism or delivered an editing mechanism that had already been expressed by the first plant during pollination via the pollen tube. Alternatively, if the singular inducer line is female in the cross, the egg cell of the singular inducer plant contains an editing mechanism that was present and possibly already expressed at the time of fertilization with a “wild-type” or non-singular inducer pollen grain. Through one of these pathways, the genome of the phantom offspring obtained through hybridization is also edited.
[0359] One embodiment of the present invention provides a method for editing plant genomic DNA, comprising (i) to (iv) below: (i) providing a first plant, the first plant being a plant singular inducer strain according to the present invention as described herein, wherein the first plant contains, expresses, or can express a DNA modifying enzyme and optionally a guide RNA as described elsewhere herein; (ii) providing a second plant containing plant genomic DNA to be edited; (iii) crossing the first plant with the second plant or pollinating the second plant with pollen from the first plant; and (iv) selecting at least one singular offspring produced by the pollination of step (c), wherein the singular offspring contains the genome of the second plant but does not contain the genome of the first plant, and the genome of the singular offspring is modified by a DNA modifying enzyme and optionally a guide nucleic acid delivered by the first plant.
[0360] In one embodiment, the present invention is a method for editing or modifying plant genomic DNA or RNA, comprising the following a) to d): a) providing a first plant which is a plant according to the present invention as described herein and which contains, expresses, or can express site-specific DNA or RNA-binding proteins as described elsewhere herein; b) providing a second plant (containing plant genomic DNA or RNA to be modified); c) pollinating the second maize plant with pollen from the first plant; and d) selecting at least one singular offspring produced by the pollination in step c) (wherein the offspring is singular, digenomic singular, or trigenomic singular, which contains the genome of the second plant rather than the genome of the first plant, and the genome of the singular, digenomic singular, or trigenomic singular offspring is modified by site-specific DNA or RNA-binding proteins delivered by the first plant).
[0361] The methods of the present invention described herein may further include a step of harvesting plant material, for example, seeds (preferably obtained from cross-pollination or fertilization).
[0362] The methods of the present invention described herein may further include the step of selecting singular offspring obtained from cross or pollination. It will be understood that singular offspring include diploid, triploid, and other offspring as described elsewhere herein.
[0363] The methods of the present invention described herein may further include a step of crossbreeding the offspring, for example, backcrossing the offspring (obtained from crossbreeding or pollination). The methods of the present invention described herein may further include a step of self-pollinating the offspring (obtained from crossbreeding or pollination).
[0364] The methods of the present invention described herein may further include the step of regenerating a plant or plant part (from an embryo obtained from cross-pollination or pollination).
[0365] The methods of the present invention described herein may further include the step of converting a singular plant or plant part (obtained from hybridization or pollination) into a doubled singular plant or plant part. It will be understood that singular offspring include diploid, triploid, and other offspring as described elsewhere in this specification. Methods for producing doubled singular plants are known in the art and are described elsewhere in this specification.
[0366] Preferably, the second plant is not a plant according to the present invention. Preferably, the second plant is not a singular-induced plant.
[0367] Preferably, the second plant is of the same species as the first plant. In some embodiments, the first and second plants are from the genus Zea, preferably Zea mays. In some embodiments, the first and second plants are from the genus Sorghum, preferably Sorghum bicolor. In some embodiments, the first and second plants are from the genus Brassica, preferably Brassica napus.
[0368] In one embodiment, the present invention relates to a descendant plant or plant part obtained by the method according to the present invention as described herein.
[0369] Polynucleic acids encoding mutated undetermined gametophyte (ig) proteins or ig proteins that induce or enhance singularities, and polynucleic acids encoding mutated centromere or kinetochore proteins or singularities that induce or enhance centromere or kinetochore proteins, are operably ligated to one or more regulatory sequences, particularly promoter sequences, of a plant or plant part, thereby enabling protein expression. Such promoters may be endogenous or exogenous (heterogeneous) promoters. Such promoters may be located at their intrinsic genomic location or elsewhere. Such promoters may enable constitutive, transient, or conditional expression, such as developmental level-dependent expression, tissue-specific expression, or inducible expression. The same applies to site-specific DNA or RNA-binding proteins encoding polynucleic acids as described elsewhere in this specification.
[0370] As used herein, the term “regulatory sequence” refers to a nucleotide sequence that influences specificity and / or expression intensity, for example, by mediating a defined tissue specificity. Such a regulatory sequence may be located upstream of the transcription start site of a minimal promoter, but may also be located downstream, for example, within a transcribed but untranslated leader sequence or intron.
[0371] In one embodiment, the polynucleic acid sequence according to the present invention described herein may be introduced into a plant or plant part by a transformation known in the art, such as Agrobacterium tumefaciens-mediated transformation. Here, the polynucleic acid may be provided on a suitable vector.
[0372] As used herein, “vector” has the common sense in the art and may be, for example, a plasmid, cosmid, phage or expression vector, transformation vector, shuttle vector or cloning vector, and may be double-stranded or single-stranded, linear or circular, or may transform a prokaryotic or eukaryotic host either by its integration into the genome or chromosome. The nucleic acid according to the present invention is preferably operably ligated in a vector having one or more regulatory sequences that enable transcription and optionally expression in a prokaryotic or eukaryotic host cell. The regulatory sequences, preferably DNA, may be homogeneous or heterogeneous to the nucleic acid according to the present invention. For example, the nucleic acid is under the control of a suitable promoter or terminator. Suitable promoters may be constitutively induced promoters (e.g., the 35S promoter from "cauliflower mosaic virus" (Odell et al., 1985), tissue-specific promoters are particularly suitable (e.g., pollen-specific promoters, Chen et al. (2010), Zhao et al. (2006), or Twell et al. (1991)), or development-specific (e.g., flower-specific promoters)). Suitable promoters may also be synthetic promoters or chimeric promoters that do not occur in nature, and consist of multiple elements, including a minimal promoter and at least one cis-regulatory element upstream of the minimal promoter that functions as a binding site for a specific transcription factor. Chimeric promoters can be designed according to desired specifications and are induced or repressed via different factors. Examples of such promoters can be found in Gurr & Rushton (2005) or Venter (2007). For example, a suitable terminator is the nos-terminator (Depicker et al., 1982). The vector may be introduced via conjugation, recruitment, gene gun transformation, agrobacteria-mediated transformation, transfection, transduction, vacuum infiltration, or electroporation.
[0373] In one embodiment, the vector is a conditional expression vector. In another embodiment, the vector is a constitutive expression vector. In another embodiment, the vector is a tissue-specific expression vector, such as a pollen-specific expression vector. In yet another embodiment, the vector is an inducible expression vector. All such vectors are well known in the art.
[0374] The vector manufacturing method described is common to those skilled in the art (Sambrook et al., 2001).
[0375] Host cells, such as plant cells, containing nucleic acids described herein, preferably induction-promoting nucleic acids or nucleic acids encoding double-stranded RNA, or vectors described herein, are also assumed herein. The host cell may contain nucleic acids as extrachromosomal (episome) replication molecules, or nucleic acids incorporated into the nuclear or plastid genome of the host cell, or nucleic acids as introduced chromosomes, such as microchromosomes.
[0376] The host cell may be a prokaryotic cell (e.g., a bacterium) or a eukaryotic cell (e.g., a plant cell or a yeast cell). For example, the host cell may be Agrobacterium, such as Agrobacterium tumefaciens or Agrobacterium rhizogenes. Preferably, the host cell is a plant cell.
[0377] The nucleic acids or vectors described herein may be introduced into host cells by well-known methods that may depend on selected host cells, including, for example, conjugation, recruitment, gene gun transformation, Agrobacterium-mediated transformation, transfection, transduction, vacuum infiltration, or electroporation. In particular, methods for introducing nucleic acids or vectors into Agrobacterium cells are well-known to those skilled in the art and may include conjugation or electroporation. Methods for introducing nucleic acids or vectors into plant cells are also known (Sambrook et al., 2001) and may include a variety of transformation methods, such as gene gun transformation and Agrobacterium-mediated transformation.
[0378] In certain embodiments, the present invention relates to transgenic host cells comprising nucleic acids, particularly those encoding induction-promoting nucleic acids or double-stranded RNA, as transgenes or vectors as described herein, or vectors as described herein. In further embodiments, the present invention relates to transgenic plants or parts thereof, comprising transgenic plant cells.
[0379] For example, such transgenic cells or transgenic plants are preferably plant cells or plants stably transformed with the nucleic acids described herein, particularly the induction-promoting nucleic acids or double-stranded RNA encoding nucleic acids described herein, or the vectors described herein.
[0380] Preferably, the nucleic acid in the transgenic plant is operably linked to one or more regulatory sequences in the plant cell that enable transcription and optionally expression. The regulatory sequences may be homogeneous or heterogeneous to the nucleic acid. The overall structure composed of the nucleic acid and its regulatory sequences according to the present invention may represent a transgene.
[0381] The transgenic plant part may be, for example, a fertilized or unfertilized seed, embryo, pollen, tissue, organ, or plant cell, where a fertilized or unfertilized seed, embryo, or pollen is produced in the transgenic plant and the nucleic acid described herein, in particular the induction-promoting nucleic acid or double-stranded RNA encoding nucleic acid described herein, is incorporated into its genome as a transgene or vector. The term transgenic plant as used herein also includes the offspring of the transgenic plant described herein with the genome of the nucleic acid described herein, in particular the induction-promoting nucleic acid or double-stranded RNA encoding nucleic acid described herein, which is incorporated as a transgene or vector described herein.
[0382] As used herein, the terms “operable ligation” or “operable ligation” mean that the ligated elements are ligated to a common nucleic acid molecule in such a way that they can be positioned and oriented relative to one another, resulting in the transcription of the nucleic acid molecule. DNA operable ligated to a promoter is under the transcriptional control of that promoter.
[0383] As used herein, the term "transformation" means the transfer of an isolated and loaned gene into the DNA, usually chromosomal DNA or genome, of another organism.
[0384] As used herein, the term “sequence identity” refers to the degree of identity between any given nucleic acid sequence and a target nucleic acid sequence. As used herein, unless expressly specified, sequence identity is preferably determined over the entire sequence length. The percentage of sequence identity is calculated by measuring the number of matching positions in the aligned nucleic acid sequence, dividing the number of matching positions by the total number of aligned nucleotides, and multiplying by 100. A matching position is a position in the aligned nucleic acid sequence where identical nucleotides exist at the same position. The percentage of sequence identity can also be determined for any amino acid sequence. To measure the percentage of sequence identity, the target nucleic acid or amino acid sequence is compared to the identified nucleic acid or amino acid sequence using the BLAST2 sequencing (Bl2seq) program from a standalone version of BLASTZ, which includes BLASTN and BLASTP. This standalone version of BLASTZ is available from the Fish & Richardson website (worldwideweb fr.com / blast) or the U.S. National Center for Biotechnology Information website (worldwideweb ncbi.nlm.nih.gov). Instructions on how to use the Bl2seq program can be found in the readme file included with BLASTZ. Bl2seq performs a comparison between two sequences using either the BLASTN or LASTP algorithm.
[0385] BLASTN is used to compare nucleic acid sequences, while BLASTP is used to compare amino acid sequences. To compare two nucleic acid sequences, set the options as follows: - Set i to the file containing the first nucleic acid sequence to compare (e.g., C:\seq l .txt); - Set j to the file containing the second nucleic acid sequence to compare (e.g., C:\seq2.txt); - Set p to blastn; - Set o to any filename (e.g., C:\output.txt); - Set q to -1; - Set r to 2; Leave all other options at their default settings. The following command will generate an output file containing the comparison between the two sequences: C:\B12seq -ic:\seql .txt -jc:\seq2.txt -p blastn -oc:\output.txt -q -1 -r 2. If the target sequence shares homology with any part of the identified sequence, the configured output file will present the sequences as aligned sequences of those homologous regions. If the target sequence does not share homology with any portion of the identified sequence, the configured output file will not present the aligned sequence. Alignment determines the length by counting the number of consecutive nucleotides from the target sequence presented in the alignment, which has a sequence from the identified sequence that begins at any matching position and ends at any other matching position. Matching positions are any positions where identical nucleotides exist in both the target sequence and the identified sequence. Gaps are not nucleotides, so gaps present in the target sequence are not counted. Similarly, gaps presented in the identified sequence are not counted because nucleotides from the target sequence, not from the identified sequence, are counted. The percentage of identity over a given length is determined by counting the number of matching positions over that length, dividing that number by the length, and then multiplying the resulting value by 100.For example, (i) when a 500-base nucleic acid target sequence is compared to a target nucleic acid sequence, (ii) the Bl2seq program presents 200 bases from the target sequence aligned in the region of the target sequence where the first and last bases of that 200-base region match, and (iii) the number of matches across those 200 aligned bases is 180, and the 500-base nucleic acid target sequence contains sequence identity over a length of 200 and 90% of that length (i.e., 180 / 200 × 100 = 90). It will be understood that different regions within a single nucleic acid target sequence that align with the identified sequence can each have their own percentage of identity. Note that the percentage of identity value is rounded to the nearest tenth. For example, 78.11, 78.12, 78.13 and 78.14 are rounded down to 78.1, while 78.15, 78.16, 78.17, 78.18 and 78.19 are rounded up to 78.2. Note that the length value will always be an integer.
[0386] "Isolated nucleic acid sequence" or "isolated DNA" means a nucleic acid sequence that no longer exists in the natural environment in which it was isolated, for example, a nucleic acid sequence in a bacterial host cell or in the nuclear or plastid genome of a plant. Where "sequence" is referred to herein, it is understood that a molecule having such a sequence means, for example, a nucleic acid molecule. "Host cell" or "recombinant host cell" or "transformed cell" is a term that refers to a new individual cell (or organism) resulting from the introduction of at least one nucleic acid molecule into the cell. The host cell is preferably a plant cell or a bacterial cell. The host cell may contain nucleic acids as extrachromosomal (episome) replication molecules, or nucleic acids incorporated into the nuclear or plastid genome of the host cell, or nucleic acids as introduced chromosomes, for example, microchromosomes.
[0387] In one embodiment, the nucleic acid molecule described herein contains fewer than 50,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 40,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 30,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 25,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 20,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 15,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 10,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains fewer than 5,000 nucleotides. In one embodiment, the nucleotide molecule described herein contains at least 100 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and fewer than 50,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and fewer than 40,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and less than 30,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and less than 25,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and less than 20,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and less than 15,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and less than 10,000 nucleotides. In one embodiment, the nucleic acid molecule described herein contains at least 100 nucleotides and less than 5,000 nucleotides.
[0388] In one embodiment, when a nucleic acid sequence (e.g., DNA or genomic DNA) is referenced that has "substantial sequence identity" with respect to a reference sequence, or has at least 80% sequence identity with respect to a reference sequence, e.g., at least 85%, 90%, 95%, 98%, or 99%, then the nucleotide sequence is considered substantially identical to the given nucleotide sequence and can be identified using stringent hybridization conditions. In other embodiments, the nucleic acid sequence contains one or more mutations compared to the given nucleotide sequence, but can still be identified using stringent hybridization conditions. Nucleotide sequences that are substantially identical to the given nucleotide sequence can be identified using "stringent hybridization conditions." Stringent conditions are sequence-dependent and vary from situation to situation. Generally, stringent conditions are selected to be about 5°C lower than the thermal melting point (Tm) for a particular sequence at a defined ionic strength and pH. Tm is the temperature at which 50% of the target sequence hybridizes to a probe that is a perfect match (based on defined ionic strength and pH). Typically, stringent conditions are selected, where the salt concentration is approximately 0.02 mol at pH 7 and the temperature is at least 60°C. Lowering the salt concentration and / or raising the temperature increases stringency. Stringent conditions for RNA-DNA hybridization (e.g., Northern blotting with a 100 nt probe) are, for example, conditions including at least one wash in 0.2×SSC for 20 minutes at 63°C, or equivalent conditions. Stringent conditions for DNA-DNA hybridization (e.g., Southern blotting with a 100 nt probe) are, for example, conditions including at least one (usually two) washes in 0.2×SSC for 20 minutes at a temperature of at least 50°C, typically around 55°C, or equivalent conditions. See also Sambrook et al. (1989), and Sambrook and Russell (2001).
[0389] RNA interference, or RNAi, is a biological process in which RNA molecules inhibit gene expression or translation by neutralizing target mRNA molecules. Two types of small ribonucleic acid (RNA) molecules (microRNA (miRNA) and small interfering RNA (siRNA)) are central to RNA interference. RNA is the direct product of genes, and these small RNAs can bind to other specific messenger RNA (mRNA) molecules, increasing or decreasing their activity, for example, by preventing mRNA from being translated into proteins. The RNAi pathway is found in many eukaryotes, including animals, and is initiated by the enzyme Dicer, which cleaves long double-stranded RNA (dsRNA) molecules into short double-stranded fragments of approximately 21 nucleotides of siRNA (small interfering RNA). Each siRNA is unwound into two single-stranded RNAs (ssRNA): a passenger strand and a guide strand. The passenger strand is degraded, and the guide strand is incorporated into the RNA-induced silencing complex (RISC). Mature miRNAs are structurally similar to siRNAs generated from exogenous dsRNAs, but before reaching maturity, miRNAs must first undergo extensive post-transcriptional modifications. miRNAs are expressed as primary transcripts known as pri-miRNAs from much longer RNA-coding genes and, in the cell nucleus, are processed by a microprocessor complex into a 70-nucleotide stem-loop structure called pre-miRNA. This complex consists of an RNase III enzyme called Drosha and the dsRNA-binding protein DGCR8. The dsRNA portion of this pre-miRNA is bound and cleaved by Dicer, generating a mature miRNA molecule that can be incorporated into the RISC complex; therefore, miRNAs and siRNAs share the same downstream cellular mechanisms. Short hairpin RNAs, or small hairpin RNAs (shRNAs / hairpin vectors), are artificial RNA molecules with tight hairpin turns that can be used to repress the expression of target genes via RNA interference. The most well-studied result is post-transcriptional gene silencing, which occurs when the guide strand pairs with the complementary sequence of the messenger RNA molecule, inducing cleavage by Argonaute 2 (Ago2), a catalytic component of RISC.It will be understood that RNAi molecules may be applied to / within plants themselves, or encoded by a suitable vector on which the RNAi molecule is expressed. Delivery and expression systems for RNAi molecules, such as siRNA, shRNA, or miRNA, are well known in the art.
[0390] The mutations described herein may be introduced by mutagenesis, which may be carried out according to any technique known in the art. As used herein, “mutagenization” or “mutagenesis” includes both conventional mutagenesis and site-directed mutagenesis or “genome editing” or “gene editing.” In conventional mutagenesis, DNA-level modifications are not produced in a targeted manner. Plant cells or plants are exposed to mutagenic conditions, such as Tilling, through exposure to ultraviolet light or the use of chemicals (Till et al., 2004). An additional method of random mutagenesis is transposon-assisted mutagenesis. Site-directed mutagenesis allows for the introduction of DNA-level modifications in a targeted manner at predefined locations in DNA. For example, TALENS, meganucleases, homing endonucleases, zinc finger nucleases, or the CRISPR / Cas system further described herein may be used for this purpose.
[0391] The mutations described herein may be introduced by random mutagenesis. Those skilled in the art will understand that the identification and selection of suitable mutations may involve appropriate selection assays, such as functional selection assays (including genotype or phenotype selection assays). In random mutagenesis, cells or organisms may be exposed to mutagenic substances, such as UV rays, X-rays, or gamma rays, or mutagenic chemicals (e.g., ethyl methanesulfonate (EMS), ethyl nitrosourea (ENU), or dimethyl sulfate (DMS)), and mutants with desired properties are selected. Mutants can be identified, for example, by tilling (targeting local lesions induced in the genome). The method combines mutagenesis, such as using a chemical mutagen like ethyl methanesulfonate (EMS), with a highly sensitive DNA screening technique for identifying single nucleotide / point mutations in target genes. The tilling method relies on the formation of DNA heteroduplexes, which are formed when multiple alleles are amplified by PCR, heated, and slowly cooled. A mismatch between two DNA strands forms a "bubble," which is then cleaved by a single-stranded nuclease. The resulting products are then separated by size, for example, by HPLC. See also McCallum et al., “Targeted screening for induced mutations”; Nat Biotechnol. 2000 Apr;18(4):455-7 and McCallum et al., “Targeting induced local lesions IN genomes (TILLING) for plant functional genomics”; Plant Physiol. 2000 Jun;123(2):439-42 (both are incorporated in their entirety by reference).As further examples, but not limited to these, methodologies described in the following publications, which are incorporated in their entirety by reference, may be employed in accordance with the present invention, such as EMS mutagenesis: Till et al., “Discovery of induced point mutations in maize genes by TILLING”; BMC Plant Biol. 2004 Jul 28;4:12; and Weil & Monde, “Getting the point-mutations in maize” Crop Sci 2007; 47 S60-S67. Those skilled in the art will understand that the (average) mutation density can be altered or fixed depending on the dose of mutagen (chemical irradiation). In some embodiments, random mutagenesis is single-nucleotide mutagenesis. In some embodiments, random mutagenesis is chemical mutagenesis, preferably EMS mutagenesis.
[0392] "Genes editing," "genome editing," "genetic modification," or "genome modification" refers to genetic engineering in which DNA or RNA is inserted, deleted, modified, or replaced in the genome (or transcriptome) of an organism. Therefore, gene editing encompasses DNA editing and RNA editing. Gene editing may include targeted or non-targeted (random) mutagenesis. Targeted mutagenesis may be achieved, for example, with designer nucleases, such as meganucleases, zinc finger nucleases (ZFNs), transcriptional activator-like effector-based nucleases (TALENs), and clustered, regularly arranged short palindromic sequence repeats (CRISPR / Cas9) systems. These nucleases create site-specific double-strand breaks (DSBs) at desired locations in the genome. The induced double-strand breaks are repaired by non-homologous end joining (NHEJ) or homologous recombination (HR), resulting in targeted mutations or nucleic acid modifications. The use of designer nucleases is particularly suitable for generating gene knockouts or knockdowns. In some embodiments, as described elsewhere in this specification, designer nucleases have been developed that specifically induce mutations in ig and / or centromere or kinetochore genes to produce, for example, gene mutations or knockouts. Alternatively, RNA-specific CRISPR / Cas systems (e.g., Cas13) enable site-specific cleavage of (single-stranded) RNA, and therefore knockdown can be achieved by, for example, RNA-specific CRISPR / Cas systems. Accordingly, in some embodiments, designer nucleases, particularly RNA-specific CRISPR / Cas systems, have been developed that specifically target mRNA to cleave mRNA and produce gene / mRNA / protein knockdowns. Delivery and expression systems for designer nuclease systems are well known in the art.
[0393] In some embodiments, the nuclease or targeted / site-specific / homing nuclease is a (modified) CRISPR / Cas system or complex, a (modified) Cas protein, a (modified) zinc finger, a (modified) zinc finger nuclease (ZFN), a (modified) transcription factor-like effector (TALE), a (modified) transcription factor-like effector nuclease (TALEN), or a (modified) meganuclease, or includes, is essentially derived from, or consists of. In some embodiments, the (modified) nuclease or targeted / site-specific / homing nuclease is a (modified) RNA-inducing nuclease, or includes, is essentially derived from, or consists of a (modified) RNA-inducing nuclease. In some embodiments, it is understood that the nuclease may be a codon optimized for expression in plants. As used herein, the term “targeting” of a selected nucleic acid sequence means that a nuclease or nuclease complex acts specifically on a nucleotide sequence. For example, in the description of the CRISPR / Cas system, a guide RNA can hybridize with a selected nucleic acid sequence. As used herein, “hybridization” or “hybridizing” means a reaction in which one or more polynucleotides react to form a complex that is stabilized via hydrogen bonds between the bases of nucleotide residues. Hydrogen bonds may be generated by Watson-Crick base pairs, Hugstein bonds, or any other sequence-specific method. The complex may include two strands forming a double-stranded structure, three or more strands forming a multi-stranded complex, a single self-hybridized strand, or any combination thereof. Hybridization is the process by which a single-stranded nucleic acid molecule attaches itself to a complementary nucleic acid strand, i.e., the process of matching this base pair.Standard procedures for hybridization are described, for example, in Sambrook et al. (Molecular Cloning. A Laboratory Manual, Cold Spring Harbor Laboratory Press, 3rd edition 2001). Preferably, this is understood to mean that at least 50%, more preferably at least 55%, 60%, 65%, 70%, 75%, 80%, or 85%, and more preferably 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% of the bases of the nucleic acid chain form base pairs with the complementary nucleic acid chain. The hybridization reaction may constitute a broader process step, such as the initiation of PGR or enzymatic cleavage of polynucleotides. A sequence that can hybridize with a given sequence is called a "complement" of the given sequence.
[0394] Gene editing may include transient, inducible, or constitutive expression of gene editing components or systems. Gene editing may include genomic integration of gene editing components or systems or the presence of episomes. Gene editing components or systems may be provided on vectors, such as plasmids, which may be delivered by suitable delivery vehicles known in the art. Preferred vectors are expression vectors.
[0395] Gene editing may include providing a recombinant template to achieve homology repair (HDR). For example, a gene element may be replaced by gene editing that provides a recombinant template. The DNA may be cut upstream and downstream of the sequence that needs to be replaced. In this way, the sequence to be replaced is excised from the DNA. HDR replaces the excised sequence with the template.
[0396] In one embodiment, the modification or mutation of nucleic acid is brought about by a (modified) transcriptional activator-like effector nuclease (TALEN) system. The transcriptional activator-like effector (TALE) can be engineered to bind to substantially any target DNA sequence. Exemplary methods of genome editing using TALEN systems are found, for example, in Cermak T. Doyle EL., Christian M., Wang L., Zhang Y., Schmidt C et al. (Efficient design and assembly of custom TALEN and other TAL effector-based constructs for DNA targeting. Nucleic Acids Res. 2011;39:e82); Zhang F., Cong L., Lodato S., Kosuri S., Church GM., Arlotta P (Efficient construction of sequence-specific TAL effectors for modulating mammalian transcription. Nat Biotechnol. 2011;29:149-153), and in U.S. Patents No. 8,450,471, 8,440,431, and 8,440,432, all of which are incorporated herein by reference. For further guidance, naturally occurring TALE or “wild-type TALE,” though not limited to these, are nucleic acid-binding proteins secreted by many species of proteobacteria. TALE polypeptides contain a nucleic acid-binding domain composed of tandem repeats of highly conserved monomeric polypeptides, primarily 33, 34, or 35 amino acids long, and differing mainly at amino acid positions 12 and 13. In a favorable embodiment, the nucleic acid is DNA. As used herein, the terms “polypeptide monomer” or “TALE monomer” are used to refer to the highly conserved repeat polypeptide sequence within the TALE nucleic acid-binding domain, and the terms “repeat variable two residues” or “RVD” are used to refer to the highly variable amino acids at positions 12 and 13 of the polypeptide monomer.As provided through this disclosure, the amino acid residues of RVD are indicated using the IUPAC single-letter codes for amino acids. A common representation of the TALE monomer contained within the DNA-binding domain is X1-11-(X12X13)-X14-33 or 34 or 35, where the subscript indicates the amino acid position and X represents any amino acid. X12X13 represents RVD. In some polypeptide monomers, the variable amino acid at position 13 is missing or absent, and in such polypeptide monomers, RVD consists of a single amino acid. In such cases, RVD can be represented instead as X*, where X represents X12 and (*) indicates the absence of X13. The DNA-binding domain contains several repeats of the TALE monomer, which can be represented as (X1-11-(X12X13)-X14-33 or 34 or 35)z, where in a favorable embodiment z is at least 5 to 40. In a more favorable embodiment, z is at least 10 to 26. TALE monomers have nucleotide binding affinity determined by the amino acid identity of their RVD. For example, polypeptide monomers with NI RVD preferentially bind to adenine (A), polypeptide monomers with NG RVD preferentially bind to thymine (T), polypeptide monomers with HD RVD preferentially bind to cytosine (C), and polypeptide monomers with NN RVD preferentially bind to both adenine (A) and guanine (G). In yet another embodiment of the present invention, polypeptide monomers with IG RVD preferentially bind to T. Therefore, the number and order of polypeptide monomer repeats in the nucleic acid binding domain of TALE determine its nucleic acid target specificity. In yet another embodiment of the present invention, polypeptide monomers with NS RVD may recognize all four base pairs and bind to A, T, G, or C. The structure and function of TALE are further described, for example, by Moscou et al. (Science 326:1501 (2009)), Boch et al. (Science 326:1509-1512 (2009)), and Zhang et al. (Nature Biotechnology 29:149-153 (2011)), and all of these are incorporated by reference.
[0397] In one embodiment, the modification or mutation of nucleic acids is brought about by a (modified) zinc finger nuclease (ZFN) system. The ZFN system uses an artificial restriction enzyme produced by fusing a zinc finger DNA-binding domain to a DNA-cleaving domain that can be manipulated to target a desired DNA sequence. Exemplary methods of genome editing using ZFNs can be found, for example, in U.S. Patent Nos. 6,534,261, 6,607,882, 6,746,838, 6,794,136, 6,824,978, 6,866,997, 6,933,113, 6,979,539, 7,013,219, 7,030,215, 7,220,719, 7,241,573, 7,241,574, 7,585,849, 7,595,376, 6,903,185, and 6,479,626, all of which are incorporated herein by reference. Further guidance may limit this, but artificial zinc finger (ZF) technology includes arrays of ZF modules targeting novel DNA binding sites within the genome. Each finger module in a ZF array targets three DNA bases. Customized arrays of individual zinc finger domains are assembled into ZF proteins (ZFPs). ZFPs may contain functional domains. The first synthetic zinc finger nucleases (ZFNs) were developed by fusing ZF proteins to the catalytic domain of the IIS-type restriction enzyme FokI (Kim, YG et al., 1994, Chimeric restriction endonuclease, Proc. Natl. Acad. Sci. USA 91, 883-887; Kim, YG et al., 1996, Hybrid restriction enzymes: zinc finger fusions to Fok I cleavage domain. Proc. Natl. Acad. Sci. USA 93, 1156-1160).By using paired ZFN heterodimers, each targeting different nucleotide sequences separated by short spacers, off-target activity can be reduced and cleavage specificity can be enhanced (Doyon, Y. et al., 2011, Enhancing zinc-finger-nuclease activity with improved obligate heterodimeric architectures. Nat. Methods 8, 74-79). ZFPs can also be designed as transcription activators and repressors and have been used to target many genes in diverse organisms.
[0398] In one embodiment, nucleic acid modification is brought about by a (modified) meganuclease, which is an endodeoxyribonuclease characterized by a large recognition site (a 12-40 base pair double-stranded DNA sequence). Exemplary methods using meganucleases can be found in U.S. Specifications 8,163,514, 8,133,697, 8,021,867, 8,119,361, 8,119,381, 8,124,369, and 8,129,134.
[0399] In some embodiments, nucleic acid modification is brought about by a (modified) CRISPR / Cas complex or system. General information regarding CRISPR / Cas systems, their components, and the delivery of such components includes methods, materials, delivery vehicles, vectors, particles, and their manufacture and use, quantities and formulations, and eukaryotic cells expressing Cas9CRISPR / Cas, eukaryotes expressing Cas-9CRISPR / Cas, including mice, can be referred to as: U.S. Patent Nos. 8,999,641, 8,993,233, 8,697,359, 8,771,945, 8,795,9 Specifications 65, 8,865,406, 8,871,445, 8,889,356, 8,889,418, 8,895,308, 8,906,616, 8,932,814, 8,945,839, 8,993,233 and 8,999,641; U.S. Patent Application Publication No. 2014-0310830 (U.S. Patent Application No. 14 / 105,031), U.S. Patent Application Publication No. 2014-0287938 (U.S. Patent Application No. 14 / 213,991), U.S. Patent Application Publication No. 2014-0273234 (U.S. Patent Application No. 14 / 293,674), U.S. Patent Application Publication No. 2014-0273232 (U.S. Patent Application No. 14 / 290,575), U.S. Patent Application Publication No. 2014-0273231 (U.S. Patent Application No. 14 / 259,420), U.S. Patent Application Publication No. 2014-0256046 (U.S. Patent Application No. 14 / 226,274), U.S. Patent Application Publication No. 201 Specifications 4-0248702 (US Patent Application No. 14 / 258,458), US Patent Publication No. 2014-0242700 (US Patent Application No. 14 / 222,930), US Patent Publication No. 2014-0242699 (US Patent Application No. 14 / 183,512), US Patent Publication No. 2014-0242664 (US Patent Application No. 14 / 104,990), US Patent Publication No. 2014-0234972 (US Patent Application No. 14 / 183,471),U.S. Patent Application Publication No. 2014-0227787 (U.S. Patent Application No. 14 / 256,912), U.S. Patent Application Publication No. 2014-0189896 (U.S. Patent Application No. 14 / 105,035), U.S. Patent Application Publication No. 2014-0186958 (U.S. Patent Application No. 14 / 105,017), U.S. Patent Application Publication No. 2014-0186919 (U.S. Patent Application No. 14 / 104,977), U.S. Patent Application Publication No. 2014-0186843 (U.S. Patent Application No. 14 / 104,900), United States Published Patent Application No. 2014-0179770 (US Patent Application No. 14 / 104,837), and Published Patent Application No. 2014-0179006 (US Patent Application No. 14 / 183,486), Published Patent Application No. 2014-0170753 (US Patent Application No. 14 / 183,429), Published Patent Application No. 2015-0184139 (US Patent Application No. 14 / 324,960), Published Patent Application No. 14 / 054,414, Published European Patent Application No. 2771468 (EP138185 70.7), European Patent Application Publication No. 2764103 (EP13824232.6), and European Patent Application Publication No. 2784162 (EP14170383.5), and PCT Patent Publications International Publication No. 2014 / 093661 (PCT / US2013 / 074743), International Publication No. 2014 / 093694 (PCT / US2013 / 074790), International Publication No. 2014 / 093595 (PCT / US2013 / 074611), International Publication No. 2014 / 093718 (PCT / US2013 / 074825), International Publication No. 2014 / 0 International Publication No. 93709 (PCT / US2013 / 074812), International Publication No. 2014 / 093622 (PCT / US2013 / 074667), International Publication No. 2014 / 093635 (PCT / US2013 / 074691), International Publication No. 2014 / 093655 (PCT / US2013 / 074736), International Publication No. 2014 / 093712 (PCT / US2013 / 074819), International Publication No. 2014 / 093701 (PCT / US2013 / 074800), International Publication No. 2014 / 018423 (PCT / US2013 / 051418),International Publication No. 2014 / 204723 (PCT / US2014 / 041790), International Publication No. 2014 / 204724 (PCT / US2014 / 041800), International Publication No. 2014 / 204725 (PCT / US2014 / 041803), International Publication No. 2014 / 204726 (PCT / US2014 / 041804) International Publication No. 2014 / 204727 (PCT / US2014 / 041806), International Publication No. 2014 / 204728 (PCT / US2014 / 041808), International Publication No. 2014 / 204729 (PCT / US2014 / 041809), International Publication No. 2015 / 089351 (PCT / US2014 / 069897) International Publication No. 2015 / 089354 (PCT / US2014 / 069902), International Publication No. 2015 / 089364 (PCT / US2014 / 069925), International Publication No. 2015 / 089427 (PCT / US2014 / 070068), International Publication No. 2015 / 089462 (PCT / US2014 / 070127) ), International Publication No. 2015 / 089419 (PCT / US2014 / 070057), International Publication No. 2015 / 089465 (PCT / US2014 / 070135), International Publication No. 2015 / 089486 (PCT / US2014 / 070175), PCT / US2015 / 05169, PCT / US2015 / 051830. See also U.S. Provisional Patent Applications 61 / 758,468, 61 / 802,174, 61 / 806,375, 61 / 814,263, 61 / 819,803, and 61 / 828,130, filed on March 15, 2013, March 28, 2013, April 20, 2013, May 6, 2013, and May 28, 2013, respectively. See also U.S. Provisional Patent Application 61 / 836,123, filed on June 17, 2013. Furthermore, refer to U.S. Provisional Patent Applications No. 61 / 835,931, 61 / 835,936, 61 / 835,973, 61 / 836,080, 61 / 836,101, and 61 / 836,127, each filed on June 17, 2013. Also refer to U.S. Provisional Patent Applications No. 61 / 862,468 and 61 / 862,355, filed on August 5, 2013, and U.S. Provisional Patent Application No. 61 / 871,301, filed on August 28, 2013.See also U.S. Provisional Patent Application No. 61 / 960,777, filed September 25, 2013, and U.S. Provisional Patent Application No. 61 / 961,980, filed October 28, 2013. See also: PCT / US2014 / 62558 filed on 28 October 2014, and U.S. Provisional Patent Applications 61 / 915,148, 61 / 915,150, 61 / 915,153, 61 / 915,203, 61 / 915,251, 61 / 915,301, 61 / 915,267, 61 / 915,260 and 61 / 915,397 filed on 12 December 2013; 61 / 757,972 and 61 / 768,959 filed on 29 January 2013 and 25 February 2013, respectively; and 62 / Nos. 010,888 and 62 / 010,879; Nos. 62 / 010,329, 62 / 010,439 and 62 / 010,441 filed on 10 June 2014, respectively; Nos. 61 / 939,228 and 61 / 939,242 filed on 12 February 2014, respectively; No. 61 / 980,012 filed on 15 April 2014; No. 62 / 038,358 filed on 17 August 2014; Nos. 62 / 055,484, 62 / 055,460 and 62 / 055,487 filed on 25 September 2014; No. 62 / 069,243 filed on 27 October 2014. Refer to PCT application filed on June 10, 2014, specifically designating the United States, application number PCT / US14 / 41806. Refer to U.S. Provisional Patent Application No. 61 / 930,214, filed on February 22, 2014. Refer to PCT application filed on June 10, 2014, specifically designating the United States, application number PCT / US14 / 41806. Also listed are: U.S. Application No. 62 / 180,709, June 17, 2015, PROTECTED GUIDE RNAS (PGRNAS); U.S. Application No. 62 / 091,455, PROTECTED GUIDE RNAS (PGRNAS), filed December 12, 2014; U.S. Application No. 62 / 096,708, December 24, 2014, PROTECTED GUIDE RNAS (PGRNAS); U.S. Application No. 62 / 091,462 (December 14, 2014), U.S. Application No. 62 / 096,324 (December 23, 2014), U.S. Application No. 62 / 180,681 (June 17, 2015), and U.S. Application No. 62 / 237,496 (October 5, 2015), DEAD GUIDES FOR CRISPR TRANSCRIPTION FACTORS; U.S. Patent Application No. 62 / 091,456 (December 12, 2014), U.S. Patent Application No. 62 / 180,692 (June 17, 2015), ESCORTED AND FUNCTIONALIZED GUIDES FOR CRISPR-CAS SYSTEMS; U.S. Patent Application No. 62 / 091,461 (December 12, 2014), DELIVERY, USE AND THERAPEUTIC APPLICATIONS OF THE CRISPR-CAS SYSTEMS AND COMPOSITIONS FOR GENOME EDITING AS TO HEMATOPOETIC STEM CELLS (HSCs); U.S. Patent Application No. 62 / 094,903 (December 19, 2014), UNBIASED IDENTIFICATION OF DOUBLE-STRAND BREAKS AND GENOMIC REARRANGEMENT BY GENOME-WISE INSERT CAPTURE SEQUENCING; U.S. Patent Application No. 62 / 096,761 (December 24, 2014), ENGINEERING OF SYSTEMS, METHODS AND OPTIMIZED ENZYME AND GUIDE SCAFFOLDS FOR SEQUENCE MANIPULATION;U.S. Patent Application No. 62 / 098,059 (December 30, 2014), U.S. Patent Application No. 62 / 181,641 (June 18, 2015), and U.S. Patent Application No. 62 / 181,667 (June 18, 2015), RNA-TARGETING SYSTEM; U.S. Patent Application No. 62 / 096,656 (December 24, 2014), and U.S. Patent Application No. 62 / 181,151 (June 17, 2015), CRISPR HAVING OR ASSOCIATED WITH DESTABILIZATION DOMAINS; U.S. Patent Application No. 62 / 096,697 (December 24, 2014), CRISPR HAVING OR ASSOCIATED WITH AAV; U.S. Patent Application No. 62 / 098,158 (December 30, 2014), ENGINEERED CRISPR COMPLEX INSERTIONAL TARGETING SYSTEMS; U.S. Patent Application No. 62 / 151,052 (April 22, 2015), CELLULAR TARGETING FOR EXTRACELLULAR EXOSOMAL REPORTING; U.S. Patent Application No. 62 / 054,490 (September 24, 2014), DELIVERY, USE AND THERAPEUTIC APPLICATIONS OF THE CRISPR-CAS SYSTEMS AND COMPOSITIONS FOR TARGETING DISORDERS AND DISEASES USING PARTICLE DELIVERY COMPONENTS; U.S. Patent Application No. 61 / 939,154 (February 12, 2014), SYSTEMS, METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION WITH OPTIMIZED FUNCTIONAL CRISPR-CAS SYSTEMS; U.S. Patent Application No. 62 / 055,484 (September 25, 2014), SYSTEMS, METHODS AND COMPOSITIONS FOR SEQUENCE MANIPULATION WITH OPTIMIZED FUNCTIONAL CRISPR-CAS SYSTEMS;U.S. Patent Application No. 62 / 087,537 (December 4, 2014), SYSTEMS, METHODS, AND COMPOSITIONS FOR SEQUENCE MANIPULATION WITH OPTIMIZED FUNCTIONAL CRISPR-CAS SYSTEMS; U.S. Patent Application No. 62 / 054,651 (September 24, 2014), DELIVERY, USE, AND THERAPEUTIC APPLICATIONS OF THE CRISPR-CAS SYSTEMS AND COMPOSITIONS FOR MODELING COMPETITION OF MULTIPLE CANCER MUTATIONS IN VIVO; U.S. Patent Application No. 62 / 067,886 (October 23, 2014), DELIVERY, USE, AND THERAPEUTIC APPLICATIONS OF THE CRISPR-CAS SYSTEMS AND COMPOSITIONS FOR MODELING COMPETITION OF MULTIPLE CANCER MUTATIONS IN VIVO; U.S. Patent Application No. 62 / 054,67 (September 24, 2014) and U.S. Patent Application No. 62 / 181,002 (June 17, 2015), Delivery, Use and Therapeutic Applications of the CRISPR-CAS Systems and Compositions in Neuronal Cells / Tissues; U.S. Patent Application No. 62 / 054,528 (September 24, 2014), Delivery, Use and Therapeutic Applications of the CRISPR-CAS Systems and Compositions in Immune Diseases or Disorders; U.S. Patent Application No. 62 / 055,454 (September 25, 2014), Delivery, Use and Therapeutic Applications of the CRISPR-CAS Systems and Compositions for Targeting Disorders and DISEASES USING CELL PENETRATION PEPTIDES (CPP);U.S. Patent Application No. 62 / 055,460 (September 25, 2014), MULTIFUNCTIONAL-CRISPR COMPLEXES AND / OR OPTIMIZED ENZYME-LINKED FUNCTIONAL-CRISPR COMPLEXES; U.S. Patent Application No. 62 / 087,475 (December 4, 2014) and U.S. Patent Application No. 62 / 181,690 (June 18, 2015), FUNCTIONAL SCREENING WITH OPTIMIZED FUNCTIONAL CRISPR-CAS SYSTEMS; U.S. Patent Application No. 62 / 055,487 (September 25, 2014), FUNCTIONAL SCREENING WITH OPTIMIZED FUNCTIONAL CRISPR-CAS SYSTEMS; U.S. Patent Application No. 62 / 087,546 (December 4, 2014) and U.S. Patent Application No. 62 / 181,687 (June 18, 2015), MULTIFUNCTIONAL CRISPR COMPLEXES AND / OR OPTIMIZED ENZYME LINKED FUNCTIONAL-CRISPR COMPLEXES; and U.S. Patent Application No. 62 / 098,285 (December 30, 2014), CRISPR MEDIATED IN VIVO MODELING AND GENETIC SCREENING OF TUMOR GROWTH AND METASTASIS. U.S. Patent Applications No. 62 / 181,659 (June 18, 2015) and No. 62 / 207,318 (September 19, 2015) are cited as examples, relating to engineering and optimization of systems, methods, encephalome, and guide scaffolds of CAS9 ortholologs and variants for sequence manipulation. U.S. Patent Application No. 62 / 181,663 (June 18, 2015) and U.S. Patent Application No. 62 / 245,264 (October 22, 2015), Novel Crisp Enzymes and Systems; U.S. Patent Application No. 62 / 181,675 (June 18, 2015) and Patent Attorney Registration No. 46783.01.2128 (filed October 22, 2015), Novel Crisp Enzymes and Systems;U.S. applications No. 62 / 232,067 (September 24, 2015), No. 62 / 205,733 (August 16, 2015), No. 62 / 201,542 (August 5, 2015), No. 62 / 193,507 (July 16, 2015), No. 62 / 181,739 (June 18, 2015), Novel Crisp Enzymes and Systems, and No. 62 / 245,270 (October 22, 2015), Novel Crisp Enzymes and Systems. U.S. Patent Application No. 61 / 939,256 (February 12, 2014) and International Publication No. 2015 / 089473 (PCT / US2014 / 070152), ENGINEERING OF SYSTEMS, METHODS AND OPTIMIZED GUIDE COMPOSITIONS WITH NEW ARCHITECTURES FOR SEQUENCE MANIPULATION, PCT / US2015 / 045504 (August 15, 2015), U.S. Patent Application No. 62 / 180,699 (June 17, 2015); Also cited is U.S. Patent Application No. 62 / 038,358 (August 17, 2014), GENOME EDITING USING CAS9 NICKASES. European Patent Application No. 3009511. Furthermore, multiple genome engineering using CRISPR / Cas systems is referenced. Cong, L., Ran, FA, Cox, D., Lin, S., Barretto, R., Habib, N., Hsu, PD, Wu, X., Jiang, W., Marraffini, LA, & Zhang, F. Science Feb 15;339(6121):819-23 (2013); RNA-guided editing of bacterial genomes using CRISPR-Cas systems. Jiang W., Bikard D., Cox D., Zhang F, Marraffini LA. Nat Biotechnol Mar;31(3):233-9 (2013); One-Step Generation of Mice Carrying Mutations in Multiple Genes by CRISPR / Cas-Mediated Genome Engineering. Wang H., Yang H., Shivalila CS., Dawlaty MM., Cheng AW., Zhang F., Jaenisch R. Cell May 9;153(4):910-8 (2013); Optical control of mammalian endogenous transcription and epigenetic states. Konermann S, Brigham MD, Trevino AE, Hsu PD, Heidenreich M, Cong L, Platt RJ, Scott DA, Church GM, Zhang F. Nature. 2013 Aug 22;500(7463):472-6. doi: 10.1038 / Nature12466.Epub 2013 Aug 23; Double Nicking by RNA-Guided CRISPR Cas9 for Enhanced Genome Editing Specificity. Ran, FA., Hsu, PD., Lin, CY., Gootenberg, JS., Konermann, S., Trevino, AE., Scott, DA., Inoue, A., Matoba, S., Zhang, Y., & Zhang, F. Cell Aug 28. pii: S0092-8674(13)01015-5. (2013); DNA targeting specificity of RNA-guided Cas9 nucleases. Hsu, P., Scott, D., Weinstein, J., Ran, FA., Konermann, S., Agarwala, V., Li, Y., Fine, E., Wu, X., Shalem, O., Cradick, TJ., Marraffini, LA., Bao, G., & Zhang, F. Nat Biotechnol doi:10.1038 / nbt.2647 (2013); Genome engineering using the CRISPR-Cas9 system. Ran, FA., Hsu, PD., Wright, J., Agarwala, V., Scott, DA., Zhang, F. Nature Protocols Nov;8(11):2281-308. (2013); Genome-Scale CRISPR-Cas9 Knockout Screening in Human Cells. Shalem, O., Sanjana, NE., Hartenian, E., Shi, X., Scott, DA., Mikkelson, T., Heckl, D., Ebert, BL., Root, DE., Doench, JG., Zhang, F. Science Dec 12. (2013).[Epub ahead of print]; Crystal structure of cas9 in complex with guide RNA and target DNA. Nishimasu, H., Ran, FA., Hsu, PD., Konermann, S., Shehata, SI., Dohmae, N., Ishitani, R., Zhang, F., Nureki, O. Cell Feb 27. (2014). 156(5):935-49; Genome-wide binding of the CRISPR endonuclease Cas9 in mammalian cells. Wu X., Scott DA., Kriz AJ., Chiu AC., Hsu PD., Dadon DB., Cheng AW., Trevino AE., Konermann S., Chen S., Jaenisch R., Zhang F., Sharp PA. Nat Biotechnol. (2014) Apr 20. doi: 10.1038 / nbt.2889; CRISPR-Cas9 Knockin Mice for Genome Editing and Cancer Modeling, Plattら, Cell 159(2): 440-455 (2014) DOI: 10.1016 / j.cell.2014.09.014; Development and Applications of CRISPR-Cas9 for Genome Engineering, Hsu ら, Cell 157, 1262-1278 (June 5, 2014) (Hsu 2014); Genetic screens in human cells using the CRISPR / Cas9 system, Wang ら, Science. 2014 January 3; 343(6166): 80-84. doi:10.1126 / science.1246981; Rational design of highly active sgRNAs for CRISPR-Cas9-mediated gene inactivation, Doench ら, Nature Biotechnology 32(12):1262-7 (2014) published online 3 September 2014; doi:10.1038 / nbt.3026, and In vivo interrogation of gene function in the mammalian brain using CRISPR-Cas9, Swiech ら, Nature Biotechnology 33, 102-106 (2015) published online 19 October 2014; doi:10.1038 / nbt.3055, Cpf1 Is a Single RNA-Guided Endonuclease of a Class 2 CRISPR-Cas System, Zetsche ら, Cell 163, 1-13 (2015); Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems, Shmakov ら, Mol Cell 60(3): 385-397 (2015); C2c2 is a single-component programmable RNA-guided RNA-targeting CRISPR effector, Abudayyeh ら, Science (2016) published online June 2, 2016 doi: 10.1126 / science.aaf5573. Each of these publications, patents, patent publications, and applications, and all documents cited therein or during examination ("Cited Documents"), and all documents cited or referenced in the Cited Documents, together with any manufacturer's instructions, descriptions, product uses, and product sheets for any product in any document cited or incorporated herein by reference, are incorporated herein by reference and may be used in the practice of the Invention. All documents (e.g., these patents, patent publications, and applications, and filed cited documents) are incorporated herein by reference to the same extent as when the incorporation of individual documents by reference is specifically and individually indicated.
[0400] In one embodiment, the CRISPR / Cas system or complex is a Class 2 CRISPR / Cas system. In one embodiment, the CRISPR / Cas system or complex is a Type II, Type V, or Type VI CRISPR / Cas system or complex. The CRISPR / Cas system does not need to generate a customized protein to target a specific sequence, but can be programmed with an RNA guide (gRNA) to recognize a specific nucleic acid target; in other words, the Cas enzyme protein may be recruited using the short RNA guide to a specific nucleic acid target locus of interest (which may include or consist of RNA and / or DNA).
[0401] Generally, CRISPR / Cas or a CRISPR system, as used in the aforementioned literature herein, refers to a transcript and other elements involved in or instructing the activity of a CRISPR-related ("Cas") gene, and includes one or more sequences encoding the Cas gene, and tracr (trans-activated CRISPR) sequences (e.g., tracrRNA or active partial tracrRNA), tracr-mate sequences (including "direct repeats" and tracrRNA-treated partial direct repeats in the description of endogenous CRISPR systems), guide sequences (also called "spacers" in the description of endogenous CRISPR systems), or "RNA" as the term is used herein (e.g., RNA guiding Cas, e.g. Cas9, e.g. CRISPR RNA, and, where applicable, trans-activated (tracr)RNA or single guide RNA (sgRNA) (chimeric RNA)), or other sequences and transcripts from the CRISPR locus. Generally, the CRISPR system is characterized by elements that promote the formation of the CRISPR complex at the site of the target sequence (also called protospacers in descriptions of the endogenous CRISPR system). In descriptions of CRISPR complex formation, the "target sequence" refers to a sequence designed to be complementary to the guide sequence, and hybridization between the target sequence and the guide sequence promotes the formation of the CRISPR complex. The target sequence may include any polynucleotide, such as DNA or RNA polynucleotides.
[0402] In some embodiments, the gRNA is a chimeric guide RNA or a single guide RNA (sgRNA). In some embodiments, the gRNA comprises a guide sequence and a tracr mate sequence (or direct repeat). In some embodiments, the gRNA comprises a guide sequence, a tracr mate sequence (or direct repeat), and a tracr sequence. In some embodiments, the CRISPR / Cas system or complex described herein does not contain a tracr sequence and / or does not depend on the presence of a tracr sequence (for example, when the Cas protein is Cpf1).
[0403] As used herein, the terms “crRNA,” “guide RNA,” “single guide RNA,” “sgRNA,” or “one or more nucleic acid components” of a CRISPR / Cas locus effector protein include, where applicable, any polynucleotide sequences that hybridize with the target nucleic acid sequence and have sufficient complementarity with the target nucleic acid sequence to direct sequence-specific binding of the nucleic acid target complex to the target nucleic acid sequence. In some embodiments, the degree of complementarity is about 50%, 60%, 75%, 80%, 85%, 90%, 95%, 97.5%, 99%, or higher when optimally aligned using a suitable alignment algorithm. The optimal alignment can be determined using any suitable algorithm for aligning sequences, and without limitation, include the Smith-Waterman algorithm, the Needleman-Wunsch algorithm, algorithms based on the Burrows-Wheeler Transform (e.g., Burrows WheelerAligner), ClustalW, ClustalX, BLAT, Novoalign (available from Novocraft Technologies, www.novocraft.com), ELAND (Illumina, San Diego, CA), SOAP (available from soap.genomics.org.cn), and Maq (available from maq.sourceforge.net). The ability of the guide sequence (within the nucleic acid target guide RNA) to direct sequence-specific binding of the nucleic acid target complex to the target nucleic acid sequence may be evaluated by any suitable assay.
[0404] The guide sequence, and therefore the nucleic acid target guide RNA, may be selected to target any target nucleic acid sequence. The target sequence may be DNA. The target sequence may be genomic DNA. The target sequence may be mitochondrial DNA. The target sequence may be any RNA sequence. In some embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of messenger RNA (mRNA), pre-mRNA, ribosomal RNA (rRNA), transfer RNA (tRNA), microRNA (miRNA), small interfering RNA (siRNA), small nuclear RNA (snRNA), small nucleolar RNA (snoRNA), double-stranded RNA (dsRNA), non-coding RNA (ncRNA), long non-coding RNA (lncRNA), and small cytoplasmic RNA (scRNA). In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of mRNA, pre-mRNA, and rRNA. In some preferred embodiments, the target sequence may be a sequence within an RNA molecule selected from the group consisting of ncRNA and lncRNA. In some more preferred embodiments, the target sequence may be a sequence within an mRNA molecule or a premRNA molecule.
[0405] In one embodiment, the gRNA includes a stem loop, preferably a single stem loop. In one embodiment, the direct repeat sequence forms a stem loop, preferably a single stem loop. In one embodiment, the spacer length of the guide RNA is 15 to 35 nt. In one embodiment, the spacer length of the guide RNA is at least 15 nucleotides. In one embodiment, the spacer length is 15–17 nt, e.g., 15, 16, or 17 nt; 17–20 nt, e.g., 17, 18, 19, or 20 nt; 20–24 nt, e.g., 20, 21, 22, 23, or 24 nt; 23–25 nt, e.g., 23, 24, or 25 nt; 24–27 nt, e.g., 24, 25, 26, or 27 nt; 27–30 nt, e.g., 27, 28, 29, or 30 nt; 30–35 nt, e.g., 30, 31, 32, 33, 34, or 35 nt; or 35 nt or greater. In one embodiment, the CRISPR / Cas system requires tracrRNA. The “tracrRNA” sequence or similar term includes any polynucleotide sequence having sufficient complementarity with a crRNA sequence for hybridization. In some embodiments, the degree of complementarity between the tracrRNA sequence and the crRNA sequence is about 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 97.5%, 99%, or more, along the shorter of the two, when optimally aligned. In some embodiments, the tracr sequence is a nucleotide of about 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, or more in length. In some embodiments, the tracr sequence and the gRNA sequence are contained within a single transcript such that hybridization between the two produces a secondary structure, such as a transcript having hairpins. In one embodiment of the present invention, the transcript or the transcribed polynucleotide sequence has at least two or more hairpins. In a preferred embodiment, the transcript has two, three, four, or five hairpins. In a further embodiment of the present invention, the transfer has up to five hairpins.In a hairpin structure, the upstream sequence portion of the loop at the 5' end of the last "N" may correspond to the tracrmate sequence, and the 3' end of the loop may correspond to the tracr sequence. Alternatively, in a hairpin structure, the upstream sequence portion of the loop at the 5' end of the last "N" may correspond to the tracr sequence, and the 3' end of the loop may correspond to the tracrmate sequence. In an alternative embodiment, the CRISPR / Cas system does not require tracrRNA, as is known to those skilled in the art.
[0406] In one embodiment, the guide RNA (which can guide Cas to a target locus) may include (1) a guide sequence that can hybridize to the target locus, and (2) a tracr mate or direct repetition sequence (as known to those skilled in the art, depending on the type of Cas protein, in the 5' to 3' direction or the 3' to 5' direction). In a particular embodiment, the CRISPR / Cas protein is characterized by utilizing a guide RNA that includes a guide sequence and a direct repetition sequence that can hybridize to the target locus, and does not require tracrRNA. In a particular embodiment characterized by the CRISPR / Cas protein utilizing tracrRNA, the guide sequence, tracr mate, and tracr sequence may be present in a single RNA, i.e., sgRNA (located in the 5' to 3' direction or the 3' to 5' direction), or the tracrRNA may be a different RNA from the RNA containing the guide and tracr mate sequences. In these embodiments, tracr hybridizes to the tracr mate sequence, directing the CRISPR / Cas complex to the target sequence.
[0407] Typically, in descriptions of endogenous nucleic acid targeting systems, the formation of a nucleic acid target complex (including a guide RNA that hybridizes to a target sequence and forms a complex with one or more nucleic acid targeting effector proteins) results in alteration (e.g., cleavage) of one or both DNA or RNA strands in or near the target sequence (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from the target sequence). As used herein, the term “sequence related to the target locus of interest” means a sequence located near the target sequence (e.g., within 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 50, or more base pairs from the target sequence, where the target sequence is contained within the target gene locus of interest). A person skilled in the art will notice, in comparison to the target sequence, a specific cleavage site of a selected CRISPR / Cas system, which may be located within the target sequence or within the 3' or 5' of the target sequence, as is known in the art.
[0408] In some embodiments, unmodified nucleic acid-targeting effector proteins may have nucleic acid cleavage activity. In some embodiments, the nucleases described herein may direct the cleavage of one or both nucleic acid strands (which may be single-stranded or double-stranded, DNA, RNA, or hybrid) at or near the target sequence, for example, within the target sequence and / or within the complement of the target sequence or at a sequence associated with the target sequence. In some embodiments, nucleic acid-targeting effector proteins may direct the cleavage of one or both DNA or RNA strands within about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 50, 100, 200, 500, or more base pairs from the first or last nucleotide of the target sequence. In some embodiments, the cleavage may be blunt-ended (e.g., for Cas9, e.g., SaCas9 or SpCas9). In some embodiments, the cleavage may be twisted (e.g., for Cpf1), i.e., may produce sticky ends. In some embodiments, the cleavage is a twisted cleavage with a 5' overhang. In some embodiments, the cleavage is a twisted cleavage with a 5' overhang of 1 to 5 nucleotides, preferably 4 or 5 nucleotides. In some embodiments, the cleavage site is upstream of the PAM. In some embodiments, the cleavage site is downstream of the PAM. In some embodiments, the nucleic acid-targeting effector protein includes a target sequence that may be mutated with respect to the corresponding wild-type enzyme such that the mutant nucleic acid-targeting effector protein lacks the ability to cleave one or both of the DNA or RNA strands of the target polynucleotide. As a further example, two or more catalytic domains of the Cas protein (e.g., the RuvC I, RuvC II, and RuvC III domains, or HNH domain of the Cas9 protein) may be mutated to produce a mutant Cas protein that substantially lacks all DNA cleavage activity.In some embodiments, if the nucleic acid cleavage activity of the mutant enzyme is less than or equal to about 25%, 10%, 5%, 1%, 0.1%, or 0.01% of the nucleic acid cleavage activity of the unmutated enzyme (for example, if the DNA cleavage activity of the mutant is ineffective or negligible compared to the unmutated type), the nucleic acid-targeting effector protein may be considered to lack substantially all DNA and / or RNA cleavage activity. As used herein, the term “modified” Cas generally refers to a Cas protein that has one or more modifications or mutations (including point mutations, cleavages, insertions, deletions, chimeras, fusion proteins, etc.) compared to the wild-type Cas protein from which it is derived. With respect to origin, this is largely based on the meaning that the enzyme from which it is derived has a high degree of sequence homology with the wild-type enzyme, but is either publicly known in the art or has been mutated (modified) in any way described herein.
[0409] In one embodiment, the target sequence should be associated with a PAM (protospacer flanking motif) or a PFS (protospacer flanking sequence or site), i.e., a short sequence recognized by the CRISPR complex. The exact sequence and length requirements for the PAM vary depending on the CRISPR enzyme used, but a PAM is typically a 2-5 base pair sequence flanking the protospacer (i.e., the target sequence). Examples of PAM sequences are described in the following paragraphs of examples, and those skilled in the art will be able to identify further PAM sequences for use with a given CRISPR enzyme. Furthermore, manipulation of the PAM interaction (PI) domain allows for the programming of PAM specificity, improving the fidelity of target site recognition and increasing the versatility of the Cas, e.g., Cas9, genome engineering platform. Cas proteins, such as the Cas9 protein, may be designed to alter their PAM specificity, as described, for example, in Kleinstiver BP et al. (Engineered CRISPR-Cas9 nucleases with altered PAM specificities. Nature. 2015 Jul 23;523(7561):481-5. doi: 10.1038 / nature14592). In some embodiments, the method involves attaching a CRISPR complex to a target polynucleotide to cause cleavage of the target polynucleotide, thereby modifying the target polynucleotide, where the CRISPR complex comprises a CRISPR enzyme complexed with a guide sequence hybridized to a target sequence within the target polynucleotide, the guide sequence being linked to a tracr mate sequence that hybridizes to a tracr sequence. Those skilled in the art will understand that other Cas proteins may be similarly modified.
[0410] The Cas proteins referred to herein, for example but not limited to Cas9, Cpf1 (Cas12a), C2c1 (Cas12b), C2c2 (Cas13a), C2c3, and Cas13b proteins may originate from any suitable source and, as well as well documented in the art, may include different orthologues from a variety of (prokaryotes) organisms. In some embodiments, the Cas protein is (modified) Cas9, preferably (modified) Staphylococcus aureus Cas9 (SaCas9) or (modified) Streptococcus pyogenes Cas9 (SpCas9). In one embodiment, the Cas protein is (modified) Cpf1, preferably from an Acidaminococcus species, such as Acidaminococcus species BV3L6 Cpf1 (AsCpf1), or from a Lachnospiraceae bacterium Cpf1, such as Lachnospiraceae bacterium MA2020 or Lachnospiraceae bacterium MD2006 (LbCpf1). In one embodiment, the Cas protein is (modified) C2c2, preferably from Leptotrichia wadei C2c2 (LwC2c2) or Listeria newyorkensis FSL M6-0635C2c2 (LbFSLC2c2). In one embodiment, the (modified) Cas protein is C2c1. In one embodiment, the (modified) Cas protein is C2c3. In one embodiment, the (modified) Cas protein is Cas13b.
[0411] A doubling singular plant or plant part arises from the doubling of singular sets of chromosomes. A plant or seed obtained from a self-pollinating doubling singular plant in any generation can still be identified as a doubling singular plant. Doubling singular plants are considered homozygous plants. A plant is considered a doubling singular plant if it is fertile, even if its entire vegetative part does not consist of cells with a double set of chromosomes. For example, a chimera can be considered a doubling singular plant if it contains viable gametes.
[0412] Somatic phantom cells, phantom embryos, phantom seeds, or phantom seedlings produced from phantom seeds may be treated with chromosome doubling agents. Homozygous plants may be regenerated from phantom cells by contacting phantom cells, e.g., embryonic cells or callus produced from such cells, with chromosome doubling agents, e.g., colchicine, pronamide, dichipyl, trifluralin, or other known microtubule inhibitors or microtubule inhibitory herbicides, or nitrous oxide to create homozygous doubling phantom cells. Treatment of phantom seeds or the resulting seedlings generally produces chimeric plants that are partially phantom and partially doubling phantom. It may be beneficial to cut the seedlings before treatment with colchicine. Doubling phantom seeds are produced when the reproductive tissue contains doubling phantom cells.
[0413] In one embodiment, the present invention relates to a method for identifying a plant or plant part as described elsewhere in this specification, for example, a plant or plant part according to the present invention. Accordingly, in one embodiment, the present invention relates to a method for identifying a plant or plant part having singular-inducing activity or enhanced singular-inducing activity (as described elsewhere in this specification). In one embodiment, the present invention relates to a method for identifying a plant or plant part that includes, or expresses, a mutated undetermined gametophyte (ig) allele, gene, or protein (a polynucleic acid encoding) and a mutated centromere or kinetochore allele, gene, or protein, preferably CENH3 (a polynucleic acid encoding) (as described elsewhere in this specification). In one embodiment, the present invention relates to a method for identifying plants or plant parts that include, or express, alleles, genes, or proteins (coding polynucleic acids) of an undetermined gametophyte (ig) that confer or enhance singular-inducing activity or ability thereof, and alleles, genes, or proteins of a centromere or kinetochore, preferably CENH3 (coding polynucleic acid), that confer or enhance singular-inducing activity or ability thereof. In one embodiment, the present invention relates to a method for identifying plants or plant parts that have reduced expression, stability, and / or activity of alleles, genes, or proteins of an undetermined gametophyte and alleles, genes, or proteins of a mutated centromere or kinetochore, preferably CENH3 (coding polynucleic acid). In one embodiment, the present invention relates to a method for identifying a plant or plant part comprising a centromere or kinetochore allele, gene, or protein, preferably a polynucleic acid encoding CENH3, which has reduced expression, stability, and / or activity of an undetermined gametophyte allele, gene, or protein and confers or enhances singular-inducible activity or its ability.
[0414] In one embodiment, such a method includes detecting an allele, gene, or protein of a mutated indeterminate gametophyte (as described elsewhere in this specification), and detecting an allele, gene, or protein of a mutated centromere or kinetochore, preferably CENH3. In one embodiment, such a method includes detecting an allele, gene, or protein of an indeterminate gametophyte having singular-inducing activity or enhanced singular-inducing activity (as described elsewhere in this specification), and detecting an allele, gene, or protein of a centromere or kinetochore, preferably CENH3. In one embodiment, such a method includes detecting reduced expression, stability, and / or activity of an allele, gene, or protein of an indeterminate gametophyte, and detecting an allele, gene, or protein of a mutated centromere or kinetochore, preferably CENH3. In one embodiment, such a method includes detecting reduced expression, stability, and / or activity of an undetermined gametophyte allele, gene, or protein (as described elsewhere in this specification), and detecting alleles, genes, or proteins of a centromere or kinetochore, preferably CENH3, that have singularity-inducing activity or enhanced singularity-inducing activity. In one embodiment, such a method includes providing a sample containing (genomic) DNA from a plant or plant part. In one embodiment, such a method includes assaying for the presence of allele, gene, or protein mutations of ig and alleles, genes, or protein mutations of a centromere or kinetochore, or assaying for singularities that induce or enhance allele, gene, or protein mutations of ig, and assaying for singoploids that induce or enhance allele, gene, or protein mutations of a centromere or kinetochore.Those skilled in the art will understand that assays for mutations may be direct or indirect; that is, mutations may be detected directly (by a suitable assay, as described elsewhere in this specification) or indirectly (by the detection of, for example, linked or associated molecular or genetic markers, as described elsewhere in this specification).
[0415] In one embodiment, the present invention relates to a method for producing a plant or plant part, comprising mutagenes of one or more (endogenous) Ig alleles, genes, or proteins, and one or more (endogenous) centromere or kinetochore alleles, genes, or proteins, preferably CENH3, and / or introducing one or more mutated Ig alleles, genes, or proteins, and one or more mutated centromere or kinetochore alleles, genes, or proteins, preferably CENH3. Those skilled in the art will understand that a single allele may be mutated and homozygosity may be achieved in subsequent generations. Those skilled in the art will understand that Ig and the centromere or kinetochore proteins may be mutated simultaneously or subsequently, in any order. For example, in the first stage, ig (or the polynucleic acid encoding the ig protein) may mutate, and in a subsequent stage, the centromere or kinetochore protein (or the polynucleic acid encoding the centromere or kinetochore protein) may mutate, either in the same plant or plant part, or in one or more subsequent generations of plants or plant parts, or vice versa.
[0416] Any mutagenic means may be applied, as described elsewhere in this specification, including, for example, random mutagenesis and site-directed mutagenesis.
[0417] Aspects and embodiments of the present invention are further supported by the following non-limiting examples.
[0418] [Table 3-1] [Table 3-2]
[0419] Examples Example 1 A mutation in CenH3(E35K), which itself exhibits low maternal induction in maize, was introduced into ig-Alvey, a maize line possessing the singular inducer ig allele (see Sequence ID No. 1). After four generations of backcrossing, the ig-Alvey genomic background was reconstructed to 99%. The main difference lies in the exchange of the CenH3 allele. This line was tested for maternal and paternal induction using glossy mutants as testers, with marker analysis and flow cytometry for ploidy confirmation. The maternal induction rate was approximately 0.5%. However, independently of the backcross version, the paternal induction rate increased to an average of 5.7–7.5%, which was much higher than expected for ig-Alvey alone (1–3%).
[0420] [Table 4]
[0421] [Table 5]
[0422] [Table 6]
[0423] In induction tests with different mutations only in the CenH3 gene, no true paternal singularities were observed. However, the maternal induction rate can be used to indicate that the tested mutation may increase the induction rate when combined with other mutations.
Claims
1. A plant or plant part comprising a polynucleic acid encoding a mutated undetermined gametophyte (ig) protein and a polynucleic acid encoding a mutated centromere or kinetochore protein.
2. The plant or plant part according to claim 1, wherein the polynucleic acid encoding the mutated ig protein comprises one or more nucleic acid insertions compared to the polynucleic acid encoding the wild-type undetermined gametophyte (ig) protein, or the polynucleic acid encoding the mutated ig protein comprises a knockout mutation or a knockdown mutation.
3. The plant or plant part according to claim 1 or 2, wherein the polynucleic acid encoding the mutated ig protein includes the insertion of one or more nucleic acids in an ig codon corresponding to a codon selected from codons 118, 119, or 120 of a wild-type maize (Zea mays) ig protein as shown in SEQ ID NO: 7 or 8, or a codon selected from codons 191, 192, or 193 of a wild-type sorghum bicolor (Sorghum bicolor) ig protein as shown in SEQ ID NO: 22, or a codon selected from codons 143, 144, or 145 of a wild-type sorghum bicolor (Sorghum bicolor) ig protein as shown in SEQ ID NO: 25, or a codon selected from codons 94, 95, or 96 of a wild-type rapeseed (Brassica napus) ig protein as shown in SEQ ID NO: 28 or 31.
4. The plant or plant part according to any one of claims 1 to 3, wherein the protein of the mutated centromere or kinetochore is selected from the group comprising CENH3, CENP-C, KNL2, SCM3, SAD2, and SIM3, and is preferably CENH3.
5. The plant or plant part according to claim 4, wherein the mutated CENH3 protein contains one or more mutated amino acids corresponding to positions 3, 17, 32, 35, 9, 24, 29, 40, 42, 50, 55, 57, 61, 74, 82, 104, 109, 120, 148, 175, 130, 151, 157, 158, 164, 166, 83, 86, 124, 127, 132, 136, 152, 155, or 172 of the reference Arabidopsis thaliana CENH3 protein, and preferably the Arabidopsis thaliana CENH3 protein has an amino acid sequence that is at least 90%, preferably at least 95%, more preferably at least 98%, identical to the sequence shown in SEQ ID NO:
12.
6. The plant or plant part according to any one of claims 1 to 5, wherein the plant or plant part is selected from the group including the genera Zea, Sorghum, and Brassica, preferably Zea mays, Sorghum bicolor, and Brassica napus.
7. The plant in question is from the genus Zea, preferably from maize (Zea mays), and the mutated undetermined gametophyte (ig) protein is a) Encoded by a polynucleic acid comprising the nucleotide sequence of SEQ ID NO: 1, or a sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 1, b) Derived from the nucleotide sequence of SEQ ID NO: 2 or 3, or a coding sequence that is at least 90% identical, preferably at least 95%, and more preferably at least 98% identical to SEQ ID NO: 2 or 3, or c) Having the amino acid sequence of SEQ ID NO: 4 or 5, or an amino acid sequence that is at least 90% identical, preferably at least 95% identical, and more preferably at least 98% identical to SEQ ID NO: 4 or 5 The plant or plant part according to any one of claims 1 to 6.
8. The plant or plant part according to any one of claims 1 to 7, wherein the plant is corn (Zea mays), and the protein of the mutated centromere or kinetochore is a mutated CENH3 protein having an amino acid substitution at position 35, preferably corresponding to position 35 of SEQ ID NO: 14 or having an amino acid substitution at position 35 of SEQ ID NO: 14, preferably the amino acid substitution is 35K, for example E35K.
9. The plant according to any one of claims 1 to 8, further comprising a polynucleic acid encoding a site-specific DNA or RNA-binding protein.
10. The site-specific DNA or RNA-binding protein is a meganuclease (MN), zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), (mutated) Cas nuclease / effector protein, such as Cas9 nuclease, Cfp1 nuclease, MAD7 nuclease, dCas9-FokI, dCpf1-FokI, dMAD7 nuclease. The plant according to claim 9, wherein the nuclease is selected (mutated) from the group comprising crease-FokI, chimeric Cas9-cytidine deaminase, chimeric Cas9-adenine deaminase, chimeric FENI-FokI, and mega-TAL, nickase Cas9 (nCas9), chimeric dCas9 non-FokI nuclease, dCpf1 non-FokI nuclease, and dMAD7 non-FokI nuclease.
11. A method for producing a plant or plant part, comprising: preparing a singular, digenomic singular, or trigenomic singular plant obtained from a cross between a first plant which is a plant according to any one of claims 1 to 10 and a second plant; and converting a singular, digenomic singular, or trigenomic singular plant or plant part into a doubled singular, doubled digenomic singular, or doubled trigenomic singular plant or plant part.
12. A method for producing the plant according to any one of claims 1 to 10, A) (i) A step of preparing a plant or plant part, and (ii) the step of mutating one or more polynucleic acids encoding an allele, gene or protein of an (endogenous) IgG as defined in any of claims 2, 3 or 7, and the step of mutating one or more polynucleic acids encoding an allele, gene or protein of an (endogenous) centromere or kinetochore protein as defined in any of claims 4, 5 or 8, and / or the step of introducing (genomes) one or more mutated polynucleic acids encoding an allele, gene or protein of an IgG as defined in any of claims 2, 3 or 7, and one or more mutated polynucleic acids encoding an allele, gene or protein of an centromere or kinetochore protein as defined in any of claims 4, 5 or 8, or B) (i) The steps of preparing one or more polynucleic acids encoding an (endogenous) mutant IgG allele, gene, or protein as defined in any of claims 2, 3, or 7, and / or one or more (genomically) introduced mutant IgG alleles, genes, or proteins as defined in any of claims 2, 3, or 7, and (ii) the step of mutating one or more (endogenous) centromere or kinetochore protein alleles, genes or protein-coding polynucleic acids as defined in any of claims 4, 5 or 8, and / or the step of introducing one or more mutated centromere or kinetochore protein alleles, genes or protein-coding polynucleic acids as defined in any of claims 4, 5 or 8, or C) (i) A step of preparing one or more polynucleic acids encoding an allele, gene, or protein of a (endogenous) mutant centromere or kinetochore as defined in any of claims 4, 5, or 8, and / or one or more polynucleic acids encoding an allele, gene, or protein of a (genomically) introduced mutant centromere or kinetochore as defined in any of claims 4, 5, or 8, and (ii) the step of mutating one or more polynucleic acids encoding an (endogenous) IgG allele, gene, or protein as defined in any of claims 2, 3, or 7, and / or the step of introducing one or more mutated IgG alleles, genes, or protein-coding polynucleic acids as defined in any of claims 2, 3, or 7 into a (genome). A method for producing a plant or plant part according to any one of claims 1 to 10, including the method described in any one of claims 1 to 10.
13. A method for identifying a plant or plant part according to any one of claims 1 to 10, comprising detecting a mutated unspecified gametophyte protein as defined in any one of claims 2, 3, or 7 and a mutated centromere or kinetochore protein as defined in any one of claims 4, 5, or 8, or detecting a polynucleic acid encoding an unspecified gametophyte protein containing a mutation as defined in any one of claims 2, 3, or 7 and a polynucleic acid encoding a centromere or kinetochore protein containing a mutation as defined in any one of claims 4, 5, or 8.
14. A method for modifying plant genomic DNA, comprising: a) preparing a first plant which is the plant described in claim 11 or 12; b) preparing a second plant (containing the plant genomic DNA to be modified); c) pollinating the second maize plant with pollen from the first plant; and d) selecting at least one phantom, digenomic phantom, or trigenomic phantom offspring produced by the pollination in step (c) (wherein the phantom, digenomic phantom, or trigenomic phantom offspring contains the genome of the second plant rather than the genome of the first plant, and the genome of the phantom, digenomic phantom, or trigenomic phantom offspring is modified by a site-specific DNA or RNA-binding protein delivered by the first plant).
15. Use of a plant or plant part according to any one of claims 1 to 12 as a singular inducer, preferably a paternal singular inducer.
16. Corn (Zea mays) seeds deposited under NCIMB deposit number NCIMB43772.
17. (igEIN) corn (Zea Mays) seeds, a representative sample deposited under NCIMB deposit number NCIMB43772.
18. A corn (Zea mays) plant grown from or obtained from the seeds described in claim 16 or 17.
19. Plant parts of corn (Zea mays) grown from or obtained from seeds according to claim 16 or 17, or obtained from the plant according to claim 18.
20. A method for identifying or selecting plants or plant parts, for example, plants or plant parts having (enhanced) singular induction activity or the ability thereof, i) Prepare a plant or plant part having reduced expression, stability, and / or activity of the gene, mRNA, or protein of an undetermined gametophyte (ig). ii) Mutating a gene encoding a protein in the centromere or kinetochore, preferably CENH3, and iii) Analyze the singular induction activity or ability in the plant or plant part, or their offspring. This includes, and optionally further iv) Selecting a plant or plant part that has (enhanced) singular induction activity or the ability thereof. A method for identifying or selecting a plant or plant part, including the method described above.
21. A method for identifying or selecting plants or plant parts, for example, plants or plant parts having (enhanced) singular induction activity or the ability thereof, i) Prepare a first plant or plant part having reduced expression, stability, and / or activity of the gene, mRNA, or protein of an undetermined gametophyte (ig). ii) Crossing the first plant with a second plant having a gene encoding a mutated centromere or kinetochore protein, preferably CENH3, and iii) Analyze the singular induction activity or ability in the obtained offspring, This includes, and optionally further iv) Selecting a plant or plant part that has (enhanced) singular induction activity or the ability thereof. A method for identifying or selecting a plant or plant part, including the method described above.
22. Use of plants or plant parts having reduced expression, stability, and / or activity of undetermined gametophyte (ig) genes, mRNA, or proteins for screening or identifying mutations in centromere or kinetochore proteins, preferably CENH3, that confer or enhance singular-inducing activity or its ability.