Methods of hybrid crop breeding
The GSE3 gene is used to create male-sterile lines with smaller grains that can be mechanically separated from restorer lines, enhancing hybrid seed production efficiency and enabling full mechanization by increasing F1 hybrid seed yield.
Patent Information
- Application Number
- PCT/EP2025/060634
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-17
- Publication Date
- 2025-10-23
AI Technical Summary
Current hybrid seed production in plants is labor-intensive and time-consuming, limiting the mechanization of breeding processes due to the need for manual separation of restorer lines from male-sterile lines based on grain size, leading to significant seed waste and reduced efficiency.
Utilizing the GSE3 gene, which encodes a histone acetyltransferase, to create male-sterile lines with reduced grain size that can be mechanically separated from restorer lines, combined with genome editing technologies to enhance hybrid seed production efficiency.
The method increases F1 hybrid seed production by 21.2%-38.3% and enables fully mechanized hybrid seed production without affecting yield, using genetically altered plants with reduced or abolished GSE3 activity.
Smart Images

Figure EP2025060634_23102025_PF_FP_ABST
Abstract
Description
[0001]M&C PC933520WOA 1 METHODS OF HYBRID CROP BREEDING FIELD OF THE INVENTION 5 The present invention relates to methods of improving the production of hybrid seeds comprising reducing or abolishing the expression or activity of GRAIN SIZE ON CHROMOSOME 3 (GSE3) in a male-sterile plant, which in turn reduces seed size andincreases the number of hybrid seeds. Also covered are male-sterile plants characterisedby reduced or abolished expression or activity of GSE3, as well as methods of using the10 plants of the invention in mechanised hybrid seed sorting. BACKGROUND OF THE INVENTION Crop hybrid technologies have contributed to significant yield improvement worldwide. 15 For example, rice yield has experienced 20%-30% improvement due to the utilization of hybrid rice in the past decades thereby contributing greatly to food security. Three-linesystems (the restorer line, the maintainer line, and cytoplasmic male-sterile [CMS]“female” line) and two-line system (the restorer line and thermosensitive-or photoperiod-sensitive genetic male-sterile [T / PGMS] female line) are widely used for hybrid rice20 breeding. However, currently, F1 hybrid seed production in many plant species requiresseveral hand-labour-intensive and time-consuming steps, which prevent the fullmechanization of hybrid breeding and seed production. In particular, the restorer lineneeds to be grown in very close proximity to the male-sterile line in the alternative row to allow transmission of enough pollen for maximum fertilization of the MS / female line. To25 avoid seed contamination of hybrid seed with seed of the restorer line, manual labour isnecessary to remove the restorer line before harvesting hybrid seeds. Theconsequences of this step are huge, in China alone more than 150,000 tons of rice seedsare wasted every year because of the requirement to pre-remove the restorer lines.Thus, there is an urgent and unmet need to improve hybrid seed production, and30 particularly for methods that allow mechanised hybrid seed production. SUMMARY OF THE INVENTION An ideal male-sterile line for mechanised hybrid seed production should form small35 grains that can be mechanically separated from the grains of the restorer line, but that M&C PC933520WOA 2 do not show negative effects on F1 hybrid seed number and hybrid rice yield in field trials. Importantly, the small-grain phenotype should be determined by a recessive allele, and this recessive allele should act maternally to influence grain size. For these reasons, finding an ideal grain-size gene for breeding ideal small-grain male-sterile lines is very 5difficult, thereby limiting the practical application of a strategy to sort hybrid and non-hybrid seed based on grain size. Several pathways that control grain size have been reported; these include MAPK signalling pathways, G protein signalling pathways, phytohormone signalling pathways, ubiquitin-related pathways, and transcriptional regulatory pathways. Although mutations in several genes have been described to 10 reduce seed size, they are associated with negative effects on other important agronomic traits. Here we discover a new grain-size gene – referred to herein as GSE3 (GRAIN SIZE ONCHROMOSOME 3), that can be used to significantly improve elite male-sterile lines15 using both traditional breeding and genome editing technologies and achieve full mechanization of hybrid seed production in both three-line and two-line systems. Importantly, the number of F1 hybrid seeds, a key economic determinant of commercialhybrid seed production, produced by male-sterile lines with reduced gse3 activity or expression was increased by 21.2%-38.3% in field trials. The major QTL gene GSE3 20 encodes a histone acetyltransferase that binds histones and influences histone acetylation levels. GSE3 is also recruited by the transcription factor GS2 to the promoters of their co-regulated genes involved in seed size control and influences the histone acetylation status of their co-regulated genes. Our findings demonstrate that genome editing or use of existing genetic variation of GSE3 enables fully-mechanized hybrid seed 25 production, dramatically enhances hybrid seed number, and can be used to immediately improve current elite male-sterile lines of hybrid rice for fully mechanized hybrid rice breeding. In one aspect of the invention there is provided a genetically altered plant, part thereof 30 or plant cell characterised by reduced or abolished activity or expression of the GRAIN SIZE ON CHROMOSOME 3 (GSE3) protein in said plant, part thereof or plant cell, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or homolog thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1. 35 M&C PC933520WOA 3 The homolog may be selected from a sequence comprising SEQ ID NO: 16, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41 and 43 or a functional variant thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 16, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41 or 43. 5 In an embodiment, the plant comprises at least one mutation in at least one GSE3 gene, wherein the mutation is preferably a loss or partial loss of function mutation, and wherein the GSE3 gene comprises a sequence as defined in SEQ ID NO: 3 to 4 or a functional variant or homolog thereof, wherein the functional variant has at least 50% overall 10 sequence identity to SEQ ID NO: 3 to 4, and wherein the loss or partial loss of function mutation reduces or abolishes the activity or expression of GSE3. In another embodiment, the plant, part thereof or plant cell comprises at least one RNAi that reduces or abolishes the expression of GSE3. 15 In another aspect of the invention, there is provided a genetically altered plant, partthereof or plant cell characterised by reduced or abolished activity or expression of a GRF4 (GROWTH REGULATING 4) protein and reduced or abolished expression of a GRF3 (GROWTH REGULATING 4) protein in said plant, part thereof or plant cell, 20 wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to25 SEQ ID NO: 14. The homolog may be selected from a sequence comprising SEQ ID NO: 46, 48, 50, 52,54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 87, 88, 89, 91, 93, 95, 97, 99 and 101or a functional variant thereof, wherein the functional variant has at least 30 50% overall sequence identity to SEQ ID NO: 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 87, 88, 89, 91, 93, 95, 97, 99 or 101. In one embodiment, the plant comprises at least one mutation in at least one gene encoding GRF4 and at least one mutation in at least one gene encoding GRF3, wherein 35 the mutation in GRF4 and the mutation in GRF3 is a loss or partial loss of function M&C PC933520WOA 4 mutation, and wherein the gene encoding GRF4 comprises a sequence as defined in SEQ ID NO: 6, 7, 8 or 11 or a functional variant or homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 6, 7, 8 or 11, and wherein the gene encoding GRF3 comprises a sequence as defined in SEQ ID 5 NO: 12 or 13 or a functional variant or homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 12 or 13, and wherein the loss or partial loss of function mutations reduce or abolishes the activity or expression of GRF4 and GRF3. 10 In another embodiment, the plant, part thereof or plant cell comprises at least one RNAi that reduces or abolishes the expression of a gene encoding GRF4 and at least one RNAi that reduces or abolishes the expression of a gene encoding GRF3. The plant may be a male-sterile plant, preferably a cytoplasmic male-sterile plant or an15 environmental genic sterility male (EGMS) plant. In particular, the EGMS plant may beselected from a thermo-sensitive genic male-sterility (TGMS) plant or photoperiod- sensitive genic male-sterility (PGMS) plant. Alternatively, the plant is a maintainer plantline.20 The genetically altered plant, part thereof or plant cell may be a crop plant. FDorexample, the plant may be selected from soybean, rice, wheat, maize, sorghum, soybean, rapeseed, cotton, sunflower, millet, barley, sugar-beet, rye, oats, ryegrass, beans, Brassica, sorghum and pea.25 In particular, the plant may be rice and may be selected from the Xiaoligeng (XLG) orY58S rice line. Alternatively, the plant may be from the Tianfeng B (TFB) rice line. In one embodiment, the plant is characterised by a reduced grain size relative to the grain size in a control or wild-type plant. 30 In one embodiment, the plant part is a grain or seed. Alternatively, the plant part is pollen.In another aspect of the invention, there is provided a grain or seed, wherein said grain or seed is characterised by reduced or abolished activity or expression of a GRAIN SIZE M&C PC933520WOA 5 ON CHROMOSOME 3 (GSE3) protein in said grain or seed, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1. 5 In another aspect of the invention, there is provided a grain or seed, wherein said grain or seed is characterised by reduced or abolished activity or expression of GRF4 and GRF3 protein in said grain or seed, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional variant or fragment thereof, and wherein the 10 functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 14.15 In another aspect of the invention, there is provided a method of hybrid breeding, themethod comprising a. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof or20 plant cell, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1; or b. reducing or abolishing the activity or expression of GRF4 and GRF325 protein in a male-sterile plant, part thereof or plant cell, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functionalvariant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional30 variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 14. The method may further comprise increasing the expression or activity of a proteinencoded by the GS2 gene in a fertility restorer plant, plant part thereof or plant cell, wherein the GS2 gene comprises a sequence as defined in SEQ ID NO: 6, 7, 8 or 11 or M&C PC933520WOA 6 a functional variant or homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 6, 7, 8 or 11. In one embodiment, the method comprises 5 a. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof orplant cell, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional10 variant has at least 50% overall sequence identity to SEQ ID NO: 1; or b. reducing or abolishing the activity or expression of GRF4 and GRF3protein in a male-sterile plant, part thereof or plant cell, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional15 variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functionalvariant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 14; and 20 c. increasing the expression or activity of a protein encoded by the GS2 genein a fertility restorer plant, plant part thereof or plant cell, wherein the GS2 gene comprises a sequence as defined in SEQ ID NO: 6, 7, 8, or 11 or a functional variant or homolog thereof, wherein the functional variant has at25 least 50% overall sequence identity to SEQ ID NO: 6, 7, 8, or 11; and d. hybridising the plant of steps a or b with the plant of step c to obtain an F1hybrid plant30 The male-sterile plant may be a cytoplasmic male-sterile plant or an environmental genicsterility male (EGMS) plant, preferably wherein the EGMS plant is selected from a thermo-sensitive genic male-sterility (TGMS) plant or photoperiod-sensitive genic male- sterility (PGMS) plant. M&C PC933520WOA 7 In another aspect of the invention, there is provided a method of hybrid breeding, the method comprising hybridising the genetically altered plants of the invention with a genetically altered fertility restorer plant, a part thereof or plant cell, wherein the restorer line is characterised by increased activity or expression of the protein encoded by the 5 GS2 (GRAIN SIZE ON CHROMOSOME 2) gene, and obtaining F1 hybrid plants. The methods of the invention may further comprise harvesting seeds obtained orobtainable from the male-sterile plant simultaneously with seeds obtained or obtainable from the fertility restorer plant, wherein preferably harvesting is mechanized harvesting. 10 In another aspect of the invention, there is provided a F1 hybrid plant obtained orobtainable by the methods of the invention. In another aspect of the invention, there is provided a method of obtaining the genetically15 altered plant cell of the invention, the method comprising introducing at least one mutation in at least one GSE3 gene, preferably all copies of the GSE3 gene, wherein the at least one mutation in each copy of GSE3 is a loss-of-function or partial loss of function mutation. Alternatively, the method comprises introducing at least one mutation in at least one GRF4 and GRF3 gene, preferably all copies of the GRF4 and GRF3 gene, 20 wherein the at least one mutation in each copy of GRF4 and GRF3 is a loss-of-function or partial loss of function mutation. In one embodiment, the seed obtained or obtainable from a pollinated male-sterile plant is characterised by a reduced grain thickness and / or grain length relative to a fertility 25 restorer plant. Preferably, the reduced grain thickness and / or grain length is a result of reducing or abolishing the expression and / or activity of GSE3 and / or GRF3 and / or GRF3, for example, by introducing at least one mutation into the GSE3 gene and / or at least one mutation into GRF4 and GRF3 gene.30 In another aspect of the invention, there is provided a method of mechanized hybrid seedproduction, the method comprising a. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof or plant cell,35 wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional M&C PC933520WOA 8 variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1; or b. reducing or abolishing the activity or expression of GRF4 and GRF3 protein in a5 male-sterile plant, part thereof or plant cell, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional variant or fragment thereof,and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional variant or fragment thereof, and wherein the10 functional variant has at least 50% overall sequence identity to SEQ ID NO: 14. The method may further comprise increasing the expression or activity of a proteinencoded by the GS2 gene in a fertility restorer plant, plant part thereof or plant cell, wherein the GS2 gene comprises a sequence as defined in SEQ ID NO: 6, 7, 8 or 11 or15 a functional variant or homolog thereof, wherein the functional variant has at least 50%overall sequence identity to SEQ ID NO: 6, 7, 8 or 11. In one embodiment, the method comprises20 a. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof or plantcell, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1; or 25 b. reducing or abolishing the activity or expression of GRF4 and GRF3 protein ina male-sterile plant, part thereof or plant cell, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional variant orfragment thereof, and wherein the functional variant has at least 50% overall 30 sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises asequence as defined in SEQ ID NO: 14 or a functional variant or fragmentthereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 14; and M&C PC933520WOA 9 c. increasing the expression or activity of a protein encoded by the GS2 gene ina fertility restorer plant, plant part thereof or plant cell, wherein the GS2 gene comprises a sequence as defined in SEQ ID NO: 6, 7, 8 or 11 or a functional variant or homolog thereof, wherein the functional variant has at least 50% 5 overall sequence identity to SEQ ID NO: 6, 7, 8 or 11; and d. hybridising the plant of steps a or b with the plant of step c to obtain an hybridseed. 10 In another aspect of the invention, there is provided hybrid seed obtained or obtainable by the method of the invention. In another aspect of the invention, there is provided a method of hybrid breeding, themethod comprising identifying and selecting plant that has a small-grain phenotype, the 15 method comprising detecting in the plant or plant germplasm at least one polymorphism or mutation in the GRAIN SIZE ON CHROMOSOME 3 (GSE3) gene and / or promoter and selecting said plant or progeny thereof, wherein the polymorphism or mutation reduces the expression and / or activity of GRAIN SIZE ON CHROMOSOME 3 (GSE3).20 In another aspect of the invention, there is provided a method of increasing grain size ina plant, part thereof or plant cell, the method comprising increasing the expression and / or activity of GRAIN SIZE ON CHROMOSOME 3 (GSE3) in said plant, part thereof or plant cell. 25 DESCRIPTION OF THE FIGURES The invention is further described in the following non-limiting figures: Figure 1 shows that the ideal male sterile line (XQA) with small grains and the restorer30 line (DHZ) with large grains allow for fully mechanized hybrid rice breeding.Figure 2 shows that GSE3 encodes a N-acetyltransferase-like protein that influenceshistone acetylation levels to contribute to grain size control. M&C PC933520WOA 10 Figure 3 shows that GSE3 interacts physically and genetically with GS2 to control grainsize and weight. Figure 4 shows that GS2 recruits GSE3 to influence the histone status of target genes5 including XIAO-P1, XIAO-P2, GW6-P1, GW6-P2, BZR1-P1 and BZR1-P2. Figure 5 shows the rapid generation of male sterile lines for mechanized production ofhybrid seeds using genome editing technology in three-line and two-line systems.10 Figure 6 shows a strategy to rapidly improve current parental lines of hybrid rice for fullymechanized hybrid seed production Figure 7 shows that loss-of-function mutations of GSE3 lead to small grainsa) The target site in GSE3. Red letters represent the target sequence and blue letters15 represent the PAM sequence. The knockout lines gse3-cri1 and gse3-cri2 are shown. b) Validation of the knockout line gse3-cri1 (single base T insertion in the second exonof GSE3) by Sanger sequencing. C) Validation of the knockout line gse3-cri2 (30 bpdeletion in the second exon of GSE3) by Sanger sequencing. n ≥ 40 (e-g) and n = 3 (h)i-j, The tilling number (i) and grain number per plant (j) of ZH11, gse3-cri1 and gse3-cri2. 20 n ≥ 12 (i-j). P < 0.01 compared with parental line (ZH11) by Student’s t-test. Bars: 3 mm (d). Figure 8 shows the phenotypes of ideal m238. j, A single base substitution on the firstexon of the GSE3 gene results in an amino acid change (Ser / Pro) in the m238. The red25 arrow indicates the location of the single base replacement. **P < 0.01 compared with ZH11 by Student’s t-test. Bars: 3 mm (a), 10 cm (b) and 5 cm (c). Figure 9 shows a phylogenetic analysis of GSE3. The protein sequences of GSE3 andits homologs were aligned using Clustal W. A phylogenetic tree was constructed using the neighbor-joining method in MEGA7.0 software with 1000 bootstrap replicates. 30 Figure 10 shows the alignment of the GSE3 homology protein in plants. Proteinsequences of GSE3 and its homologs in plants were aligned using Clustal W. The red box and red arrow indicate the conserved amino acid (Ser) at position 68 of GSE3. M&C PC933520WOA 11 Figure 11 shows the overexpression of GSE3 forms large grains. a, Mature rice grainsof ZH11 and pActin:GSE3 transgenic lines. b, Relative expression level of GSE3 in ZH11and pActin:GSE3 transgenic lines. RT–qPCR was used to measure expression levels ofGSE3 and normalized to ACTIN1. n ≥ 40 (c-e) and n = 3 (f). **P < 0.01, *P < 0.055 compared with ZH11 by Student’s t-test. Bars:3 mm (a). Figure 12 shows the phenotype of pActin:GSE3 transgenic plants Values in (c-h) aregiven as mean ± SD. **P < 0.01, *P < 0.05 compared with parental line (ZH11) by Student’s t-test. Bars:10 cm (a), 5 cm (b). 10 Figure 13 shows that a simultaneous knockout of GRF3 and GRF4 forms small grains.a, Sanger sequencing confirmed that the knockout lines grf3-cri1 grf4-cri2 and grf3-cri1grf4-cri3 have a single T base insertion in the third exon of GRF3. The sequencinganalysis also revealed a single T base deletion in the first exon of the GRF4 gene in the15 grf3-cri1 grf4-cri2 line, as well as a single C base deletion in the first exon of the GRF4gene in the grf3-cri1 grf4-cri3 line. The gene structure of GRF3 and GS2 / GRF4 and thetarget sites (red arrow) are shown. n ≥ 50 (c-d), n = 3 (e), n ≥ 10 (f-g). Values in (c-g)are given as mean ± SD. **P < 0.01 compared with parental line (ZH11) by Student’s t-test. Bars: 3 mm (b). 20 Figure 14 shows the genome editing of GSE3 in TFA and TFB background usingCRISPR / Cas9 system. The knockout line TFBgse3-cri3 and TFAgse3-cri3 were validated to have a single base T insertion in the second exon of GSE3 through Sanger sequencing. The gene structure of GSE3 and the target sites (red arrow) are shown. 25 Figure 15 shows the phenotypes of TFB and TFBgse3-cri3. n ≥ 40 (d-f), n =3 (g), n ≥ 10(h-i) n ≥ 10 (j>m), n ≥ 20 (n). Values in (d-n) are given as mean ± SD. **P < 0.01 comparedwith TFB by Student’s t-test. ns, not significant (Student's unpaired t-test). Bars: 10 cm (a), 5 cm (b), and 3 mm (c) 30 Figure 16 shows the grain phenotypes of TFA, TFAgse3-cri3, HZ, and DHZ. a-c, Thegrain length (a), grain width (b), and 1,000 grain weight (c) of parent lines TFA (pollinated with HZ), TFAgse3-cri3(pollinated with DHZ), HZ, and DHZ. n ≥ 40 (a-b) and n =3 (c). Values in (a-c) are given as mean ± SD. **P < 0.01 compared with TFA (c) and HZ (c)35 by Student’s t-test. M&C PC933520WOA 12 Figure 17 shows the grain yield per plant of hybrid rice TFA × HZ (TYHZ) and TFAgse3-cri3 × HZ. The grain yield per plant of hybrid rice TFA × HZ (TYHZ) and TFAgse3-cri3 × HZ,and the plant density was 20 cm × 20 cm. 5 Figure 18 shows genome editing of GSE3 in Y58S background using CRISPR / Cas9technology. a, Target sites in the second exon of the GSE3 gene. Red arrows indicate the location. b, The knockout line Y58Sgse3-cri4 was validated to have a single base Ainsertion in the second exon of GSE3 through Sanger sequencing. 10 Figure 19 shows the grain phenotypes of Y58S, Y58Sgse3-cri4, and R900. a-c, The grainlength (a), grain width (b), and 1,000 grain weight (c) of parent lines Y58S, Y58Sgse3-cri4,and R900. n ≥ 40 (a-b) and n =3 (c) Values in (a-c) are given as mean ± SD. **P < 0.01compared with Y58S (c) and R900 (c) by Student’s t-test. 15 Figure 20 shows the phenotypes of Y58S and Y58Sgse3-cri4. n ≥ 10 (c-d), n ≥ 10 (e-h), n≥ 20 (i) Values in (c-i) are given as mean ± SD. **P < 0.01, *P < 0.05 compared withY58S by Student’s t-test. ns, not significant (Student‘s unpaired t-test). Bars: 10 cm (a), 5 cm (b). 20 Figure 21 shows the grain phenotypes of Y58S (pollinated with R900) and Y58Sgse3-cri4 (pollinated with R900). b-e: n ≥ 10. Values in (b-e) are given as mean ± SD. **P <0.01, compared with Y58S (pollinated with R900) or Y58Sgse3-cri4 (pollinated with R900) by Student’s t-test. Bars: 3 mm (a) 25 DETAILED DESCRIPTION OF THE INVENTIONIn the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly 30 indicated to the contrary. In particular, any feature indicated as being preferred or advantageous may be combined with any other feature or features indicated as being preferred or advantageous. The practice of the present invention will employ, unless otherwise indicated, 35 conventional techniques of botany, microbiology, tissue culture, molecular biology, M&C PC933520WOA 13 chemistry, biochemistry and recombinant DNA technology, bioinformatics which are within the skill of the art. Such techniques are explained fully in the literature. The following features apply to all aspects and embodiments of the invention.5 As used herein, the words "nucleic acid", "nucleic acid sequence", "nucleotide", "nucleic acid molecule" or "polynucleotide" are intended to include DNA molecules (e.g., cDNA or genomic DNA), RNA molecules (e.g., mRNA), natural occurring, mutated, synthetic DNA or RNA molecules, and analogs of the DNA or RNA generated using nucleotide 10 analogs. It can be single-stranded or double-stranded. Such nucleic acids or polynucleotides include, but are not limited to, coding sequences of structural genes, anti-sense sequences, and non-coding regulatory sequences that do not encode mRNAs or protein products. These terms also encompass a gene. The term "gene" or “gene sequence” is used broadly to refer to a DNA nucleic acid associated with a biological 15 function. Thus, genes may include introns and exons as in the genomic sequence, or may comprise only a coding sequence as in cDNAs, and / or may include cDNAs in combination with regulatory sequences. The terms "polypeptide" and "protein" are used interchangeably herein and refer to20 amino acids in a polymeric form of any length, linked together by peptide bonds. The aspects of the invention involve recombinant DNA technology and exclude embodiments that are solely based on generating plants by traditional breeding methods. 25 The term "plant" as used herein encompasses whole plants, ancestors and progeny of the plants and plant parts, including seeds, fruit, shoots, stems, leaves, roots (including tubers), flowers, tissues and organs, wherein each of the aforementioned comprise at least one genetic modification as described herein (e.g. in GSE3 and / or GRF3 and / or GRF4). The term "plant" also encompasses plant cells, suspension cultures, callus 30 tissue, embryos, meristematic regions, gametophytes, sporophytes, pollen and microspores, again wherein each of the aforementioned comprises at least one genetic modification as described herein (e.g. in GSE3 and / or GRF3 and / or GRF4). The invention also extends to harvestable parts of a plant of the invention as described herein, but not limited to seeds, leaves, fruits, flowers, stems, roots, rhizomes, tubers 35 and bulbs. M&C PC933520WOA 14 In a preferred embodiment, the plant part is a grain or a seed. Such terms may be used interchangeably herein. In another preferred embodiment, the plant part is pollen. 5In any aspects of the invention described herein, the plant may be a crop plant; forexample, rice, wheat, maize, sorghum, soybean, rapeseed, cotton, sunflower, millet,barley, sugar-beet, rye, oats, ryegrass, beans and pea.A control plant as used herein according to all of the aspects of the invention is a plant10 which has not been modified according to the methods of the invention. Accordingly, inone embodiment, the control plant does not have altered expression of a GSE3,GS2 / GRF4 or GRF3 nucleic acid and / or altered activity of a GSE3, GS2 / GRF4 or GRF3polypeptide, as described herein. In an alternative embodiment, the plant has been genetically modified, as described above. In one embodiment, the control plant is a wild- 15 type plant. The control plant is typically of the same plant species, preferably having the same genetic background as the modified plant. In a traditional hybrid seed production process, the male-sterile line needs to be grown near the restorer line in the alternative row for pollination (Fig.1a). Before large-scale20 mechanized harvesting of hybrid seeds, the restorer (i.e. male / pollinator) line needs tobe removed from the field by hand labour. These are limiting steps for fully mechanized hybrid rice breeding. If hybrid seeds can be mechanically separated from the seeds of the restorer line on the basis of the difference in their seed sizes (as shown in Fig. 1a),seeds of male-sterile and restorer lines can be mixed in a certain ratio for mechanized25 planting and harvesting in the field, which enables fully mechanized hybrid rice breeding. As male-sterile lines produce the hybrid seeds, the key is to discover ideal grain-size genes / alleles for breeding ideal small-grain male-sterile lines with the recessive small- grain alleles, which should form small grains without penalties in F1 hybrid seed number and hybrid rice yield in field trials. 30 We have discovered an ideal grain size gene – referred to herein as GSE3 formechanized hybrid seed production. We bred an ideal small-grain male-sterile line XQA by introgressing the GSE3small grainallele into an elite male-sterile line TFA. We also crossed Kuangsijiadi (an indica variety with large grains) with the elite restorer line HZ35 and produced a new large-grain restorer line DHZ. A new hybrid rice variety XQHZ from M&C PC933520WOA 15 the XQA × DHZ combination was achieved. A sifter with a suitable sieve aperture can successfully separate the small F1 hybrid seeds from the mixed harvest (as shown in Fig. 1l). It is worth noting that the quantity of F1 hybrid seed of the XQA × DHZ combination was increased by 22.2% compared with that of the original TFA × HZ 5 combination in field test (Fig. 1m), thereby dramatically increasing F1 hybrid seed production efficiency. Meanwhile, the grain yield of XQHZ plants was comparable with that of the hybrid rice TYHZ (Fig.1p-q). These field trials support that the utilization of the GSE3small grainallele can successfully improve the elite male-sterile line TFA for mechanized production of F1 hybrid seeds, thereby enabling fully-mechanized hybrid10 rice breeding. Critically, this small-grain phenotype was achieved without penalties in F1hybrid seed number and hybrid rice yield in field trials. In fact, this small-grain phenotypeis shown to increase F1 hybrid seed number (Figure 1m, Figure 5d). In one aspect of the invention, there is provided a genetically altered plant, part thereof 15 or plant cell characterised by reduced or abolished activity or expression of the GRAIN SIZE ON CHROMOSOME 3 (GSE3) protein in said plant, part thereof or plant cell,wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variantor fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1. 20 GSE3 encodes a nuclear-localized N-acetyltransferase-like protein that influences histone acetylation levels. GSE3 encodes a N-acetyltransferase-like protein with the conserved GNAT motif (Fig.2o).25 As used herein, the term “reducing” or “reduced” expression or activity means a decreasein the levels of GSE3 expression and / or activity by up to or more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% when compared to the level in a wild-type or controlplant. In one embodiment, reducing means a decrease in at least 50% compared to the level in a wild-type or control plant. Reducing also may or may not encompass abolishing 30 expression. The term “abolish” expression means that no expression of GSE3 is detectable (no transcript) or that no functional GSE3 polypeptide is produced. In one embodiment, GSE3 comprises a sequence as defined in SEQ ID NO: 1 or afunctional variant or homolog thereof. In another embodiment, the GSE3 nucleic acid M&C PC933520WOA 16 sequence comprises or consists of a sequence as defined in SEQ ID NO: 4 or afunctional variant or homolog thereof. Afunctional variant may also encompass a ‘fragment’ of the GSE3 amino or nucleic acid5 sequence. By “fragment” is meant a portion of a whole length variant. A fragment may comprise a certain percentage of a full length variant, for example 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90% or 95%, or may refer to specific domains, such as the GNAT domain, of a whole length variant.10 The term “functional variant” (or “variant”) as used herein with reference to any of the sequences described herein refers to a variant sequence or part of the sequence which retains the biological function of the full non-variant sequence. Accordingly, a functional variant of GSE3 will retain the ability to co-regulate target genes with GS2 and GRF3, as described herein. In one embodiment, the functional variant will retain the ability to co-15 regulate the expression of a target gene selected from XIAO, OsBZR1, RGG2, OsER1, OsGRF1, GL6, SRS5, DEP1, PGL1, OsGA20ox1, OsMADS5, and OsMADS15. A functional variant may be able to mediate histone acetylation, through activity as a histone acetyltransferase.20 It is possible to determine the functional activity of a variant of GSE3 by measuring the transcript abundance of the genes that are under GSE3 regulation, for example using real-time PCR or western blots. Alternatively, it is possible to analyse the histone acetyltransferase function of a functional variant by performing CHIP-qPCR analysis to identify the relative enrichment of histone H4Ac (pan-acetyl) at the promoter regions of 25 genes. Example genes that could be studied to analyse GSE3 activity include OsXIAO, GW6 and OsBRZR1 or a homolog thereof. This approach is shown, for example, at Figures 4f-i.A functional variant also comprises a variant of the gene of interest which has sequence 30 alterations that do not affect function, for example in non-conserved residues. Also encompassed is a variant that is substantially identical, i.e. has only some sequence variations, for example in non-conserved residues, compared to the wild type sequences as shown herein and is biologically active. Alterations in a nucleic acid sequence which result in the production of a different amino acid at a given site that do not affect the 35 functional properties of the encoded polypeptide are well known in the art. For example, M&C PC933520WOA 17 a codon for the amino acid alanine, a hydrophobic amino acid, may be substituted by a codon encoding another less hydrophobic residue, such as glycine, or a more hydrophobic residue, such as valine, leucine, or isoleucine. Similarly, changes which result in substitution of one negatively charged residue for another, such as aspartic acid 5 for glutamic acid, or one positively charged residue for another, such as lysine for arginine, can also be expected to produce a functionally equivalent product. Each of the proposed modifications is well within the routine skill in the art, as is determination of retention of biological activity of the encoded products. 10 In one embodiment, a functional variant has at least 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 15 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% overall sequence identity to the non-variant nucleic acid or amino acid sequence. As shown in Figure 10, homologsof GSE3 are found in all plants, including soybean, wheat, maize, sorghum, tomato, potato and brassicas. 20 Accordingly, in one embodiment, the homolog of GSE3 is selected from a sequence comprising SEQ ID NO: 16, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41 or 43 or a functional variant thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 16, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41 or43.25 In one embodiment, the promoter of the homolog of GSE3 is selected from a sequencecomprising SEQ ID NO: 15, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42 or 44 or afunctional variant thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 15, 18, 20, 22, 24, 26, 28, 30,32, 34, 36, 38, 40, 42 or 44. 30 The term homolog, as used herein, also designates a GSE3, GFR3 or GS2 / GFR4 geneorthologue from another plant, preferably a crop plant. A homolog may have, in increasing order of preference, at least 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 35 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, M&C PC933520WOA 18 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or at least 99% overall sequence identity to the amino acid represented by any of SEQ ID NO: 1, 10, or 14 to the nucleic acid sequences as5 shown by SEQ ID NOs: 4, 11 or 12, respectively. Functional variants of GSE3 (and also GRF3 and GS2 / GRF4) homologs as defined above are also within the scope of the invention. 10 Two nucleic acid sequences or polypeptides are said to be "identical" if the sequence of nucleotides or amino acid residues, respectively, in the two sequences is the same when aligned for maximum correspondence as described below. The terms "identical" or percent "identity," in the context of two or more nucleic acids or polypeptide sequences, refer to two or more sequences or subsequences that are the same or have a specified 15 percentage of amino acid residues or nucleotides that are the same, when compared and aligned for maximum correspondence over a comparison window, as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. When percentage of sequence identity is used in reference to proteins or peptides, it is recognised that residue positions that are not identical often differ by 20 conservative amino acid substitutions, where amino acids residues are substituted for other amino acid residues with similar chemical properties (e.g., charge or hydrophobicity) and therefore do not change the functional properties of the molecule. Where sequences differ in conservative substitutions, the percent sequence identity may be adjusted upwards to correct for the conservative nature of the substitution. Means for 25 making this adjustment are well known to those of skill in the art. For sequence comparison, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. 30 Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters. Non-limiting examples of algorithms that are suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.035 algorithms. M&C PC933520WOA 19 We have identified exemplar homologs for GSE3 using the NBCI BLAST program. Input data: Program BLASTP, Database = All non-redundant GenBank CDS, SEQ ID NO: 1. Our search identified the following sequences to be homologs of GSE3: Table 1: Sequence of homologs of GSE3, GS2 and GRF3 5 Thus, the nucleotide sequences of the invention and described herein can also be used to isolate corresponding sequences from other organisms, particularly other plants, for example crop plants. In this manner, methods such as PCR, hybridization, and the like 10 can be used to identify such sequences based on their sequence homology to the sequences described herein. Topology of the sequences and the characteristic domains structure can also be considered when identifying and isolating homologs. Sequences may be isolated based on their sequence identity to the entire sequence or to fragments thereof. In hybridization techniques, all or part of a known nucleotide sequence is used 15 as a probe that selectively hybridizes to other corresponding nucleotide sequences present in a population of cloned genomic DNA fragments or cDNA fragments (i.e., genomic or cDNA libraries) from a chosen plant. The hybridization probes may be M&C PC933520WOA 20 genomic DNA fragments, cDNA fragments, RNA fragments, or other oligonucleotides, and may be labelled with a detectable group, or any other detectable marker. Methods for preparation of probes for hybridization and for construction of cDNA and genomic libraries are generally known in the art and are disclosed in Sambrook, et al., (1989) 5 Molecular Cloning: A Library Manual (2d ed., Cold Spring Harbor Laboratory Press, Plainview, New York). Hybridization of such sequences may be carried out under stringent conditions. By "stringent conditions" or "stringent hybridization conditions" is intended conditions under10 which a probe will hybridize to its target sequence to a detectably greater degree than toother sequences (e.g., at least 2-fold over background). Stringent conditions are sequence dependent and will be different in different circumstances. By controlling the stringency of the hybridization and / or washing conditions, target sequences that are 100% complementary to the probe can be identified (homologous probing). Alternatively, 15 stringency conditions can be adjusted to allow some mismatching in sequences so that lower degrees of similarity are detected (heterologous probing). Generally, a probe is less than about 1000 nucleotides in length, preferably less than 500 nucleotides in length. Typically, stringent conditions will be those in which the salt concentration is lessthan about 1.5 M Na ion, typically about 0.01 to 1.0 M Na ion concentration (or other 20 salts) at pH 7.0 to 8.3 and the temperature is at least about 30°C for short probes (e.g., 10 to 50 nucleotides) and at least about 60°C for long probes (e.g., greater than 50 nucleotides). Duration of hybridization is generally less than about 24 hours, usually about 4 to 12. Stringent conditions may also be achieved with the addition of destabilizing agents such as formamide. 25 It will be most valuable to introduce GSE3small graininto current elite male-sterile lines for mechanized production of F1 hybrid seeds and develop new hybrid rice varieties, but it is a time-consuming process using traditional breeding approaches. We explored the genome editing technology to knockout the GSE3 gene in the male-sterile line TFA and 30 the maintainer line TFB to develop TFAgse3-cri3and TFBgse3-cri3. Small F1 hybrid seeds from mixed planting TFAgse3-cri3and DHZ were mechanically separated (Fig. 5c). F1 hybrid seed number from the TFAgse3-cri3× DHZ combination also dramatically increased by 21.2% compared with the original TFA × HZ combination, boosting hybrid seed production efficiency (Fig.5d). It is plausible to improve current elite male-sterile lines by M&C PC933520WOA 21 editing the GSE3 gene and their respective restorer lines by editing grain- thicknessgenes for fully mechanized hybrid rice breeding (Fig.6a). We further achieved mechanized F1 hybrid seed production in a famous two-line hybrid rice YLY900 only by editing the GSE3 gene in thermosensitive male-sterile line Y58S. 5The reason is that the difference between the restorer line R900 and Y58S grainthickness was relatively large, (Fig.5j). Importantly, the Y58Sgse3-cri4× R900 combination also dramatically increased F1 hybrid seed number by 38.3% compared with the original Y58S × R900 combination, therefore dramatically elevating hybrid seed production efficiency (Fig.5m). 10 Accordingly, in another embodiment, the plant comprises at least one mutation in at least one endogenous GSE3 gene or GSE3 promoter of the GSE3, wherein the mutation reduces or abolishes the expression or activity of GSE3.15 As used throughout, by “GSE3 promoter” is meant a region extending at least or approx.3256 bp upstream of the ATG codon of the GSE3 ORF. In one embodiment, thesequence of the GSE3 promoter comprises or consists of a nucleic acid sequence asdefined in any one of SEQ ID NO: 3 or a functional variant or homolog thereof. In oneembodiment, the GSE3 promoter may also include 5’ UTR sequences.20 In the above embodiments an ‘endogenous’ nucleic acid may refer to the native or natural sequence in the plant genome. Also included in the scope of this invention are functional variants (as defined herein) and homologs of the above identified sequences. 25 By “at least one mutation” is meant that where the target gene (and promoter) is present as more than one copy or homeolog (with the same or slightly different sequence) thereis at least one mutation in at least one gene and / or promoter. In a preferred embodiment,all genes / promoters are mutated – that is, all copies or homeologs of the gene compriseat least one mutation. 30 Preferably, at least one mutation is introduced into each GSE3 gene present in a plant,part thereof or plant cell. That is, all homeologs or copies of GSE3 comprise at least onemutation. M&C PC933520WOA 22 In one embodiment, the mutation may be introduced using genome editing, such as CRISPR. Alternatively, the at least one mutation may be introduced using targeted mutagenesis. 5 To achieve effective genome editing via introduction of site-specific DNA DSBs, four major classes of customisable DNA binding proteins can be used: meganucleases derived from microbial mobile genetic elements, ZF nucleases based on eukaryotic transcription factors, transcription activator-like effectors (TALEs) from Xanthomonas bacteria, and the RNA-guided DNA endonuclease Cas9 from the type II bacterial 10 adaptive immune system CRISPR (clustered regularly interspaced short palindromic repeats). Meganuclease, ZF, and TALE proteins all recognize specific DNA sequences through protein-DNA interactions. Although meganucleases integrate nuclease and DNA-binding domains, ZF and TALE proteins consist of individual modules targeting 3 or 1 nucleotides (nt) of DNA, respectively. ZFs and TALEs can be assembled in desired 15 combinations and attached to the nuclease domain of FokI to direct nucleolytic activity toward specific genomic loci. In a preferred embodiment, the genome editing method that can be used according to the various aspects of the invention is CRISPR. CRISPR technologies can be used to 20 introduce precise edits into a genome or nucleic acid. CRISPR technologies comprise three essential components: a guide RNA (gRNA), CRISPR and a CRISPR associatedendonuclease or enzyme. The guide RNA (single guide RNA, sgRNA) is a type of RNA molecule that binds to the endonuclease and specifies, based on the sequence of the gRNA, the location at which endonuclease will cut DNA. One major advantage of the 25 CRISPR system, as compared to conventional gene targeting and other programmable endonucleases is the ease of multiplexing, where multiple genes can be mutated simultaneously simply by using multiple sgRNAs each targeting a different gene. In addition, where two sgRNAs are used flanking a genomic region, the intervening section can be deleted or inverted. 30 Cas9 is thus the hallmark protein of the type II CRISPR-Cas system, and is a large monomeric DNA nuclease guided to a DNA target sequence adjacent to the PAM (protospacer adjacent motif) sequence motif by a complex of two noncoding RNAs: CRISPR RNA (crRNA) and trans-activating crRNA (tracrRNA). The Cas9 protein contains two nuclease domains homologous to RuvC and HNH nucleases. The HNH 35 nuclease domain cleaves the complementary DNA strand whereas the RuvC-like M&C PC933520WOA 23 domain cleaves the non-complementary strand and, as a result, a blunt cut is introduced in the target DNA. Heterologous expression of Cas9 together with an sgRNA can introduce site-specific double strand breaks (DSBs) into genomic DNA of live cells from various organisms. For applications in eukaryotic organisms, codon optimized versions 5 of Cas9, which is originally from the bacterium Streptococcus pyogenes, have been used. Alternatively, Cpf1, which is another Cas protein, can be used as the endonuclease. Cpf1 differs from Cas9 in several ways: Cpf1 requires a T-rich PAM sequence (TTTV) for target recognition, Cpf1 does not require a tracrRNA, (i.e. only a crRNA is required) and the Cpf1-cleavage site is located distal and downstream to the10 PAM sequence in the protospacer sequence (Li et al., 2017). Furthermore, after identification of the PAM motif, Cpf1 introduces a sticky-end-like DNA double-stranded break with several nucleotides of overhang. As such, the CRISPR / Cpf1 system consists of a Cpf1 enzyme and a crRNA. 15 The single guide RNA (sgRNA) is the second component of the CRISPR / Cas(Cpf) system that forms a complex with the Cas9 / Cpf1 nuclease. sgRNA is a synthetic RNAchimera created by fusing crRNA with tracrRNA. The sgRNA guide sequence located atits 5′ end confers DNA target specificity. Therefore, by modifying the guide sequence,it is possible to create sgRNAs with different target specificities. The canonical length of 20 the Cas9 guide sequence is 20 bp. Cas9 and Cpf1 expression plasmids for use in the methods of the invention can be constructed as described in the art. Indeed, this present invention used CRISPR / Cas9 technology to generate loss-of-function mutants of GSE3 in the male sterile line TFA, the 25 maintainer line TFB and the male sterile line Y58S. Cas9 or Cpf1 and the one or more sgRNA molecules may be delivered as separate or as single constructs. Where separate constructs are used for the delivery of the CRISPR enzyme (i.e. Cas9 or Cpf1) and the sgRNA molecule (s), the promoters used to drive 30 expression of the CRISPR enzyme / sgRNA molecule may be the same or different. In one embodiment, RNA polymerase (Pol) II-dependent promoters or the CaMV35S promoter can be used to drive expression of the CRISPR enzyme. In another embodiment, Pol III-dependent promoters, such as U6 or U3, can be used to drive expression of the sgRNA. 35 M&C PC933520WOA 24 Accordingly, using techniques known in the art it is possible to design sgRNA molecules that target a GSE3 gene or promoter sequence as described herein. In a furtherembodiment, it is also possible to design sgRNA molecules that alternatively or additionally target a GS2 and / or GRF3 gene or promoter sequence.5 In one embodiment, the method uses a CRISPR construct defined herein to introduce atargeted mutation into a GSE3 gene and / or promoter. Preferably, the CRISPR constructcomprises a sequence defined in SEQ ID NO: 106.10 In a further embodiment, the method uses a CRISPR construct to additionally introducea mutation into a GS2 and / or GRF3 gene and / or promoter.Thus, aspects of the invention involve targeted mutagenesis methods, specifically genome editing, and in a preferred embodiment exclude embodiments that are solely15 based on generating plants by traditional breeding methods. The genome editing constructs may be introduced into a plant cell using any suitable method known to the skilled person (the term “introduced” can be used interchangeably with “transformation”, which is described below). In an alternative embodiment, any of 20 the nucleic acid constructs described herein may be first transcribed to form a preassembled Cas9-sgRNA ribonucleoprotein and then delivered to at least one plant cell using any of the above described methods, such as lipofection, electroporation, bolistic bombardment or microinjection. 25 Specific protocols for using the above described CRISPR constructs would be well known to the skilled person. As one example, a suitable protocol is described in Ma & Liu (“CRISPR / Cas-based multiplex genome editing in monocot and dicot plants”) incorporated herein by reference. 30 The invention also extends to a plant obtained or obtainable by any method described herein. Preferably, the mutation is a loss of function or partial loss of function mutation, wherein preferably the mutation reduces or abolishes the histone acetylase activity of GSE3.35 Alternatively, the mutation reduces or abolishes the expression of GSE3. M&C PC933520WOA 25 Using routine methods, and also the methods disclosed below, the skilled person will readily be able to identify suitable loss or partial loss of function mutations. We have also identified a number of appropriate, and preferred mutations that achieve a loss or partial 5 loss of function in the GSE3 gene. In a preferred embodiment, the mutation that is introduced into the endogenous GSE3 gene or promoter thereof can be selected from the following mutation types10 1. a "missense mutation", which is a change in the nucleic acid sequence thatresults in the substitution of one amino acid for another amino acid; 2. a "nonsense mutation" or "STOP codon mutation", which is a change in thenucleic acid sequence that results in the introduction of a premature STOP codon and, thus, the termination of translation (resulting in a truncated 15 protein); in plants, the translation stop codons may be selected from "TGA" (UGA in RNA), "TAA" (UAA in RNA) and "TAG" (UAG in RNA); thus any nucleotide substitution, insertion, deletion which results in one of these codons to be in the mature mRNA being translated (in the reading frame) will terminate translation.20 3. an "insertion mutation" of one or more nucleotides or one or more aminoacids, due to one or more codons having been added in the coding sequence of the nucleic acid; 4. a "deletion mutation" of one or more nucleotides or of one or more aminoacids, due to one or more codons having been deleted in the coding 25 sequence of the nucleic acid; 5. a "frameshift mutation", resulting in the nucleic acid sequence beingtranslated in a different frame downstream of the mutation. A frameshift mutation can have various causes, such as the insertion, deletion or duplication of one or more nucleotides.30 6. a “splice site” mutation, which is a mutation that results in the insertion,deletion or substitution of a nucleotide at the site of splicing. Preferably the mutation is any mutation that reduces or abolishes the expression or activity of GSE3. 35 M&C PC933520WOA 26 In one embodiment, the mutation is a STOP codon mutation. In another embodiment, the at least one mutation in at least one GSE3 gene is a 10 bpdeletion in the third exon of the GSE3 gene, preferably at a deletion of up to 10 or more 5nucleotides at positions 960 to 970 of SEQ ID NO: 4 or corresponding positions in ahomologous sequence. A deletion of residues 960 to 970 of SEQ ID NO: 4 is referred toas DEL1 in the examples and is exemplified by SEQ ID NO: 110.In another embodiment the at least one mutation in a least one GSE gene is a single10 base insertion that results in a frameshift mutation. The single base insertion may be inthe second exon of GSE3 gene. More preferably, the mutation is a single base insertion of a ‘T’ at position 271 of SEQ ID NO: 4 (referred to as gse3-cri1 in examples, Figure 7).This gse3-cri1 mutation is exemplified in SEQ ID NO: 108. Alternatively, the mutation isa single A insertion at position 313 of SEQ ID NO: 4 (referred to as gse3-cri4 herein).15 This gse3-cri4 mutation is exemplified in SEQ ID NO: 112. In another embodiment, the at least one mutation is a deletion mutation of up to 30bp or more at positions 275 to 304 of SEQ ID NO: 1 (referred to as gse3-cri2 in examples,Figure 7). A deletion of residues 275 to 304 of SEQ ID NO: 4 is referred to as gse3-cri220 in the examples, and is exemplified by SEQ ID NO: 109. In another embodiment, the at least one mutation is in the first exon of the GSE3 gene, and preferably is a substitution mutation that results in a non-conservative mutation inSEQ ID NO: 1. More preferably, the mutation is a non-conservative mutation at position25 68 of SEQ ID NO: 1 or a corresponding position in a homologous sequence. Even morepreferably, the mutation is a serine to proline mutation (referred to as m238 allele, seeFigure 8). This S>P mutation (m238 allele) is exemplified in SEQ ID NO: 111. In apreferred embodiment, the mutation is a non-conservative mutation at position 202 ofSEQ ID NO: 4 or a corresponding position in a homologous sequence, as exemplified 30 by SEQ ID NO: 110. We have identified that this serine residue in GSE3, corresponding to position 68 of SEQ ID NO: 4, is conserved across the identified homologs. This serine residue may be identified in a homolog by alignment, as shown in Figure 10. 35 M&C PC933520WOA 27 Accordingly, in one embodiment, at least one mutation is introduced into a homolog of GSE3, preferably wherein the at least one mutation is a non-conservative mutation of aserine residue. That is, the serine residue may be substituted with a different residue with different chemical properties to the serine. 5 In a most preferred embodiment, the serine corresponds to position 202 of SEQ ID NO: 4. In another preferred embodiment, the at least one mutation is a non-conservative mutation of a serine residue within a RHSP or RNSP motif. As shown in Figure 11, thismotif is highly conserved around the conserved serine, and so will allow the skilled 10 person to easily identify a corresponding position. From the sequences of homologs in the sequence listing, the skilled person would routinely consider using a reverse translate tool – for example, the Expasy translate tool - to translate the DNA sequence into aprotein sequence. Using an alignment tool, such as EMBOSS Needle, the skilled person would be able to identify the conserved serine for mutations. 15 We have identified exemplar corresponding positions in select homologs, by way of example. In one embodiment, the corresponding position to position 202 of SEQ ID NO: 4 is selected from: 1. Position 57 of SEQ ID NO 116 (corresponding to homolog in Zea mays);20 2. Position 58 of SEQ ID NO: 117 (corresponding to homolog in Glycine Max);3. Position 48 of SEQ ID NO: 118 (corresponding to homolog in Brassica napus);4. Position 58 of SEQ ID NO: 119 (corresponding to homolog in Camelina sativa);5. Position 60 of SEQ ID NO: 120 (corresponding to homolog in Sorgum Bicolour)6. Position 99 of SEQ ID NO: 121 (corresponding to homolog in Brassica rapa); or25 7. Position 59 of SEQ ID NO: 122 (corresponding to homolog in Arabidopsisthaliana) By “corresponding or homologous position in a homologous sequence” is meant an equivalent position in a similar protein or nucleic acid sequence, due to common evolutionary origin or structural conservation. Homologous positions or as used herein 30 “corresponding positions in homologous sequences” can thus be determined by performing sequence alignments once the homologous sequence has been identified. For example, homologs can be identified using a BLAST search of the plant genome of interest. M&C PC933520WOA 28 Alternatively, the expression or activity of GSE3 may be reduced or abolished using a number of gene silencing methods known to the skilled person, such as, but not limited to, RNA interference, and in particular small interfering nucleic acids (siNA) against endogenous GSE3. Accordingly, in one embodiment, the plant, part thereof or plant cell 5 comprises at least one RNAi that reduces or abolishes the expression of GSE3. In one embodiment, the siNA may include, short interfering RNA (siRNA), double- stranded RNA (dsRNA), micro-RNA (miRNA), antagomirs and short hairpin RNA (shRNA) capable of mediating RNA interference. Silencing or reducing expression levels10 of GSE3 nucleic acid may also be achieved using virus-induced gene silencing.In another aspect of the invention, there is provided a genetically altered plant, part thereof or plant cell, wherein the plant, part thereof or plant cell comprises an RNA interference construct that reduces or abolishes the expression of at least one GSE3 15 nucleic acid sequence. Thus, in one embodiment of the invention, the plant comprises and expresses a nucleic acid construct comprising a RNAi, shRNA snRNA, dsRNA, siRNA, miRNA, ta-siRNA, amiRNA or co-suppression molecule that targets the GSE3 nucleic acid sequence as described herein and reduces expression of the endogenous GSE3 nucleic acid sequence. A gene is targeted when, for example, the RNAi, snRNA,20 dsRNA, siRNA, shRNA miRNA, ta-siRNA, amiRNA or cosuppression molecule selectively decreases or inhibits the expression of GSE3 compared to a control plant.Alternatively, a RNAi, snRNA, dsRNA, siRNA, miRNA, ta-siRNA, amiRNA or cosuppression molecule targets a GSE3 nucleic acid sequence when the RNAi, shRNAsnRNA, dsRNA, siRNA, miRNA, ta-siRNA, amiRNA or co-suppression molecule25 hybridises under stringent conditions to the gene transcript. The silencing RNA molecule may be introduced into the plant using conventional methods, for example a vector and Agrobacterium-mediated transformation. Stably transformed plants are generated and expression of the GSE3 gene compared to a wild30 type control plant is analysed. We further identified the transcription factor GS2 as a GSE3-interacting protein. The gain of function allele GS2AAformed large grains, indicating that GS2 is a positive regulator of grain size. Simultaneous disruption of GS2 / OsGRF4 and OsGRF3 caused35 small grains, like those observed in the loss-of-function of GSE3 alleles (Fig.2a, d-g). M&C PC933520WOA 29 Our genetic analyses further support that GSE3 functions genetically with GS2 / OsGRF4 and OsGRF3 to control grain size and weight (Fig.3d-h). Our RNA-seq, qRT-PCR, and CUT-Tag data reveal that the expression levels of several grain size genes, such as XIAO, GW6, and OsBZR1, were reduced in both ZH11-GSE3small grainand grf3-cri1 grf4- 5 cri2 young panicles (Fig.4g-h). GS2 directly binds to the promoter regions of those genes to promote their expression (Fig.4e-f). GSE3 also associates with the promoter regions of those genes, and the associations of GSE3 with the promoters of these genes depend on the functional GS2 / OsGRF4 and OsGRF3 (Fig.4i). We also found that the enrichment of H4Ac (pan- acetyl) in ZH11-GSE3small grain young panicles was reduced in the promoter10 regions of those selected genes compared with that in ZH11 young panicles. Thus, these findings reveal that GS2 directly binds to the promoter regions of several grain size genes and recruits GSE3 to regulate their expression by influencing the histone H4 acetylation levels, thereby regulating grain size. In addition, the gse3 alleles15 increased grain number as well as tiller number. In contrast, grf3-cri1 grf4-cri2 did notshow any difference in grain number and tiller number compared to ZH11 (Fig.13f-g). Thus, it is unlikely that GSE3 interacts with GS2 to promote expression of common grain number and tiller number-related genes. It is possible that GSE3 may interact with other transcription factors involved in the regulation of grain number and / or tiller number to20 promote expression of grain number and tiller number-related genes. Accordingly, in a further aspect, there is provided a genetically altered plant, part thereof or plant cell characterised by reduced or abolished activity or expression of a GRF4 (GROWTH REGULATING 4) protein and reduced or abolished expression of a GRF325 (GROWTH REGULATING 3) protein in said plant, part thereof or plant cell. In one embodiment, the GRF4 protein comprises a sequence as defined in SEQ ID NO:10 or a functional variant or homolog thereof. In another embodiment, GRF4 is encodedby the GS2 (GRAIN SIZE ON CHROMOSOME 2) gene. The sequence of GS2 may 30 comprise a sequence as defined in SEQ ID NO: 5 or 11 or a functional variant or homolog thereof. In one embodiment, the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional variant or homolog thereof. In a further embodiment, the GRF3 protein M&C PC933520WOA 30 is encoded by a nucleic acid sequence comprising a sequence as defined in SEQ ID NO: 12 or a functional variant or homolog thereof. As used herein, the term “reducing” or ‘reduction’ of expression or activity means a5 decrease in the levels of GS2 / GRF4 or GRF3 expression and / or activity by up to or morethan 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% when compared to the level in a wild-type or control plant. In one embodiment, reducing means a decrease in at least 50% compared to the level in a wild-type or control plant. Reducing also may or may not encompass abolishing expression. The term “abolish” expression means that10 no expression of GS2 / GRF4 or GRF3 is detectable (no transcript) or that no functionalGS2 / GRF4 or GRF3 polypeptide is produced.A functional variant is defined above. Nonetheless, a functional variant of GFR3 or GFR4 will retain the same transcriptional function activity, and will be able to mediate15 transcription of the genes under their respective controls. Transcript abundance of targetgenes can be measured using rt-PCR and western blots. A functional variant of GFR3or GFR4 should also retain the ability to interact with GSE3. This interaction can bedetermined by routine methods, including pull-down assays and the yeast two-hybrid (Y2H) system. 20 In one embodiment, the homolog of GS2 / GRF4 is selected from a sequence comprisingSEQ ID NO: 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72 or 74, or a functionalvariant thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72 or 74. 25 In one embodiment, the promoter of the homolog of GS2 / GRF4 is selected from asequence comprising SEQ ID NO: 47, 49, 51, 53, 55, 57, 59, 31, 63, 65, 67, 69, 71, 73 or 75 or a functional variant thereof, wherein the functional variant has at least 40% or atleast 50% overall sequence identity to SEQ ID NO: 47, 49, 51, 53, 55, 57, 59, 31, 63, 65,30 67, 69, 71, 73 or 75. In one embodiment, the homolog of GRF3 is selected from a sequence comprising SEQID NO: 76, 78, 80, 82, 84, 86, 87, 88, 89, 91, 93, 95, 97, 99 or 101 or a functional variant thereof, wherein the functional variant has at least 40% or at least 50% overall sequence35 identity to SEQ ID NO: 76, 78, 80, 82, 84, 86, 87, 88, 89, 91, 93, 95, 97, 99 or 101. M&C PC933520WOA 31 In one embodiment, the promoter of the homolog of GRF3 is selected from a sequencecomprising SEQ ID NO: 77, 79, 81, 83, 85, 90, 82, 94, 96, 98, 100, 102, 103, 104, or 105 or a functional variant thereof, wherein the functional variant has at least 50% overall 5 sequence identity to SEQ ID NO: 77, 79, 81, 83, 85, 90, 82, 94, 96, 98, 100, 102, 103, 104, or 105. In one embodiment, the plant comprises at least one mutation in at least one endogenous gene encoding GRF4 and / or at least one (GS2) GRF4 promoter and at least10 one mutation in at least one gene encoding GRF3 and / or at least one GRF3 promoter. As used throughout, by “GRF4 / GS2 promoter” is meant a region extending at least orapprox. 3 Kbp upstream of the ATG codon of the GS2 ORF. In one embodiment, thesequence of the GS2 promoter comprises or consists of a nucleic acid sequence as15 defined in any one of SEQ ID NO: 6, 7, 8 or a functional variant or homolog thereof. Inone embodiment, the GS2 promoter may also include 5’ UTR sequences.As used throughout, by “GRF3 promoter” is meant a region extending at least or approx.3 Kbp upstream of the ATG codon of the GRF3 ORF. In one embodiment, the sequence20 of the GRF3 promoter comprises or consists of a nucleic acid sequence as defined inSEQ ID NO: 13 or a functional variant or homolog thereof. In one embodiment, the GRF3promoter may also include 5’ UTR sequences. In a preferred embodiment, the at least one mutation is an insertion mutation in the third25 exon of GRF3. Preferably, the mutation is a T insertion mutation at position 545 of SEQID NO: 12. This is known as the Grf3-cri1 mutation and is exemplified in SEQ ID NO:113. In a preferred embodiment, the at least one mutation is a deletion of at least one residue 30 in the first exon of the GRF4 gene, wherein the single base deletion causes a frameshift position. Preferably, the mutation is the deletion of position 247 or 248 of SEQ ID NO: 11. These mutations are referred to as the Grf4-cri2 and Grf4-cri3 mutations, respectivelyand are exemplified in SEQ ID NO: 114 and 115 respectively. M&C PC933520WOA 32 The mutation in GRF4 and / or the mutation in GRF3 may be a loss or partial loss of function mutation, wherein the loss or partial loss of function mutations reduces orabolishes the activity or expression of GRF4 and / or GRF3. Most preferably, themutation(s) reduces or abolishes the activity or expression of GRF4 and GRF3.5 In one embodiment, the at least one mutation in GRF3 is selected from an insertionmutation in the third exon of GRF3, preferably a T insertion at position 545of SEQ ID NO: 12 or a corresponding position in a homologous sequence; and further a single basedeletion mutation in the first exon of the GRF4 gene, wherein the single base deletion10 causes a frameshift mutation, preferably selected from the Grf4-cri2 or Grf4-cri3mutation. This embodiment is reflected in the examples as Grf3-cri1 Grf4-cri2 and Grf3-cri1 Grf4-cri3 (Fig 13a), and the effect is shown in Fig.13b-g.In another aspect of the invention, there is provided a genetically altered plant, part 15 thereof or plant cell, wherein the plant, part thereof or plant cell comprises an RNA interference construct that reduces or abolishes the expression of at least one GRF4 nucleic acid sequence and a RNA interference construct that reduces or abolishes the expression of at least one GRF3 nucleic acid sequence. Preferably, at least one RNAi reduces or abolishes the expression of all copies of GRF4 and GFR3. 20 In one embodiment, RNA interference comprises the use of a siNA which may include, short interfering RNA (siRNA), double-stranded RNA (dsRNA), micro-RNA (miRNA), antagomirs and short hairpin RNA (shRNA) capable of mediating RNA interference of endogenous GRF4 and / or GRF3 nucleic acids. 25 We have also shown, in Example VI and in Figure 12, that the overexpression of GRAINSIZE ON CHROMOSOME 3 (GSE3) results in a large-grain phenotype. Accordingly, in one aspect of the invention, there is provided a genetically altered plant,30 part thereof or plant cell, wherein the plant, part thereof or plant cell is characterised by increased expression of GRAIN SIZE ON CHROMOSOME 3 (GSE3). In another aspect of the invention, there is provided a genetically altered plant, part thereof or plant cell, wherein the plant, part thereof or plant cell is characterised by M&C PC933520WOA 33 increased expression of a polypeptide encoded by the GSE3 (GRAIN SIZE ONCHROMOSOME 3) gene. As used herein, the term “increased” or “increasing” expression means an increase in 5 the levels of the GRAIN SIZE ON CHROMOSOME 3 (GSE3) gene or encoded polypeptide expression by up to or more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% when compared to the level in a wild-type or control plant. In one embodiment, increasing means an increase of at least 50% compared to the level in a wild-type or control plant. 10 In one embodiment, the method of increasing expression of the GRAIN SIZE ONCHROMOSOME 3 (GSE3) gene and / or a polypeptide encoded by GSE3 comprisesintroducing and expressing a nucleic acid construct comprising a nucleic acid sequence encoding GRAIN SIZE ON CHROMOSOME 3 (GSE3), where the nucleic acid sequence15 encoding GRAIN SIZE ON CHROMOSOME 3 (GSE3) is operably linked to a regulatorysequence, such as a promoter. In one embodiment, the plant comprises at least one mutation in at least one endogenous gene encoding GRAIN SIZE ON CHROMOSOME 3 (GSE3) and / or at least20 one GRAIN SIZE ON CHROMOSOME 3 (GSE3) promoter.The mutation in GRAIN SIZE ON CHROMOSOME 3 (GSE3) may be a gain-of-functionmutation, wherein the gain-of-function function mutations increases the expression of GRAIN SIZE ON CHROMOSOME 3 (GSE3). 25 In a preferred embodiment, the mutation is the insertion of at least one or more additional copy(ies) of a nucleic acid encoding a GRAIN SIZE ON CHROMOSOME 3 (GSE3) polypeptide or a homolog or variant thereof such that said sequence is operably linked to a regulatory sequence and wherein said mutation is introduced using targeted genome 30 editing. Preferably, said mutation results in an increase in the expression of a GRAIN SIZE ON CHROMOSOME 3 (GSE3) nucleic acid compared to a control or wild-type plant. That is, in a preferred embodiment, the mutation is a knock-in mutation of a GRAIN SIZE ON CHROMOSOME 3 (GSE3) nucleic acid. M&C PC933520WOA 34 A knock-in mutation refers to a CRISPR knock-in, in which a DNA sequence is used as a template to insert a target DNA sequence into a CRISPR-mediated Double-Stranded Break. This relies on a cell’s endogenous homology-directed repair system. Such mutations are well-known in the art, and methods are widely available in literature. 5 There is also provided a method of obtaining any of the genetically altered plant, part thereof or plant cells described above. In one embodiment, the method comprises introducing at least one mutation in at least one GRAIN SIZE ON CHROMOSOME 3 (GSE3) gene, preferably all copies of the GRAIN SIZE ON CHROMOSOME 3 (GSE3)10 gene. In another embodiment, the method may comprise introducing and expressing in a plant or plant cell a nucleic acid construct comprising a nucleic acid sequence encoding an GRAIN SIZE ON CHROMOSOME 3 (GSE3) polypeptide as defined in SEQ ID NO: 1 or15 a homolog or functional variant thereof, as defined herein. Preferably, the nucleic acid sequence is operably linked to a regulatory sequence, preferably a promoter. In one embodiment, the progeny plant is stably transformed with the nucleic acid construct described herein and comprises the exogenous polypeptide or polypeptides 20 that are heritably maintained in the plant cell. The method may also comprise the additional step of collecting seeds from the selected progeny plant. In a further aspect of the invention, there is provided a plant, part thereof or plant cell characterised by an increase in grain weight and / or size compared to a wild-type or 25 control pant, wherein preferably, the plant comprises increased expression of GRAIN SIZE ON CHROMOSOME 3 (GSE3). In preferred embodiments, the plant is further characterised by: i) at least one mutation in the GRAIN SIZE ON CHROMOSOME 3 (GSE2)30 gene and / or its promoter, preferably an insertion mutation of at least one or more additional copy(ies) of a nucleic acid encoding a GRAIN SIZE ON CHROMOSOME 3 (GSE3) polypeptide or a homolog or variant thereof; or ii) expression of at least one nucleic acid construct comprising a nucleic acidsequence encoding an GRAIN SIZE ON CHROMOSOME 3 (GSE3) M&C PC933520WOA 35 polypeptide as defined in SEQ ID NO: 1 or a homolog or functional variantthereof, as defined herein. Nucleic acid constructs, including genome editing constructs, can be introduced into a 5 plant cell using any suitable method known to the skilled person (the term “introduced” can be used interchangeably with “transformation”, which is described below). Specifically, methods of transforming a plant with a nucleic acid construct to achieve GSE3 overexpression is described at Example IV. 10 Transformation of plants is now a routine technique in many species. Any of several transformation methods known to the skilled person may be used to introduce the nucleic acid construct of interest into a suitable ancestor cell. The methods described for the transformation and regeneration of plants from plant tissues or plant cells may be utilized for transient or for stable transformation. 15 Transformation methods include the use of liposomes, electroporation, chemicals that increase free DNA uptake, injection of the DNA directly into the plant (microinjection), gene guns (or biolistic particle delivery systems (biolistics)) as described in the examples, lipofection, transformation using viruses or pollen and microprojection. Methods may be20 selected from the calcium / polyethylene glycol method for protoplasts, ultrasound- mediated gene transfection, optical or laser transfection, transfection using silicon carbide fibres, electroporation of protoplasts, microinjection into plant material, DNA or RNA-coated particle bombardment, infection with (non-integrative) viruses and the like. Recombinant plants can also be produced via Agrobacterium tumefaciens mediated 25 transformation, including but not limited to using the floral dip / Agrobacterium vacuum infiltration method. Accordingly, in one embodiment, at least one nucleic acid construct molecule or CRISPR construct as described herein can be introduced to at least one plant cell using any of 30 the above described methods. Optionally, to select transformed plants, the plant material obtained in the transformation is, as a rule, subjected to selective conditions so that transformed plants can be distinguished from untransformed plants. For example, the seeds obtained in the above- 35 described manner can be planted and, after an initial growing period, subjected to a M&C PC933520WOA 36 suitable selection by spraying. A further possibility is growing the seeds, if appropriate after sterilization, on agar plates using a suitable selection agent so that only the transformed seeds can grow into plants. As described in the examples, a suitable marker can be DsRed. Alternatively, the transformed plants are screened for the presence of a 5 selectable marker, such as, but not limited to, GFP, GUS (β-glucuronidase). Other examples would be readily known to the skilled person. Following DNA transfer and regeneration, putatively transformed plants may also be evaluated, for instance using PCR to detect the presence of the gene of interest, copy 10 number and / or genomic organisation. Alternatively, or additionally, integration and expression levels of the newly introduced DNA may be monitored using Southern, Northern and / or Western analysis, both techniques being well known to persons having ordinary skill in the art. 15 The method may further comprise the step of regenerating a transgenic plant from a plant cell described above. Preferably, the transgenic plant comprises in its genome a nucleic acid sequence selected from SEQ ID NO: 4 or a homolog or functional variantthereof, and obtaining progeny derived from the transgenic plant, where the progeny exhibits an alteration in grain or seed size, preferably an increase in grain size. 20 It is further envisaged that any of the mutations or genetic alterations described for any of GSE3 or GRF3 and GS2 / GRF4 may be combined. That is, a method or plant described herein may comprise a first mutation in at least one GSE3 gene and a second or third mutation in at least one GRF3 and / or at least one GS2 / GRF4 gene.25 In one embodiment, the plant is a male-sterile plant. A male-sterile plant is a plant thatis unable to generate functional anthers, pollen or male gametes. A male-sterile plant line may also be known as an A line. In one embodiment, the male-sterile plant is acytoplasmic male-sterile plant or an environmental genic sterility male (EGMS) plant. 30 By ‘cytoplasmic male-sterile’ or ‘CMS’ plant is meant a plant which is unable to producefunctional pollen (i.e. the plant is male-sterile) due to a maternally inherited trait that isencoded in the mitochondrial genome. Such CMS plants possess a male-sterilecytoplasm arising from the mitochondrial-encoded CMS-causing gene (hereafter termed35 a CMS gene) and lack a functional nuclear RESTORER OF FERTILITY (Rf, or fertility M&C PC933520WOA 37 restorer) gene or genes. In the F1 plants, the Rf gene restores male fertility, and thecombination of nuclear genomes from the CMS line and the restorer line produces hybridvigour. 5CMS plants have been well reviewed in the literature, for example in Melonek et al.(2021) and Xu et al. (2023). Therefore, the exemplar CMS plants listed below are meant to be representative only – the invention is intended to be applied to all CMS plants.CMS plant lines and the genetic basis for their sterility are identifiable from the literature,10 for example, Chen and Liu (2014), Xu et al. (2023) and Melonek et al. (2021). In one embodiment, the plant is rice, and the CMS line is selected from Xiaoligeng (XLG),Lead Rice-type (LD-CMS), Boro II type (BT-CMS), Java-type, Dongxiang-type CMS type(D1-CMS), D-type CMS, K52 (K-CMS), Dissi (D-CMS), Dwarf Abortive (DA-CMS), Gang15 type, Hong Lian type (HL-CMS), Indonesian paddy valley type, Chinese wild type (CW-CMS) RT98-type CMS (RT98-CMS), RT102-type CMS (RT102-CMS), and Wild-abortive(WA-CMS). More preferably, the CMS line in rice is selected from LD-CMS, BT-CMS,WA-CMS, HL-CMS, Indonesian paddy valley type, D-type CMS, CW-CMS and TFA1. Inone embodiment, the CMS plant is from the XLG line. 20 In one embodiment, the plant is wheat, and the CMS line is T-CMS (Melonek et al.2021).In one embodiment, the plant is a Brassica spp, and the CMS line is selected from Ogura (CMS-Ogu), pol, Gülzow’ (G)-type CMS, Pol-type CMS (CMS-Pol), CMS-Nao, and C-25 type CMS. Most preferably the plant is from the P-type CMS.In one embodiment, the plant is maize, and the CMS line is selected from CMS-T, CMS-C, and CMS-S. 30 In a preferred embodiment, the CMS line is selected from a CMS line listed within Table 2. In a preferred embodiment, the CMS line is characterised by a CMS-gene listed within Table 2. Table 2. Exemplar characterised CMS lines and CMS-associates genes. Based on table35 produced in Toriyama K.2021 and Chen and Liu (2014). M&C PC933520WOA 38 M&C PC933520WOA 39 The EGMS plant may be selected from a thermo-sensitive genic male-sterility (TGMS)plant, a photoperiod-sensitive genic male-sterility (PGMS) plant, a photo-thermosensitive genic male sterility (PTGMS) plant and a humidity-sensitive genic male sterility5 (HGMs) plant. By ‘environment-sensitive genic male’ or ‘EGMS’ plant is meant a plant that is unable toproduce functional pollen (is sterile) due to genetic and environmental factors. EGMSplants have been well reviewed in the literature, for example at Ashraf et al. (2020) and10 Xu et al (2023). Therefore, the exemplar EGMS plants listed below are meant to berepresentative only – the invention can be applied to all EGMS plants.In a preferred embodiment, the EGMS plant is selected from a line listed within Table 3. In another preferred embodiment, the EGMS plant is characterised as having a gene15 responsive for EGMS, selected from a gene listed within Table 3. Thermo-sensitive-genic-male-sterility (TGMS) system refers to the ability of the male gametes to be sterile or fertile at a higher or lower temperature than the critical point. The TGMS system induces male sterility to male fertility through temperature variations20 at the critical anther developmental stage of the crop. TGMS lines include Y58S, 5460S,Zhu1S, Norin PL 12, Annong 810 S, Hennong S, Sokcho-MS, SA2, J207S and G20S inrice. In one embodiment, the plant is rice and the preferred TGMS plant is from the Y58Sline.25 Photoperiod-sensitive-genic-male-sterility (PGMS) refers to plants where the malegametes can be sterile or fertile in accordance with length of exposure to light. PGMSplants can be further divided into long photoperiod sensitive genic male-sterile type plants and short photoperiod sensitive genic male-sterile type plants. The longphotoperiod sensitive genic male-sterile type is male sterile in the long day environment30 and fertile in the short day environment, while the short photoperiod sensitive genic male- sterile type is male sterile in the short day environment and fertile in the long day M&C PC933520WOA 40 environment. PGMS lines include Nongken 58S (NK58S), EGMS, 201, CIS 28-10,Zhenong S, X-88, 7001 S, Mian9S and Yi D1S.In one embodiment, the plant is wheat and the PGMS line is Norin26. 5 Photo-thermo sensitive genic male sterility (PTGMS) refers to lines controlled by the interaction of both photoperiod and temperature. Table 3: Characterised EGMS lines in rice. 10 Alternatively, the plant may be a maintainer plant line. A maintainer line may also beknown as the B line. A maintainer plant line refers to a plant developed by crossing the malesterile line to a fertile plant. A maintainer line is characterised by a recessive Rf gene, and does not possess the CMS-causing gene or allele (i.e., the mitochondrial-15 encoded allele of sterility). Therefore, the maintainer line pollen exhibits normal fertility.The maintainer plant may be functionally defined as a plant that can pollinate a CMS plant to produce more CMS plants. M&C PC933520WOA 41 Example of suitable maintainer lines include Tianfeng B (TFB) in rice.In Figure 5a-b, it is further shown that introducing a loss of function mutation (gse3-cri3mutation) into GSE3 of the TFA CMS line results in a reduction in grain size (a) and grain 5thickness (5b) relative to the non-mutated TFA line. Figure 15 also shows that the samemutation can be applied to a maintainer line (TFB) and a significant reduction in grain length, width and thickness will be observed, relative to the non-mutated maintainer line. Similarly, introducing a loss of function mutation (gse3-cri4) into the TGMS plant line 10 Y58S using gene editing, resulted in decreased grain size (Fig 5i) and grain thickness (Fig 5j). Finally, it is also shown that the introduction of a loss of function mutation in GRF3 andGRF4 (grf3-cri1 mutation and grf4-cri2 mutation) results in a significant reduction in grain15 size (Fig.13b), grain length (Fig.13c) and grain width (Fig.13c). Accordingly, in one embodiment, the genetically altered plant, part thereof or plant cell of the invention is characterised by a reduced grain or seed size relative to the grain sizein a control or wild-type plant. 20 By “reduced grain size” or “reducing seed size” it is meant a reduction in mean grain or seed size by at least 5%, 10%, 12%, 15%, 17%, 20%, 22%, 25%, 27%, 30%, 32%, 35%, 37%, 40%, 42%, 45%, 47% or 50%, relative to the grain or seed size of a control or wild-type of a plant. Grain (or seed) size may encompass one or more of grain (or seed) 25 length, width or thickness. Preferably, at least grain (or seed) thickness is reduced, although as shown in the examples, reducing the expression or activity of GSE3 reduces grain / seed length, grain / seed width and grain / seed thickness. Nonetheless, grain / seed thickness may be considered the most important characteristic of the genetically altered plants that allows for mechanised hybrid seed sorting. 30 Figure 5d further shows that the seed number per plot of a TFA (GSE3 mutant) x DHZ combination was increased by 21.2% compared with that of the TFA (WT) X HZ combination. It is also shown in Figure 5l-m that the F1 hybrid seed number per plot of the thermos-sensitive genic sterile male Y58S (GSE3 mutant) x R900 combination was 35 also increased by 38.3% compared to a Y58S (WT) x R900 cross. M&C PC933520WOA 42 Accordingly, in one embodiment, the genetically altered plant, part thereof or plant cell of the invention is characterised by an increased grain / seed number relative to thegrain / seed number in a control or wild-type plant. By an increase in grain / seed numberis meant the number of seeds / grains produced per plot or per plant. The increase may5 be an increase of at least 5%, 10%, 12%, 15%, 17%, 20%, 22%, 25%, 27%, 30%, 32%, 35%, 37%, 40%, 42%, 45%, 47% or 50%, relative to the seed / grain number of a controlor wild-type of a plant. In one embodiment, the plant shows increased efficiency of F1 seed production. By 10 efficiency of F1 seed production, it is meant the number of F1 seeds produced per plot following hybridisation. An increase in efficiency of F1 seed production may be an increase of seed number at least 5%, 10%, 12%, 15%, 17%, 20%, 22%, 25%, 27%, 30%, 32%, 35%, 37%, 40%, 42%, 45%, 47% or 50%, relative to the seed number of a control or wild-type of a plant. 15 In one embodiment, the plant does not show a significantly altered yield of F1 grain / seeds. Yield is normally defined as the measurable produce of economic value from a crop. This may be defined in terms of quantity and / or quality. We have shown that the grain yield of the plants with the small-grain phenotype is comparable with that of20 control hybrid rice (Fig.1p-q). Achange in seed or grain yield may be measured by one or more of the following: a) achange in seed or grain biomass (total seed weight) which may be on an individual seed basis and / or per plant and / or per hectare or acre; b) number of (filled) seeds; c) seed 25 filling rate, which is expressed as the ratio between the number of filled seeds divided by the total number of seeds; d) harvest index, which is expressed as a ratio of the yield of harvestable parts, such as seeds, divided by the total biomass; and e) thousand grain weight (TGW), which is extrapolated by dividing the total weight of a collection of grain by the total number of grain and then multiplying by 1000. An altered TGW may result 30 from an increased seed size and / or seed weight, and may also result from an increase in embryo and / or endosperm size. In another aspect of the invention, there is provided a grain or seed, wherein said grainor seed is characterised by reduced or abolished activity or expression of a GRAIN SIZE35 ON CHROMOSOME 3 (GSE3) protein in said grain or seed, as described above. M&C PC933520WOA 43 Preferably the seed or grain is F1 hybrid seed or grain from an F1 hybrid plant. Furthermore, the seed may be characterised by a reduced seed size (i.e. seed lengthand / or width and / or thickness) compared to the seed size (seed length and / or widthand / or thickness) in a wild-type or control plant (i.e. a plant without reduced or abolished 5 expression or activity of GSE3). In another aspect of the invention, there is provided a grain or seed, wherein said grain or seed is characterised by reduced or abolished activity or expression of GRF4 and GRF3 protein in said grain or seed, as described above. Preferably the seed or grain is 10 F1 hybrid seed or grain from an F1 hybrid plant. Furthermore, the seed may be characterised by a reduced seed size (i.e. seed length and / or width and / or thickness)compared to the seed size (seed length and / or width and / or thickness) in a wild-type orcontrol plant (i.e. without reduced or abolished expression or activity of GS2 / GRF4 and GRF3). 15 In another aspect of the invention, there is provided there is provided a grain or seed, wherein said grain or seed is characterised by reduced or abolished activity or expression of GSE3, GRF4 and GRF3 protein in said grain or seed, as described above. Preferably the seed or grain is F1 hybrid seed or grain from an F1 hybrid plant.20 Furthermore, the seed may be characterised by a reduced seed size (i.e. seed lengthand / or width and / or thickness) compared to the seed size (seed length and / or widthand / or thickness) in a wild-type or control plant (i.e. without reduced or abolished expression or activity of GSE3 and GS2 / GRF4 and GRF3).25 We have also shown that a gain of function mutation in GS2 / GRF4 can be used toachieve a large grain phenotype. We have also shown that overexpression of GRAIN SIZE ON CHROMOSOME 3 (GSE3) can be used to achieve a large grain phenotype. These large-grain phenotypes can contribute to achieving full mechanisation of hybridseed production in three-line and two-line systems, if used in tandem with a small-grain30 producing male-sterile line. The Applicant also envisages that a gain-of-function mutations known in the art to form large grans in rice, such as GRAIN WEIGHT3 (GS3), GRAIN WEIGHT2 (GW2), GRAIN SIZE ON CHROMOSOME 5 (GSE5) / GRAIN WIDTH5 (GW5), GRAIN LENGTH 3 (GL3), M&C PC933520WOA 44 LARGE1, LARGE2, and MKP1 18-31 may be used in tandem with a small-grain producing male-sterile line as described herein. Accordingly, in another aspect of the invention, there is provided a genetically altered 5 fertility restorer plant, a part thereof or plant cell, wherein the restorer line is characterised by increased activity or expression of the protein encoded by the GS2 (GRAIN SIZE ON CHROMOSOME 2) gene. In a further aspect of the invention, there is provided a genetically altered fertility restorer 10 plant, a part thereof or plant cell, wherein the restorer line is characterised by increased expression of the protein encoded by the GSE3 (GRAIN SIZE ON CHROMOSOME 3)gene. In a preferred embodiment, the method of increasing expression comprises introducing 15 a mutation into the plant genome, where said mutation is the insertion of at least one or more additional copy(ies) of a nucleic acid encoding a GRAIN SIZE ON CHROMOSOME 3(GSE3) polypeptide or a homolog or variant thereof such that said sequence isoperably linked to a regulatory sequence and wherein said mutation is introduced using targeted genome editing. 20 In another preferred embodiment, the method of increasing expression comprises introducing and expressing in a plant or plant cell a nucleic acid construct comprising a nucleic acid sequence encoding a GRAIN SIZE ON CHROMOSOME 3 (GSE3) polypeptide as defined in SEQ ID NO: 1 or a homolog or functional variant thereof, as25 defined herein. Preferably, the nucleic acid sequence is operably linked to a regulatory sequence, preferably a promoter. In a further embodiment of the invention, there is provided a genetically altered fertility restorer plant, a part thereof or plant cell, wherein the restorer line is characterised by30 increased expression of a protein selected from GRAIN SIZE ON CHROMOSOME 2(GS2 / OsGRF4), GS3, GRAIN WEIGHT2 (GW2), GRAIN SIZE ON CHROMOSOME 5 (GSE5) / GRAIN WIDTH5 (GW5), GRAIN LENGTH 3 (GL3), LARGE1, LARGE2, and MKP118-31. Preferably, this genetically altered fertility restorer plant is used in a method with a small-grain producing male-sterile lin. For example, in a method of hybrid35 breeding; and / or M&C PC933520WOA 45 i. a method of mechanised hybrid seed sorting; and / orii. a method of reducing seed size of a male-sterile plant, where preferablyseed size is selected from one or more of seed length, seed width and seed thickness, most preferably seed thickness; and / or 5iii. a method of increasing the purity of F1 hybrid seeds,as described below. By a “fertility restorer plant” is meant a plant characterised by homozygous dominantalleles of a nuclear RESTORER OF FERTILITY (Rf) gene and an absence of a10 mitochondrial sterility gene (i..e, a male sterile-causing gene). A fertility restorer geneblocks or compensates for specific mitochondrial dysfunctions resulting from male-sterilegenes that are phenotypically expressed during pollen development. A fertility restorerplant is preferably functionally defined, as any plant can pollinate a male-sterile plant to produce fertile F1 (hybrid) progeny. These plants may be referred to as “restorer” plants15 herein. Alternatively, these plants may also be referred to as “male plants” or“pollinators”. Examples of suitable fertility restorer plants include HZ, R900, Minghui63, 93-11, Shuhui527, Chenhui727, Fuhui838, Ce64-7, IR26 and IR24 in rice.20 In one embodiment, the fertility restorer plant is characterised by a dominant Rf alleleselected from Rf1, Rfa1, Rf1b (Rf5), RF2, Rf4, Rf5 (RF1b), Rf-A619, Rf17 or Rf102. That is, the plant is characterised as having a functional RESTORER OF FERTILITY (Rf)gene selected from Rf1, Rfa1, Rf1b (Rf5), RF2, Rf4, Rf5 (RF1b), Rf-A619, Rf17 or Rf102. 25 Table 4. Preferred RF genes, based on table produced in Ashraf et al. (2020). M&C PC933520WOA 46 Literature has shown that a restorer line may be generated by introgressing a restorer gene into parental lines of a hybrid. Interestingly, hybrids produced from these lines showed equivalent or better agronomic performance relative to their counterparts. 5 Comprehensive methods are provided in Jiang et al. (2022). As used herein, the term “increasing” or “increased” expression or activity means anincrease in the levels of GS2 / GRF4 expression and / or activity by up to or more than10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% when compared to the level in a10 wild-type or control plant. In one embodiment, increasing means an increase of at least50% compared to the level in a wild-type or control plant. In one embodiment, the plant comprises at least one mutation in at least one GS2 (endogenous) gene and / or at least one GS2 promoter, where the mutation increases the15 expression and / or activity of GS2 / GRF4. As such, preferably the mutation is a gain of function mutation. Preferably, at least one mutation is introduced into each GS2 genepresent in a plant, part thereof or plant cell. That is, all homologs of GS2 comprise atleast one mutation.20 Accordingly, in one embodiment, the genetically altered plant, part thereof or plant cellis characterised by a GS2 protein comprising a sequence as defined in SEQ ID NO: 10,or a functional variant or fragment thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10. In a preferred embodiment, the GS2protein is encoded by a sequence as defined in SEQ ID NO: 5 or 11 or a functional 25 variant or fragment thereof. In one embodiment the mutation is a gain-of-function mutation in the GS2 gene or polypeptide. By gain-of-function it is meant a mutation that results in an increase in the M&C PC933520WOA 47 expression or activity of a gene or polypeptide. In this case, a gain-of-function mutation is applied with respect to GS2 or a functional variant or homolog. Increasing the expression or activity of the GS2 gene or an encoded polypeptide(s) 5 (preferably GRF4) results in increased grain size and weight, and also increased grain yield. As such a gain-of-function mutation may be defined as mutation that increases grain size or grain weight, or increases grain yield, and / or results in an increase in gene activity and / or results in an increase in protein synthesis, wherein preferably said geneactivity or protein is involved in grain initiation and development, most preferably wherein10 the protein is GRF4. A change in the activity or expression of the GS2 gene / polypeptidemay be determined by measuring transcript abundance following from the activity of the transcription factor may be measured using routine methods, for example, qt-PCR. Alternatively, the mean grain size, weight or yield of grain can be measured and compared to the grain of control (i.e., non-mutated) plant. Alternatively, the any change 15 in expression or activity of the GRF4 product (a transcription factor) of GS2 can be determined by measuring protein-protein interactions, using techniques standard in the art, such as, but not limited to, interaction assays using recombinant proteins, yeast-2- hybrid, immunoprecipitation or bimolecular fluorescence. 20 For any of the above embodiments, the mutation may be introduced into the GS2 gene or the promoter thereof. Preferably, the mutation is a substitution mutation. That is, the substitution of one basefor another, different base. More preferably, the mutation is the substitution mutation25 TC487- 488AA in the GS2 gene. Most preferably, the mutation is a substitution mutation,TC>AA at position 487 / 488 of the GS2 gene, wherein GS2 comprises a sequence asdefined in SEQ ID NO: 5 or 11. This mutation has been shown to perturb miRNA regulation (OsmiR396c) of GS2, resulting in a gain-of-function effect that produces large and heavy grains and also an increased grain yield. 30 The mutation in the endogenous gene can comprise at least one mutation in any one of the following sites: the coding region of the GS2 gene, preferably exon 3; a micro RNA(miRNA) binding site, preferably at the miR396 binding site; an intronic sequence, preferably intron 2 and / or intron 3; and / or at a splice site, in the 5’UTR, the 3’UTR, the35 termination signal, the splice acceptor site or the ribosome binding site. M&C PC933520WOA 48 In one example the miR396 binding or recognition site comprises or consists of the following sequence or a variant thereof, as defined herein: 5 CCGTTCAAGAAAGCCTGTGGAA: SEQ ID NO: 9 Preferably the mutation is any mutation that prevents the cleavage of the sequence by microRNA and thus its subsequent degradation. This results in an increase in the levels of both GS2 mRNA and protein. In one embodiment, the mutation is a substitution.10 In a specific embodiment, the mutation is one or both of the following: -a T to A at position 4 of SEQ ID NO: 9 or a homologous position thereof;- a C to A at position 5 of SEQ ID NO: 9 or a homologous position thereof.15 In an additional or alternative embodiment, the mutation is in intron 2 and / or intron 3 at least one of the following: -an A to G at position 724 or 725 of SEQ ID NO: 5 or a homologous positionthereof; -a T to C at position 1672 of SEQ ID NO: 5 or a homologous position thereof.20 Alternatively, or in addition to at least one of the above described mutations in theendogenous gene, the mutation is in the GS2 promoter. Preferably said mutation is anymutation that increases the expression of GS2. In one example, the mutation is at leastone of or any combination thereof of the following mutations. The former positions are 25 positions in the haplotype A promoter (for example, a promoter that comprises or consists of SEQ ID NO: 6 or a variant thereof). The latter positions are positions in thehaplotype C promoter (for example, a promoter that comprises or consists of SEQ ID NO: 7 or a variant thereof).30 - a C to T substitution at position -941 or -935 from the GS2 start codon or atposition 60 of SEQ ID NO: 6 or position 66 of SEQ ID NO: 7; or a homologousposition thereof; -a T to A substitution at position -884 or position -878 from the GS2 start codon orat position 118 of SEQ ID NO: 6 or position 124 of SEQ ID NO: 7; or a35 homologous position thereof; M&C PC933520WOA 49 -a C to T substitution at position -855 or -849 from the GS2 start codon or atposition 148 of SEQ ID NO: 6 or position 154 of SEQ ID NO: 7; or a homologousposition thereof; -a C to T substitution at position -847 or -841 from the GS2 start codon or at5 position 157 of SEQ ID NO: 6 or position 163 of SEQ ID NO: 7; or a homologousposition thereof; -a C to T substitution at position -801 or -795 from the GS2 start codon or atposition 204 of SEQ ID NO: 6 or position 210 of SEQ ID NO: 7; or a homologousposition thereof;10 - a C to T substitution at position -522 or -516 from the GS2 start codon or atposition 484 of SEQ ID NO: 6 or position 489 of SEQ ID NO: 7; or a homologousposition thereof; -a G to C substitution at position -157 from the GS2 start codon or at position 850of SEQ ID NO: 6 or position 516 of SEQ ID NO: 7; or a homologous position15 thereof; In one embodiment, the mutation is -a T to A substitution at position -884 or position -878 from the GS2 start codonor at position 118 of SEQ ID NO: 6 or position 124 of SEQ ID NO: 7; or a20 homologous position thereof; and -a C to T substitution at position -847 or -841 from the GS2 start codon or atposition 157 of SEQ ID NO: 6 or position 163 of SEQ ID NO: 7; or a homologousposition thereof; -a C to T substitution at position -801 or -795 from the GS2 start codon or at25 position 204 of SEQ ID NO: 6 or position 210 of SEQ ID NO: 7; or a homologousposition thereof. In a preferred embodiment, the mutation increases the expression or activity of GS2 or an encoded polypeptide. 30 In a further aspect of the invention, there is provided the use of the male-sterile plant and / or the maintainer plant and / or the fertility restorer plant of the invention in hybrid breeding. In another aspect of the invention there is provided the use of male-sterile plant and / or the fertility restorer plant of the invention in mechanised hybrid seed sorting. 35 M&C PC933520WOA 50 In a further aspect of the invention, there is provided iv. a method of hybrid breeding; and / orv. a method of mechanised hybrid seed sorting; and / orvi. a method of reducing seed size of a male-sterile plant, where preferably5 seed size is selected from one or more of seed length, seed width and seed thickness, most preferably seed thickness; and / or vii. a method of increasing the purity of F1 hybrid seeds,where the method comprises10 a. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein, as defined above, in a male-sterileplant, part thereof or plant cell; and / or b. reducing or abolishing the activity or expression of GRF4 and GRF315 protein, as defined above, in a male-sterile plant, part thereof or plant cellto obtain a male-sterile plant line that is characterised by a reduced seed size, as described above, compared to a control or wild-type plant. This plant can then be used in hybrid breeding. 20 In a further embodiment of any the above methods, the method further comprises obtaining and / or breeding the male sterile plant or plant line described above with a restorer line characterised by increased expression or activity of GS2. In a preferred embodiment, the method of increasing expression or activity of GS2 comprises introducing at least one mutation into a GS2 gene or promoter. 25 In a further aspect of the invention, there is provided i. a method of hybrid breeding; and / orii. a method of mechanised hybrid seed sorting; and / oriii. a method of increasing the purity of F1 hybrid seeds,30 where the method comprises selecting a plant, part thereof or plant cell, wherein the plant, part thereof or plant cell comprises reduced or abolished expression and / or activity of GRAIN SIZE ON CHROMOSOME 3 (GSE3). In a further embodiment, the methodcomprises detecting in the plant or plant germplasm at least one polymorphism or mutation in at least one GSE3 gene and / or promoter and selecting said plant or progeny, M&C PC933520WOA 51 wherein the polymorphism or mutation reduces the expression and / or activity of GSE3. In a further embodiment, the method may comprise crossing or hybridising said selected plant with a restorer plant, for example as defined herein. The restorer plant may be characterised by a large-grain phenotype, as described herein. 5 In a further aspect of the invention, there is a method of increasing (i.e. improving) theF1 hybrid seed number (i.e. increasing the number of hybrid seeds produced) in a(hybrid) plant, the method comprising: a. reducing or abolishing the activity or expression of a GRAIN SIZE ON10 CHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof or plantcell as defined above; and / orb. reducing or abolishing the activity or expression of GRF4 and GRF3 protein, asdefined above, in a male-sterile plant, part thereof or plant cell.15 By an increase / improvement in grain / seed number is meant the number of seeds / grainsproduced per plot or per plant. The increase may be an increase of at least 5%, 10%, 12%, 15%, 17%, 20%, 22%, 25%, 27%, 30%, 32%, 35%, 37%, 40%, 42%, 45%, 47% or 50%, relative to the seed / grain number of a control or wild-type of a plant.20 As discussed above, the term “reducing” or “reduced” expression or activity means adecrease in the levels of GSE3 and / or GRF3 and GRF4 expression and / or activity by up to or more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% when compared to the level in a wild-type or control plant. In one embodiment, reducing means a decreasein at least 50% compared to the level in a wild-type or control plant. Reducing also may 25 or may not encompass abolishing expression. The term “abolish” expression means that no expression is detectable (no transcript) or that no functional polypeptide is produced. The levels of GSE3 and / or GRF3 and GRF4 expression and / or activity can be reducedor abolished in male-sterile plants using targeted mutagenesis techniques, such as30 CRISPR, or by using RNAi methods, as described herein. In a further embodiment, the method may further comprise increasing the expressionand / or activity of a protein encoded by the GS2 gene (GRF4), as defined above, in afertility restorer plant, plant part thereof or plant cell. 35 M&C PC933520WOA 52 As discussed above, the term “increasing” or “increased” expression or activity meansan increase in the levels of GS2 / GRF4 expression and / or activity by up to or more than 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90% when compared to the level in a wild-type or control plant. In one embodiment, increasing means an increase of at least 5 50% compared to the level in a wild-type or control plant. The level of GS2 / GRF4 expression and / or activity can be increased in a fertility restorerplant using targeted mutagenesis techniques, such as CRISPR. Alternatively,GS2 / GRF4 expression may be increased by introducing and expressing a nucleic acid 10 construct comprising a nucleic acid sequence encoding GS2, where the nucleic acid sequence encoding GS2 is operably linked to a regulatory sequence, such as a promoter. The method may optionally further comprise the step of hybridising the male-sterile plant 15 described above with the fertility restorer plant, also described above to obtain an F1 hybrid plant and / or F1 hybrid seed. In an embodiment, the hybrid plant has an increasedseed number (compared to the number of seeds from a hybrid plant that does not contain one or more of the above-described genetic alterations (i.e. a wild-type or control plant as described above)). 20 In a further embodiment, the method may further comprise harvesting seeds obtained orobtainable from the male-sterile plant simultaneously with seeds obtained or obtainable from the fertility restorer plant. A particular advantage of the invention is that the seeds can be harvested mechanically. Accordingly, the method may further comprise25 mechanically harvesting the seeds of the invention, where hybrid seeds are separated from restorer plant seeds based on seed size, such as seed length, seed width and seedthickness. Preferably, hybrid seeds are mechanically sorted based on seed thickness. In one embodiment, a sieve or sifter may be used to mechanically separate hybrid and 30 non-hybrid seeds. The size of the sieve may be at least 2.5mm or less. As shown in Figure 1l, when the width of the sieve was less than 2.08mm, the purity of F1 hybrid seeds was more than 96.01%, which is above the purity standard required for commercial seed production. M&C PC933520WOA 53 In a further aspect of the invention, there is provided an F1 hybrid plant obtained orobtainable by the methods of the invention. In a further aspect of the invention, there is provided F1 hybrid seed obtained or 5 obtainable by the methods of the invention. In another aspect of the invention, there is provided a method of increasing grain size in a plant, part thereof or plant cell, the method comprising increasing the expression and / or activity of GRAIN SIZE ON CHROMOSOME 2 (GS2) in said plant, part thereof or plant10 cell. Appropriate and preferred methods to increase the expression and / or activity of GS2have been described elsewhere, and are intended to apply here. In another aspect of the invention, there is provided a method of increasing grain size in aplant, the method comprising increasing the expression of GSE3 in said plant. In a15 preferred embodiment, the method of increasing expression comprises introducing at least one or more additional copy(ies) of a nucleic acid encoding a GRAIN SIZE ON CHROMOSOME 3 (GSE3) polypeptide or a homolog or variant thereof such that saidsequence is operably linked to a regulatory sequence and wherein said mutation is introduced using targeted genome editing. Alternatively, in another preferred 20 embodiment, the method comprises introducing and expressing in a plant or plant cell a nucleic acid construct comprising a nucleic acid sequence encoding an GRAIN SIZE ON CHROMOSOME 3 (GSE3) polypeptide as defined in SEQ ID NO: 1 or a homolog orfunctional variant thereof, preferably wherein the nucleic acid sequence is operablylinked to a regulatory sequence, most preferably a promoter. 25 By “increasing grain size” or “increase seed size” it is meant an increase in mean grainor seed size by at least 5%, 10%, 12%, 15%, 17%, 20%, 22%, 25%, 27%, 30%, 32%, 35%, 37%, 40%, 42%, 45%, 47% or 50%, relative to the grain or seed size of a controlor wild-type of a plant. Grain (or seed) size may encompass one or more of grain (or 30 seed) length, width or thickness. Preferably, at least grain (or seed) thickness is increased, although as shown in Figure 11, increasing the expression or activity of GSE3 alters grain / seed length, grain weight and grain / seed thickness. Nonetheless, grain / seedthickness may be considered the most important characteristic of the genetically altered plants that allows for mechanised hybrid seed sorting. 35 M&C PC933520WOA 54 In a further final aspect of the invention, there is provided a method of screening a population of plants and identifying and / or selecting a plant that will have altered expression and / or activity of at least one gene or polypeptide, and therefore an alteration in grain or seed size in a plant, compared to a control or wild-type plant, the method 5 comprises detecting at least one polymorphism or mutation in a gene and / or promoter, wherein said mutation or polymorphism leads to an alteration in the level of expression and / or activity of the corresponding protein compared to the level in a plant not carrying said mutation or polymorphism or nucleic acid construct (e.g. a control or wild-type plant). Said mutation or polymorphism may comprise at least one insertion and / or at least one10 deletion and / or substitution. Preferably, the at least one gene is selected from GSE3,GS2, GRF3 and / or GRF4. In a most preferred embodiment, the method is a method of screening a population of plants and identifying and / or selecting a plant that will have decreased expression and / or 15 activity of at least one GSE3 gene or polypeptide, and a decrease in grain or seed size in a plant, compared to a control or wild-type plant, comprising detecting at least one polymorphism or mutation in said at least one GSE3 gene and / or promoter, wherein said mutation or polymorphism leads to a decrease in the level of expression and / or activity of the corresponding protein compared to the level in a plant not carrying said mutation20 or polymorphism or nucleic acid construct (e.g. a control or wild-type plant).In a further preferred embodiment, the method is a method of screening a population of plants and identifying and / or selecting a plant that will have decreased expression and / or activity of at least one GRF3 and GRF4 gene or polypeptide, and a decrease in grain or 25 seed size in a plant, compared to a control or wild-type plant, comprises comprising detecting at least one polymorphism or mutation in said at least one GSE3 gene and / or promoter, wherein said mutation or polymorphism leads to a decrease in the level of expression and / or activity of the corresponding protein compared to the level in a plant not carrying said mutation or polymorphism or nucleic acid construct (e.g. a control or30 wild-type plant). In a further preferred embodiment, the method is a method of screening a population of plants and identifying and / or selecting a plant that will have increased expression of at least one GSE3 gene or polypeptide, and an increase in grain or seed size in a plant, 35 compared to a control or wild-type plant, comprising detecting at least one polymorphism M&C PC933520WOA 55 or mutation in said at least one GSE3 gene and / or promoter, wherein said mutation or polymorphism leads to an increase in the level of expression of the corresponding protein compared to the level in a plant not carrying said mutation or polymorphism or nucleicacid construct (e.g. a control or wild-type plant).5 In a further preferred embodiment, the method is a method of screening a population of plants and identifying and / or selecting a plant that will have increased expression of at least one GS2 gene or polypeptide, and an increase in grain or seed size in a plant, compared to a control or wild-type plant, comprising detecting at least one polymorphism 10 or mutation in said at least one GS2 gene and / or promoter, wherein said mutation or polymorphism leads to an increase in the level of expression of the corresponding protein compared to the level in a plant not carrying said mutation or polymorphism or nucleic acid construct (e.g. a control or wild-type plant). 15 Suitable tests for assessing the presence of a polymorphism would be well known to the skilled person, and include but are not limited to, Isozyme Electrophoresis, Restriction Fragment Length Polymorphisms (RFLPs), Randomly Amplified Polymorphic DNAs (RAPDs), Arbitrarily Primed Polymerase Chain Reaction (AP-PCR), DNA Amplification Fingerprinting (DAF), Sequence Characterized Amplified Regions (SCARs), Amplified 20 Fragment Length polymorphisms (AFLPs), Simple Sequence Repeats (SSRs-which are also referred to as Microsatellites), and Single Nucleotide Polymorphisms (SNPs). In one embodiment, Kompetitive Allele Specific PCR (KASP) genotyping is used. The method may also comprise the step of assessing whether the polymorphism has an25 effect on grain or seed size distribution as described herein. Methods to screen for aneffect on grain or seed size distribution are provided in the examples below. The method may further comprise selecting one or more plant cells or plants for propagation. The selected plants may be propagated by a variety of means, such as by 30 clonal propagation or classical breeding techniques. For example, a first generation (or T1) transformed plant may be selfed and homozygous second-generation (or T2) transformants selected, and the T2 plants may then further be propagated through classical breeding techniques. The generated transformed organisms may take a variety of forms. For example, they may be chimeras of transformed cells and non-transformed35 cells; clonal transformants (e.g., all cells transformed to contain an expression cassette); M&C PC933520WOA 56 grafts of transformed and untransformed tissues (e.g., in plants, a transformed rootstock grafted to an untransformed scion). The invention is now described in the following non-limiting examples: 5 EXAMPLES Example 1: The male-sterile line (XQA) with an ideal small-grain size allele allows for fully mechanized hybrid rice breeding 10 We have previously developed a super hybrid rice TYHZ (Tianyouhuazhan), which has been widely cultivated in China over more than 1.5 million hectares annually during the past decades. Tianfeng A (TFA), Tianfeng B (TFB), and Huazhan (HZ) are the cytoplasmic male-sterile line, the maintainer line, and the restorer line of hybrid rice15 TYHZ, respectively. To discover an ideal grain size gene / allele for breeding ideal small-grain maintainer / male-sterile lines, we crossed the maintainer line TFB with a range ofrice varieties with small grains and tried to breed ideal small-grain maintainer / sterile lines. Only by crossing a small-grain japonica / geng variety Xiaoligeng (XLG) with TFB, we successfully bred an ideal small-grain maintainer line Xiaoqiao B (XQB) through the 20 pedigree method, indicating that XLG contains an ideal small-grain allele. The grains of XQB were obviously small compared with those of TFB while the tiller number and grain number per panicle of XQB were significantly increased in comparison to those of TFB (Fig.1b-d), resulting in the increased grain number per plant (Fig.1e). By contrast, the plant height, panicle length and growth period of XQB were similar to those of TFB 25 (Fig.1b and Extended Data Fig.1b, g-h). We then backcrossed XQB with TFA and bred a new cytoplasmic male-sterile line (XQA, Xiaoqiao A). When pollinated with XQB, XQA exhibited similar phenotypes to XQB. These results suggest that XQA is an ideal small-grain male-sterile line for mechanized F1 seed production of hybrid rice because XQA showed obviously small grains and increased grain number per plant.30 We previously crossed a large-grain indica variety Kuangsijiadi with the restorer line HZand bred a restorer line DHZ (Da huazhan). The grain length, grain width, grain thickness, and grain weight of DHZ were significantly increased compared with those of HZ (Fig. 1f-j). By contrast, plant height, tiller number, grain number per panicle, and grain number M&C PC933520WOA 57 per plant of DHZ are comparable with those of HZ. These results suggest that DHZ can be used as the restorer line with large grains for mechanized hybrid seed production.We tested whether F1 seeds from the male-sterile line XQA pollinated with DHZ can 5be separated from DHZ seeds. The male-sterile line XQA pollinated with the restorer lineDHZ formed smaller and lighter grains than TFA pollinated with HZ. Grain length, grain width, grain thickness, and grain weight of F1 hybrid seeds from the XQA × DHZ combination were decreased compared with those of F1 seeds from the TFA × HZ combination (Fig. 1f-j). These results suggested that XQA maternally influences grain 10 size, and the small-grain phenotype of XQA is determined by a recessive allele(s). Considering that the value of grain thickness was smaller than the values of grain length and width, the grain thickness is the determinant for separating F1 seeds from mixed harvest (Fig.1k). We compared the grain thickness of TFA (pollinated with HZ), XQA (pollinated with DHZ), HZ and DHZ and found that there is an obvious gap between XQA15 and DHZ in grain thickness (Fig. 1j), suggesting that small F1 hybrid seeds of the XQA× DHZ combination could be separated from the mixed harvest grains of XQA and DHZ.We therefore planted DHZ in the mixture with XQA and conducted separation experiments. We designed the grain thickness-based sorting sifter as follows. The holes20 of the sifter are rectangular, the theoretical width of sifter aperture is between the grain thickness of the restorer line and the grain thickness of the male-sterile line, and thetheoretical length of sifter aperture is more than the length of the male-sterile line. Grains from mixed planting DHZ and XQA were harvested and then separated using a sifter with 2cm aperture length (far exceeding grain length) and different aperture widths. The25 purity of F1 hybrid seeds gradually declined with increasing sieve aperture width (Fig.1l). When the width of the sieve aperture was less than 2.08 mm, the purity of F1 hybrid seeds was more than 96.01%, which meets the purity standard of F1 hybrid seeds forcommercial production (Fig.1l). According to the Hybrid Rice Seed Standard (GB4404.1- 2008, China, https: / / openstd.samr.gov.cn / bzgk / gb / ), the minimum purity for hybrid seeds 30 is 96%. Hybrid seeds produced by traditional methods usually have a purity range of 96%-98% 32,33. When the width of the sieve aperture was larger than 1.96 mm, the loss ratio of hybrid seeds was less than 3.21% (Fig.1l). These results indicated that a sieve with less than 2.08 mm aperture width can be used to efficiently separate small F1 hybrid seeds from the bulked seed harvested from a mixed planting XQA and DHZ during the 35 F1 seed production. Thus, these field trials demonstrate that the ideal male-sterile line M&C PC933520WOA 58 (XQA) and the restorer line (DHZ) allow fully-mechanized hybrid rice breeding andeconomically effective fully-mechanised hybrid seed production. Example II: The XQA × DHZ combination dramatically increases hybrid seed 5 number Considering that the ideal small-grain male-sterile line should not have negative effects on F1 hybrid seed number and hybrid rice yield, we investigated the F1 hybrid seed number and yield per plot of the XQA × DHZ combination and the TFA × HZ combination.10 As shown in Fig.1m, the F1 seed number per plot of the XQA × DHZ combination was increased by 22.22% compared with that of the original TFA x HZ combination, although the F1 hybrid seed yield of the XQA × DHZ combination was decreased by 18.7% in comparison to that of the TFA × HZ combination (Fig.1n). Considering that seed numberis a determinant for commercial hybrid seed production, the XQA × DHZ combination 15 dramatically increases the efficiency of hybrid seed production. We also found that the F1 hybrid seed yield per plot of the XQA × DHZ combination using mixed planting and traditional approach was similar, indicating that the mixed planting does not affect the pollination efficiency of the population in the field. 20 We further investigated the grain yield of the elite hybrid rice TYHZ from the TFA × HZ combination and the improved hybrid rice variety (XQHZ) from the XQA × DHZ combination. As shown in Fig.1o-q, the grain yield of hybrid rice XQHZ (heterozygous plants) was comparable with that of original hybrid rice TYHZ, although the male-sterile line XQA has small grains and the decreased grain yield, and DHZ had large grains.25 Similarly, we found no significant difference in the grain yield of hybrid rice from XQA x HZ combination and TYHZ (TFA x HZ combination). It is plausible that the male-sterile line XQA possesses the recessive small-grain mutation that did not influence grain size in hybrid rice (heterozygous plants). Considering that the genetic basis of crop heterosis is complicated, including dominance, overdominance and epistasis, it is reasonable that 30 DHZ did not increase the yield of hybrid rice XQHZ, although DHZ formed large grains. Taken together, these field trials demonstrated that XQA is an ideal male-sterile line for mechanized hybrid seed production, because XQA can increase hybrid seed number and does not decrease hybrid rice yield. M&C PC933520WOA 59 Example III: Identification of ideal small-grain size alleles of GSE3 in XLG, XQA, and XQB varieties and the mutant from large scale mutagenesis screens. To identify the mutation that is responsible for the small-grain phenotype in the male- 5 sterile line XQA and the maintainer line XQB, we crossed a japonica variety Zhonghua 11 (ZH11) with XQB and mapped a major locus for GRAIN SIZE ON CHROMOSOME 3(GSE3). We then generated the near-isogenic line ZH11-GSE3small grain in ZH11background. The grains of ZH11-GSE3small grainwere obviously shorter, narrower, and thinner than ZH11, resulting in the decreased grain weight (Fig.2a, d-g). The tiller 10 number and grain number per panicle of ZH11-GSE3small grainwere increased compared with those of ZH11 (Fig.2b-c and h-i). Seed number per plant and per plot of ZH11-GSE3small grainwere significantly increased by 25.3% and 24.3%, respectively, compared with those of ZH11, although the grain yield per plot of ZH11-GSE3small grainwas reduced by 5.7% 247 in comparison to ZH11 (Fig.2j-l). Thus, the effects of ZH11-15 GSE3small grainon seed number and yield were similar to those observed in XQA and XQB. To identify the GSE3 gene, we crossed ZH11-GSE3small grainwith ZH11 and generated an F2 population. The genomic DNAs from F2 plants with small grains were pooled for whole genome resequencing, and the ZH11 was sequenced as a control. The20 sequencing data and MutMap analyses were conducted. The SNP / INDEL-index in thepooled F2 plants was calculated in the whole genome. Five SNPs and one INDEL onchromosome 3 were linked to the small-grain phenotype ofZH11-GSE3small grain. Among them, only the DEL1 variant occurs within the exonic region and had an SNP / INDEL-index = 1 (Supplementary table 1). We further confirmed this deletion using PCR25 amplification (Fig.2m). This DEL1 (10 bp deletion) happens in the third exon of theLOC_Os03g55530 gene, which leads to the frame shift mutation (Fig.2n-o and ExtendedData Fig. 5), suggesting that the LOC_Os03g55530 is the candidate gene for GSE3.To confirm that LOC_Os03g55530 is the GSE3 gene, we then conducted the genomic30 complementation test. The genomic fragment of LOC_Os03g55530 (gGSE3) includingthe 3256 bp promoter, the GSE3 gene, and the 1208 bp 3’ sequence was transformed into ZH11-GSE3small grain. We measured the grain length, grain width, grain thickness,thousand grain weight, the tiller number, and grain number per panicle of gGSE3;ZH11-GSE3small graintransgenic plants. Our results showed that the genomic 35 fragment of GSE3 can complement the phenotypes of ZH11-GSE3small grain(Fig. 2p-t). M&C PC933520WOA 60 These genomic complementation tests also supported that the phenotypes of ZH11-GSE3small grain were solely caused by the GSE3 mutation. We further generated the lossof function mutants of GSE3 in the ZH11 background using the CRISPR / Cas9technology. The gse3-cri1 had a single base T insertion that causes a frame-shift, and5 gse3-cri2 had 30 bp base pair deletion (Fig.7a-c). The gse3-cri1 and gse3-cri2 producedsmall grains compared with ZH11, while the tilling number and grain number per plant ofgse3-cri1 and gse3-cri2 were increased (Fig. 7d-j). Taken together, these resultsdemonstrated that the GSE3 gene is LOC_Os03g55530. We further sequenced theGSE3 gene in XLG, XQA, XQB, TFA, and TFB and found that XLG, XQA, and XQB10 contain the same deletion as detected in ZH11-GSE3small grain, while TFA and TFB do not own this deletion (Fig. 2n). We also examined indica varieties and 80 japonica varietiesand did not detect this deletion, indicating that the GSE3small grain allele is a rare allele.Consistent with this, we analyzed 3024 rice genome (www.rmbreeding.cn) and foundnone of these rice varieties contain the GSE3small grain mutation. Considering that the15 small-grain phenotype of GSE3small grain is the adverse trait for conventional rice breedingalthough it is beneficial for mechanized hybrid seed production, it is reasonable that theGSE3 small-grain allele has not been selected by rice breeders.We simultaneously adopted the large-scale mutagenesis strategy to identify ideal grain20 size genes for breeding ideal small-grain male-sterile lines. Since 2008, we treated japonica varieties Zhonghua11 (ZH11) and Zhonghuajing (ZHJ) with ethylmethanesulfonate (EMS) for nine times and totally grew M2 populations about 20 hectares. In these screens, we isolated 609 small-grain mutants, and most of them showed defects in key agronomic traits, such as small and short plants, small panicles,25 the reduced grain number, the decreased tiller number or small leaves. Fortunately, only m238 mutant exhibited obviously small grains, increased tiller number, and elevatedgrain number (Fig.8a-i), suggesting that m238 is an ideal small-grain mutant for breedingideal male-sterile lines in mechanized hybrid seed production. We further identified themutation in m238 using the MutMap approach.5 SNPs and 3 INDEL on Chromosome 330 were linked to the small-grain phenotype of m238.SNP5 happens in the first exon of the LOC_Os03g55530 / GSE3 gene, resulting in a substitution of serine with proline atposition 68 of GSE3 (Fig.8j). In addition, the m238 phenotypes were similar to those ofgse3 alleles (Fig.2a-l and Extended Data Fig.8a-i). To further confirm that the m238 isa new allele of GSE3, we crossed the m238 allele with ZH11- GSE3small grain and35 investigated the phenotypes of ZH11-GSE3small grain / m238 F1 plants. M&C PC933520WOA 61 The grains of ZH11-GSE3small grain / m238 F1 plants showed similar phenotypes to ZH11-GSE3small grain and m238 mutants, supporting that m238 is allelic to ZH11-GSE3small grain.The result of large-scale mutagenesis screens also suggested that there is no too many5 genes whose single gene mutations can be used for breeding ideal small-grain male-sterile line. GSE3 encodes a nuclear-localized N-acetyltransferase-like protein that influences histone acetylation levels. GSE3 encodes a N-acetyltransferase-like protein with the10 conserved GNAT motif (Fig.2o). The last 93 amino acids in the C terminus were changedin GSE3 small-grain, indicating that the GSE3small grain is a loss of function allele. Asshown in Fig. 9, homologs of GSE3 were found in rice, maize, soybean, Arabidopsis,Brachypodium distachyon, Setaria italica and other plant species. We aligned GSE3 andits homologs and observed that the serine (Ser) residue at position 68 of GSE3 was15 conserved (Fig. 10). A substitution of serine with proline of GSE3 occurred at thisconserved position in m238 (Fig. 8j). We then examined the transcript levels of GSE3using quantitative real-time PCR. The GSE3 gene was expressed in roots, leaves, anddeveloping panicles. To further observe tissue-specific expression patterns of GSE3, wegenerated pGSE3:GUS construct and transformed it into ZH11 background. We20 detected the GUS (β-glucuronidase) activity in the developing spikelet hulls, panicles, young seedlings, and roots of pGSE3:GUS transgenic plants. To determine thesubcellular localization of GSE3, we generated p35S:GFP-GSE3 construct andtransiently expressed in rice protoplast. GFP fluorescence was observed in the nuclei and co-localized with DAPI staining signals, indicating that GSE3 is a nuclear-localized25 protein. Considering that GSE3 encodes an GCN5-related N-acetyltransferase-like protein that is localized in nuclei, and GCN5 related proteins have been described to influence histone acetylation, we asked whether GSE3 could bind histones and influence histone30 acetylation levels in rice. The pull-down assay showed that GSE3 strongly interactedwith biotin-labeled histone H2A, H3, and H4 peptides (Fig. 2u), indicating SE3 can bindhistones. We then detected the overall acetylation levels of histone H4 in ZH11 andZH11-GSE3small grain panicles using the Cleavage Under Targets and Tagmentation(CUT&Tag) method. As shown in Fig.2v, the levels of H4Ac (pan-acetyl) in ZH11-GSE3 M&C PC933520WOA 62 small-grain panicles were obviously lower than those in ZH11 panicles, indicating thatGSE3 influences overall acetylation levels of H4. Example IV: Overexpression of GSE3 increases grain size 5 To further uncover the functions of GSE3 in the control of grain size, we overexpressed GSE3 in the ZH11 background and generated the pActin:GSE3 transgenic plants.Expression levels of GSE3 in pActin:GSE3 transgenic lines were higher than that inZH11. The pActin:GSE3 transgenic lines formed long and thick grains compared withZH11, while the width of pActin:GSE3 grains was similar to that of ZH11 (Figure 11a, c-10 e). The grain weight of pActin:GSE3 transgenic lines obviously increased compared withthat of ZH11 plants (Figure 11). The plant height, leaf length, and leaf width of pActin:GSE3 transgenic lines were comparable to those of ZH11. The tilling number andgrain number per panicle of pActin:GSE3 transgenic lines were significantly reducedcompared with those of ZH11, while the panicle length of pActin:GSE3 transgenic lines15 was increased compared with that of ZH11 (Fig. 12b,f- h). These results revealed thatGSE3 positively regulates grain size and negatively influences grain number, suggestinga compensation mechanism between grain size and grain number.Example V: GSE3 physically interacts with GS2 20 To understand how GSE3 regulates grain size, we performed yeast two-hybrid screen to identify GSE3-interacting proteins. GSE3 was fused with the GAL4 DNA-binding domain (BD). Interestingly, we identified several clones corresponding to the transcription factor GS2 / OsGRF4 in this screen. The GS2AA allele has been reported to increase grain size and weight as well as nitrogen utilization efficiency. We further25 confirmed that the full length GSE3 can interact with the full length GS2 in yeast cells(Fig. 3a). To verify which regions of GSE3 are responsible for the interaction of GSE3with GS2, GSE3 was divided into three fragments, including GSE3-N (1-84), GSE3-GNAT (85-177) and GSE3-C (178-414). GSE3-GNAT (85-177) and GSE3-C (178-414) interacted with GS2, but not the GSE3-N (1-84). Similarly, the human ERG1 (ETS-30 related gene 1) can interact with both N- terminal region and C-terminal region of ETS2GS2 has two conserved and functional domains including the QLQ domain that mediatesprotein-protein interaction and the WRC domain that is involved in transcriptionalregulation. GS2 was divided into five segments, including GS2-N-QLQ (1-112), GS2-QLQ (50-112), GS2-WRC (113-181), GS2-C (169-394) and GS2-QLQ-WRC-C(50-394).35 Only the segments that contain the QLQ domain can interact with GSE3. We then tested M&C PC933520WOA 63 whether GSE3 could interact with GS2 in vitro. GST-GS2 and MBP-GSE3 fusion proteins were expressed in E. coli, respectively. MBP-GSE3 was pulled down by GS2-GSTimmobilized on glutathione sepharose beads (Fig. 3b), indicating that GSE3 physicallyinteracts with GS2. We further performed co- immunoprecipitation assay by transiently5 expressing p35S:GFP-GSE3 and p35S:MYC- GS2 in N. benthaminana leaves. Totalproteins were immunoprecipitated with GFP- Trap-Agarose, and the immunoblot wasdetected with anti-GFP and anti-MYC antibodies. MYC-GS2 was detected in theimmunoprecipitated GFP-GSE3 complex, but not in the negative control GFP (Fig. 3c).These results demonstrated that GSE3 interacts with GS2 using several approaches.10 Example VI: GSE3 acts genetically with OsGRF3 / 4 to control grain size and weight We have previously identified the gain-of-function allele GS2AA, which produces large grains. Here we asked whether loss-of-function mutant GS2 / OsGRF4 could form small15 grains. We therefore generated the loss-of-function mutant grf4-cri1 in the ZH11background using CRISPR / Cas9 technology, but the size of the grf4-cri1 mutant grainswas similar to that of ZH11 grains. We then knocked out both GRF4 and its closesthomolog GRF3 in the ZH11 background. The grf3-cri1 grf4-cri2 had a single base Ainsertion in GRF3 and a single base T deletion in GRF4 that cause frame-shift in both20 genes (Fig. 13a). The grf3-cri1 grf4-cri3 had a single base A insertion in GRF3 and asingle base C deletion in GRF4 that cause a frame-shift in both genes (Fig. 13a). Wefound that grf3-cri1 grf4-cri2 and grf3- cri1 grf4-cri3 double mutants produced smallgrains compared with ZH11 (Fig.13b-e), indicating that OsGRF4 and OsGRF4 functionredundantly to control grain size.25 Considering that GSE3 interacts with GS2 / OsGRF4, and OsGRF4 and OsGRF3 functionredundantly to regulate grain size, we asked whether GSE3 can also physically interactwith OsGRF3. As we expected, GSE3 interacted with OsGRF3 in yeast cells. We then investigated whether GSE3 can function genetically with GS2 / OsGRF4 and OsGRF3 to30 control grain size. To test this, we crossed pActin:GSE3 #1 with grf3-cri1 grf4-cri2 andisolated pActin:GSE3 #1; grf3-cri1 grf4- cri2 plants (Fig. 3d). Expression level of GSE3in pActin:GSE3 #1; grf3-cri1 grf4-cri2 plants was similar to that in pActin:GSE3 #1 (Fig.3e). The grain length and weight of pActin:GSE3 #1 were increased in comparison tothat of ZH11, while the length, width and weight of pActin:GSE3 #1; grf3-cri1 grf4-cri235 grains were similar to those of grf3- cri1 grf4-cri2 grains (Fig. 3f-h). These genetic M&C PC933520WOA 64 analyses indicated that the increased grain size and weight of pActin:GSE3 #1 almostdepend on the functional OsGRF3 / 4 and also suggested that GSE3 act genetically withOsGRF3 / 4 to control grain size and weight.5 Previous studies showed that the gain-of-function allele GS2AApromotes cell expansion but also increases cell proliferation in the grain hull. Considering that GSE3 acts genetically with OsGRF3 / 4 to control grain size, we performed morphological and cellularanalysis for the outer surface of ZH11, the loss of function of grf3-cri1 grf4- cri2, andZH11-GSE3small grain spikelet hulls. The results showed that GRF3 and GS2 / GRF410 influence both cell expansion and cell proliferation in spikelet hulls. Consistent with this, GSE3 also influences both cell expansion and cell proliferation.Considering that the transcriptional regulator GSE3 interacts genetically and physically with the transcription factor GS2 to control grain size, we speculated that GSE3 and GS215 could have co-regulated genes involved in grain growth. To test this, we performed RNA-sequencing of young panicles from ZH11, ZH11-GSE3small grain and grf3-cri1 grf4-cri2.RNA-seq data showed that 39.1% of the down-regulated genes in ZH11- GSE3small grainexhibited decreased expression in grf3-cri1 grf4-cri2 (Fig. 3i). Similarly, 16.7% of up-regulated genes in ZH11-GSE3small grainshowed increased expression in grf3-cri1 grf4-20 cri2 (Fig. 3j; Supplementary data 1). These results further supported that GSE3 andGRF3 / GRF4 have overlapped functions. Example VII: GS2 / OsGRF4 recruits GSE3 to influence histone acetylation levels of grain size genes25 Considering that GSE3 regulates the histone acetylation levels, and GS2 is a transcription factor, we asked whether the interaction between GS2 and GSE3 could regulate the expression of downstream grain size genes by influencing histone acetylation. To test this, we firstly conducted a Cleavage Under Targets and30 Tagmentation (CUT&Tag) assay using young panicles from pGS2:GS2AA-GFP transgenic lines to identify the target genes of GS2. CUT&Tag data showed that a largenumber of target genes (about 60%) are overlapped in two biological replicates, indicating the reproducibility of the method. The GS2- binding sites were mainly locatedin the promoter regions, which accounted for about 60% of all the peaks in both35 replicates. We then identified the GS2 binding motifs using the meme-chip method. The M&C PC933520WOA 65 CTGACA motif was the most enriched, and Arabidopsis and rice GRF transcriptionfactors have been reported to bind the CTGACA motif 47-49. The additional sequencesof GTGGGNCC were also enriched (Fig. 4a, Extended Data Fig. 23). The sequencesGTGGGNCC have been reported to be the binding motif of TCP family transcription5 factors. We further performed DNA affinity purification sequencing (DAP-seq) in vitro toidentify the binding motifs of GS2. DAP-seq data showed that about 63% of GS2 bindinglocated in the promoter regions and GS2 was mainly bound to the TSS region consistentwith CUT&Tag results. Five putative binding motifs of GS2 were identified Among them,the top one is the most conserved sequence CTGACA (Fig.4b), which is identical to the10 sequence identified from the CUT&Tag analysis. Toconfirm these putative binding sequences of GS2, we conducted electrophoretic mobility shift assay (EMSA) in vitro.As shown in Fig. 4c-d, GST-GS2 fusion proteins bond to the probe containing theCTGACA motif, but not bond to GTGGGCCC sequences. These results demonstratedthat GS2 can bind the CTGACA motif. 15 We then performed RNA-sequencing of young panicles from ZH11 and ZH11- GSE3smallgrain and found that expression of several genes involved in grain size and grain numbercontrol were changed, such as XIAO, OsBZR1, RGG2, OsER1, OsGRF1, GL6, SRS5,DEP1, PGL1, OsGA20ox1, OsMADS5, and OsMADS15 (data not shown). We further20 compared this RNA-seq data and the CUT-Tag data of GS2 and identified 429overlapping genes (data not shown). Histone acetylation is related to transcriptionalactivation in plants and animals. To confirm this, we selected genes that have beenreported to regulate grain size in rice, including XIAO, GW6, and OsBZR1. Mutations inXIAO, GW6, or OsBZR1 caused small grains. 25 As shown in Fig. 4e, the CUT&Tag data showed that the binding peaks of GS2 were located in the promoter regions of XIAO, GW6, and OsBZR1. We then used pGS2:GS2AA-GFP lines to perform ChIP-qPCR assay and found that GS2 canassociate with the promoter regions of XIAO, GW6, and OsBZR1 that contain the30 CTGACA motif (Fig. 4f). The RNA levels of those genes were significantly decreased inyoung panicles in grf3-cri1 grf4-cri2 compared with those in ZH11 (Fig.4g). These resultsindicated that GS2 associates the promoter regions of these genes and activates theirexpression. M&C PC933520WOA 66 We also examined the expression levels of those genes in ZH11-GSE3small grain. As shownin Fig.4h, expression levels of those genes were relatively reduced in ZH11- GSE3smallgrain in comparison to those in ZH11. Previous studies have reported that histonemodifiers can be recruited by transcription factors to the promoter regions of their target5 genes and regulate gene expression. We therefore asked whether GSE3 could associatewith the promoter regions of those genes through GS2 because GSE3 has no predictedDNA-binding domain. As shown in Fig.4i, ChIP-qPCR results confirmed the associationsof GSE3 with the promoter fragments of those genes in rice protoplasts. Furthermore,we detected the reduced associations of GSE3 with the promoter regions of those genes10 in grf3-cr1 grf4-cri2 rice protoplasts compared with those in ZH11 protoplasts, indicatingthat the associations of GSE3 with the promoter regions of those genes, at least in part,depend on the functional GS2 / OsGRF4 and OsGRF3.We further detected the histone H4 acetylation levels in the promoter regions of those15 selected genes in ZH11 and ZH11-GSE3small grainyoung panicles using ChIP-qPCR. The enrichment of H4Ac (pan-acetyl) in ZH11-GSE3small grain young panicles was reduced inthe promoter regions of those selected genes compared with that in ZH11 young panicles. Thus, these results suggested that the transcription factors GS2 / OsGRF4 andOsGRF3 recruit GSE3 to the promoter regions of their targetgenes and regulates their20 expression by influencing H4Ac (pan-acetyl) levels in the promoter regions of targetgenes. Example VII: Genome editing of GSE3 immediately improves elite male-sterile lines (TFA and Y58S) for mechanized production of hybrid seeds in three-line and25 two-line systemsConsidering that the male-sterile line XQA with the GSE3small grainallele can be used for the mechanical sorting of F1 hybrid seeds, we speculated that GSE3 could be utilized to improve other elite male-sterile lines for the mechanized production of F1 hybrid seeds. 30 However, it is time-consuming and labour-intensive using the conventional rice breedingapproach to introgress GSE3small grain to these elite male-sterile lines. To rapidly improvemale-sterile lines, we developed a one-step method to generate the loss-of-function mutants of GSE3 in the male-sterile line TFA and the maintainer line TFB using35 CRISPR / Cas9 technology. This strategy can simultaneously knock out the GSE3 gene M&C PC933520WOA 67 in any rice male-sterile lines and maintainer lines. The TFBgse3-cri3 and TFAgse3-cri3 have asingle T insertion that causes a frame-shift, respectively (Fig.14). The TFBgse3-cri3produced small grains compared with TFB, while the tiller number and grain number perpanicle of TFBgse3-cri3 were significantly increased in comparison to those of TFB (Fig.5 15a-i). By contrast, the plant height, panicle length, leaf length and leaf width of TFBgse3-cri3were similar to those of TFB (Fig.15a-b, j-m). The grain yield per plant of TFBgse3-cri3slightly decreased compared with that of TFB (Fig. 15n). We tested whether F1 seedsfrom the male-sterile line TFAgse3-cri3 pollinated with DHZ could be separated from DHZ seeds. The male-sterile line TFAgse3-cri3 pollinated with the restorer line DHZ formed10 smaller and lighter grains than TFA pollinated with HZ (Fig. 5a-b, Fig.16). Grains frommixed-grown DHZ and TFAgse3-cri3 were harvested and then separated using a sieve withdifferent aperture widths. When the width of the sieve aperture was 2.04 mm, the purityof F1 hybrid seeds was 96.04%, whereas the loss ratio of hybrid seeds was 5.76%, whichcan meet the purity standard of F1 hybrid seeds for commercial production (Fig. 5c).15 Then, we tested the F1 hybrid seed number per plot of the TFAgse3-cri3× DHZ combination and the TFA × HZ combination. The F1 seed number per plot of the TFAgse3-cri3 × DHZcombination was increased by 21.2% compared with that of TFA × HZ combination (Fig.5d), indicating that the TFAgse3-cri3 × DHZ combination dramatically increases theefficiency of F1 hybrid seed production. 20 Interestingly, we found that the F1 hybrid seed yield of the TFAgse3-cri3 × DHZ combinationwas similar to that of the TFA × HZ combination (Fig. 5e), but different from that of theXQA × DHZ combination (Fig.1n). This difference may attribute to the fact that grains ofXQA, which is a variety derived from TFA, were lighter than those of TFAgse3-cri3 (Fig. 1g).25 It is possible that XQA contains the GSE3small grain mutation as well as other mutations thatinfluence grain size. In the actual F1 hybrid seed production process, the seed settingrate of the sterile line also influences the yield of F1 hybrid seed. The seed setting rateof the male-sterile line is usually between 30% and 50%, rather than 100%. We foundthat the seed setting rate of the male-sterile lines TFAgse3-cri3 in the TFAgse3-cri3 × DHZ30 combinations was higher than their original combinations TFA × HZ. Next, weinvestigated the yield of hybrid rice TYHZ and the improved hybrid rice variety TYHZ(gse3-cri3) derived from the TFAgse3-cri3 × HZ combination. The grain yield of TYHZ (gse3-cri3) was comparable with that of TYHZ (Fig.5f-h). Similarly, the grain yield of hybrid rice from the TFAgse3-cri × HZ combination was comparable to that of TYHZ (Fig. 17). These35 results indicated that the disruption of the GSE3 gene by genome editing technology can M&C PC933520WOA 68 rapidly generate the male-sterile line with small grains for mechanized production of F1 hybrid seeds. We then asked whether it is still effective to achieve mechanized F1 hybrid rice seed 5 production in two-line hybrid rice varieties by editing of GSE3 in male-sterile lines. YLY900 (Y liangyou 900) is one of the top super hybrid rice varieties in China, and its yield is more than 15 t / hm269. Y58S and R900 were the thermosensitive male-sterile line and the restorer line of YLY900, respectively. We generated the loss-of-function10 mutant of GSE3 in the Y58S background using the CRISPR / Cas9 technology. The Y58Sgse3-cri4 has a single A insertion in the GSE3 gene, which causes a frame-shift (Fig.18). The Y58Sgse3-cri4 produced smaller and lighter grains than Y58S (Fig. 5i-j, Fig. 19),while the tiller number and grain number per panicle of Y58Sgse3-cri4 were significantlyincreased in comparison to those of Y58S (Fig. 20a-d). By contrast, the plant height,15 panicle length, leaf length and leaf width of Y58Sgse3-cri4 were similar to those of Y58S(Fig. 20a, e-h). The grain yield per plant of Y58Sgse3-cri4 was decreased compared withthat of Y58S (Fig. 20i). We found that there were dramatic differences in the thicknessof Y58Sgse3-cri4 and R900 grains (Fig. 5j). The thermosensitive male-sterile line Y58Sgse3-cri4 pollinated with the restorer line R900 also formed smaller and lighter grains than Y58S20 pollinated with R900, as gse3 alleles act maternally to influencing grain size (Fig.21).Then, we tested whether F1 hybrid seeds from the thermosensitive male-sterile lineY58Sgse3-cri4 pollinated with R900 could be separated from R900 seeds using the sieve.Grains from R900 and Y58Sgse3-cri4 were harvested and then separated using a sifter with25 different aperture widths. When the width of the sieve aperture was 1.96 mm, the purityof F1 hybrid seeds was 96.65%, whereas the loss ratio of hybrid seeds was 3.2%, whichcan meet the purity standard of F1 hybrid seeds for commercial production (Fig. 5k).These results indicated that the sieve with a 1.96 mm aperture width can be used toefficiently separate small F1 hybrid seeds from R900 seeds during the F1 seed30 production. Then, we tested the F1 hybrid seed number per plot of the Y58Sgse3-cri4 ×R900 combination and the Y58S × R900 combination. The F1 hybrid seed number perplot of the Y58Sgse3-cri4 × R900 combination was increased by 38.3% compared with thatof the Y58S × R900 combination, while the F1 hybrid seed yield of the Y58Sgse3-cri4 ×R900 combination was comparable with that of the Y58S × R900 combination (Fig. 5l-35 m). The seed setting rate of the male-sterile lines Y58Sgse3-cri4 in the Y58Sgse3-cri4 x R900 M&C PC933520WOA 69 combination was higher than their original combination Y58S x R900. The grain yield ofhybrid rice YLY900 (gse3-cri4) from the Y58Sgse3-cri4 × R900 combination wascomparable with that of YLY900 from the original Y58S × R900 combination (Fig. 5n-p).We also computed the standard deviation and coefficient of variance of grain area of5 YLY900 and YLY900 (gse3-cri4). The standard deviation and coefficient of variance ofgrain area of YLY900 (gse3-cri4) were similar with those of YLY900, indicating that gse3alleles does not influence the grain size uniformity in hybrid rice plant YLY900 (gse3-cri4) (heterozygous plant). Considering that gse3-cri4 is a recessive small-grain allele,hybrid rice YLY900 (gse3-cri4) (heterozygous plant) showed similar phenotypes to10 YLY900, indicating that the editing of GSE3 does not affect hybrid rice performance inthe field. Therefore, these results revealed that editing GSE3 can rapidly improve thecurrent elite male-sterile line Y58S for mechanized F1 hybrid seed production and alsodramatically increase the efficiency of F1 hybrid seed production.15 Example VII: Methods and Materials Plant materials and growth conditions Rice plants used in this study were listed in supplementary table 3. All rice plants were grown in standard paddy fields with a density of 20 cm × 20 cm in Changping (Beijing,20 China), Fuyang (Zhejiang province, China), or Lingshui (Hainan province, China). For rice protoplast assay, ZH11 and grf3-cri1 grf4-cri2 seedlings were grown in a growthchamber under dark conditions at 28℃ for 10 days. N. benthamiana plants were grownin the greenhouse at the Institute of Genetics and Developmental Biology for a duration25 of 16 hours of light and 8 hours of darkness. Morphological and cellular analysis Agronomic traits were measured at the maturity stage, including plant height, tilling number, panicle length, the number of primary and secondary branches, leaf length, leaf 30 width, and grain number per panicle. Mature grains from the main panicles were utilized for measuring grain length, width, thickness, and weight. A WSEEN grain analysis software (WSeen, China) was used to measure the grain length and width. Grain thickness was measured using a digital caliper (Jianye Tools, China). To calculate the grain weight, 100 mature grains were weighted using an electronic analytical balance 35 (Sartorius, Germany). Three replicates were weighted for each sample. M&C PC933520WOA 70 For outer epidermal cell size and cell number measurement of spikelet hulls, mature grains were observed with a scanning electron microscope (SEM) after gold spraying treatment. The cell length and cell width of central cells in the spikelet hulls were 5 measured by using ImageJ software, the grain number in the grain-length and grain- width direction were counted. Hybrid rice seed production and the separation test Hybrid seed production experiments were conducted in the paddy field at Lingshui 10 (Hainan province, China) where Chinese breeders usually produce F1 hybrid seeds. Briefly, hybrid seed production experiments consisted of several steps, including seeding, transplanting, pollinating, and harvesting. The sterile lines TFA, XQA, TFAgse3-cri3, Y58S, and Y58Sgse3-cri4, restore lines HZ, DHZ, and R900 were used, 25 days old seedlings were transplanted into the experiment field. The sterile line TFA, XQA, TFAgse3-15cri3, Y58S, Y58Sgse3-cri4and restore line HZ, DHZ, and R900 have different heading date in Lingshui. There are 15 days seeding interval for TFA after HZ seeded in TFA×HZ combination, 25 days seeding interval for TFAgse3-cri3after DHZ seeded in TFAgse3-cri3×DHZ combination, 7 days seeding interval for Y58Sgse3-cri4after R900 seeded in Y58Sgse3-cri4×R900 combination. The parents lines of XQA×DHZ and Y58S×R900 20 combination were seeded at the same time. The final spacing between rice plants were maintained at 20 cm. The TFA X HZ and Y58S X R900 combinations used a row ratio of 1:6 for restore lines and male sterile lines. The male sterile line needs to be grown near the restorer line in the alternative row for pollination. For the XQA × DHZ, TFAgse3-cri3× DHZ, and Y58Sgse3-cri× R900 combinations, restore lines and male sterile lines were 25 mixed transplanted at a ratio of 1:6. The plots are 3 m in length, and 2.4 m in width, with three replicates. Sieves with different diameters of sieve aperture, ranging from 1.84 mm to 2.12 mm with a 0.04 mm spacing, were utilized to separate hybrid rice seeds from mixed harvest.500 30 g mixed harvested seeds were used for the separation test. The purity and loss of ratio were counted. According to the separation test, When the width of the sieve aperture was 2.08 mm (XQA × DHZ combination), 2.04 mm (TFAgse3-cri3×DHZ combination) and 1.96 mm (Y58Sgse3-cri×R900 combination), respectively, the purity of F1hybrid seeds were more than 96% in those combinations, which can meet the purity standard of F135 hybrid seeds for commercial production. In the XQA × DHZ, TFAgse3-cri3×DHZ, and M&C PC933520WOA 71 Y58Sgse3-cri×R900 combination, the mixed harvested seeds were sifted using a sieve with an aperture size of 2.08 mm, 2.04 mm and 1.96 mm, respectively. The seeds that passed through the sifter were weighted to determine the grain yield per plot of F1hybrid rice seeds, and the number of F1hybrid rice seeds per plot were calculated. For the TFA × 5 HZ and Y58S × R900 combinations, the seeds from sterile lines were harvested and weighted to determine the grain yield per plot of F1hybrid rice seeds, and the number of F1 hybrid rice seeds per plot were calculated. To test the grain yield per plot of the F1 hybrid rice, the rice plants were grown in standard paddy fields with a space of 20 cm. Each plot had three replicates and contained 36 plants. 10 Constructs and plant transformation All constructs were generated either using the ClonExpress II One Step Cloning Kit (C112-1, Vazyme) or T4 DNA ligase (M0202S, NEB). The GSE3 genomic sequence wasamplified from ZH11 genomic DNA using primers C99-gGSE3-F and C99-gGSE3-R and15 cloned into pMDC99 vector to generate plasmid pGSE3:GSE3. The GSE3 promotersequence was amplified from ZH11 genomic DNA using primers C164-pGSE3-GUS-F and C164-pGSE3-GUS-R and cloned it into the pMDC164 vector. To generatep35S:GFP-GSE3 plasmid, the coding sequence of GSE3 was amplified with the primersC43-GSE3-F and C43-GSE3-R from cDNAs of ZH11 panicles and cloned into the20 pMDC43 vector. The GSE3 coding sequence was amplified using primers 1266-GSE3-F and 1266-GSE3-R and cloned into pw1266 vector to generate p35S:FLAG-GSE3 plasmid. To generate the MBP-GSE3 plasmid, the coding sequence of GSE3 wasamplified with the primers MBP-GSE3-F and MBP-GSE3-R and cloned into the pMAL- c2 vector. To generate the GST-GSE3 plasmid, the coding sequence of GSE3 was25 amplified with the primers GST-GSE3-F and GST-GSE3-R and cloned into the pGEX4T- 1vector. To generate BD-GSE3 plasmid, the coding sequence of GSE3 was amplifiedwith the primers BD-GSE3-F and BD-GSE3-R and cloned into the pGBKT7 vector. Togenerate p35S:MYC-GS2 plasmid, the coding sequence of GS2 was amplified with theprimers MYC-GS2-F and MYC-GS2-R and cloned into the pCAMBIA1300-221-MYC30 vector. To generate the GST-GS2 plasmid, the coding sequence of GS2 was amplifiedwith the primers GST-GS2-F and GST-GS2-R and cloned into the pGEX4T-1 vector. Togenerate the AD-GS2 plasmid, the coding sequence of GS2 was amplified with theprimers AD-GS2-F and AD-GS2-R and cloned into the pGADT7 vector. To generatepETnT-FLAG-GS2 plasmid, the coding sequence of GS2 was amplified with the primers35 pETnT-FlAG-GS2-F and pETnT-FlAG-GS2-R and cloned into the pETnT vector. To M&C PC933520WOA 72 generate SK-gRNA-GSE3-1 and SK-gRNA-GSE3-2 intermediate vectors, the primersGSE3-gRNA1-F and GSE3-gRNA1-R, GSE3-gRNA2-F and GSE3-gRNA2-R were annealed and insert into the linearized intermediate vector SK-gRNA (digest with AarI)using T4-DNA ligase. To generate pC1300-Cas9-GSE3-1 and pC1300-Cas9-GSE3-25 plasmids, The intermediate vectors SK-gRNA-GSE3-1 and SK-gRNA-GSE3-2 weredigested with Kpn I and Bgl II, the gRNA segments were recovered and insert intolinearized pC1300-Cas9 (digest with Kpn I and BamH I), respectively. To generate SK-gRNA-GRF3 and SK-gRNA-GRF4 vectors, the primers GRF3-gRNA-F and GRF3- gRNA-R, GRF4-gRNA-F and GRF4-gRNA-R were annealed and insert into the10 linearized intermediate vector SK-gRNA (digest with AarI) using T4-DNA ligase,respectively. To generate pC1300-Cas-GRF3-GRF4 plasmid, The gRNA segments from SK-gRNA-GRF3 (digested with Kpn I and Sal I) and SK-gRNA-GRF4 (digested with XhoI and Bgl II) were recovered and insert into the linearized pC1300-Cas9 (digest with KpnI and BamH I). The primers used for construct are shown in Supplementary table 4. The15 plasmids pC1300-Cas9-GSE3-1, pC1300-Cas-GRF3,GRF4, were transferred to Agrobacterium strain GV3101 and then transformed into ZH11 to generatecorresponding transgenic lines. The plasmid pC1300-Cas9-GSE3-2 was transformed into TFA, TFB, and Y58S. 20 RNA extraction, qRT-PCR, and RNA-seq analysis Total RNA was extracted from young panicles using the RNAprep pure kit (TIANGEN, DP439), according to the corresponding instructions. The HiScript II 1st Strand cDNA Synthesis Kit (Vazyme, R211) was used to produce the first strand cDNA, and followed the manufacturer's instructions. For qRT-PCR, we used 2×RealStar Fast SYBR qPCR 25 Mix (Genstar, A301-10) on a Lightcycler 480 (Roche, Switzerland). Data were normalized with ACTIN1. The primers used for qRT-PCR are shown in Supplementary table 4. For RNA-seq analysis, we extracted total RNA from young panicles (5 cm) of ZH11, 30 ZH11-GSE3small grainand grf3-cri1 grf4-cri2, each with three replicates. The libraries were constructed, and sequencing was performed by BerryGenomics Corporation on the Illumina NovaSeq 6000 platform using 150-bp double-end sequencing. The resulting clean reads were aligned to the Nipponbare reference genome (MSU7.0) using TopHat and Bowtie 2 software. 35 M&C PC933520WOA 73 GUS staining and subcellular localization analysis Seedlings, panicles and seeds at various developmental stages from pGSE3:GUS transgenic lines were soaked in GUS staining solution (each 100 mL contain X-Gluc: 0.075g, 1 M NaH2PO4: 4.23 mL, 1 M Na2HPO4: 5.77 mL , K3Fe(CN)6: 0.1g, 0.5 M EDTA: 5 2 mL, Nonibet–P40: 0.1 mL ). The samples were subjected to vacuum conditions for 30 minutes, followed by staining at 37°C for 12 hours. Chlorophyll was subsequently eliminated using 70% ethanol, and the process was repeated several times. Finally, the stained samples were photographed using either a digital camera or a microscope. p35S:GFP-GSE3 was transiently expressed in rice protoplast and used to observe the10 GFP fluorescence with Confocal microscopy (Zeiss LSM 710, Germany). The nuclei were stained with DAPI (1 μg mL−1). Yeast two-hybrid assay Matchmaker Gold Yeast Two-Hybrid System (Clontech) was used to conduct the yeast 15 two-hybrid assay following the manufacturer’s instructions. To construct an expression cDNA library for yeast two-hybrid screening, we utilized young panicles of ZH11 that were shorter than 5 cm. BD-GSE3 was used to screen the library. To check theinteraction region between GSE3 and GS2, GSE3 and GS2 were divided into several parts. To generate BD-GSE3-N, BD-GSE3-GNAT, and BD-GSE3-C plasmids, we20 amplified the corresponding segments of GSE3 by primers (Supplementary table 4), andclone to vectors pGBKT7, respectively. To generate AD-GS2-N-QLQ, AD-GS2-QLQ, AD-GS2-WRC, AD-GS2-C, and AD-GS2-QLQ-WRC-C plasmids, we amplified thecorresponding segments of GS2 by primers (Supplementary table 4), and cloned to vectors pGADT7, respectively. The yeast strain AH109 was used for transformation. 25 Pull-down assays The plasmids pGEX4T-1, pMAL-C2, GST-GS2, and MBP-GSE3 were each transferredinto BL21 cells. The proteins were then expressed in BL21 with 0.4 mM IPTG at 28°C for 3 hours. The BL21 cells were collected by centrifugation, the precipitates were30 resuspended in RS buffer (each 40 mL contains, 1.5 M NaCl: 4 mL, 1 M HEPES (pH7.5): 2mL, 10% TritonX-100:4 mL, 50% Glycerol: 8 mL, EDTA-free protease inhibitor cocktail: 1 tablet) and sonicated for 3 minutes (5 s on and 5 s off) at 20 amplitudes. Equal amounts of proteins were added in different combinations in 1 mL RS buffer. Then, 20 μL of Glutathione Sepharose 4B beads (GE Healthcare) were added to the above 35 samples and incubated at 4°C for 45 minutes with gentle rotation. The beads were then M&C PC933520WOA 74 collected through centrifugation at 700 rpm for 3 minutes and washed with RS buffer at least five times. Next, 40 μL of 2×SDS loading buffer were added to the beads. The beads were boiled at 95°C for 10 minutes to wash the proteins and prepare for western blotting. The proteins were detected using an anti-GST antibody (Abmart, MA9025, 5 dilution, 1:5000) and an anti-MBP antibody (NEB, E8032S, dilution, 1:10000) through immunoblotting. Histone pull-down assay To perform the experiment, 2 µg of GST-GSE3 fused proteins and 1 µg of biotinylated 10 histone peptides (including H2A, H2B, H3, and H4 peptides) were incubated together in peptide-binding buffer (50 mM Tris-HCl 7.4, 200 mM NaCl, 0.05% NP-40) for 30 minutes at 4°C with gentle rotation. Next, 10 µL of Streptavidin coupled agarose beads (Sigma, 69203-3) were added to the mixture and incubated at 4°C for another 30 minutes with gentle rotation. The beads were then washed with peptide-binding buffer three times at 15 4°C. Finally, the bound proteins were detected using an antibody to GST (Abmart, MA9025, dilution, 1:5000) through immunoblotting. Co-immunoprecipitation The plasmids pMDC43, p35S:GFP-GSE3, p35S:MYC-GS2 were transferred to20 Agrobacterium strain GV3101, respectively. The transgene Agrobacterium GV3101 werecollected and activated using an activation buffer (each 100 mL contains, 2 M MgCl2: 0.5 mL, 1 M MES (pH 5.7): 1 mL, 100 mM AS:150 μL) based on the different combinations. N.benthaminana leaves (4-5 weeks) were transformed by injection of AgrobacteriumGV3101 cells with different combinations. N.benthaminana leaves were collected and25 ground in liquid nitrogen, the total proteins were extracted using extraction buffer (each 20 mL contains, 1 M Tris-HCl (pH 7.4): 1 mL, 5 M NaCl: 600 µL, 50% glycerol: 8 mL, 10% Triton X-100: 4 mL, Complete protease inhibitor Tablets: 1 / 2 tablet, 100 mM PMSF:200 µL) and immunoprecipitated with 20µL GFP-Trap-Agrose (Chromotek, gta-20). The beads were then collected through centrifugation at 700 rpm for 3 min, and subsequently30 washed with the Wash buffer (each 20 mL contains, 1 M Tris-HCl (pH 7.4): 1 mL, 5 MNaCl: 600 µL, 50% glycerol: 8 mL, 10% Triton X-100: 200 µL, Complete protease inhibitor Tablets: 1 / 2 tablet) at least for four times. 40 μL 2×SDS loading buffer wereadded to the beads and boiled at 95°C for 10 min to wash the proteins and prepared for western blot. The proteins were then detected through immunoblotting using an anti- M&C PC933520WOA 75 GFP antibody (Abmart, M20004, dilution, 1:5000) and an anti-MYC antibody (Abmart, M20002, dilution, 1:5000), respectively. DAP-seq analysis 5 . We expressed and purified the FLAG-GS2 fused protein in BL21.2 μg purified fused FLAG-GS2 protein and 25 μL anti-Flag M2 Magnetic Beads (Sigma, M8803) were incubated in 400 μL PBS buffer for 1 hour at 25°C with slow rotation. Then, the beads were washed four times with PBS (containing 0.005% NP40) and then washed twice with PBS alone. The beads were resuspended with 40 μL PBS and incubated with 100 ng 10 rice genomic library for 1 hour at 25°C with slow rotation. Next, the beads were washed with PBS (containing 0.005% NP40) four times, followed by two washes with PBS only. The beads were resuspended in 25 μL Elution Buffer (10 mM Tris-HCl pH 8.5), and the DNA segments were eluted by incubating at 98°C for 10 min. Next, we use the eluted DNA as templates for PCR amplification (Primer A (forward primer) and Primer B 15 (reverse primer) listed in Supplementary table 4) to construct the library and for sequencing. The BerryGenomics Corporation sequenced the library on the Illumina NovaSeq 6000 platform using 150-bp double-end sequencing. The resulting clean reads were aligned to the Nipponbare reference genome (MSU7.0) using Bowtie 2 software. Peak calling was performed using MACS2 software, while motif enrichment was carried20 out using MEME-ChIP. CUT&Tag assay To identify genes with differential acetylation of H4Ac (pan-acetyl) between ZH11 and ZH11-GSE3small grain, nuclei were extracted from young panicles of ZH11 and ZH11- 25 GSE3small grainfor the CUT&Tag assay. 0.1 g of young panicles were ground in liquid nitrogen and resuspended in 2 mL of precooled nuclei extract solution I (10 mM Tris-Cl (pH 8.0), 0.44 M sucrose, 10 mM MgCl2, 0.1 mM PMSF, 5 mM β- Mercaptoethanol, andComplete protease inhibitor tablets (1 tablet for each 50 mL solutions)). The sampleswere filtered by using 40 μm cell sieve and nuclei were collected through centrifugations.30 The nuclei were then washed three times with nuclei extract solution II (10 mM Tris-Cl (pH 8.0), 0.25 M sucrose, 10 mM MgCl2, 0.1 mM PMSF, 5 mM β- Mercaptoethanol, andComplete protease inhibitor Tablets (1 tablet for each 50 mL solutions)). The nuclei wereresuspended with 1 mL nuclei extract solution I containing 0.15% Triton X-100 and counted using a blood counting chamber. About 10,000 nuclei were collected and 35 resuspended in 50 μL Antibody buffer (150 mM NaCl, 20 mM HEPEs (pH7.5), 2 mM M&C PC933520WOA 76 EDTA, 0.5 mM Spermine, 0.1% BSA). Then, 1 ul of H4Ac (pan-acetyl) (Active Motif, 39925) antibody was added to the nuclei and slowly rotated at 4°C overnight. The nuclei were washed with CUT-Wash buffer (0.5 mM Spermine, 150 mM NaCl, 20 mM HEPEs (pH7.5)) for twice, and then resuspended in 50 μL CUT-wash Buffer. Next, 0.5 μL of 5 Goat anti-Rabbit secondary antibody was added to the samples, and the mixture was rotated slowly at 4°C for 1 hour. The nuclei were washed with CUT-Wash buffer twice and resuspended in 100 μL of CT-300 Buffer (0.5 mM Spermine, 300 mM NaCl, 20 mM HEPEs (pH7.5)).1 μL of pG-Tn5 (4 pmol) was added to the nuclei, and the mixture was rotated slowly at 4°C for 1-2 hours. The nuclei were washed twice with CT-300 Buffer10 and resuspended in 300 μL of Tagmentation Buffer (10 mM MgCl2, 0.5 mM Spermine, 300 mM NaCl, 20 mM HEPEs (pH7.5)), and the samples were rotated slowly at 37°C for 1 hour. Finally, 2.5 μL 20 mg / mL protease K, 3 μL 10% SDS and 10 μL 0.5 M EDTA were added to the solutions in order, and incubated at 50°C for 1 hour. The DNA was extracted using the Phenol-chloroform method and dissolved in 25 μL of 1 × TE 15 solutions. The above DNA was used as templates for PCR amplification (Primer P5 (forward primer) and Primer P7 (reverse primer) listed in Supplementary table 4) to construct the library and for subsequent sequencing. The library was sequenced by BerryGenomics Corporation using Illumina NovaSeq 6000 platform with 150-bp double- end sequencing. The clean reads were aligned to the Nipponbare reference genome20 (MSU7.0) using Bowtie 2 software, and peak calling was performed using MACS2 software. To identify the target genes of GS2, 1 million nuclei from young panicles of pGS2:GS2AA-GFP transgenic plants were extracted for CUT&Tag assay. The anti-GFP antibody (Invitrogen, A11122 ) was used in the pGS2:GS2AA-GFP CUT&Tag assay.MEME-ChIP was used for motif enrichment. 25 Electrophoretic mobility shift assay GST-GS2 fusion proteins were expressed in Escherichia coli BL21 (DE3) strain and thenpurified using Glutathione Sepharose4B beads (GE Healthcare, 17-0756-01). The Light Shift Chemiluminescent EMSA kit (Thermo Fisher Scientific, 20148) was used to perform 30 electrophoretic mobility shift assay according to the manufacturer’s instructions. Chromatin immunoprecipitation and quantitative real-time PCR analysis The young panicles of transgenic plants pGS2:GS2AA-GFP and rice protoplasts withtransient expression of p35S:FLAG-GSE3 were collected for ChIP assay, respectively. 35 We ground samples in liquid nitrogen and isolated nuclei. The chromatin complexes were M&C PC933520WOA 77 extracted and then subjected to sonication (Bioruptor pico, Diagenode) to obtain an average size of about 300 bp. The anti-FLAG antibodies (Abmart, M20008) / anti-GFP antibodies (Invitrogen, A11122) and protein A+G beads (Millipore, 16-663) were used for Immunoprecipitations. The precipitated DNA was recovered by using the QIAGEN DNA 5 purification kit, and acted as the template for the subsequent qRT-PCR. References 10 Ashraf, M.F. Peng, G.; Liu, Z. Noman, A. Alamri, S. Hashem, M. Qari, S.H. Mahmoud al Zoubi, O. (2020) Molecular Control and Application of Male Fertility for Two-Line HybridRice Breeding. Int. J. Mol. Sci. , 21, 7868.Chen L, Liu YG (2014). Male sterility and fertility restoration in crops. Annu Rev Plant 15 Biol.65:579-606. Jiang H, Lu Q, Qiu S, Yu H, Wang Z et al (2022). Fujian cytoplasmic male sterility and the fertility restorer gene OsRf19 provide a promising breeding system for hybrid rice.Proc Natl Acad Sci U S A.23;119(34) 20 Melonek, J., Duarte, J., Martin, J. et al (2021). The genetic basis of cytoplasmic malesterility and fertility restoration in wheat. Nat Commun 12, 1036.Toriyama K. Molecular basis of cytoplasmic male sterility and fertility restoration in rice. 25 Plant Biotechnol (Tokyo).2021 Sep 25;38(3):285-295. Xu Y, Yu D, Chen J, Duan M (2023). A review of rice male sterility types and their sterility mechanisms. Heliyon. Jul 13;9(7). 30 SEQUENCE LISTING M&C PC933520WOA 78 5 M&C PC933520WOA 79 SEQ ID NO: 1 GSE3 wild-type. Underlined = region of difference with SEQ ID NO: 2.Bold = serine of m238 allele, conserved within many plants.MVVPVPTTFEGAREMEVVEVREYREDRDRAAVEEVERECEVGSSGGGEAKMCLFTD LLGDPLCRIRNSPAYLMLVAETANGGGGGNGREIIGLIRGCVKTVVSGGSVQAGKDPI 5 YSKVAYILGLRVSPRYRRKGVGKKLVGRMEEWFRQSGAEYSYMATEQDNEASVRLF TGRCGYSKFRTPSVLVHPVFGHALQPSRNAAIRKLEPREAELLYRWHFAAVEFFPADI DAVLSKELSLGTFLAVPAGTRWESVEAFMDAPPASWAVMSVWNCMDAFRLEVRGAP RLMRAAAVATRLVDRAAPWLKIPSIPNLFAPFGLYFLYGVGGAGPASPRLVRALCRHA HNMARKGGCGVVATEVSACEPVRAGVPHWARLGAEDLWCIKRLADGYNHGPLGD 10 WTKAPPGRSIFVDPREF SEQ ID NO: 2 GSEsmall grain.Underlined = region of difference with SEQ ID NO: 1 MVVPVPTTFEGAREMEVVEVREYREDRDRAAVEEVERECEVGSSGGGEAKMCLFTD LLGDPLCRIRNSPAYLMLVAETANGGGGGNGREIIGLIRGCVKTVVSGGSVQAGKDPI 15 YSKVAYILGLRVSPRYRRKGVGKKLVGRMEEWFRQSGAEYSYMATEQDNEASVRLF TGRCGYSKFRTPSVLVHPVFGHALQPSRNAAIRKLEPREAELLYRWHFAAVEFFPADI DAVLSKELSLGTFLAVPAGTRWESVEAFMDAPPASWAVMSVWNCMDAFRLEVRGAP RLMRAAAVATRLVDRAAPWLKIPSIPNLFAPFGLYFRRRRPGLPAARPRAVPPRPQH GPQGRLRRGRHRGLRLRARPRRRAALGAPRRRGPLVHQASRRRLQPRPARRLDQ 20 GAAGPLHLRRPKRVLGGF SEQ ID NO: 3 GSE3 promoter GCCAGGCATCGTGTTGAGTGCGTGGGGAGCTTGTATACACGGCGCCCTATCCCG AAATTTCCCCTTGTAACTCTCACCCTAGCGATTCAAAATCCCCCCATAGCCAGCGG 25 CCGCCGGCTCCTCGTCCTCATACGCCGCGCCGCTACCGCCGGCGCCAATGGCG CAAACCCTAATCACCGTCGGCGATGGGAGAGAGAGGGAAAGGGAGAGGGAGAG AGAGAGAGGGGAGGGAGAGGCGAACGGCGGCGGCGAGAGGGAGAGGGGAGAG GGAGAGAAGAAGGCGACGGCACTATCACCGTTGGAGGCGAGCGCCACCACTGTC TAGCTATAAAATCCCATTTGGGTAGCCCCTTAATTAACTATGTCTAGCTATACAAAT 30 GCCAAAGTGAAATCTGAGATCAATTCAGAAAACCACCGGGTCCCTTTTTCCATTTT GGAACAAGGTTTTTGTCAAACTATTTATTGATATATTTTGTTAGTTACTACTCAACAA GTAACAATATAATAATGTACATTAACTCACGAGTAAAACAAATAATCCATCGTTTCA TATTATAAGTCGCTTTAATTTTTTTCTGGTTAAATATTTTTAAGTTTGCAACATCTAC AACATTAAACTAGTTTCATTAAATCTGACATACATTAAATATATTTTGATAATATGTT 35 TGTTCGCGTGGAAACAATACTATATTTTTTCCGATAAACTTGCATTGGTTAAACTCA AAGAAGTTTGCATAGAAAAAAACAAAACATCTTATAATATTAACTTTTGGATACTCG TGGCACGCTTTTCAAATTACAAAACGGTGTGTTTCGTGCTATAACTTTCTATATGAA AATTGTTCTAAAATATCAGATTAATCCATTTTTCAAGTTTATAATAATCAAAAGGAAA AGGCTATACGTACTAACTTTTTCTTAATAAATTCTTACTAATACTGATGTGTCATGC 40 TATGAGTGCATCTAGATTTTTTTGTAACTTTCATCTACTTAAATCAAACGGTTGAGA TGATTGTTAGTAAAAGTTTAGTAAGAAAAATTTAATACGTGTAGCATTGCTCTAATC AAAACTCAATTAATCACATGTTATTACCACCTCGTTTTGCGTGAAAAACTTAATTTT M&C PC933520WOA 80 CATCTTCAGTACATTCAACCACCACCTAAACGGAAGGAGTATGTGTCTATACCATA CAGTTAATTTGTATCGTGGTTTTTTCAATTTGACACTGATAAAGTAGCATAGACATT GCATATATTGTAATGTCTAACTATGCCATGTAAGCATATAATACCAAACAAAGTTTG GCATGTTTCAGCGAAACTAAAACAAAACCCATTATATCAGCTTACCTTCGCGGATA 5 AAGATCCTCTCTAGAAAACTGATAACTTTTCTATCCTCGTAACAGATAAGCTCTCCC AAAACCAAAGTAAACTGGGAAAAAAACAATCGTTTTCAGAGGGGATAAAACGATAT GATGGACCAATCCAAAAGTCTATTTTGGATATTTTTTTTCCTAATATATTGACGTGC AATCCTTTTACACGTTCGTGAAAAAGATATCCTTTTTCATAAGAGAAATGAGTTAAC AAAATGATTATGATCATCAGCTAGCTAGTTCTACCAATCCATTATGTAAGTCCTCAC 10 TAAGGGAGTGTTTAGATTCAGGGGTGTAAAGTTTTGGCGTGTCATATCGGGTATTA TATAGGGTGTCACATGGGGTGTTCGGGCACTGATAAAAAAATAATTACAAAATCTG TCAGTAAACCACAAGACGAATTTATTAAGCCTAATTAATTCATCATTAGTAAATATTT AATGTAACACCACATTGTCAAATTATCGAGCAATTATGCTTGAAAGATTCGTCTCG CAAATTAGTCGCAATATGTTCAATTAGTTATTTTTAGCATATATTTAATACTTCATGC 15 AAGTGTTCAAACGTTCGATGCGATATGGTGAAAAATTTTGGGGTGGGATCTAAATA GGGCCTAATATTAGGAGGTAGTAAAAGTACATCGTATAACGATAAAACAATAGTAG TAAATGATAGTTCGACTCCATTACATGTCACATAGAAATATTAAATGACATTCAATT CAACTCCCAGGTGTTTGTCGTGTTGTGAACGGCATATAGCGGTTCGACGGCACTA CTTGCCACAATAATATTTGGCTATTTGTCGCGTAGTTAATAACATGTGGTGGTTCA 20 ATTTTATTGCTTGCCGTACAATAATATTTGGCGCCATTTTTACATAGATATATTTAAA CCGCCATACCAATATTTTAGCGTTATTAATTTTACGTAGAGATATTTAAACCACCAT AGCAACATTTTCGCATCATCATATTTTTACGTAGAGATATTTGAACCACTATACTAA CATTTTAGCATTCTTCTGTTTAAACGTACGCTTTCCCGTTGTCCTGTGGTTTTGTGA GCACGGATCGACGATGAAAACAGACGAAGAACATTGTTCCTGAGAGCAACAAAAC 25 ACAAGAAAATCCCATGGCGCAACCGAAACCAGAGATGGGACCAGCAACCCTTTTC CTCTGTGGCCCCGCCAAGTGTAGGGTGAAAAAATTCAGGTATAAATAGCGCAGCA GCCATTGCTAGCCCTCCCACACACAGCAGGTAGCTAGCTTAGCCCCTCACTTTGT GGCAGATCATACTCGCTGGCTACACCACCACCCACATCCATCTGAGCAGTGGAAT CATCATCTCCTCCTATATCAGTTGTCATCACCCGGCCTGTAGTGTGCGCTCTGATC 30 CTAGTTAGCTGCTGCTCTTCGTTCCGATCCTAGCTCGTTGCGAATTGAGCCAACG GATCAATCGCGCGCTCCGGCCGCCACAGCCTTTATTACCCAAGCTGTTTCTTTTC CAGAATTTAAAGAAGTAGCAGCTGTAGACAGTAGTTATCCTCGTCGCGACCCCTT GTGTGCAGAGCAGGCGGCGGCGGCGGCGTTAGCTAGCTCGTCGAGCAACATATA TGGCAGCTAGTAGTAGCACGTAGTTACTGCAGTGAGGAATTCGTCCGGAATCGCG 35 TTGGGAGCGCGCGGAGTAGCGGACGGCTGAGCGGAGCGCTCTCGGTTTGATAA GGGTTCGCGTGTGTGTGTTCGGTGTGGGGGAGAGTAATTGCACGCATCACAAGA ACACTCCTCGCCTCTACTTGGGCTCTCCTGCTCACGTGATTCATACTCTCCCTACT CCCCTCTCCCCACATCCGTTGCGCCTTGCATCTACTCTGTGCCAATCACCGGGCA TATATAC 40 SEQ ID NO: 4 GSE3 coding sequence (annotated)ATGGTTGTGCCTGTGCCGACGACGTTTGAGGGAGCGAGAGAGATGGAGGTGGTG GAGGTGCGTGAGTACAGGGAGGACCGGGACCGCGCCGCCGTCGAGGAGGTGGA GCGGGAGTGCGAGGTGGGCTCCTCCGGCGGCGGCGAAGCCAAGATGTGCCTGT 45 TCACGGATCTCCTCGGCGACCCGCTCTGCCGCATTCGCAACTCGCCGGCCTACC TCATGCTGGTAGCGGAGACAGCGAACGGCGGCGGCGGCGGCAATGGCAGGGAG ATCATCGGCCTCATCCGCGGTTGCGTCAAGACCGTCGTCTCCGGCGGCAGCGTC CAGGCCGGCAAGGACCCCATCTACTCCAAGGTCGCCTACATCCTCGGCCTTCGC GTCTCGCCTCGTTACCGGAGGAAGGGGGTGGGGAAGAAGCTGGTGGGCAGGAT 50 GGAGGAGTGGTTCCGGCAGAGCGGGGCGGAGTACTCGTACATGGCGACGGAGC AGGACAACGAGGCGTCGGTGCGCCTCTTCACCGGCCGCTGCGGCTACTCCAAGT TCCGGACGCCGTCGGTGCTCGTGCACCCGGTGTTCGGCCACGCGCTCCAGCCCT M&C PC933520WOA 81 CGCGCAACGCCGCCATCAGGAAGCTCGAGCCGCGCGAGGCCGAGCTGCTGTAC CGGTGGCACTTCGCCGCCGTCGAGTTCTTCCCCGCCGACATCGACGCCGTGCTG TCCAAGGAGCTCTCGCTCGGGACGTTCCTGGCCGTGCCGGCCGGGACGCGGTG GGAGAGCGTCGAGGCGTTCATGGACGCGCCACCGGCGTCGTGGGCAGTGATGA 5 GCGTGTGGAACTGCATGGACGCCTTCCGCCTCGAGGTGCGGGGCGCCCCGCGC CTGATGCGCGCCGCGGCGGTCGCGACGCGGCTGGTGGACCGCGCGGCGCCGT GGCTCAAGATCCCGTCCATCCCGAACCTCTTCGCGCCCTTCGGCCTCTACTTCctct acggcgTCGGCGGCGCCGGCCCGGCCTCCCCGCGGCTCGTCCGCGCGCTGTGCC GCCACGCCCACAACATGGCCCGCAAGGGCGGCTGCGGCGTGGTCGCCACCGAG 10 GTCTCCGCCTGCGAGCCCGTCCGCGCCGGCGTGCCGCACTGGGCGCGCCTCGG CGCCGAGGACCTCTGGTGCATCAAGCGTCTCGCCGACGGCTACAACCACGGCCC GCTCGGCGACTGGACCAAGGCGCCGCCGGGCCGCTCCATCTTCGTCGACCCAA GAGAGTTTTAG 15 Preferred positions of mutations marked: ^Bold = position for a gse3-cri1 mutation (insertion)^ Underlined = position of gse3-cri2 mutation (deletion of residues 275 to 304).^ Bold and underlined = position of gse3-cri4 mutation (A insertion at position313)20 ^ Lower case, italics = position of DEL1 mutation (10bp deletion at positions 960to 970) ^Bold italicised = position of m238 allele (T>C substitution at position 202).SEQ ID NO: 5 OsGS2NGR2 (wild-type genomic sequence)25 ATGACGATGCCGTATGCCTCCCTGTCTCCGGCGGTGGCCGACCACCGCTCGTCC CCGGCAGCCGCGACCGCCTCCCTCCTCCCCTTCTGCCGCTCCACCCCGCTCTCC GCGTAAGCAACGCGAACCCGCGGCTACAACCCATTTTCTTGGCTCCAGTGGTGCA TGTGACAACACGGTGAGACGTTGTGTGTGGGTGGGTGGGTGCAGGGGCGGTGG 30 TGTTGTCGCGATGGGGGAGGACGCGCCGATGACCGCGAGGTGGCCGCCGGCGG CGGCGGCGAGGCTGCCGCCGTTCACCGCGGCGCAGTACGAGGAGCTGGAGCAG CAGGCGCTCATATACAAGTACCTGGTGGCAGGCGTGCCCGTCCCGCCGGATCTC GTGCTCCCCATCCGCCGCGGACTCGACTCCCTCGCCGCCCGCTTCTACAACCAT CCCGCCCGTACGTCGTGTTCCTATTTCTTGCCTCTCCTCTACCATCGCTGCATTGC 35 TTTTGGATGCTTGTTTAGTGTCGGCCTCTTTGTTTATTCCGATCAGGCGTACTTTG CTTCCATTTGTTAATTGGCTCCGGGTCATTTGTTAATCCGGGTTACGCGATTCAAG AAACATGCGTGTGGTTTTTATGCTATCCTCCGGATTTGGTTATAAAAAGGCTTGTTT TTAAATCCAAAACTCGTGCTCGCTTCACGATTAGCGCATCATTTTTTTTTTATGGGG GGGGGGGGGGGAGAGTTTGCCCATCATTCTGTCTCTGTTTGATCTGATAGAGGAC 40 GTGCACACGCTCTTGTCTGAAATAAAATCTTTTGTTTATCAGTATGCCCATGGGAT AAGCCATTTTCTCTGTGAACCAACACCCTGGCAAACTGTTTTTTTGCTCGCCATTTT TGAGCGATTGCTAAGAACAGATAACTATGCCCTGCATATGGATCGGATATGGACTT CTCAAATATTCAAATGCCATTCTATTAGGAACTCAAAATGCATTACCAACAAATGCA TTCTTGTGTGTAACACGGTTGCTACGATGTGCCTGTTTTTGTACAGTTGGATATGG 45 TCCGTACTTCGGCAAGAAGCTGGACCCAGAGCCAGGGCGGTGCCGGCGTACGG ACGGCAAGAAATGGCGGTGCTCGAAGGAGGCCGCGCCGGATTCCAAGTACTGCG AGCGCCACATGCACCGCGGCCGCAACCGTTCAAGAAAGCCTGTGGAAACGCAGC TGGTCGCCCAGTCCCAACCGCCCTCATCTGTTGTCGGTTCTGCGGCGGCGCCCC TTGCTGCTGCCTCCAATGGCAGCAGCTTCCAAAACCACTCTCTTTACCCTGCTATT M&C PC933520WOA 82 GCCGGCAGCAATGGCGGGGGCGGGGGGAGGAACATGCCCAGCTCATTTGGCTC GGCGTTGGGTTCTCAGCTGCACATGGATAATGCTGCCCCTTATGCAGCTGTTGGT GGTGGAACAGGCAAAGATCTCAGGTGATTGTTCATTTCTTTTTTTTTAATCAAACG CCATATTTACTTGTTTAGCACTGTCTTGAATCATGATATGTATCCTTCCGTTGTCTA 5 AAAAAAAGGTGTCATGCTCTAACTGATTGGTGTCAGGTGGATGCAGTTATGAATCT GTATTTTTCTTTGTGATCGGTTAATAACTGTGTCCCATTTGTTTGCATTGGTGGCAA TCGAACCAGCTGTCCATGCTCAGTAGTACTACTTCGATTTGGTGCTGCAATCACTG AAAGTCTGAAACTTTACTCTCTGCACTGCAAAAATTTGTGTTATGTTTAGGTTTCCA GAGTGCTGCCTCTTTGCCCTTCCCATACTTTCTGGTATCAGTTTTCAGCCCCAGAA 10 GCCGGGGACAGTCTCCATAAGAGATTTCTGCTCAGGTGAAACTGGGGTGCAGGG TCTTAACATGGCTTTGGCCCAGTAGTTTGAAACATGTACTGTCCATAAAGATGATA CTACTACATATTTGTGTCTGCCCTCGCAGTGCTTGTGCCTGCTGGTAGCTGATCAT GGCTTCCCTTGGCATTTACTCCACTTCTTTATTCCTCCACAGAATCCAGTTGTTTCT GTCTCTGCTCTTCAGGGGCAGTCAATTATTTGGCCCTTGCAAAATACTATCTCTGA 15 AGATGTCTCACCGATCACCACTATACCTGAAACATTTTCCAGTGGCCAGCGTGAG CTGCATGATGCTCCAAGTCAACTCTATACTCATCCAATGTTGATGATTAGATTTTAA CAATGCAACTCTTTGATTTATCTTCCCTACAAAAAAAAAGGAACTCTTTGATTTATC TTCGGTGAATCTCAGTCTGACCTTAGTACCTAGCCTCATTATTTACTTCACCAAATG TATAACTCTACAGTGCTTGTTCGTGTTGATTTGGTTTAGTTTAGTTATTGAATTATTC 20 GGTCACCTTAGTCTTTGATTGTTTTTTTCTTTCTGCTCTTGTCATCAACTGTTTAGG GTTCAGCTGACTTGCTGCTGCAACTAAACTGTCTTCTGGTTTTACTGCAAAATAGA ATGTTTCTTGGGCCATGATCTGCTGCTATATATGATTAGTTAAACCATGGTTCTATG TTTTCTTATATGAATTCATGACAAGAATACTAACTTTTGGAAAAGGTAATTTTATTTT TTTTGTATGATAATAATGCTTTGGATTCTTTCTAGTTTATCTGTCGGACTTAGGTTA 25 ACTACATTTCCTCCGGTACATGGATTTATTTCATTCTTACAATTGAGCCCTTATGAA TATTTTCTTCCTAATTCTGTTCTAAAAAGTTAGAATTGACATATTTTCGATAGGTACA TGCCTAGCACTTGCATTCGTGTTTCCTACTAATTCCCAATCACTGTATCTTCTCAAA TTCAGGTATACTGCTTATGGCACAAGATCTTTGGCGGATGAGCAGAGTCAACTCAT TACTGAAGCTATCAACACATCTATTGAAAATCCATGGCGGCTGCTGCCATCTCAGA 30 ACTCGCCATTTCCCCTTTCAAGCTATTCTCAGCTTGGGGCACTAAGTGACCTTGGT CAGAACACCCCCAGCTCACTTTCAAAGGTTCAGAGGCAGCCACTTTCGTTCTTTG GGAACGACTATGCGGCTGTCGATTCTGTGAAGCAAGAGAACCAGACGCTGCGTC CCTTCTTTGATGAGTGGCCAAAGGGAAGGGATTCATGGTCAGACCTCGCTGATGA GAATGCTAATCTTTCGTCATTCTCAGGCACCCAACTGTCGATCTCCATACCAATGG 35 CATCCTCTGACTTCTCGGCGGCCAGTTCTCGATCAACTAATGGTACGACTACTTGA TCTCCCCCCAATTACTTCGTGCGTGTTTATGTCTGTATCCTGCAATGTCTGAAGAT TTCTTACTGAAAACGTCATCTGGTCTGTGTGCAGGTGACTGA 40 SEQ ID NO: 6: OsGS2 haplotype A promoter TATCGATGGCAACAGTGCATGAGCATATATTTATTTCATTGACCTACGGTTGCATG TCTTCGATCTCTATGGAGTAGTACCGAGGCTAAGTTTAGTTTCAAACTTTTCCTTCA AACTTACAGCTTTTTTATCACATTAAAACTTTCCTACATACAAACTTTCAACTTTTCC 45 ATCACATCTTTTAATTTCAACCAAACTTCTAATTTTAACGTGAACTAAAAACACCCT GAATTCAAAACTCTTTTTATTTTCCTTCAAGATGTCCGATGCACACGCTCTATGTAG ACGCAAGAAGATGTTGGAGCAGCAGACTAACAGTAGCAAAAAAATGGCAGGTCGA AAAGCAACTGCGACGGTTGCTCCGTCATCCTCTCATCGCCTTTTTATTGCTCCGG CGTTGGGAACCGCAACAATGGAACAGCCCAAATCGACAGTCCCCTCCCCCCCCC 50 TCCCCCATCCTCTCTCTCCCCACGCAATACTTGTCACTACTCGCGCTGCTCACTAC AGCGTCTCTGCATGTATATCCATCTATCCATCCATTCCCCCATTTTCCAAATAAAAA TACAGCAAACCAAACACAAACGCAGCCTCGCACTGTACTCGAAGAAAAATCGGTG CTGTACGTACTACGCCACGAGATAACGAGAGAGAGAGAGAGAGAGAGAGAGAGA M&C PC933520WOA 83 GAGGAGAAAATGGAAATGCTTCTGCTCGTACCACGCCGCTACGTCCGCTAGGTC GACAGGCCCGGGCGGAGGCAGGTGTTTGTCGTCTAGCTCGGGTCGGAGCGCGC CTTCTCGTGTCGGGCTCGACGTCCGCGACTCCTCGCCCCTGGTCGAGAGCTCGC AGGCGCAGCGGGAGAGAGAGAGAGAGAGAGAGAGAGAGAGACAAGCCGCGCAA 5 TAAAGGCGCGCGCGCGAGCGAGCGAAGCAAAGCACCATTACTAAAGACCGCGGC GTGTGCTTGCGTTGCGAGCGAGCGAGAGCGAGAGAGAGATTGAGAGAGAGAGA GGGAAGGG (the -941, -884, -855, -847, -801, -522 and -157 SNPs are highlighted in bold) 10 SEQ ID NO: 7: OsGS2 haplotype C promoter CTAAATTATCGATGGCAACAGTGCATGAGCATATATTTATTTCATTGACCTACGGTT GCATGTCTTCGATCTCTATGGAGTAGTACCGAGGCTAAGTTTAGTTTCAAACTTTT 15 CCTTCAAACTTACAGCTTTTTTATCACATTAAAACTTTCCTACATACAAACTTTCAAC TTTTCCATCACATCTTTCAATTTCAACCAAACTTCTAATTTTAGCGTGAACTAAACA CACCCTGAATTCAAAACTCTTTTTATTTTCCTTCAAGATGTCCGATGCACACGCTCT ATGTAGACGCAAGAAGATGTTGGAGCAGCAGACTAACAGTAGCAAAAAAATGGCA GGTCGAAAAGCAACTGCGACGGTTGCTCCGTCATCCTCTCATCGCCTTTTTATTG 20 CTCCGGCGTTGGGAACCGCAACAATGGAACAGCCCAAATCGACAGTCCCCTCCA CCCCCCTCCCCCATCCTCTCTCCCCCCACGCAATACTTGTCACTACTCGCGCTGC CCACTACAGCGTCTCTGCATGTATATCCATCTATCCATCCATTCCCCCATTTTCCA AATAAAAATACAGCAAACCAAACACAAACGCAGCCTCGCACTGTACTCGAAGAAAA ATCGGTGCTGTACGTACTACGCCACGAGATAACGAGAGAGAGAGAGAGAGAGAG 25 AGAGGAGAAAATGGAAATGCTACTGCTCGTACCACGCCGCTACGTCCGCTAGGTC GACAGGCCCGGGGGGAGGCAGGTGTTTGTCGTCTAGCTCGGGTCGGAGCGCGC CTTCTCGTGTCGGGCTCGACGTCCGCGACTCCTCGCCCCTGGTCGAGAGCTCGC AGGCGCAGCGGGAGAGAGAGAGAGAGAGAGAGAGAGAGAGACAAGCCGCGCAA TAAAGGCGCGCGCGCGAGCGAGCGAAGCAAAGCACCATTACTAAAGACCGCGGC 30 GTGTGCTTGCGTTGCGAGCGAGCGAGAGCGAGAGAGAGATTGAGAGAGAGAGA GGGAAGGG (the -935, -878, -849, -841, -795, -516 and -157 SNPs are highlighted in bold) 35 SEQ ID NO: 8 GS2 / GRF4-promoter sequence (full) CGTCGTTGCAACAAATCACCCACGTATTATGGAAACGCTCTATTAAGGGGATCATT TCAACTTTTGTTTCACCAACAATATTTCATCAAGTGTACTCGTGATGTTTCACTATG TATAGATCAAATGTTGCAGTGAATTAAAATATCCTTTTGCTATTTGCTGAAACATTA TTTTATATATGGTAAAACAACACCTGATTTTTTGGAAACCGTTTGTTTCATACTTTGA 40 AGTAAAATGTTCCATGAGGTGGTTTGGTTGAGTTTAAGCTTTTTAAACAATTGAAAC GTTTTCAATCTATTTAATGAAACAATTCCGGTCTACTTGATGGAACAGCGCCCGAT TTTTGAAAAAAATTCAGGTGCCTCGCCTTTTTTATCAGTTTATACGTTGATTGAAGG TAAAACAAATCGACCAAAATGGTAAATGCATTAATGGACATGAGTCTCCATACAGT TCGACCACTAATAAACTAGTCAATTCGTTGAACATAGGTCTAACAATTTACTTATTG 45 GCTTATAAGGCCTTGTTTAGTTCGCGAAAAAGAAAATTTTGGGTGTCACATCAGAC GTTTGACTGGATGTCGGAAGGGGTTTTCGGACATGAATGAAAAAACTAATTTCATA ACTCGACTGGAAACCGCGAGACGAATTTATTAAGCCTAATAAATCCGACATTAGCA CATGTGGGTTACTGTAGCACTTATGGCTAATCATCGACTAATTAGGTTCGAAAGGT TCGTCTCGCGATTTTCATACAAACTGTGCAATTAGTTTTTCATTTTATCTATATTTAG 50 TGCTCCATGCATGTGTCCAAAGATTCGATGTAATGTTTTTGAGAAAAAAAATTGGA AACTACGCAAGGCCTAATTTAAGTTTTAGAACTTAAAGTGTTTATTCTAAGTTTTCT TTTCATCGTAGTTTTTCTTGCAGCCGGTTTTCAAACCACTAGTATTATAGATATAAT M&C PC933520WOA 84 TTTTTTATTTGCAATTTATTTTTTACGATTTATCAACCGCAGTTTATCCGATTAGTCT TTGGAAAATGTTACTGGAAGAACTAAAACCCAGATAACCCACACCAAATAATATTA ATAAAAAATCTGCATCTGCATTATAGTACGTCAATCCAGCTCCACATGTTACTGATC TTGTGGTACTGTAGTAGTAATAGTACTCCCTCCGTCCTAGTATAGTGGACGTTTCC 5 ATTCGTTTTATTTGAAAAATTAGTGCAAATATAAAAAAAGATAAGTCATATGGAAAG TATTTTGATAATAAAGCAATTGACAAACAAAATAATTAATAATTCCAAAATTTTTTAA ATAAGACGAATGATCAAAAATTATAAACAAAAACTCAAGAAGACAAGTAATATGGG ACAGAGGTAGTAATAGTATATTAGTACACTGTTTTGAGGTATTCCAACGTCGAATA AAACAGAGAAGTACCACACCATTATTGTGAACTCGAAACCTGGTAGTGATTAATTG 10 CCTCATGGCAAGCAAACTGAAACAAACTACTATTACTACTGCTCTCCCGTTTTATAT TGTTCATCGTATAACCCAAAATCAGAATTTCCAAATTATATCTTGTAATCTTGACTG CATCGTTTGACATACAATACTATTAAATCTATACATATAAATTGAGTCTGTATACGT ATATACAAGCACTTGAGATGGTTAAGCACTTTTTTTAGCATTCTAAGTTTCTTATTTT GTAGGATTTTTAGTACGAGGTAAGACATACTTGAAAAAAATTATAAGAACTAGAGT 15 GCATGTGACCACCTAACTCCTTGCAATTTTTATTCTTATAATTTGAAAATCCTATAA ACCAAATAAGCCCTTCAAAGGAAATTAAATCATGAGGTTTGAGGTTAGGTTTGAAT TCTCTAAAAAGTGGAGGAAAGGACTCAACAGAAAAAAAAATCCTATAGAATTTCGA TCCTATAAAATTTTAGTTAAAAATACTTTGTTCCAAAATTGCCATGGATAAAATGTAA TTTCTATGCATACAACTAAATTATCGATGGCAACAGTGCATGAGCATATATTTATTT 20 CATTGACCTACGGTTGCATGTCTTCGATCTCTATGGAGTAGTACCGAGGCTAAGTT TAGTTTCAAACTTTTCCTTCAAACTTACAGCTTTTTTATCACATTAAAACTTTCCTAC ATATAAACTTTTAACTTTTCCATCACATCTTTCAATTTCAACCAAACTTCTAATTTTA GCGTGAACTAAACACACCCTGAATTCAAAACTCTTTTTATTTTCCTTCAAGATGTCC GATGCACACGCTCTATGTAGACGCAAGAAGATGTTGGAGCAGCAGACTAACAGTA 25 GCAAAAAAATGGCAGGTCGAAAAGCAACTGCGACGGTTGCTCCGTCATCCTCTCA TCGCCTTTTTATTGCTCCGGCGTTGGGAACCGCAACAATGGAACAGCCCAAATCG ACAGTCCCCTCCACCCCCCTCCCCCATCCTCTCTCCCCCCACGCAATACTTGTCA CTACTCGCGCTGCCCACTACAGCGTCTCTGCATGTATATCCATCTATCCATCCATT CCCCCATTTTCCAAATAAAAATACAGCAAACCAAACACAAACGCAGCCTCGCACTG 30 TACTCGAAGAAAAATCGGTGCTGTACGTACTACGCCACGAGATAACGAGAGAGAG AGAGAGAGAGAGAGAGGAGAAAATGGAAATGCTACTGCTCGTACCACGCCGCTA CGTCCGCTAGGTCGACAGGCCCGGGGGGAGGCAGGTGTTTGTCGTCTAGCTCG GGTCGGAGCGCGCCTTCTCGTGTCGGGCTCGACGTCCGCGACTCCTCGCCCCTG GTCGAGAGCTCGCAGGCGCAGCGGGAGAGAGAGAGAGAGAGAGAGAGAGAGAG 35 ACAAGCCGCGCAATAAAGGCGCGCGCGCGAGCGAGCGAAGCAAAGCACCATTAC TAAAGACCGCGGCGTGTGCTTGCGTTGCGAGCGAGCGAGAGCGAGAGAGAGATT GAGAGAGAGAGAGGGAAGGG SEQ ID NO: 9: miR396 binding or recognition site in GS240 CCGTTCAAGAAAGCCTGTGGAA SEQ ID NO: 10 GS2 amino acid sequence (GRF4)MTMPYASLSPAVADHRSSPAAATASLLPFCRSTPLSAGGGVVAMGEDAPMTARWPP AAAARLPPFTAAQYEELEQQALIYKYLVAGVPVPPDLVLPIRRGLDSLAARFYNHPALG 45 YGPYFGKKLDPEPGRCRRTDGKKWRCSKEAAPDSKYCERHMHRGRNRSRKPVETQ LVAQSQPPSSVVGSAAAPLAAASNGSSFQNHSLYPAIAGSNGGGGGRNMPSSFGSA LGSQLHMDNAAPYAAVGGGTGKDLRYTAYGTRSLADEQSQLITEAINTSIENPWRLLP SQNSPFPLSSYSQLWALSDLGQNTPSSLSKVQRQPLSFFGNDYAAVDSVKQENQTL RPFFDEWPKGRDSWSDLADENANLSSFSGTQLSISIPMASSDFSAASSRSTNGD* 50 M&C PC933520WOA 85 SEQ ID NO: 11 GS2 / GRF4-coding sequenceATGGCGATGCCGTATGCCTCCCTGTCTCCGGCGGTGGCCGACCACCGCTCGTCC CCGGCAGCCGCGACCGCCTCCCTCCTCCCCTTCTGCCGCTCCACCCCGCTCTCC 5 GCGGGCGGTGGTGGCGTCGCGATGGGGGAGGACGCGCCGATGACCGCGAGGT GGCCGCCGGCGGCGGCGGCGAGGCTGCCGCCGTTCACCGCGGCGCAGTACGA GGAGCTGGAGCAGCAGGCGCTCATATACAAGTACCTGGTGGCAGGCGTGCCCGT CCCGCCGGATCTCGTGCTCCCCATCCGCCGCGGACTCGACTCCCTCGCCGCCCG CTTCTACAACCATCCCGCCCTTGGATATGGTCCGTACTTCGGCAAGAAGCTGGAC 10 CCAGAGCCAGGGCGGTGCCGGCGTACGGACGGCAAGAAATGGCGGTGCTCGAA GGAGGCCGCGCCGGATTCCAAGTACTGCGAGCGCCACATGCACCGCGGCCGCA ACCGTTCAAGAAAGCCTGTGGAAACGCAGCTGGTCGCCCAGTCCCAACCGCCCT CATCTGTTGTCGGTTCTGCGGCGGCGCCCCTTGCTGCTGCCTCCAATGGCAGCA GCTTCCAAAACCACTCTCTTTACCCTGCTATTGCCGGCAGCAATGGCGGGGGCG 15 GGGGGAGGAACATGCCCAGCTCATTTGGCTCGGCGTTGGGTTCTCAGCTGCACA TGGATAATGCTGCCCCTTATGCAGCTGTTGGTGGTGGAACAGGCAAAGATCTCAG GTATACTGCTTATGGCACAAGATCTTTGGCGGATGAGCAGAGTCAACTCATTACTG AAGCTATCAACACATCTATTGAAAATCCATGGCGGCTGCTGCCATCTCAGAACTCG CCATTTCCCCTTTCAAGCTATTCTCAGCTGGGGGCACTAAGTGACCTTGGTCAGA 20 ACACCCCCAGCTCACTTTCAAAGGTTCAGAGGCAGCCACTTTCGTTCTTTGGGAA CGACTATGCGGCTGTCGATTCTGTGAAGCAAGAGAACCAGACGCTGCGTCCCTTC TTTGATGAGTGGCCAAAGGGAAGGGATTCATGGTCAGACCTCGCTGATGAGAATG CTAATCTTTCGTCATTCTCAGGCACCCAACTGTCGATCTCCATACCAATGGCATCC TCTGACTTCTCGGCGGCCAGTTCTCGATCAACTAATGGTGACTGA 25 Bold and underlined = position 247, grf4-cri2 mutation Bold and italicised = positon 248, grf4-cri3 mutation30 SEQ ID NO: 12 GRF3 coding sequenceATGGCGATGCCCTTTGCCTCCCTGTCGCCGGCAGCCGACCACCGGCCCTCCTTC ATCTTCCCCTTCTGCCGCTCCTCCCCTCTCTCCGCGGTCGGGGAGGAGGCGCAG CAGCACATGATGGGCGCGAGGTGGGCGGCGGCGGTGGCCAGGCCGCCGCCCTT 35 CACGGCGGCGCAGTACGAGGAGCTGGAGCAGCAGGCGCTCATATACAAGTACCT CGTCGCCGGCGTGCCCGTCCCGGCGGATCTCCTCCTCCCCATCCGCCGTGGCCT CGACTCACTCGCCTCGCGCTTCTACCACCACCCTGTCCTTGGATACGGTTCCTAC TTCGGCAAGAAGCTGGACCCGGAGCCCGGACGGTGCCGGCGTACGGACGGCAA GAAGTGGCGGTGCTCCAAGGAGGCCGCGCCGGACTCCAAGTACTGTGAGCGAC 40 ACATGCACCGCGGCCGCAACCGTTCAAGAAAGCCTGTGGAAGCGCAGCTCGTCG CCCCCCACTCGCAGCCCCCCGCCACGGCGCCGGCCGCCGCCGTCACCTCCACC GCCTTCCAGAACCACTCGCTGTACCCGGCGATTGCTAATGGCGGCGGCGCCAAC GGAGGCGGTGGTGGTGGTGGCGGTGGCGGCAGCGCGCCTGGCTCGTTCGCCTT GGGGTCTAATACTCAGCTGCACATGGACAATGCTGCGTCTTACTCGACTGTTGCT 45 GCTGGTGCCGGAAACAAAGATTTCAGGTATTCTGCTTATGGAGTGAGACCATTGG CAGATGAGCACAGCCCACTCATCACTGGAGCTATGGATACCTCTATTGACAATTC GTGGTGCTTGCTGCCTTCTCAGACCTCCACATTTTCAGTTTCGAGCTACCCTATGC TTGGAAATCTGAGTGAGCTGGACCAGAACACCATCTGCTCGCTGCCGAAGGTGG AGAGGGAGCCATTGTCATTCTTCGGGAGCGACTATGTGACCGTCGACTCCGGGA 50 AGCAGGAGAACCAGACGCTGCGCCCCTTTTTCGACGAGTGGCCAAAGGCAAGGG ACTCCTGGCCTGATCTAGCTGATGACAACAGCCTTGCCACCTTCTCTGCCACTCA M&C PC933520WOA 86 GCTCTCGATCTCCATTCCAATGGCAACCTCTGACTTCTCGACCACCAGCTCACGA TCACACAACGATGAGTGA Bold = positon 545, site of insertion of grf3-cri1 mutation 5 SEQ ID NO: 13 GRF3 promoter sequenceAAAGAAAGAAAGAGCCCATCGCTGCCATGCGTGCAGGGCAGGCAGTTCGTTCCC 10 GGCTTTATTCCTCTTTTTTTCTTCGTCCCCTTCGGCTGCACCGCCTCAGCTGAGCT CCGCGAATCGAGTCTGGTCAATCTGGTCGGGTGTACTTCAGGTCTGGTTGATGCT GCGATGAACAGTGCTTTACATACTGGACTATAGACACAGCTACCCACTTGTTGTAC ATCAAACAACGGGAGAAAATCCTGTGTATTTCAACCAACGGCGCCGTGTGTACTA ATGTGATCAAACAGGGCAGGACAGAACTTGGCGTATCGATCGGTGAGCATCCTGT 15 TCTTGCACATGTTCATTCCCAGCATGCAAAGCATGGAACCACCAAAAAGCAAGGG CTTAGCTTTAGCCGTTGCATAAGCCGTAAGGCAACGGAAAGTTACAATGCATGCT CGCCTGCCTGCCTCTGCTTGGTCGATATCTCGCATGCCTTCGCCTGCGGGTGTTT TGTCGCTGGGATGACCTGACCGTCCGCCTGCGTTGCCTGCGCGCACAAGCGCAG CAGCAGCAGCAGCAGCAGCTTGTGTGTGAAAGAAACAACGGTAGGAGTCGTGCG 20 CTTGTGATCACTGCCGGATTGACGGTGAGTGGGCTGCCGCTGCACGTGAAGCCC GTACGCGCGCCGATTTGGCTTTCTCGACAGGTCTGTAAAATTAAAGACCGAGCGG TGAACAGATAAAATAATAGGAGGAGTTCCGCCCAACCGTTGATGAAAACATCGAC AACGTCCTTTTACAGGTGAAGAATAGGCGTGCGCCATGCATGTTCAGGTAACGTA ACGTCAAGCCTTCTTGAACCCAACACAGCTAGCCACGAGCCCACGACACAAGAGT 25 TTAAAAAAAGAACGACACAGTTTACTCTGATGTGTTAATATGACGCTCACTGACGG GACCTTTTTTATTTATAGAAAAAATAGAAATATTTAGAGAATTTCAATAAAGATTTAA AATGATAGAATTATATCCTATCATTTAAAGTTCTTACATAACGAACAATCATATATG GATTTTGAATAAATTTAGTAAGAGTTTCAACCTCTTGAAATTTTTTAGTTTATCACTC TCATCTAATTCATTTGTTTTTTTCCTATGTTCCAAAGAAACGGTTATTTCTATGTTAT 30 TTATGTATTTTGTAATCATCTATTTTGCACTTGCATTCCGTCAGAGTTATTAGTTTGT TCATGTTTTTTCGTTTTTCCACTCCAACGATTAAGTGGGGGCCCTGAAAACGAAAA CACAACACCGACTCATGTGGTAATGTCCCTTGGTCCTACTTGTCATCCTCACAGAT CCTTTTCTAATTTCAATAAGTTTGTTGGTGGTTAGGGTTTTGCGTGGACATTGGCG AGCTAAGTGTGTGGAACTTAGGCCATTTAAGTCTCCACCCCAATACAGACATGCA 35 CTATCACCTACTAATAAACATGAAGCCAAGGGAGTAGAGATTTTCATGATTTTATTT CAAGATTGGTGACTAATAAAAGAACAACTTGTTAAAAGCATTTCACACGCTCATCA AAGAATATTGTGTTAAACATGACCTCCACCCTAGTTTCATGTACATGTATGGTCAAA TCTTTAACTCCATATCTGATCCAAATATGTCCCATAAATATCATAAACCGTGGCATA GCGATCGGCACGGTAGGTTCTATCTTAAGATAACATAGAGTTTGAAATCATCAATC 40 TACATAATGCTTATGTCCCCGTTGAAAATCGGCCTTATCTGAATCGTTGGATGATT CATTGACTCACCTTATATCAATCCTTGCAGGGAATCCTATGCTAACGAAATATGAC ACCGATGGGTTGTATGTTGATGGTGACAAGAAACACAACTTAGGGTGTGTTTGAG GAGGAGATTGGAGAGGTTGGGAAGATACGTAAAACGAGGTGAGCCATTAGCGCA TGATTAATTGAGTATTAAATATTTTAAATTTCAAAAATGGATTAATATGACTTTTTAA 45 AGTAATTTTCTTATAGAATATTTTTTAAAAAAACACACCGTTTACTAGTTTGGGAAG CGTGCGCGCGGAAAACGAGGTGCTTTCTCACCCTATATCACACAAACGAACGCAA CCTTAGGCTCTATATCGTCGACTTAGGGGATGATGCTAAGTGCACCATCGAAAAC TTATACATAGGTGTACATTATGTGATGTGTCTTTGATTCAGTAATTTTTTTGTTTGAT TTTCTTCTCATTTTTTTTAATAAAATCTTAAACCGCCGGCTGTTAAAAAGCTAACAA 50 GACATGTAATGTATTACTCCGCAATGATCAATACAACTCTCTATGTTGTAAGAAAAA AAAACAAATATGTGGTGTATATTAATATATGCATAGTAATCAAAGTTTAATTTATCTT M&C PC933520WOA 87 TTTGCGACAACATGTAGAAGTCATATTTAGACTTCTATCATATTTAGACTTCTATCT TTCTAGAGTACTTTATCTTATTAAAAATATTAAAATTTTTAACAACTATTCAGATGAG ATATAAGATACTAGAGGACATTCTTAAGTGTTTATAATCCACTCCCACAAAACCGAA TACTGAAAGCGAGCAGCACGCGAACCCAAAGGTAAGAAACCGATCAAGACGCAG 5 GTGCTTGTCCCTTAGCTGGAACCCTCCCGTGTCGGCCTCCTCGCCCACCCCCGC GCACCGCTCCTAGCTACCCGGCCGGCTATATAGCCCCATCGTTATTCGTCCCCGT CTCCCCCCCTCCTCCCTCCCTTCATCAGGCAGAGCCGGCCTGCAGCCTCGCTCG CATCAGCTTGCACACAGCCCACGGCGCCACCACCAGCGAGGCGGGACAGAGGG AAGGAGACACAGACCAGCCAAGTAAAAGGCAAAAGCACAGCACATTAAAAGAGAG 10 AGGGCGCAAGCAAGCGGCAGAGAGGAGAGAGAGAGAGAGAGAGGCGCACAGAG AGAGTGTGTGTGTGTGTGAGGGAGAGAGAGATAGCGGCAGCATAT SEQ ID NO: 14 GRF3 amino acid MAMPFASLSPAADHRPSFIFPFCRSSPLSAVGEEAQQHMMGARWAAAVARPPPFTA 15 AQYEELEQQALIYKYLVAGVPVPADLLLPIRRGLDSLASRFYHHPVLGYGSYFGKKLD PEPGRCRRTDGKKWRCSKEAAPDSKYCERHMHRGRNRSRKPVEAQLVAPHSQPPA TAPAAAVTSTAFQNHSLYPAIANGGGANGGGGGGGGGGSAPGSFALGSNTQLHMD NAASYSTVAAGAGNKDFRYSAYGVRPLADEHSPLITGAMDTSIDNSWCLLPSQTSTF SVSSYPMLGNLSELDQNTICSLPKVEREPLSFFGSDYVTVDSGKQENQTLRPFFDEW 20 PKARDSWPDLADDNSLATFSATQLSISIPMATSDFSTTSSRSHNDE GSE3 homologs SEQ ID NO: 15 Secale cereal>SECCE5Rv1G0354590-promoter sequence: 25 TGAAAATGGTGTGCCTGAGCACCGAGAACCGACAACGTTTGAACAGTCCATCCAA TTTCACCATGCAATGCGCTATTGGCATACTCACGCGCAACTCCAAAATGACTTGTT TGAGCACATGTGGACTCACATTGGCAAGCAGTAGATGTATCAGTTCATTTTCTTTC TATGTATTAGTGACAATTTAAATTTAGTTGTAAACTTCATGTTTTTTATTCAGTTGTG AAACTATTTATTATTATGGCACTATGTAATTTATTTGGTTGTAAACAATATGCAATGC 30 TTTATTTTAGATTAAAAAATGTCTTTGTTTGGCCTCCGACCGCGGATCAAATGCATC GACTGCGTTGGGCACACTGCCAATGCATCCTCGACGGCACCCTTTGCCCGACTC AAATGAACAAAAAACGCATAAAATGAATATCGGGTCGCCTTGTTGGAGTTGTCCTA AGAGCATCTCCAGCCGTTGGCCCCCCCAGGACGCGTAAAAATCGCCCCTGGGGG CGAACCGGCGATACAATCGGCGCTGGGGGCGGTTTTGCGCCCAATCGTCGCCCC 35 CAACTCGCCCCCAGGCGCCGAAATTGGCCCACTTTGCAGCCCAATTTCGGCGAAT AAAGGGCCCATATGGGCGAGAATAGGCCCATATTCGGCGTGGTTTCGCCGTGTC TCGGCATTATCAACACAATTATTTCTTATCACATATTTCATCACAGAAAAATCAAATA CTTCAATAAAATAGTACAACAACAAATAGTTCAATACAAATTATATAGTTCAACAAAT AAAAACTCGTATTTCATCACACGGCGTCCTCCTTGAGCCTCCATAGGTGCTCAATC 40 AGATCTTTCTGCAGTTGATGATGCACCTGTGGGTCTCGGATCTCATGACGCATACT GAGATAGGCAGTCCAACTTGCCGGTAGCTGGTGATCAACTTCGGTTAGAGGACCC TGCCTGTAGTATGGTTCAGTGTCAAACACTGGGTCTTCTTGCTCGCTCTCGATGAT CATGTTGTGCAAGATGACACAACAAGTCATAATCTCCCACATTTGATCTTTGGACC AGGTCTGAGCGGGGTACCGAACAACAGCGAATCGAGATTGGAGCACACCAAATG 45 CCCGCTCGACATCCTTCCTCCAAGCCTCCTGAACTTTCTCAAACCATGCGTTCTTG CCTCCTGCTGCAGGGTTTGAGATCGTCTTCACAAATGTCGACCATCTCGGATAGA TGCCATCTACTAGGTAGTACCCCTTGTTGTATTGGTGCCCATTGATCTCGAAGTTC ACCGGAGGAGAATGGCCCTCAACGAGCTTGGCAAAAACAGGAGAGCACTGCAGC ACGTTGATGTCATTGTGAGTTCCTGGCATACCAAAGAAGGAGTGCCAAATTCAGA 50 GGTCCTGTGTGGCTACCGCCTCAAGCACCACACTACAACCGCCTTTGGCGCCTTT GTACATCCCCTGCCAACCAAATGGGCAATTCTTCCATTTCCAATGCATGCAGTCGA TGCTTCCAAGTATCCCAGGAAATCATCTTGCTGCATTCTGGGCTAGGATCCGAGC M&C PC933520WOA 88 AGTGTCTTTCGCATTGGGTGTTCTCAAGTATTGTGGCCCAAACACTGCCATCACTG CCCGACAGAACTTGTAGAAACACTCTATGCTGGTGGACTCGGCCATGCGCCCATA GTCATCGAGTGAATCACCGGGAGCTCCATATGCAAGCATCCTCATCGCTATCGTG CACTTCTGGATGGAGGTGAATCCAAGAGCGCCAGTGCAATCCATCTTGCACTTGA 5 AGTAGTTGTCGAACTTCCGGATGGAATTCACAATCCTGAGGAAGAGCTTTCGGCT CATCCGATAACGGCGCCGAAATGTTTTCCCGCCGTGAAATGGACCATCGGCGAA GTAGTCGGAGTAGAGCATGCAGTAGCCTTCAAGACGAAGCCGGTTCTTTGCTTTG AACCGCCCCGGCGCCGAGCCACCTCGCTGCGGCTTTCCATGGCTCGCCAGCAG CTGGGCGAGAGCGGCGAGCACCATGAGATGCTCTTCTTCCTGGACGTCGGCCGC 10 GGCTTGCTCCTCCAGCAGCGCGGCGAGCTCTTCCTCGTCATCCGAGTCCATCGC CGAGACAGGCAAAACGCCGAACACCTTGCGCTCGGTGGGCGCGGGGGGTGCGA GGCGAGTCGGGGGATGAAAACCTTGATTTTTCCCCTGTCGGTGTGGGCCAGGCG CGCTTTTCCCTAGCGTCGGAGCCCCCAACTGCTCCCCAGCGCGCCGGGTTCGGC CCGTAACCGCCGGGCGAAAAAAAGATCCGAACCGGCGATTTTCGGCTTCCTAGG 15 ACGCGACTGGACCGTTTTTTCGGCGCCGGCGCCGAAAAAATGGCCTGAGGGCCT GTCGGAGGCGCGGCTGGAGATGCTCTAAGCAGCAGCCACGCAGCGCCGTCAGA GCAGCAACCAACGAGTAGAGTAGGCAGTATCGCAGCACCGAAGCAGTGTAGGCA GCCACAGTCACTCTCCTCTCCGCCTTCCTCCTCGTCGGGAATTGAAAAGCCTAGG GCTACGGCCCCGGAGGGACTCGCTCCGAGCTCCACGGCCTTTACTGCCGAATTA 20 AACCCTCCGCTTCCAAAACTTTGAAGTGGCTACTCCTTCCTCTCCCAATCGGCCGT ACCGATTCCTAACCAATTCATTTCAGTGCGTGGGGGCACTGGGTCTGGGTGGGG AGTAGTAACTCAGACGGCGGCAGGCAGGCACGCAAAAGAGCGCCTCCTCCTTCC CCCGCTCTCACGTGATTCATACCCGCGCCCTCCCCTCGTCCGGGCCCCCCCTCC CCGCCTGCCCCGCCGTCTAGCTATATAGTATCCAAATCGCCCAGCCTCCACCTCC 25 TCCACCTCCCTTTCTCGCGAGCTCGAGCCTGGCCGGGCGTGTCTGTCTGTCTCTT TGCCTGCTACCCATCTTTGCACTGCTTACAGCCAGGAGAGTGAGTGCGTGCGTGA GTGAGTGAGTGAGCTCGATCGGTCCGTC SEQ ID NO: 16 Secale cereal> SECCE5Rv1G0354590.1 cds:protein_coding30 ATGGTTGAGGCGGTGGCGGAGGAGGAGCCGATGGTGGTGGTGCGGGAGTACGACGACGTCCGCGACCGCG GCGGCGTGGAGGAGGTGGAGCGGGAGTGCGAGGTGGGGTCCAGCGGCGGCGGCGGCGGCGAGATGTGCCT CTTCACGGACCTCCTCGGCGACCCGCTCTGCCGCATTCGCAACTCGCCGGACTTCCTCATGCTGGTCGCG GAGACAGCAACCGGCAGCGGCGGCGGCGCTGACGACTGCACCGAGATTATCGGCCTCGTCCGCGGCTGCG TCAAGTCCGTCGTCTCCGGCGGCTCCCACGCCAAGGACGACCCCATCTACACCAAGGTCGCCTACATCCT 35 CGGCCTCCGCGTCTCGCCCAACCATCGGAGGAAAGGGGTGGGGAGGAACCTCGTGGAGAGGATGGAGCAA TGGTTCCGGCAGAAGGGCGCCGAGTACTCGTACATGGCGACGGAGCAGGACAACGAGGCGTCCGTGCGCC TCTTCACCTCCCGCTGCGGCTACTCCAAGTTCCGCACGCCGTCGTTGCTCGTGCACCCGGTGTTCCGCCA CGCCCTCAAGCCCTCGCGCCGCGCCTCCATCGTGCGCCTCGAACCCCGCGACGCCGAGCGCCTCTACCGC TGGCACTTCGCTGCCGTCGAGTTCTTCCCCGCTGACATCGACGCCGTGCTGTCCAACGCCCTGTCGCTCG 40 GCACGTTCCTGGCGCTGCCGGCGGGCACAAGTTGGCACGGGGACGTCGACGCGTTCCTCGCCGCGCCGCC GGCGTCGTGGGCTGTGCTGAGCGTGTGGAACTGCATGGACGCCTTCCGCCTAGAGGTGCGTGGCGCACCC CGCCTGATGCGCGCCGCGGCGGGCGCAACGCGGCTGGTGGACCGAGCGGCGCCGTGGCTCCGGATACCCT CCATCCCGAACCTCTTCGCGCCGTTCGGGCTCTACTTCCTCTACGGCCTGGGCGGCGCTGGACCAAGGGC GCCGTGGCTGGTACGCGCGCTGTGCCGGCACGCGCACAACATGGCCCGACGCGGCGGGTGCGGCGTGGTG 45 GCCACCGAGGTGGCCGCACTCGAACCCGTCCGCGCCGGGGTGCCGCACTGGGCGCGCCTCGGCGCCGAGG ACCTCTGGTGCATCAAGCGGCTCGCCGACGGCTACAGCCACGGCCTGCTCGGCGACTGGACCAAGGCGCA GCCCGGGCGGTCCATCTTCGTCGACCCCAGAGAGTTTTAG M&C PC933520WOA 89 SEQ ID NO: 17 Zea mays>Zm00001eb058360_T001 cds:protein_coding ATGGTCGAGGCGCCGGCCGTCGTGGTGGTGGTGCGGGCGTACGACGCCGCGCGCGACCGC GTCGGCGTGGAGGAGGTGGAGCGCGCGTGCGAGGTCGGGTGCAGCGGCGGCGGCAAGATG TGCCTCTTCACGGACCTCCTCGGCGACCCGCTCTGCCGCATTCGCCACTCGCCGGACTCC 5 CTCATGCTGGTCGCAGAGACAGCAACCGGCCCCAACAGCACGGAGATCGCCGGAGTCGTC CGCGGCTGCGTCAAGACCGTCGTCTCTGCTGGCACCACCACCACCCAGCAGGCTAATAAG GACGACCCCATCTACACCAAGGTTGGCTACATCCTCGGCCTCCGCGTCTCGCCCAGCCAC CGGAGGAAAGGGGTGGGGAAGAAGCTGGTGGATCGTATGGAGGAGTGGTTCCGGCAAAGG GGGGCCGAGTACTCGTACATGGCGACGGAGCAGGACAACGAGGCGTCGGTGCGGCTGTTC 10 ACGGGCCGCTGCGGCTACGCCAAGTTCCGCACGCCGTCCGTGCTGGTGCACCCGGTGTTC CGGCACGCCCTGAGGCCGTCCCGGAGCGCGGCCATCGTGGCCCTGGAGCCGCGGGAGGCG GAGCTGCTGTACCGCTGGCACTTCGCCGGCGTGGAGTTCTTCCCCGCCGACATCGGCGCC GTGCTATCCAACGCGCTGTCGCTGGGCACGTTCCTGGCGCTGCCGTCGTCGCCGGCGCGG TGGGAGGGCGCGGAGGCGTTCGTGGCGGCGCCGCCGGCGTCGTGGGCCGTGCTGAGCGTG 15 TGGAACTGCATGGACGCCTTCCGGCTGGAGGTGCGCGGGGCGCCGCGCCTGATGCGCGCC GCGGCGGGCGCCACGCGGCTGGTGGACCGCGCGGCGCCGTGGCTCGGGATCCCCTCCATC CCCAACCTGTTCGCGCCGTTCGGGCTCTACTTCCTCTACGGCCTGGGCGGCGCCGGGCCC GACGCGCCCAGGCTCGCCCGCGCGCTGTGCCGCAGCGCGCACAACATGGCGCGCGACGGC GGGTGCGGCGTGGTGGCCACCGAGGTCGGCGCCTGCGAGCCCGTCCGCGCCGGGGTGCCG 20 CACTGGGCGCGGCTAGGCGCCGAGGACCTCTGGTGCATCAAGCGGCTCGCGGACGGCTAC GGCTCCGGCCCGCTCGGCGACTGGACCAAGGCGCCGGCCAGACGCTCCATATTCATCGAC CCGAGGGAGTTTTAG SEQ ID NO: 18 Zea mays>Zm00001eb058360-promoter 25 AGTCACATGATCATGCGATTGCTTAACCCCAATATTTAGGAGGTGAGGTTGAAGAAGGCA TATTTCCTGTTGGAGCCACTACCTAGTTGTGGAATGTAGTTCGGCAAAAGGAAAAATAAA CTCATCAAAGGTAATATCCCGTAATATATAAACTAGTCCAAAATAGGTACTTGAAGCCTT TGTAAAACATGTCGAATCCTAAGACTCCAAAATATGGAACAAAATAAAAGTTTGTGAGTG TTGTATGAGCAAACGTTGGGCCAACAAACTAGTCCAAAGATACATAAAGGCTCATAATCA 30 AGTGTTTCATGAAACAAGAGTTCCAACAATTTCATAATTGATGACTTTGCTTGGTGTACG ACTGATCAAGTAAATAATAGCAAGAAAACTATCATCCCAAAATTCTAATTGCATATAGGC TTGTGCAAGGAGGGAAAGACCAACTTCAACTGTATGGTGATGCTTTCCTTCAGCTAACCC ATTTTGTTGGTGTTTGTAAGGACACAATACATGATGGGCAACGCCACTACCCATCTCGCT CGACCCCAACCTCTTCTAAAAGGGTCCACTCCTGCCTTAGCATGTAAAATGATGGGGAGA 35 GGCTCATGCTCAAGGGACAACAAACCTTGTAATCAAACACTAGAACACCCATTCACACTT GTGCACCAGATACTTGGGAACTCATCCCTCTCTCGCCCTAGTACAGACTCAGACAAGGTA ACACGAGCAACACAAACTGGACATAGGGACATTCTACTCAAACTAGTATAAATCCTTGTG TTCTATTGGCATATCATTCGCTTCTAAACACACAAATCATGATAATCACTAGATATTGGT ATGAAACATCACATATTGTGTTGAGCGTATGGTGCGACTCATCGTCATACCAGTCCAACA 40 TCCACTAAAACCTCTCGTCACTACACAGAATCCATCGGACCCATCCACTTTGGACCGACG AGGCAACTACTCACTGGGTTCGACAATCCTATAAAAGGAACACTGGCACAAATAGTGCCA CAATATGATCTAATACATGTTGTCGGACCAATCGAAAACACCTATGAAACTCTCAGGTCT M&C PC933520WOA 90 AATCACACAGATTGTCAGACTCAACGCATAGGTCTAATGAGGTACAACGATCGAGTGCTC TCAAGTTGTCTAGGTGGAGGCATCAGATCGATCCACTGTGGTTTGATGATCCTAAAACTT AGAGTTCGACACTTCGCAAACCCCTCCAACTTGTCCAGATTCAAAAATGTGGAATTGGGC TTTGACCCTTATTTCTTTTTTGCCTTTCTAAGCCACCTAAAGCTAAATACAACTAGTGTG 5 CACTATACTAAACTTGTTACACTTATCTAGATCAAGCTAATAGCTCATACCCCTTTAATA GTATAGTTAAAGAAAAAATCAGGTCCTAAACTAATATAAATGTCAAACTAGTGCATCCTG AACCTTTTTCTTTCATCCTTTAAAAAACTGCAATGAAAATCAATAAGCAGGGGCATCAGA ACCGTACATTGACCAAACCAATCTCCATCATTGTGACCTAACTCAAATTTATCTATGGCA TTAGTCACAATTATCTTATATTGTCAATAATCACCAAAATACTTGAGGGACTACACCCTT 10 TCACTCTTGAATGAGTTATGCTTTCGTAGAATAGGTTGGAGAAACATGGAGATGATGGAT CCACAGGATATGTGGGCTTTTACAAAATTGGATAATCTTATGCCCTCCGGAAAAGAAAGA GCAATTGAAGATCACTATGGAGAAAACCAAGAGTGCGACAAGGACTGTTCTGGCCACCCA ATTCAACTTGGGGACGTTGTCTAAATATATAACATGGTGTTGAAATGTTTCGTGGCGGAA AGGAAGTTGCGGTGAAATATTTGATGATCCCATGGAAGACTTCCACAGAAGCTTGGTGGT 15 GTGCCTTGGATGCTCCTTTGGGTTCTAGCTGGCTTTGGGGTTTGAGACCTTTCTTCCTTT GCTCTAGGGTATTCGTTGTTGGTCGCTGAGTTCTAACTTTTCTGTTTTGCAAGGAGAGGG TTCCCTGTCTTCGGTAGGAAATAGAGAGTGTTTTGTCTAAACTTGGTGTTGTCATGGATT CTTGTCTTTTGACAATAGTTAAAAGGCTGAGGCATGTGACATTAAACCGTAAAAAAACTA CAATAAATAAGTTTATCATAATCGGATAGAAATTAATATAGACTTCTGAATGGTATTTAT 20 TATGTTAAAGAAAGAATACAGAAGAAAGATAAACAAGCTTAATTAGCTGCACTAATTAAA TTTATGGGTTTTAGAAAACTTCCAATACCTATATATATAGCTCAAGATTTTGGACCGTGT TGAAGCTCAAGTTGATATGCTAAAAGAAGATTAAACCTGAATAGTACTGAACAAAGTGAA AATGGAAAGAGCGGGACCGATTGTTACCTTTTTCACCTGGTGGGGCCAACTCCGAGGTCC AATGGCCGGTATAAATAGCAGGGCAAGCGGCTCCCACGCCATCAACGGAGCTTAGCCTGG 25 GCAGCAGTGAGCAGCCGCCTTCCCCTGTGCTGCGGATCGCTACACTACACCACCACTCAC CACTCAGTTACTCACCACACATCCATCCAACTATCCAAGCCCACAGTGAGGCAGTGAGCA GCATCCCCTCTCCTGTAGGGGGAGTCTGTAGTAGCAACCTGTTGTCGATCCTACTCATCA GTCATCTTCCCTGCGCAAAGCCTGCAGCTTTACCACCGCCGCCGCCGCCGCCGCCTCCTC TCCTTCGAAAGCGCGCCTCCTGCTAGCTGGCTGTCCTGAGCTCGGGCCTGTCCCGCGCAC 30 GCAATCGCGCCCTGCCTAGCAAGCTAGACAGGCGCCGCCAGTTATTGCACCGGGTGCGAT ATCGCGCTAGCCGCGGAGTGCGGAGTAGTGGGCGGCGGATAGGCTGCGCCGCTACTCCTT CGTGCGTGTGTGGGGAGGATTAGTAACTGTAGAGGAAGGCAGGCACGCAAAAGGACGCCC TCGCCTCCGCCTCCCGCCCGCGCACTCACGTGATTCATACTTGCGGCCTTCTCGCTCGCT CCCTTCTTCATCACCACCCCTCTCGCGCCGCAGCGGGCGGGCAACTACTATATCAAGCAC 35 CAAATTAAACCCACCCAGCTCCTCTCTCGCCCCTTCGTTCCTGGGAGACCGTACCGGCCG GCCGCGGGAG SEQ ID NO: 19 Setaria italic>KQK86967 cds:protein_codingATGGTCGAGGCGACGGCGGCGGAGATAGTGGTGGTGCGGGAGTACGACAGCGCGCGCGAC 40 CGCCGCGGAGTGGAGGCGGTGGAGCGCGCCTGCGAGGTCGGGTCCTGCGGCGGCGGCAAG ATGTGCCTCTTCACCGACCTCCTCGGCGACCCACTCTGCCGCATTCGCCACTCGCCGGCC TCCCTCATGCTGGTCGCGGAGATAGCAACCGGCCCCAACACCAACAGCACGGAGATCGCC GGCCTCGTCCGCGGCTGCGTCAAGACCGTCGTCTCCGGCACCACCCAGGCCAACGACCCC M&C PC933520WOA 91 ATCCACACCAAGGTCGGCTACGTCCTCGGCCTCCGCGTCTCCCCGAGACACCGGAGGAAG GGTGTCGGGAAGAAGCTGGTGGACCGGATGGAGGAGTGGTTCCGGCAAACGGGGGCGGAG TACTCGTACATGGCCACGGAGCAGGACAACGAGGCGTCGGTGCGGCTCTTCACCGGCCGC TGCGGCTACGCCAAGTTCCGCACGCCGTCCGTGCTCGTGCACCCGGTGTTCGGCCACGCC 5 CTCCGGCCCTCGCGGAGCGCGGCCATCGTGCGCCTCGAGCCGCGGGAGGCCGAGCTGCTC TACCGCTGGCGCTTCGCCGGCGTCGAGTTCTTCCCCGCTGACATCGACGCCGTGCTGTCC AACGACCTCTCGCTCGGCACGTTCCTGGCCGTGCCGGCGGGGGAGCGGTGGGAGGGCGTA GAGGCGTTCCTGGCCTCGCCGCCGGCGTCGTGGGCGGTGCTCAGCGTGTGGAACTGCATG GACGCCTTCCGCCTCGAGGTGCGCGGCGCGCCGCGCCTGATGCGCGCCGCGGCGGGCGCG 10 ACGCGACTGGTGGACCGCGCGGCGCCGTGGCTCGGGATCCCCTCCATCCCCAACCTGTTC GCGCCGTTCGGGCTCTACTTCCTCTACGGCCTCGGCGGCGCCGGCGCCGACGCGCCCCGG CTCGCCCGCGCGCTGTGCCGCCACGCGCACAACATGGCCCGCGACGGCGGCTGCGGCGTC GTGGCCACCGAGGTCAGCGCCTGCGAGCCCGTCCGCGCCGGGGTGCCGCACTGGGCGCGG CTCGGTGCAGAGGACCTCTGGTGCATCAAGCGGCTCGCCGACGGCTACAGCTCCGGGCCG 15 CTCGGCGACTGGACCAAGGCGCCGGCCGGGCACTCCATATTCATCGACCCGAGGGAGTTT TAG SEQ ID NO: 20 Setaria italic>KQK86967-promoter GCCGCTCCGGCCATTTACCGCAGACCCTATCAACTCCCCCGCTGGCCCGCTGCTGCTCTA 20 TCTCTCTCCTGAAACCGGCCGCGCGCACTGCGCCGCCACGCGGCCTCATGCCATATGCTC TCTCGCTGCCCTCTCGAGGCTCGCCGTCCGTTCCTCTCTCTCGTCGAGAAAGGACGGGCC TTTTCTTTTCGCAGCTTTTTTGATACCTCGGCTCATGGCCGGTTGGTTGTGCAGGGGTCC ACAGATGTGCTCTGCAGGCGCCGGGGACACCTTGGATTTGCTCTGCCGCCGATGCTAGCG CTTTCCCCGGGCTGATCGCGTCGAGGTCCCATTCAGAACGGCGGAAAACCTCTTGGAGCT 25 AGGACTTCTCAGTTTCATGTAACGGCTCTGAGCTTTACGTCCAACTTTTATCACATATAT TCGCATGCTCGACGTTAACATATTCGCAACTCAGATATTCCTCAGACTGAAAGCCACCTG CACATTTGCGCATCATACAGTTTATCACTTTGTCCTTTTTTCTCGGTCGGTGCCGTATGG TCGTCTTGTGTTGTGCTGCAGATAACTCAGAGCTCAGTTCCGGCGTGCAGGAGTGCTGTT CTATCGCCGTATAACACAAACTTGCTGAAAAGGACACGGTTGTATCTATGTCAGTTAGCA 30 AACGTTCCTTTTGCCTGAACTCCTTGTTGGGTTGTCGAAATATTCTCTCGACTCCCCATT TGCACCTACTGAAAGTGAAATAACGCATCGATATAATTAAGCAACGATGTTTTTTTTTTC TGGATGAACGACAAAAGAAAATTCCGGCGCGCGCGGTAACATGGTTATTGGCGCTCGATA CTTTTCGCTGTGCCATGCATGGTGATCACAATATATGCGTGAAAGGCAAGAGCGTGTAAA AAGAGTGGAATTGAAATAGCCTTCATGTTCTGCTCTTTGAGTAGCTCTCTGATATTTCCT 35 TCTCTAAAAAATTCTACCGTGTGGAAGAACAAAATTTAAGCCCATTGGTACAGACAAGTG TTCCCTATGATGAAACCAAGAAAACAACCCAATGAGCGTTGGAGGTAAGTGATCCTTTTT CTTTTTCTTGCGAGGAAGGCAAGTGATCCTTTAATTTGGAGCCATGCACGCAGTGGCATG CTCTGGACAACACCACAACTCTGTACTTCATTATTGCTCATCATACAGCAGTTTTATTGG CTTATTCAAGCCTTGTAAGCATGCAACACTACTCAAGGTTCACCAGCATCATGTTCATCC 40 AAGTACCTAGAGGGCTAGAGCCACCCTCGTCACACATGATATATGCCACGGGTTAGCGTC ACACTGAATAGCAGAGCAAAAGCCTCTGGCCCGCGCCAATTGCATAGCTAGCACGTATCC CAACCCTCGTGCAACTAATTCCCTCTCATGCACTGCTCTTAAATAAATTGCTCCTCCATC GTTTTTTGTCAGTACAAGTAGCCTTGTGCAACAACGCTTGTATTCTCACTCCGGTAGCGA M&C PC933520WOA 92 CTACTACTGTCTAGGTTGATGTCAATGGCCATATGGTCATCAGGCATTCAGGCTTCTATT TTCTGGTGCATCTGCCTCAGTTATCTTCCATGACGCTAGCTATGTGTGTTGTTGCCTAGG GCTCTGCCGACGCCAAGAAGCATGTCGTCGCTAGAGGAGAGCTCCCCCCATACATTCATA ATAAACCAAAAGGATAAAGTGGGTTTTCTCGAAACATCACCAGAAGATGAGAAGATTAAC 5 CTAGGTAAAACAAAGTGGGAAAAAATGGGAGACAGCGGGGTGGATGCATGGCCCACACAG AAAACCTTTTCACCTGGTGGGGCCAGCTACGAGGTCAGGTGCAAAACCGCCGGGTGTATA TAAATAGCGCCGCAACCCGGGCTTCACGCCATCAAGCAGAGGGCAGAGCAGAGAGCAGCA GTGGCCAGCCGCCTTCCCCTGTTCTAGCTGTGGAAATCGCTACGGCGCCACTGTCACCAC TCACACCACTCGCCAGTTACTCACCACATCCATCCGCGCCCAGCTCTCGCATCCCATAGG 10 CTGCCCTGTTCTCTGATCCCACCGGTCTAGTCGTCTTCTTCCCTGCACAGCCCGCGGCTC TGCTACAACGCCCGCGCTTTCCTCCTCCGAAAGCGGCGCTGCAGCAGTGTCCTCCTGCCC TGCGCACGCAATCGAGCCCCGGCTAGCTAGGCAAGCTAGGCGCTCGCCAGTTATTGCACC GGGGATATCGCTGCCGGGTGGGCGGCGGATAGCTGCGCTGCGCCGCTACTCCGTGCGTGT TGGGGAGGAGTAATTGTGGAGGAGACAGGCACGCAAACGATCGCCCTTCGCCTCCCGCCG 15 CGCGCCCTCACGTGATTCATACTTGCCGCCTCCTCACTTCGACTCTCCTCTCGCAACCAT ACCAAGCAAGCACCAAACACACCCAGCTCCTCCCTTAGCCGGCGGCGGGGGTCTTCGTTC CTGCTCCCTGCTCGAGGGAGGCCGCGCGCGCCGCCGGGAA SEQ ID NO: 21 Hordeum vulgare>HORVU.MOREX.r3.5HG0512710.1 20 cds:protein_coding ATGGCTGAGGCGGCGGTGGAGGAGGAACCGGTGGTGATGGTGCGGGAGTACGACGACGTC CGCGACCGCGGCGGAGTGGAGGAGGTGGAGCGGGAGTGCGAGGTGGGCTCCAGCGGCGGC GGCGGCGAGATGTGCCTCTTCACGGACCTCCTCGGCGACCCGCTCTGCCGCATTCGCAAC TCGCCGGACTTCCTCATGCTGGTCGCGGAGACGGCAACCGGCGACGGCGGCGCGGAGGTA 25 ATCGGCCTCGTCCGCGGCTGCGTCAAGTCCGTCGTCTCCGGCGGCTCCCACTCCAAGGAC CCCATCTACACCAAGGTCGCCTACATCCTCGGCCTCCGCGTCTCGCCCAACCATCGGAGG AAAGGGGTGGGGAGGATGCTCGTGGAGAGGATGGAGCAATGGTTCCGGCAGAAGGGCGCC GAGTACTCGTACATGGCCACGGAGCAGGACAACGAGGCGTCCGTGCGCCTCTTCACCTCC CGCTGCGGCTACACCAAGTTCCGCACGCCGTCGTTGCTCGTCCACCCGGTGTTCCGCCAC 30 GCCCTCAAGCCCTCGCGCCGCGCCTCCATCGTGCGCCTCGAGCCCCGCGACGCAGAGCGC CTCTACCGCTGGCACTTCGCCGCCGTCGAGTTCTTCCCCGCCGACATCGACGCCGTGCTG ACCAACGCACTGTCGCTCGGCACATTCCTGGCGATTCCCGCGGGGTCCAGGTGGGACGGG GACGTCGAGGCGTTCCTCGCCGCGCCGCCGGCGTCGTGGGCGGTGCTGAGCGTGTGGAAC TGCATGGAGGCCTTCCGCTTAGAGGTACGCGGAGCGCCTCGACTGATGCGCGCCGCGGCC 35 GGCGCGACACGGCTGGTGGACAGAGCGGCGCCGTGGCTCCGGATACCCTCCATTCCGAAC CTCTTCGCGCCGTTCGGGCTCTACTTCCTCTACGGCCTGGGCGGCGCCGGGACGGACGCG CCGCGGCTGGTGCGCGCGCTGTGCCGGCACGCGCACAACATGGCACGGCACGGCGGCTGC GGCGTGGTGGCCACCGAGGTCGCCGCTCTGGAGCCCGTGCGCGCCGGCGTGCCACACTGG GAGCGCCTAGGCGCCGAGGACCTCTGGTGCATCAAGCGGCTCGCCGACGGCTACAGCCAC 40 GGCCCGCTCGGCGACTGGACCAAGGCGGAGCCCGGGCGCTCCATCTTCGTCGACCCCAGA GAGTTTTAG SEQ ID NO: 22 Hordeum vulgare>HORVU.MOREX.r3.5HG0512710-promoter M&C PC933520WOA 93 TCCAAAACCAACAAGGGGTGGACCTAGATGCACTTACAATCATCATTCTGGACTCGTTCC TTCCTTCTGTGGCCACACATCCATCTTAACATGCCATTTTCGACACACCTAACTGCTGCA CATGTCGCCTAGTCGACCAACACTTGGCGCCATACAACATTGTGAGTCAAACCGCCGTCC TATAGAACTTGCCTTTTAGCTTTTGTGGCACTCTCTTGTCACAGAGAATACCAGAAGCTT 5 GGTGCCACTTCATCCATCCGGCTTTGATTCAATGATTAACATCTTTATCGGTATCCCCGT TTTTCTGCAGCATTGACCCCAAATAATAAAAGGTGTCCTTCTGAGGTACCACCCGCCCAT CAAGGCTAACCTCCTCCTCCTCCTCCTCGTGCCTAGTAGTACTCGTTGTAGTTCTACTAA GTCTAAAACCTTTCGATTCCAAGGTTTGTCTCCATATATCTAACTTCCTATTGACCTGTA CGACTATAATCGACTAGCACCACATCACCCGCAAAGAGCATACATCATGAGATATCTCCT 10 TATATATCCCTTGCGACCTTGTCCACATCAAGGCAAAAAAGATAAGGGCTCGAGGCTGAC CTCTAGTGCAGTCCTATTTTAATCGGGAAGTCATCGATGTCGCCATCAATTGTTTGAACA CTTGTCACAATATTATTGTACATGTCCTTTATGATCGTAATGTACTTTGTTGGGACTTTG TGTTTCTCCAAGGCCCACCACATGAGATTCTGCGATATCTTATCATAGGCCTTCTCCAAG TCAATGAACACCATAAGGAGGTTCTTCTTTTGCTCCCTGTATCTCTTCATAAGTTGTCGT 15 ACCAAGAAAATGGCTTCCATGGTTGACCTCCCAGTCATGAAATCAGCTTAATTCCACGAT AATAAGTACAACTATGGACACCCCCTTGAAGATTGGTACTAATATACTTCGTCTCCATTA TTCTGGCATCTTGTTTGCTCGAAAAATGAGGTTGAAGAGTTTGGTTAGCCATTACTATCG ATATGTCTCCGAGGCCTCTCCACACCTCAACGGGGATACAATCAGGGCCCATTGCCTTGC CTCCTTTGATCCGCTTTGAAGCCTCCTTTACCTCGAACTCCTGGATTCACCGCACAAAAC 20 GCCTGCTAGTATCATCAAAGTTGTCGTCCAATTTAATCTAGAGCTCTCATTCTCCCCATT GAACAATTTGTCGAAGTACTCCTGCCATCCATGCTTAATCTCCTCATCCTTCACCAGGAG TTGGTGTGCTCCGCCCTTGATGCATTTGACTTGGTCAACATCCATCGTCTTCTTTGCTCT GATCTTGGCCATCTTATAGATGTCCCTTTCGCCTTCCTTAGTGTCTAACCTCTGGTAGAG GTCCTCATACGTCCGACCCCTTGGTTCACTCACAATTCGCTTTGCGGTCTTCTTTGTCAT 25 CTTGTACTTCTTTATATTATCCGCACTCCTAATGAGATATAGGCATCTGAAACAATCGTT CTTCTCACTGAAAGCCTTTTGGACGTTATCGTTCCACTACCAGAGAGAGGATCGATAAAG TTACAAAAATGTTAAGAATCTGACGTTACATGTTAAACTTTTCTTAGCAAACTCCTAGAT TTTCCATGCGAAAATAGATAAGTTTGCCATAATTGTGTAAAAGGATTGTTAATTTTTTGA AATGAATCGTGTATTAAGATTTGTCATTTTGTAAAGTGAAAATTGTCATGTGGTCTAATA 30 AATCGTTATAAAAATTGTCATGCTGGAGCGAGTAATTGCCACACGTTTAATCGGGATTCG ATGTACCCGAGGGCTTTAGATTTTTTAACCCTAAAAATATGATGTAAGATGTATCTCGGG GTCCCTAAAGTTACAAAATTTACACCGACAACTTACAAAACATAAAATAATAATAAAAAT TAGGGTCATGCATCTCCAATCCAATGTTTTTTTTTTCTTTCCTTTCACCACAAAATGATG TTCATACTGTCATCATCACTAAGCGGAGTAGGAGTAGGCGGACCGATGTGGAGTATCTAG 35 TGCCGAGCTCATATGCTCCCTAATGAACAGTTAAATTAAAAAAATAGTAAAAATCTGAAT TTTATTGGGTTAAAATTTTGACAAATGTTCTCAGTGTTGCACATTTTCATCCTGAAATAG GAAGTCGTGGTTAAGAAAACAGGATCGACAGTTCGACACTCCAAAAATGCTATTTTCAAA AGCAGTTTGGAGTGCTGTTTTTTTTTCACGAGTTTCCCAAATGTTATTTCGGGATGAAAA TATGAAAGCGCTGAGAACACTTGTCAAAGTTTGCACCCAATAAAATTCAGTTTTGTCGAA 40 CTTTCTTTCAGAATTCACTGTTCACGTGAGCTTGAATGAGCTGTGGGAGTAGAAATGGAC TTTGGTAGGAAGTGTTCGTTCAGCCCACCCATCCCTTTGCCCGGTGACCCCAGCTCGCTG CTCAATATAAGTAGCAGCCGCGCAGCCCCGTCACAGCAGCAACCAACGAGCAACGACCAG TGTACTCCGCAGTATCACTTCCCAGGACCAAAGCAGTGTAGGCAGCCACACTCCCGCTCT M&C PC933520WOA 94 CCTCTCCTCTCCTCTCCGTCTTCCTTCCTTCCTCATCGTCGCCAATTGAAAAGCCTACGG CCGGGCCCGATCGATGCATGCGACTCGCTCGGAGCTCCACAGCCTTTACTGCCGAATTTA AACCCTTGGCCTCCACAACTTTTAAGAAGCCATAGTGGTGGCTACTCCTCTCCGAATCGG CCGTACCGGTACCGATTCCTAACCATTTCTCAGTGCGCGGGGGCGCTGGTGGTGGTGGGG 5 AAGTAGTAACTGCGACGGCGGCAGGCACGCGAAAGAGCGCCTCCGCCTTCCCCCGCTCTC ACGTGATTCACACCCGCGCCCTCCCCCCTCCTCCGGGCCCCCCCCTCCCCGCCTTCTCCG CCGTCTAGCTATATATCCAAATCGCCAGCCTCCGCTCCTCCATCATCTCCCTTTCTACCG AGCTCGAACCTGCTGGCCATCTTTGCAGTCAGGACTCAGGATAGTGAGTGAGTGAGTGAG TGAGCTCGATC 10 SEQ ID NO: 23 Helianthus annuus>mRNA:HanXRQr2_Chr15g0682711 cds:protein_coding ATGGCGGCGGAGGGTGGTAGTACGGCGGAGATAGTGGTGGTTGTTAGAGAGTACAACCCT AAAACCGATAGTGAACGAGTTGAACAAGTTGAAAGCAGCTGTGAAGTGGGCCCCAACGGC 15 GAACTATCACTTTACACTGATCTGTTGGGTGATCCGATTTGTAGAGTTAGAAACTCACCG GCTTATCTCATGCTGGTGGCGGAGATGGTGGTTAGCGGTGGCGATGGAGCGGAGATAGTA GGGATGATAAGGGGTTGTATTAAGACGGTTACTTGTGGTAGTAAATTCTCGCGTAGTAGA TTGGGTGAATGTTCGAAACCGTTACCGGTTTTAACTAAACTCGCCTACATTTTGGGCTTG CGGGTGTCTCCGGTTCATCGGAGAATGGGAATCGGATTAAAGCTTGTTCGTAGAATGGAA 20 CAATGGTTCGAAGATAACGGTGCCGAATACTCCTACATCGCAACCGACGACGCTAATGAG CCTTCGGTCAGTCTCTTCACGGATAAATGCGGCTACGCTAAGTTCCGTAACCCTTCCGTC CTTGTCCACCCGGTATTTGCTCACCGTCTCCCCGTCAACAACCGTGTAACCATCATCAAA CTCACCCCCTCCGACGCGGAGTCACTCTACCGCCACCGCTTCTCCACCACCGAGTTTTTC CCGCGCGACATTGACGCGGTACTCAACAACCACCTCAACCTAGGGACCTTTTTGGCATTG 25 CCAAAAGGGTACATTTGGGCGGGGCCGGATAAATTTTTATCCGGCCCACCTGAAAATTGG GCCGTTATGAGCGTGTGGAATTGTAAGGATGTTTTTAAACTCGAAGTGAAAGGCGCATCG AAGTTAAGAAAAGGGATTGCTAAAACGACCCGGGTCGTGGATCGGGCCTTCCCGTTTCTA AGGTTACCCTCACTCCCCAAAATTTTTAGCCCATTTGGGCTCCACATGTTATATGGGCTA GGTGGGTCAGGCCCATTGTACACGAAGTTCATCAAGGCTCTATTTGGGTTCGCCCATAAC 30 CTAGCCAAGGAGTGCAAGTGTGGGGTGGTTGCAACTGAAGTGTCTAGTGAGGACCCGCTC AAGTTAGCAATCCCACATTGGAAGGTTCTATCATTCACGGATTTGTGGTGTATCAAGAGA CTTGGGGAAGACTACAGTGACGGATCAGTGGGTGACTGGAGAAAGTCACAGCCTGGTTTA TCCATTTTTGTTGACCCTAGAGAGTTCTAA 35 SEQ ID NO: 24 Helianthus annuus>HanXRQr2_Chr15g0682711-promoter TCAGGGATGCTAATCTAGCAATGTTGGCGAAGTGGTGGTGGCGGTTCAAAACTGAAAAAA AAGGTATGTGGCGTCGGATAGTGTGGGCGGTGCATCACAGTTCTAGATCATGGAATGATA TCCCGGTTAAGGTCTCGGTGGCAGGGCCATGGAAGAATATCCATAGCATTCGACAGACAT TAGTGCATGCAAATATCGATTTGTATCAAGACATCTCAGTGGCGGTGGGTGACGGGAAGA 40 ATGTAATGTTCTGGTTGGATTGCTGGTTGGACCAGGTGCCATTATACATTAAATTTCCAG CTCTATTCAAGGAAGAAATGGACAAATAATGCATGGTGGCAGACCGATGGGCCAACAGCG ATTCTGGGCACATCTTTTACTGGGCCTGGGCTCGGCCAGAGCTTGGATCTGAAGCGGCTC M&C PC933520WOA 95 ATCAGTTGCAGGGGTTAATTAGTATGCTAGAAGACTGGAATGGCTCAACGGATGCGGACG TTTGGAAATGGAATCATGATCCAGATGGGAATTTTTCTGTTTCGAATGTTAAAAAGTTGC TTGGTTCTGTGGATCGAAACAGGCCTGAAAGAGTATTTGAATGGAACAATTGGGTTCCGA AAAAGGTCGGTATTGTAGCTTGGCGAGCGGAGATGGAAAGATTGCCGACAAAGTGTGCAT 5 TAGCCAGAAGAAACGTCCCGGTTCCAAATCAACTGTGTGTTTTATGCGGAGAGTATGTAG AGACATCAGAGCATGTATTTGTTTCATGTCATTTTGCGCAAACGTTATGGCAAAACGTGG CGGGATGGTGCAAGATTCCGCCGGTAATTGCGTTCGGCATTGGTGATTTACTAAAGCTGC ATGAATCGAGTTCGGGTTCGAGAAAGGTGAGGAAAGTAATACATGCTTTAGTTTTGGTGA CGTTTTGGAGTATTTGGAAGACTAGAAACGAAGTGGTTTTTAGACAAGTGAATGCAAATA 10 CGACAAGGAGTTTGGACGAGATCAAGTCGGTAACGTATCTATGGGTCAAATCTAGAGCAA AAGTGGCGGCTTTGAGTTGGGAAGATTGGAGTCGGTTTAGCTTGGGTGTCTTGTGATTAT GTAACTTAATAGTAATTGATGTATTCTGGTGGTTCTTGGTGTCAAACGTTGTAATTTTGG TGAGTAGCGCCTTGCTACATTGAATAAAATTTTATGTCGGCCGTTCAAAAAAAAAATAAG AAGATATATATAAAAATACCCTAATAAATAATATGATTACATGATTCGTGCATCAATACA 15 ATCATGTTTAATTTGTAACTGTTGAATAAATATGTAATAAAGGGTCGTAATAATAAGGGG ATATATAGCTTAGTGGTATCTTGAGGGTTAAGATAAGGCTTTGGACTAATAGGTCCTGGG TTCGATTCTCACAAAGGAGGTTTTCTCATATTTATTGAGTTTCCTCCTTGGTGTATAGGC ATTATGTCTAGTGGAGATGGATATGATCGGGTGGTTCCGCTGATGGCACGATGATACTCC AGTGATCCGTTAGTGATCCAAATTTGCCGTTAAAAATAATATATATAAAAATACCCTAAT 20 AAATAATATAATTACATGATTCGTGCATCGATACTACCATGTTTAATTTGTAACTTTTGG TAATAACTTTATGGATTTAAATGGGGCAATCGTTACATAATCACGTCATGATGTTTTGAA TAAATGCATATTTAAAATCTCACATTGAAATCCTAAGATGACTTTTTTCAAAAGAATATC TGAAATTATTTCGACACTCTTTTTATGTGGAAGATATATAATAAATAATATATAATTACA TGAATATAAAGTTGATTCGTGCATCAATCTTACCATGTTTAATTTATAACTTTCTGTAAT 25 AACTTAGGGGAGGGGGTGGTCACTAGTGATAGAATTCTATCACTCGCAATATCCAATTAT GTTCCGCCATGTCAGCAACAATTTTTCCATCACTCACAATCTTTTTTAGTGGGAGTTGTC ATCACTCACCACCACAACCAACAATTTTCCCCAACCAACAACTACCCTCACAATAAAACT CATCACGACCGTGGAAAAAATAACGAAATCATATTTCGTTATGTTATAACGCGTTTTCAT GTATGGTGTCGGTGGAAATCAATTCCGTATCACGCGTTATAAATGCTCATCACGGTACCG 30 CCCCCTCCACCCTTATGGATACTGAATCAGGTCATGATCGTTTCTCAACGAAAATAAAAT GATTATAAAACTGCGCAGTCACAGCCGCCAATTTCCCTCATCTTTTTAGCTTCCAAAAAA TAAAATCGATATTTTTTATTTTATTTAAGTAAAAAGCAAAATTAAATAAGCTATATTCGA AGATGAGACCCCTATATAAATGGAGCCATGATGGAAGAACATACCACGTCAATAATAAAA CACTCACTCTCTCTCTCTCTCTCTTTCTCTCTCTAAAAACAAAACAGGCAACCTTTTCAA 35 CATCTTTTACATTAGTTACGTCAACTGAATAACATTCGAG SEQ ID NO: 25 Glycine max>KRH28367 cds:protein_coding ATGGGTGAGGAGCTATCACCTACGTTGGTTGTGAGAGAGTTCGACCTGAATAAAGACCGA GAGAGAGTAGAAACCGTTGAAAGGTCATGCGAAGTTGGACCCAGCGGCAAGCTTTCTCTC 40 TTCACCGACATGCTCGGCGACCCAATTTGCAGGGTCCGCCATTCACCTGCTTTTCTCATG CTGGTAGCGGAGATTGGCGAAGAGATAGTAGGAATGATAAGAGGTTGCATAAAAACTGTC ACATGCGGGAAAAGATTGTCCAGAAATGGAAAATACAACAACACTAATGTGAAACATGTC M&C PC933520WOA 96 CCAGTGTACACCAGAGTCGCATATATACTAGGCCTTCGCGTTGCTCCCAACCAACGTAGA ATGGGAATAGGGTTGAAACTAGTGCATAGAATGGAGTCTTGGTTCAGAGATAACGATGCG GAGTATTCTTACATGGCAACGGAAAGAGACAATTTAGCATCCATTAAACTCTTCACCGAC AAATGCGGGTATTCAAAATTTCGTAATCCGTCCATTCTTGTCAACCCAGTTTTTGCTCAC 5 CGAGCAAGGGTGTCTCCAAGGGTCACAATCGTAAGTCTTTCTCCCTCAGACGCCGAGTTT GTCTACCGTCGCCATTTTGCCACAACAGAATATTTCCCGCGCGACATTGACTCAATTCTG AACAACAAACTAAATTTGGGAACATTTTTGGCGCTCCCAAATGGGTCCTACAGTGCCGAG ACGTGGCCCGGCCCGGATCTTTTTCTGTCGGACCCACCGCACTCCTGGGCCATGGTCAGT GTGTGGAACACTAAGGAAGTGTTCACGCTCGAGGTGCGTGGTGCGTCGCGCTTGAAACGC 10 ACACTTGCAAAGACGTCCCGGTTGGTGGACCGGGCCTTGCCGTGGCTGCGGTTACCGTCG ATGCCGGACTTGTTTAGGCCGTTTGGGTTTCAGTTCATGTACGGGCTGGGAGGCGAAGGC CCAGAGGGTGTGAAGATGGTGAAGGCCCTGTGTGGGTTCGTGCACAACCTCGCTATGGAA AAAGGGTGTAGTGTGGTGGCCACGGAAGTGTCCTCAAACGAGCCCTTAAGGTTCGGAATA CCCCACTGGAAGATGCTGTCGTGTGAAGATTTATGGTGTATGAAGCGGTTGGGCGAGGAT 15 TACAGTGACGGTTCCGTTGGTGACTGGACCAAATCTCAGCCCGGCATGTCCATTTTTGTT GACCCGAGAGAGGTCTAA SEQ ID NO: 26 Glycine max>KRH28367-promoter ACAACTTTCAAAGATAAAGTCGTAAATTAATTAATATATTCTTTCCTTTTCTATTCAAGT 20 CGTAAAAATGTAAATATTTTTTATTTTTTTAATTATGACTTTGACATTATAAATATTTTT TTTAAAATAGTTATCAAAAGTCCTAATTTTGATTTAGGACTTTAAAGTTAAGTTATAATT TTTTAAAAATTATTTTTTAACGAGTTTAAATTATTTTTAATTTTTTTTAATATTTTTGTT ACTAATATTAACAAATAATATTAAATTAAAATTTATCAAATATTATTAAATCTTTTGATT TTTTAAATTAATTTTACGGTTATCATTAAGTAAATGAATAATTTCCTACATAATTTTTAA 25 GTAAATATTTTTTCACTATTAATATTTTTATTTAAATTTAAGTCTTTTATTTACATTATT TTTTAATTTTAAAAATAAATAATAATTCTAATAAAAAATATATTAAATAAAACAGTATAA AAATATTAAATAAAATTATAAATTAAAACTAATTTAAATTCGTTAAAAAGTAATTTAAAA AAATTACAAATTCTAATAAATATTTAAAAAAATACTTTGAGGTCATAAATCAATTTATGA TTTTAAGGTTATAATTTAAAAAATTAGATTTTTTTTTTTACAACTTCAATAAAACAGAAA 30 AAGAATTATGAATTAACTTATGACTTTATCTTCAAAAATCATAAGATATCTCATAACTTA TTTTTGTTAGATATTAATTTAAAATTTATTAGATATAATAACTTACAAGAATAATGACCC TTTTTTTAACTTTCCCCATAGCCATAGGCTTTAATCAGTACGTATGATAAAAGTACATAA AATCTTGGCAGCTCGTGAAGAAAGTATCCTAGCCGTCAATTTGAAACTGAGGTTAATTAT TATACGAATGGTGACGATCATTACCAGCAGCAGCTAAAGAGAAACCCGCTTTTGCAGCTC 35 ATTAGTCACTAGTCAGCGGTCTCAGCACTGAAACAACAACATAATTCAAATAAATAAAAT GCCCCCACTTTGCTTTTTAACAGCACCCGTTTCCGCAGAAAGTGCTTAGGAGTATTTTTT TATTATTATAAAGCAAATGGTAATTAGATTATTTGAAGATTCATCCATTATAATTTTACG GCACTATCCAATTAATTAATATTAATATAATATTTGGTAGAATGAGAAAAACGCATTGTT CTCATGGTAACGCTATCTAGCTACTAAATCACATCCTTTGCATGTAATTCACACTTCACA 40 GCATTGAGAGTGTGTTTGATTTAAAGAAATGAAAGAAATATAAAAGTAAATTGAAATAAA AAGATAAAATATTTAAATTAAAGTAAAATATAAAAATATAGAATTCACACTAAATTTTAA AACTTTCATCTCAAATACACTAAAATATATTTTAACTAAATACTACCTGAACATACTATC M&C PC933520WOA 97 ATCCAAAATAAACACAGTAAATAACACACGTGATGACAAAATTGAAATTAAGTGTCCAAT GAAAGATAGATATATAATATTGTACTAGTATAGTAGATATATACTAGATTCCATCCCACT CCCACATGAAATATGCATGTGCATCTCATTCCATATGATGCATATATAAACATTACACAT TATGGGGCGGAGGTAGTAGGGGGAAAGAGAGTAGTATTAAAGGTAATGATTACAAATGGG 5 GTCATGATGGCAACGCGTCCTTCAGTATAAAATAAATGATTATTAAAAAATCAATTAAGA GAAACATCTCCACAATACTGCGCAGCCATAGCCGCCACTTCCTCATCCCTTGCCCCTTCC CCTTCCTTTCGCTTCGCAATGAATATTATTGATTCATTTTCATTTTTTTTTTATTTCTTT TCCGTCATTATTATCACAAGTACCAGTCATAGAGCGAGAAAAATAAATTAAAAGTCCAAT ATCAAAGAAAAAAAAGTCGCGCAGATAAAGAAAACAAACAGATGGAAGGTTGGGGGAATA 10 GATAAAGACATTCAGAATTATTTTAACGTTATATATAATTCGAAGCGAAAGCAAAAGGAA GTGGTAGATTCTGGAGAGAGAAACTAAAAAATAATAGTCCTATTTTTATTTTTTATTAAA AAAAACAGCAAGAAAGAATAGCTAAAAAGTTCAGTATTTAAAGCCCACAAGGGACTTGTA TTGTACCTCTGCGGAAGAATAGAAGCAGAAATATCTAAATAAACAGCTCCCTTCTTTTTT CTTTATTTGCTCCTAAATAAACAGTTCTGATTTTCCAAAATTCGTGGCTCTGCTCTGGGT 15 CTCAAATTTCGAAGCTTTGGTAGGCTTAGGTTTCTTGCTTATAGCACAGTGCACAAGTTA AACCAGCCTTGAAAAAGCTGCGATTTGCTCACTCTCAAAATCCCTCCTCTATCTATAAAA CTCTACCCGTCGCTCACAATAGTTCCTCGCACTTTGTTTGCTATCCATCTTATCATCTTC TCTACCACGCATATATACCTGTTATTATACTCTCTCTCAA 20 SEQ ID NO: 27 Pisum sativum>Psat2g170520.1 cds:protein_coding ATGCCTAATCATAATGAATCCATTAGTGTGGTTGTTAGAGAGTTCCAAGTTAAGAAAGAC ACAGAAAGAGTAGTAACACTCGAAAACACATGTGAAGTTGGACCCACTAACAAACTTTCT CTCTATACCGACATGCTCGGTGACCCCATTTGCAGACTCCGTAACTCCCCTTCTTTCCTA ATGCTGGTTGCTGAAATTGGTGAAGAGATAGTTGGAATGATAAGAGGATGCATCAAAACA 25 GTCGCATGCGGTAAAAGCCTCTCAAGATCAAAAGCAGCTCTCACTAAACAAATCCCACTT TACACCAAACTCGCTTACATATTAGGCCTTCGAGTTTCTCCTAATCAACGGAGAATGGGA ATAGGGTTGAAGCTGGTAAAGAAAATGGAAGCGTGGTTTAAAGATAACGGCGCTGAATAT TCATACATGGCAACGGAAAGCGAAAATTTAGCGTCCGTTAAACTCTTCACAGAGAAATGC GGTTATTTAAAGTTCCGTACGCCGTCAATCCTCGTTAACCCCGTTTACGCTCACAGAACT 30 AAAGTTTCACGAAACGTTACTATCATTCCGTTAACTCCTTCAGACGCCGTTATTCTCTAT CGTAACCGTTTCTCCACCATCGAGTTTTTCCCTCGCGACATTGACTCAGTTCTAAATAAT AAACTCAGTCTCGGTACTTTTCTCGCCGTACCTTGTGGGTCCTACAGTGTTGAAAACTGG CCGGGCCCAGTAAGGTTTCTTTTGGGCCCACCTTGTTCTTGGGCCGTTTTAAGTGTTTGG AATTCGAAGGAGGTTTTTAAGCTTGAGGTACGTGGCGCATCGCGCGTGAAGCGGGCTTTT 35 GCTAAGACAACGCGGGTTTTGGATCGGGCTTTTCCGTGGTTGAAGATGCCGTCGGTTCCG GATTTGTTTAGGCCGTTTGGGTTTCATTTTTTGTATGGGCTTGGTGGTGAAGGCCCAAAG GTTGTTAAGATGGTTAGGGCCTTGTGTGGGTTTGCGCATAATATAGCTATGGAATATGGG TGTGGGGTTGTGGCTACTGAGGTTGCTTCTTGTGAACCGCTTAGGTTAGGAATACCGCAC TGGAAAATGCTATCTTGCGCAGAGGATTTATGGTGCATTAAAAGACTTGTGGAGGATTAT 40 AGTAATGATTCTGTTGGGGATTGGACTAAATCTGTGCCCGGAATCTCCATTTTTGTGGAC CCTAGAGATGTTTAA M&C PC933520WOA 98 SEQ ID NO: 28 Pisum sativum>Psat2g170520-promoter GGGTGTGTACCTGGTACATAGTTGCCAAAATTGCTTTCAAATTTTTAGTCTAATTGCTTT AGCGTTCTGCTATGTTTATTTATTCCTTGGAAGAAATAAAAACACAAAAGACAACTCATA CAGTTTATGGTTGCTACTACTACTTTAAGAGTCTTGATGTTAGGCATTCCCAAGGTTTCA 5 TGTATAAATGAATGGTTGATTTCTGAAGTCGTTGTGTTTGCAAGCTTTGATAAGACGTCT TCCCTGGTGTTTTGTTCTCTCAGGACTTGCGCGATCTCATGTGACTTGAGCCTGACCAAC CTTTGACATGAGATACTACTAGTTGCAAGTTCGTTCGCAACTTAATCTTATATGCTCTCA TCTAGGATGCCAACATGAACCCATCAATGAATGCCTCATACTTTATTTGATTATTGGCAT AAGGGAATTGAAATTGCAAGGGTGCCCCGATTGTTAGTCCTATATCAGTCTCCAAGATGA 10 AGCCGGTGCCACTTCCCCTACTGTTGGAAGATCTATTAGTGAACTACGCTACTGTTTTAT GGTTGAACGGGAATTGGATAGTCTGAAAAGAGCGAGCTCCTCGCTGCGATGGTAAGAGTC TTCAATGCGGTGGAACGAGGGTGAACTAGCAATATTAACATTCCAACGTGCAAGTTAGTA AAGTCAAATAAATATAGCTTATGAGTGTGAATATGATACAATATCTTTCATGGTGCATGG TAGATCTTGTATATAAGAGTCTGTTATAATTGTCTTCCTTTAAGATGCCAAGTGACACGG 15 AAGATCGTTCGATCAAATCTAAAGGTAGTTTATGCATAAATAATTAGGATCACATGTCTG ATTATTTGATAAGATTATATCTACGTCTGACTTGCCTACTCTATTGAGCCTACTCCTGAT TGGACCCTTAGGATTGGGCCATTATATTGAGCATTTTTAGCTGACTCAAAACTATTTTCT TAAATACTCCCTCTGTCCTAAATTATAAGATGTTTTGGGCATTTCACACAAATTAAGAAA TATAATTAATTTTCTATGGAAAAAAGAAATTATGTTTTACTTTACAATATTGTCCTTCAT 20 TTATTATATACGAAAGAGAAATTTAAAGAATCAGAGATAAAACTAATAATAAATAAAATA GGATATATTAGAAAAAACTATCATTAATATTGCATTAACATTATAAAATTTAATAATTTG GAACAATTTTTTTAAGTGAATTATAATTTTGAGGAATACTAAACATAATAGAACGAGGAT TATTTAAATAAATAGTAAAAAATACCTATCATTGTCTTATAATTTGGAGGAGTACTAAAC ATAATAGAATGAGGATTATTTAAATAAATTGTGAATAACGCCAATGTTATGCGGTGTAAC 25 GAGAATTTGATAGTCTAAAAAGAGTGAGCTCTTCTGATGCAGTGGTAAGAGTCTTCGGCG TGCGTAAAAGAATGAGATGAATTTATAAGGTTAGTACTCTAATATTCAAGTCAATAAATA CGAAAAAATGCTCTTTATGAGAATGAAAGTTAGATAATACTTAACGCATTATAGAGTTTA TTAGAGCTGTCAAAACAAACTATTCAACCTTAAAAGGTTCGGCCTTGACGAGCCTCAAGC TTTTTAGAGCAGGACAAAAACAACTCTACCTTGATGGGCCTGAAGCTTCTTTAGAGGTGG 30 GCAATTTTTTGTATTGTAATATGGACTTAAAAAGTATTATTCGAATTAAACGAACTTTTT TAAATAATTTTTTCAATAAAAATTAATATAAAAACTCAATAATATTTATTATAATACAAT AAAAACTATTATTTTAAATATAATACATTTTTAAATATTATATTATATTTTATAAATAAT ATAATATGTGTTTAAATATAATATATTAACTAAATTATATATAATAATAATAATAATTAA ATATCTAAAATAATAAAAACTAATATAATATGAGAAAATTATGAAATAAAACCTAGAATA 35 ACTAAATATTAAATTAGTTATTTTAATGTATAATATTAAAAAAAATATAAATAAAATTTA ATGAAGTGAAAGAAATAATGTTGTTTTGATATGTAGACAATCTAAATATTAATATAATTT TTTTTAATTTATAGTTATACAGTTGATTTATTATAGACTTTTTTAAAATCAGAAAGAAAA ACCTTTAAAAATATGGACTTCAAATTTTGACTTATTTTTGACGGCTCTAGAATTTATTTA TGAAATACTTTTAATGGTGTCTTCATCTAATCCGTTGAAGATGCGGAAAACCATCCGTTC 40 AAAGTTAAGACTAGATAATGCACGAGGAAACCCTTAGGATTGGCTCCGCTTTATTAGCTT GGAACAATTTTCTTAAAAATGAAATAAATATTTTTAATTGAGACAGAGAGTGTATTATTT AAATAAATAGTAATTAATACCTATCATTGTCTTATTCTCTCTCTCCGTCGTCTTCACTTG M&C PC933520WOA 99 TTATTTTTCTCTTGTTATTTATATAAACAGACTAGATATCTTCCAAATTAGTGCAAACTT ATCGTCGCGTATGTTGCTTACTTGACCCCTACTCCTACTTTCAGTCTCATATATTCCCAT TCCCTCGCAATTCTACCATCACCCATACCAGCACTCACCATTTCTTTCTATTTTCTTCTA CCAAAT 5 SEQ ID NO: 29 Brassica napus> CDY29565 cds:protein_coding ATGATAGTGGTTAGAGAATACGACTCGAGCAGAGACTTAGCCGATGTGGAGGCTGTGGAG CGACGTTGCGAGGTCGGACCAAGCGGCAAGCTTTCTCTCTTCACCGACCTTTTGGGTGAC CCGCTTTGTAGGATCCGACATTCACCTTCTTTCCTTATGTTGGTGGCTGAGATGGGTACG 10 GAGAAGAAGGAGATAGTGGGCATGATTAGAGGTTGCATCAAAACCGTTACATGTGGCATA AAACTCGATTTAAATCATAAATCCCAAAACGACACCGTTAAACCTCTTTACACTAAACTC GCCTACGTTTTGGGCCTCCGTGTCTCTCCTTCTCACAGGAGGCAAGGGATAGGGGTTAAG CTCGTGAAGATGATGGAAGAATGGTTTAGGCAAAACGGCGCCGAATATTCGTATATAGCA ACTGAAAACGACAATGAAGCTTCCGTTAATCTATTCACCGGTAAATGCGGTTACTCCGAG 15 TTTCGTAAACCGTCTATCTTGGTCAACCCGGTTTACGCCCACAGAGTCAACGTCTCGCGT CGTGTAACCGTCATTAAATTGGACCCGGTTGATGCAGAGTCGCTGTACCGACTCCGGTTC AGCACAATAGAGTTTTTCCCGCGGGATATCGATTCGGTACTGAATAACGAACTCTCTCTC GGGACTTTCGTGGCGGTGCCACGTAGCAGCTGTTACGGGTCTGGTTCAGGATCATGGCCC GGTTCGGCTAAGTTCCTGGAGTACCCGCCCGAGTCATGGGCCGTGTTGAGCGTTTGGAAC 20 TGCAAAGACTCGTTTCGGCTGGAGGTTCGTGGCGCGTCGCGACTGAAACGCGTAGTGGCT AAAACGACGCGTGTGGTTGATAAAACCCTGCCGTTTCTGAAACTCCCTTCGATTCCGTCG GTTTTTAAACCGTTTGGGCTTCACTTTATGTATGGGATTGGTGGAGAAGGCCCACAAGCG GCGAAGATGGTGAAGTCATTATGTGGTCACGCGCATAACCTGGCTAAGGATGGTGGTTGC GGCGTTTTGGCGACGGAAGTGGCAGGAGAAGAGCCGTTGCGGCAAGGAATACCGCACTGG 25 AAAGTGTTGTCGTGCGATGAGGATCTGTGGTGTATAAAACGGCTTGGGGAAGACTATAGT GATGGTGCAATCGGTGACTGGACTAAATCATCACCTGGTTCCTCCATTTTTGTGGACCCC AGAGAATTTTAA SEQ ID NO: 30 Brassica napus> CDY29565 promoter 30 ATTAAATAATTCAATATATTATTTAAAATTTTGTGGTTTATATTATATATTTTTATAACAATTT TATGTTTTTTAGACGTAGTTAATATACCAAAATAAAAAATAACCACCCCATAATAATAAA AGCTTGTTATATTTTTGACATTGATGAAGTATAATTTATTGTTTACCCAAATTATAGGAA ACATTTTTAAAATAAAAAAACATAATTTGTTATACTTTATTTTAAGAAAATACATTAATT AAACTATAAAAGACAAAAAAAATACTAAAAAGTAGGTAAACTTATTACTTCAGTGACATT 35 CAAATGTAAATAACTTTGAAAATTATGAGGCAATTTATATTGGTACTTCTATTTTAATAA TAGAGATAGAGAGAAAGGAGTAAGGACATTAAGAGAGATTTAATTTCATACTTATATTCT ATTTTTTTTAAATACTGCTATTTTGATTTATTTTCAACATATTACATATCGTATTTTGAT TAATTTTTATACGATGTGTTATGTTGAAAATAAATCAAATTATCTGCTGCATCATCAACC GTTAAAATATTGTCATACTTTCTTTATTCTATGAATTGATTTATAAAAAATCGGGTGGTT 40 AAAGTTTTTGTAATTCCCTAATCCAATAAATAAACAAAAAGCTCTTCAAATATCAGTTAA GGTCTACTTACTACCATTTCGAACTAACTAGTATTTCTATTAAAGCAAAAAAAAAGAGAG TAAAACAAAATAGATAATGTAGAGACTATAAATAGGATCGAGAAGAACTTTAAACCACTT M&C PC933520WOA 100 CAATCATGAACTACTCGCCTTCTCAAACTTTTAAAACTCATCATAAATAATAAAGTCTAT TTCTCTTTCCTTTTAAACCCAAATCCTATAAACTTTCATAGCTTCTCTGTTCATTACTTA TATCTCACGTTATACATATAGCTCCTATAAATGCTTCTCTTTCCTCTCGAATAATCTTCC CTCACTACTTTCTATATAAGACCCTTGTCTACTTCACTCTTCCTCTTAACTCTCTCTTCT 5 TCTCTTTGCCTCTTCTATCCGCTCTCCATATAAAAGAAAT SEQ ID NO: 31 Lolium perenne>KYUSt_chr4.9071 cds:protein_coding ATGGTCGTGGCGGTTGCTGCGGTGGTGGAGGAGGAAGAGGCGGCGGCGGTCGAGGTGTGG GTGCGAGAGTACGACGGTGGTCGTGACCGCGGTGGCGTGGAGGAGGTGGAGCGGGAGTGC 10 GAGGTGGGCTCCAGCGGCGGCGGCTCCGGCAAGATGTGCCTCTTCACTGACCTCCTCGGC GACCCACTCTGCCGCATTCGCAACTCGCCGGCCTACCTGATGCTGGTCGCGGAGATAGCA ACCGGCACCGGAACCGGCGGCGGCGGCGGTGGCACCAGGATTGTCGGCCTCGTCCGCGGC TGCGTCAAGTCTGTCGTCTCCGGCACCTCCCATGGCAAGGACCCCATCTACACCAAGGTC GCGTACATCCTCGGCCTCCGCGTCTGCCCCACCCACCGGCGGAAAGGGGTTGGGAAGAAG 15 CTCGTGGAGCGGATGGAGGAGTGGTTCCGGCAGAAGGGCGCGGAGTACTCGTACATGGCG ACGGAGCAGGACAACGAAGCGTCCGTGCAGCTCTTCACGGGCCGCTGCGGCTACTCCAAG TTCCGCACGCCGTCCGTACTCGTGCACCCGGTGTTCCGCCACGCCCTCGGGCTCTCCCGC CGCGTCTCCATCGTGAAGCTCGAGCCTCGCGACGCCGAACGGCTCTACCGCTGGCACTTC GCTGCCGTGGAGTTCTTCCCGGACGACATCGACGCCGTGCTGTCCAACGCCCTGTCGCTC 20 GGCACGTTCGTGGCGGTGCCCGCGGGGACGAGGTGGGACGGAGACGTCGAGGCGTTCATC GCGTCGCCGCCGGCGTCGTGGGCGGTGCTGAGCGTGTGGAACTGCATGGACGCCTTCCGC CTCGAGGTGCGCGGAGCTCCGCGCGTGATGCGCGCCGCCGCGGGCGCGACACGGATGGTG GACCGCGCGGTGCCGTGGCTCGGGATCCCTTCCATCCCAAACGTGTTCAGGCCGTTCGGG CTCTACTTCCTCTACGGCCTGGGCGGCGCTGGGCCCGGGGCGCCGAGGATGGTGCGCGCG 25 CTGTGCCGGCACGCGCACAACATGGCACGTCGCGGCGGGTGCGGTGTGGTGGCCACCGAG GTCGCCGCATGCGAGCCTGTCCGTACCGGTGTGCCGCACTGGGCGCGCCTGGGCGCCGAG GACCTCTGGTGCATGAAGCGGCTCGCGGACGGGTACACCCACGGCACGCTCGGTGACTGG ACCAAGGCGACGCCCGGACGCTCCATCTTCGTCGACCCTAGAGAGTTTTGA 30 SEQ ID NO: 32 Lolium perenne>KYUSt_chr4.9071-promoter TTGTCATGTCCACACCATCAGTTCCTGCTCTCTCCAATTTTGCTCTAGCTTTCACACTGG AAATAGATGCATCTGGTACTGGGTTAGGAGCAGTCCTCATGCAGCAGGGCAGGCCATTAG CATATTTTAGTAAAGCACTGGGACCAAAATCAGCCACACTATCCATATATGAAAAAAAGG CACTAGCTATTCAGGAATCACTAAGGAAATGGAGACATTATCTGCTTGGTAACCAACTCA 35 TTATCAAGACTGGTCAAAGGAGCCTCAAATATCTCTCTAGCCAAAGATTACTGGAAGGGA TTCAACACAAAATCATGCTCAAGCTATTGGAGTTTGACTACTCCATTGAATACAAAAAGG GAACAGACAACACAGCTGCAGATGCCTTGTCTAGGAAATATACTGATCCTATGGAGGAGC AATGTACTACAATTTCTGCAACTATTCCCACTTGGATGACTGAAGTGGTTGACACTTATG TCAATGAAACCAAGTGCACCCAACTACTCCAAGAACTGGCGATCTCTGCAACCAGTAACC 40 CTAAGTACACACTTACTTCAGGGATTCTCAGATACAAAAACAGAATTGTTCTGGGAACAG CTACTGACTTGAGAGATATAATCTTCAATGCATTTCACTCATCAATCTTTGGAGGATATT CTGGCAACAGGGTCACTCATCACATGATCAAGAGACTTTCCTTCTGGCCACATCTTAAAC M&C PC933520WOA 101 AGTTTACTGCTGACAAAGTAGCGCAATGCCCAGTTTGCCAGATATCAAAAATTGAACGAG TTCATTACTCCACTAAATATTCCTGATAGATAATGGGCAAAAGTGAGCATGGATTTCGTT GAAGGACTACCTAGATCCAAAGGCAAGGATGTGATACTAGTTGTCGTCGATCGACTAACC AAATATGCTCATTTCCTCACATTGGCTCGCCCATTTACTGCACATCAGGTAGCAATATTG 5 TTTATGGACAACATATTCAAGCTGCATGGACCTCCTAAAGTAATTGTGAGTGGAGACAAG ATATTCACCTAAAAAAGTTGGCAAGATATATTTACTACACTCAAAGTGGACCTACACTTT AGTACAACATATCACCCTGAATCAGATGGTCAAACTGAACGAGTCAACCAGTGCTTGGAG TAGTACCTTCGCAGCATGGCATTCAAGGAATCAAAAAAATGGGCAGAATGGTTACCAGCT GCAGAGTGGTGGTACAATACTTCATACCATACATCCCTGAAGACATCACCTTTTGAGGCC 10 CTTTATGGATACTCTCCCCCACAGCTGCATGACATAGCAGTTCCCTGTGATGTATCACCA GAAGTGCAAGTTACACTGCAAGAGAAGGATCGCATCCTCAAATCTCCGCAGCAAAATTTG ATGCAAGCTCAAAATCGAATGAAGGAGTACGACATTCCTGTCCCACAATGGCCGATCCAT TGGGAGAACATGTCACTTGAGGAAGCAACTTGGGAGGATGCCAAATTCATCGAGGCGACA TTCCCAACCTTCCAGCCTTGAAGTCTAGTCTTGCCCAGGGAGTATTGCCACAACCTGAAA 15 ACAGTGCTGCATTTCCCGACGCTGTCGACAACATGCAAAGATTAAGTTAAGCAACGGTCC AGATCGACAAGTCTTCATCGGACGGCTCCATTTCAATTTCGTTTGAACTGTCTGTCGTAT GTGACAAGTTATGTCAGTTTGTAAGGCTACTTTAATTTGAATCGCGGCTCCTCTGGGGAC TCTATATAACCAGGGAGAGTGGGTGGCTGGGTGGCGTGATAATGAAAACCCTAAAAACGC TATCTTCCAGTTTATCAGTTCTCCTTCCTGCCCAGTACTTATTTTCCTTATTTAGCTACT 20 GTTAGATCTTGATCTGAGCTAGTTCTTGCAAAGAATCCGTAGTCTCCTTTCCAAATCATG GCCGGGTCGTGACACATCCGATATTCCGATTGAACCTGACCCCGTAGTGGCGTCCGTCCG ATCCAGCCTAGCTTGCTTCTTCCTCGTTTGCCAGATAAAAGCCTACGGCCCAAAGCGCCT GGCTCTATGCTCCACAGCCTTTATTGGCTATCATCCTCGCCAATTAAACTCGTCGCTTCG AAAATTTAAAGAAGCGGCTAGTGGCTCTAGCCTTCAAATCGGCCTCTCCTCCTCCCTTCC 25 CTGTGGCTATTATTATTGCTGTGGGGATATATATCGTCCGGTAGTGCTTCGATCCGATTT CAGTGGTGTGTGGGGAGTAATTGTAATTGAGAAAGGCGGCAGGCTAGGGTAGCGAAAGAA CACCACCATCGCCCCCCTCCCCCACTACCCCGGCCCTCTCCCTCACATGATTCATACTCG CGCCCTCCCTCCGGCCCCGGACCTTCAACTGCCCGGCCACTAGCTATATATATAACCACT CGCCACACGCTCCACTCCTCTCTCTGTCTTTCGATCCTGGTATCTGGTTGGCCAACCATA 30 CATATTTCCACTCCTTGCTACCAAGAAAGTGAGCTCGATC SEQ ID NO: 33 Triticum aestivum>TraesCS5B02G405900.1 cds:protein_coding ATGGTTGAGGCGGCAGCTGTAGTGGCAGAGGAGGAGCATGACGCGGCGGTGCGGGAGTAC GACGACGTCCGCGACCGCAGCGGCGTGGAGGAGGTGGAGCGGGAGTGCGAGGTGGGGTCC 35 AGTGGCGGCGGCGGCGGCGAGATGTGCCTCTTCACGGACCTCCTCGGCGACCCGCTCTGC CGCATTCGCAACTCGCCGGACTTCCTCATGCTGGTCGCGGAGACAGCAACCGGCAGCGGC GGCGGCGGCGGCGCGGAGATAATCGGCCTCGTCCGCGGCTGCGTCAAGTCCGTCGTCTCC GGCGGCTCCCACTCCAAGGACCCCATCTACACCAAGGTCGCCTACATCCTCGGCCTCCGC GTCTCACCCAATCATCGGAGGAAAGGGGTGGGGAGGAAGCTCGTGGAGAGGATGGAGCAA 40 TGGTTCCGGCAGAAGGGCGCGGAGTACTCGTACATGGCGACGGAGCAGGACAACGAGGCC TCCGTGCGCCTCTTCACCTCCCGCTGCGGCTACTCCAAGTTCCGCACGCCGTCGTTGCTC GTGCACCCGGTGTTCCGCCACGCCCTCAAGCCCTCGCGCCGCGCGTCCATCGTGCGCCTC M&C PC933520WOA 102 GAGCCCCGCGACGCCGAGCGCCTCTATCGCTGGCACTTTGCCGCCGTCGAGTTCTTCCCC GCCGACATAGACGCCGTGCTGTCCAACGCCCTGTCGCTCGGCACGTTCCTGGCACTGCCG GCGGGCACGAGTTGGGACGGGGACGTCGAGGCGTTCCTGGCCGCGCCGCCGGCTTCGTGG GCGGTGCTGAGCGTGTGGAACTGCATGGACGCCTTCCGCCTCGAGGTACGTGGCGCACCC 5 CGCCTGATGCGCGCCGCGGCCGGCGCAACGCGGCTGGTGGACCGAGCGGCGCCGTGGCTC AGGATACCCTCCATCCCGAACCTCTTCGCGCCGTTCGGGCTCTACTTCCTCTACGGCCTG GGCGGCGCCGGGCCGGAGGCGCCGCGGCTGGTGCGCGCGCTGTGCCGGCACGCGCACAAC ATGGCCCGACGCGGCGGGTGCGGCGTGGTGGCCACCGAGGTCGCCGCCCTCGAGCCCGTC CGCGCCGGGGTGCCGCACTGGGCGCGCCTCGGCGCCGAGGACCTCTGGTGCATCAAGCGG 10 CTCGCCGACGGCTACAACCACGGGCTGCTCGGCGACTGGACGAAGGCGCCGCCCGGGAGC TCCATCTTCGTCGACCCCAGAGAGTTTTAG SEQ ID NO: 34 Triticum aestivum>TraesCS5B02G405900-promoter TCGTCAAAATATAAAGTTATTGAAGATCTTCTTGAATTTTGGCAACACCTTTAACATTCT 15 GGATACCTCTCACTAAATTAGTAATTATTTTTTCTACACAACCACAAACAATGCAAGGTG GTTAAGTTGGGAAATGCATTTGAAGGGGCTTCGACTGGTTTACCCCATTGTCAAGTTGCA CTTCTAGCATGGTGGTGAAACAACAAAGTATACGTTAAAAGCATAGCAATATTACCGTAA GCCATCACATGCAAATGTTTTTCATTGGCTCGGCTTAGTAATTCTTTCCAACTAATTTGT GCGGCTTGCTGAAATGTGAGGTTTGTTGTATCCATGTGCTCTTCAAATTACTTTACTTAA 20 CTTAGAGGTTGTTGAGAGTATCCACGGTACATGTAAGTGGGTAGCAGCAATCTCGTGTTG GTCTCAACCGTTGGTAGTAAATTAGATGGTTCGTATGTATGAATTATGAAAGCGAACCAT CACCACTACCTTGAGTGTTTATAGGAGTGAAGATATTACACTACAAAAATTAGTGATCAA AGTTGGATATGAAAATCATGTCGATATTCAAATCAAAAAACCTTTCAAGCGCGAGATGGA GTAAATAAACTAATCATGTTTTAGTGGCTAAAGTGACTAATGGGTTCCGAATGAAGCAAG 25 TGGCTAGTGTTTTTTATTGGATATTCGTCAAAATCGAGCACCCAACCCCCCCTACAACGC CACATGCGCCTGTGCGGGAGCTGCCACTTGGTAACAATGGATACTAAAGTACTAAGAAAT ACTACAATAGAAACATTCGTTGTTGTGAAAATATAGAAAACCCAACTTAATCAAATTGGG GATTATGTGCGCGGATATGTTGTTGTGAAAATATAGAAAACATTCGTTGTTGTGAAAATA TAGAAATATTCGTTGTTGTGAAAATATAGAAAACCGACTAGTTACAAAACTTACAAACAC 30 ATGCAAGTTTGCTGTTTTGGACCCGAAAAAAAGGTTTATTGTTTTTGAGAAGAGCATCAA TGGAGTTACGGAAATGTTAAAATTTTGATGTTAGATTTTTAACATTTCCCAGCTAATGCC CAGATTTGCCATGCGAAAATAGATAAATTTTGCCATGTTTGTGTATAAGGAGTTGCCATG TTCTAAAGTGAAAATTGTCATACAATGTAAAAATTTGTCAGAAAATCATCATGTTGCGTG AGTAATTGCCATGCGTTCGATCGGAATCCGACGTACCTAGGGAGTCAACCCTAAAAATAT 35 GATATCGAGTGTATATTGGAGTCCTATAAAAAGTTACAAAACTTAGACCGACAAGTTACA AAATGTAAAAAATGATCTGGCTGGTGCATCTCCAATCCGATGATAAAGTTGTTCTTTTTT TCCTTTTCTTTCACCACAAATGATGCTCTTACTCTCATCGTCACTAGGGTAGTAGTGAGG GTTCGTTTAGCCCACATATCCTTTTCCGCGATGCCCCCAACTACGTGCTCGGTCGAGGAT AAAATCTCTGATATATGTAGTAGCCCCGTCACAGCGGCAACTTGGCCACTTGCTCGCCTC 40 CTTGGAGGCAGCTTGTCCTGGGTCTGGCAAGGCGATTGCTTGCTCCTAGCATAGGAGGTT TCTACATGCACAATACAGAAGGTGAAGAAGGCTGTCAGGACCATAGGCAAGAAAAGTGGC GTCATTGGTAAGGCATCACTGACTGCTTAATGAAATATACTCAGTTCCGTCATCCTCCAT M&C PC933520WOA 103 CTCCATCTGGCTCATCTCCGTGGATTTGGTGGCATCCGATGTATTCAATCTTGATTTCGT TGAGGGTTTCTTCGAGACATTGATGTGCGAGGGTTCGTGGTTGTTACATGTTATGTATGT TTGTCGGGTTGATCGTGGGGTCATGGAGTTTGTTTCTTAGCAAATAATTTGCACGTGTTC TGTATGATGTTTTTTCGTCCGTTTTTTGTTAATTAACTTTCTCTTCTTCTTAATTAATTG 5 ATGACCCAATTTTTTTGCCACCGTTTCAAAAAAAAAAAACAGTAACCAGTACAGCATCGT AGCACTGAGAAGCAGTGTGGCCACCCAGCCACACTCCCGCTCTCCTCTCCCCCTCTCCTT CCCTGCTCTTCCTGTGGCCAGATCGCTACTCTGATCCATCCTAGTCGTAGCCGCCCTTCC TCCTCGTCGCCAATTAAAAAGCCTACGGCCCCGGAGCAACTCGCTCCGAGCTCCACAGCC TTTATTCCCCAATTAAACCTTCGATTCCAAAGCTTTTAAGATGCGGCTAGTGGCACTCCT 10 CTCCGAATCGGCCGTACCGATTCCATTCCTAACCATTTGAGTGTGTGCGTCGTGGTAACT GAGACGGCGGCAGGCAGGCAGGCACGCAAAAGAGCGCCTCCGTCTTCCCCCGCTCTCACG TGATTCACACCCGCGCCCTCCCCCCTCCTCGCCTGCCCCGCCGTCTAGCCTATATATCCA GTCTGGCCAAGCCAAGCCCTCAGCTTCTCTGCTACTACCTCTCACTTTCTCTCGAACCTG TCCGCCTGCCACCGGTCCTTACACTGCTTACAGTCAGGAGAGTGAGCTCGATCGCTCGGC 15 C SEQ ID NO: 35 Triticum aestivum>TraesCS5A02G401100.1 cds:protein_coding ATGGTTGAGGCAGTGGCGGAGGAGGAGCCGGTGGTGCGGGAGTACGACGACGTCCGCGAC CGCGCCGGCGTGGAGGAGGTGGAGCGGGAGTGCGAGGTGGGGTCCAGCGGCGGCGGCGGC 20 GAGATGTGCCTCTTCACGGACCTCCTCGGCGACCCGCTCTGCCGCATTCGCAACTCGCCG GACTTCCTCATGCTGGTCGCGGAGACAGCAACCGGCAGCGCCGCCGGCGGCGGCGCGGAG ATAATCGGCCTCGTCCGCGGCTGCGTGAAGTCCGTCGTCTCCGGCGGCTCCCACGCCAAG GACAAGGACCCCATCTACACCAAGGTCGCCTACATCCTCGGCCTCCGCGTCTCACCCAAC CATCGGAGGAAAGGGGTGGGGAGGAAGCTCGTGGAGAGGATGGAGCAATGGTTCCGGCAG 25 AAGGGCGCGGAGTACTCGTACATGGCGACGGAGCAGGACAACGAGGCGTCCGTGCGCCTC TTCACCTCCCGCTGCGGCTACTCCAAGTTCCGCACGCCGTCGTTGCTCGTCCACCCGGTG TTCCGCCACGCCCTCAAGCCCTCGCGCCGCGCCTCCATCGTGCGCCTCGAGCCCCGCGAC GCCGAGCGCCTCTACCGCTGGCACTTTGCCGCCGTCGAGTTCTTCCCCGCCGACATCGAC GCCGTGCTGTCCAACGCCCTGTCGCTCGGCACGTTCCTGGCGCTCCCGGCGGGCACAAGT 30 TGGCACGGGGACGTTGAGGCATTTCTCGCCGCGCCGCCGGCTTCGTGGGCGGTGCTGAGC GTGTGGAACTGCATGGACGCCTTCCGTCTCGAGGTGCGCGGCGCACCCCGTCTGATGCGC GCCGCGGCCGGCGCAACGCGGCTGGTGGACCGAGCGGCGCCGTGGCTCCGGATACCCTCC ATCCCGAACCTCTTCGCGCCGTTCGGGCTCTACTTCCTCTACGGCCTGGGCGGCGCCGGG CCGGGAGCGCCGCGGCTGGTGCGCGCGCTGTGCCGGCACGCGCACAACATGGCCCGACGC 35 GGCGGATGCGGCGTGGTGGCCACCGAGGTCGCCGCCCTCGAGCCCGTGCGCGCCGGGGTG CCGCACTGGGCGCGTCTCGGCGCCGAGGACCTCTGGTGCATCAAGCGGCTCGCCGACGGC TACAACCACGGGCTGCTCGGCGACTGGACCAAGGCGGCGCCCGGGCGCTCCATCTTCGTC GACCCCAGAGAGTTTTAG 40 SEQ ID NO: 36 Triticum aestivum>TraesCS5A02G401100-promoter AATATTTTGCTAGTCAAATATGTTGTCAAATTTTAACCCAAAATACAATAGGGACTAATA AACCAGTAAGAGCATCTTCAGCAGATGCGCAAAATAAGTCGTGCACTGGGTAAAATGAGT M&C PC933520WOA 104 ATATAGCGTGTGCGACCGAAAGTAGCGCTCTAGCAGCCACGCTATAATCCCGTGCACGGT AAATAAGATTCAGCGCGCTAGAGAAAACTGCATCGCACGTCCTACTACCTCCGTCTAGGT GAATAAATCATTCACGTAGTTCTAGGTCATCAATTTGAGGAATTAAATATGTGTCATGTG TCATGAAAAGTATATCACTAGATTTCTACACGGATGTAGTTTTTTAATATATATTTTTTG 5 TCACATATAATACATATTTAGATAGTTAAATCATCGACCTAAAACTACACGAATGACTTA TTCACCGAGACGGAGATAGTATATTTTGTGCGCCGGCTCCAGCGCGCTGGGCAAGACTGC ATCACATGCTAGGTATTTTGTGCGCCAGCTCCAACGGCAAAATTTTACCGCGCCTGAAAA AGCAACCACCCTCTCGTGCCTGCAAAAACGCTATTTTGACGCTGCAATTTTTTTAGCGCA CCACCCTACCGCGCGTCTGTTCTCTAAGAGCATCTCCAGTCGTGCCCCCAAGAAGACCCC 10 CCAGGCGACTTTTCAGCCGCCGGCGCTAAAAAATCGGCCCAGTCGGGCCCCCAAGGGCCC AGTTTTCGCCGGCTCGGGCCGAAATTGGCGCCGGCGGACCCAACCCGGACCCGGCGCGCT GGGGATCGCTCGGGGGCGCCGGGCGAATCTTTTTTGGTGCGAAGAGCCGCGGGCCAGCCG CGTCAGCGACTCGACTCTCTTCTCGCCGTTTCGTCGTCCTCATCACCTCATTTCCCACGG CGAATCAATGCCAAAGCTGTGCGCTGCGCGCGCTGCCGCGCCGGTCAGTCTCAATTGATG 15 CCTCACAGGCAGCGCAGTGAAGGCCGGGCGACGCGCGTCCCCTCGGCCGGCACGCCATCA AGCCACGCGTAACGCGCCGGCCAAGCCGACCACGCGGCGCCTCCGCCTCTATAAGCCGAC CACCAGCGCGCCGGAGGCACGCACAGACACTCCACCCCCGACGCGCCCATTTCTCCCCCC TTCCTCCTCTCGCCGTCTCCAGTCCATTAGAATGTCCGAGCGCTTTCCAGGCGACGGCGC GGCGGCGAATGGCTTCGGGCGCCGCCATCTTCACGAGGACGAGACTCGCCTCCTTTTCGA 20 GGTCGAGTACCCGGTCCCGCCGGACATGCGGGTGCCCGGGGCTTGGAGGATCAGCGCCGG CGGCGTGCTGGTGCCACCACCACCCACCGGGGCGGCGCAGCGTGCGGAGATCGCACGTAT CCGCGCGTCCCTGCCACGGGCGGCGAGGGAGGGGCCACGGTACGTCCCCAACAGCCCGCT CTGGGAGCCCTACTTCCGACGCCGTCACGCCGAGTAGCTCGAGGCCACCAACGGCGTCGT GCCCTCCGGCAGGCTCAGCTCCGATGGCCGGCGCCGGTGGTGGGGCGTGCCCGGCCGCAC 25 GTTGGAGGCCGTCCTCGAGTACATCGAGGGCGGCAACACGCCGCGCCTCGAGTACCCCGC TCCCCCGTCCTTCTCACGCCGTCGTGGAAGCTCCTAGACGCCGAAGCGCATGGAGCGGCC CGGGGCGTCCTCCTCGTCCGGCCGCTCGTCGGGCTCTCCCTGCCTCCGCCCCGTCAAGCC GGAGCCCCAGGACACGCCTGTCAGCGCGCGCACCCGCAGCTCCGGTGTCCGCATCGGCGA CAACGCCTCCCCCACCGGCCGCTTCGTCCTCGTCAAGCCCAAGCCGGAGCCCGGCCTCCC 30 CGCGGAGTACGAGGAGATAGCCCGGCACGGCTTCTCCGACGAGGACGCCCTGCGGTGGGC GCGGGACGACTACCTCCGGACGGAGATGACCCGACAGACCCGGGCCCTGGAGGAGATCGC CGCCCGCAAGCGTGGGCGCGAGGACGAGCACGGCATCGTCGTCCTCGACGGCGACGACGA CGACGACGCCCCCGGACCGTCCAACCCGCCGCGCCAACCCGGGGAGGGATGCAGCAGGGA CGGCGGCAGCAGCGGCGGAGACGACGACGACGACTACACGCGGTTCTACCGCCTCCTCGG 35 CATGTAGAACTGCAAAGGCGGGCGGCGGACGGCGAGCGGCGAGGTAGACGGCGAGGAGCG GCTATGGAGACGGCGGGGGCAGCCCGTGCTAGTTTATTTTTCTTTTTTTGTAAAATATGT TTAAATTTGAACGAACTCACCGATGTTTACGATAAATTTGAGCCGTGTTTGCGCCGTATT TGACTTTTTCAAGAAAACATGCACCGCGACTGGGGGGCATCACGCCCCCAGTGCGCGGTT TAGCGCCGGTGCACCCCTAAAGAGCAATTTTTAGCCCCTTCTGGAAGGCCAACAGCTGGA 40 AATGCTCTAAGTAGCAGCCGCGCAGCCCCGTCGCAGCAGCAACCAACGAGTATCACCGCA CCGAAGCAGTGTAGGCAGCCACACTCGCTCTCCTCTCCTCTCCTCTCTGCCTTCCACCTC GCGTCGCCAATTGAAAAGCCTACGGCCCCGGAGCGAGTCGCTCCGAGCTCCGCAGCCTTT ACTGCCCAATTAAAACCCTTGGCTTCCAAAACTTTTAAGATGCGCTAGTGGCTACTCCTC M&C PC933520WOA 105 CCCGAATCGGCCGTACCGATTCCTAAGCATTTCATTCCAGTGCGTGGGCGCACTGGGTGG GTCGTAGTAACTGAGACGGCGGCAGGCACGCGAAAGAGCGCCTCCTCCTTCCCGGCTCTC ACGTGATTCATACCCGCGCCCTCCCCTCGTCCGGGCCGGGCCCCCCTCCCCGCCTGCCCC GCCGTCTAGCTATATATCCAAATCGCCGAGCCTCCACCCCCTCCATCTCCCTTTCTCGCG 5 AGCTCGTCCGTTTGCCTGCTACCCATCTTTGCACTGCTTACAGTCAGAGAGTGAGTGAGT GAGCGAGCTCGATCGGTCGGTC SEQ ID NO: 37 Triticum aestivum >TraesCS5D02G411300.1 cds:protein_codingATGGTTGAGGCGGCAGCTGCAGTGGCGGAGGAGGAGCATGACGCGGCGGTGCGGGAGTAC 10 GACGACGTCCGCGACCGCGGCGGCGTGGAGGAGGTGGAGCGGGAGTGCGAGGTGGGGTCC AGCGGCGGCGGCGGCGGCGGCCAGATGTGCCTCTTCACGGACCTCCTCGGCGACCCGCTC TGCCGCATTCGCAACTCGCCGGACTTCCTCATGCTGGTCGCGGAGACAGCAACCGGCAGC GCCGGCGGCGGCGGCGGCGGCGCCGAGATAATCGGCCTCGTCCGCGGCTGCGTCAAGTCC GTCGTCTGCGGCGGCTCCCACTCCAAGGACCCCATCTACACCAAGGTCGCCTACATCCTC 15 GGCCTCCGCGTCTCACCCAACCATCGGAGGAAAGGGGTGGGGAGGAAGCTCGTGGAGAGG ATGGAGCAATGGTTCCGGCAGAAGGGCGCGGAGTACTCGTACATGGCGACGGAGCAGGAC AACGAGGCGTCCGTGCGCCTCTTCACCTCCCGCTGCGGCTACTCCAAGTTCCGCACGCCG TCCTTGCTCGTGCACCCGGTGTTCCGCCACGCCCTCAAGCCCTCGCGCCGCGCCTCCATC GTGCGCCTCGAGCCTCGCGACGCCGAGCGCCTCTACCGCTGGCACTTTGCCGCCGTCGAG 20 TTCTTCCCCGCAGACATCGACGCCGTGCTGTCCAACGCCCTGTCACTCGGCACGTTCCTG GCACTGCCGGCGGGCACGAGTTGGGACGGGGACGTCGAGGCGTTCCTCGCCGCGCCGCCG GCATCGTGGGCGGTGCTGAGCGTGTGGAACTGCATGGACGCCTTCCGCCTCGAGGTGCGT GGCGCACCCCGCCTGATGCGCGCCGCGGCCGGCGCAACGCGGCTGGTGGACCGGGCGGCG CCGTGGCTCCGGATACCCTCCATCCCGAACCTCTTCGCGCCGTTCGGGCTCTACTTCCTC 25 TACGGCCTGGGCGGCGCCGGACCGGGGGCGCCGCGGCTGGTGCGCGCGCTGTGCCGGCAC GCGCACAACATGGCCCGACGAGGCGGCTGCGGCGTGGTGGCCACCGAGGTGGCCGCCCTC GAGCCCGTGCGCGCGGGGGTGCCGCACTGGCCGCGGCTCGGCGCCGAGGACCTCTGGTGC ATCAAGCGGCTCGCCGACGGCTACAACCACGGCCTGCTCGGCGACTGGACCAAGGCGCCG CCCGGGCGCTCCATCTTCGTCGACCGGCCTGCTCCTTCCAGAGCACTCTGCCGTGCCAGT 30 GTTCCCTCCCGAACCTTCTGA SEQ ID NO: 38 Triticum aestivum>TraesCS5D02G411300-promoter TGTGCAGATCGCAGAAGTGCCGTACTTTTGGTGCTAGGATCGGTAAATCGTGAAGACGTA CGACTACATCAACCGCGTTGTCATAACGCTTCCTCTTACGGTCTACGAGGGTACGTAGAC 35 AACACTCTCCCCTCTCATTGCTATGCATCATCATGATCTTGCGTGTGCGTAGGAATTTTT TTTAAATTACTATGTTCCCCAACATTTACTGGTTCACCCCATTGTCAAGTTGCACTTCTA GCATAGTGATGAAACAACAAAGTATACGTTAAAAGCATAGCAATATTACCATAAGACGTC ACATGCAGATGTTTTTTATTGGCTAGGCTTAGTAATTCTTTCCAACTAATTTGTGCGCCT TGCTGAAATGTGAGGTGTGTTGTATCCATGTGCTCTTCAAATTACTTTACTTAACTTAGA 40 GGTTGTTGAGAGTATCCACGGTACATGTAAGTGGGTGGCAGCAATCTCGTGTCGGTCTCA ACCGTCGGTAGTAAATTAGATGGTTCGTATGTATGAGTTATGGGAAGTGAACCATCACCA CTAGCTTGACTGTTTATAGGAGTGAAGATATTACACTACAAAAATTAGTGATCAAAGTTG M&C PC933520WOA 106 GATATGAAAATCGTGTCGATATTCAAATCAAAAAACCTTTCAAGCGTGAGATGGAGTAAA TAAACTAATCATGTTTTAGTAGCTAAAGTGACTAACGGGTTCCGAATGAAGCTAAGTGGC TAGTGTTTTTTATTGGATATTCGTCAACATCGAGCACCCAAACCCCCCTACAACGCCACA GGCGCCTGTGCGGGAGCTACCACTTGGTAACCATGGATACTAGAGTACTAAGAAATACTA 5 CAATAGAAACATTCGTTGTGGTGAAAATATAGAAAAAACAACTTAATCAAATTGGGGATT ATGTGCGTGGATATTTCCAGTTTCCACCGACTAGTTACAAAACTTACAAACACATGTAAG TTTGTTGTTTTGGACCAGAAAAAAAGGTTTATTGTTTTTCAGAAGAATATCAATGGAGTT ATGGAAATGTTAAAAATTTGATCTTAGATTTTTAACATTTCCCAGCTAACGCCCATATTT GCCATGCGAAAACAGATAAATTTGCCATGTTTGTGTATAAGGATTTGCCATGTTCTAAAG 10 TGAAAATTGTCATACAGTGTAAAAATTTGTCAGAAAATTATCATGCCGTCGCGAGTAATT GCCATGCGCTCGATCGGAATCCGACGTACCTAAGGGAGTCAACCCTAAAAATATGACATC GGGTGTATATTGGAGTCCTATATAAAGTTACAAAACTTAGACCGACAAGTTACAAAACGT AAAAAATGATCTGGATGGTGCATCTCCAATCCGATGATAAAGTTGTTTTTTTCTTTTCTT TCACCACAAATGATGCTCTTACTCTCATCGTCACTAGGGTAGTAGTGAGGGTTCGTTTAG 15 CTCACATATCCCTTTCCGCGGTGCCCCCAACTCCGTGGTCGATCGAGGATAAAATCTCTG TTATAAGTAGCAGCCCTGTCAGAGCGGCAACTTGGCCGCTTGCTCGCCTCCTTGGAGGCA GCTTGTCCTGGGTCTGGCAAGGCGATTGGCTGCTTCCTAGCGGAAGAGGTTTCTACGGGC ACAATAAAGAAAAGGAAGATGGCTCTCAGGACCATAGCACGAAGAGTGACGCCATTGATA AGGCGTCAGTGGCTGCTTAATGGAGGGTACTCAGTCCCATCGTCCTCCATGTGTGTCTGG 20 CTCATCTCCATGGATTTGGTGGCGTCCTATGTATTCAACCTTAATTTTGTTGAGGGTTTC TTTGAGACATTGAAGTGCGAGGGTTCGTGGTTGTTACATATTATGTATGTTTGTCGGGTT AACCGTGAGGTCATGGAGTTTGTTTCTTAGCAAGTACTAATTAAAATACAAAAGATAATT GTTTTAGAGAGTAGTCCACGTGTGTTCTATATGTATCGTTTTTTGTCTGGTTTTCTGTTA ATTAAGTGGGTAATTCTCTTCTTCTTAATTACTAATTGATGCTTTTGCGTCCTTTTCCAA 25 TAAAAACAAACAACCAGTAGGCAGCATCGGAGCACTGAAGCAGTGTGGCCACCCAGCCAC ACTCCCGCTCGTCCTTTTTGCGCTTCCTGTGGCCAGATCGCTACTCTGATCCATCTAGTC GTAGCTGCCCTTCCTCCTCGTCGCCAATTAAAAGGCCTACGGCCCCGGAGGGCCTCGCTC CGAGCTCCACAGCCTTTATTCCCCAATTAAACCCTTCGCTTCCCAAGCTTTTAAGATGCG ACTAGTGGCCACTCCTCTCCGAATCGGCCGTACCGATTCCTAATAACCATTTCAGTGTAT 30 GCGTCGTGGTAACTGAGACGGCGGCAGGCAGGCACGCAAAAGAGCGCCTCCGCCTTCCCC CGCTCTCACGTGGTTCACACCCGCGCCCTCCCCTCCTCGCTTGCCCCGCCGCCTAGCTAT ATATCCAGTCTCGCCAAGCCAAGCCCTCAGCTCCTCTGCTACCTCTCACTTTCTCTCGAA CCTGGCCGTCTGTCTGCCTGCCTGCCACCCATCCTTGCACTGCTTACAGTGAGGAGAGTG AGCTCGATCGGTCGGCC 35 SEQ ID NO: 39 Beta vulgaris >KMT09828 cds:protein_codingATGGTGATGATGAAGAAGAAAGTAGTAGTAGTTATTAGAGAATATGAGAAAGAAAGAGAT AGTAAAGAAGTTGAAGAGGTAGAAAGAAGGTGTGAAGTTGGTCCCAGTTCCAAAGTTTCT CTATTCACTGACCTCTTAGGGGACCCTCTTTGCCGCATACGCCATTCCCCTGCTTTTCTC 40 ATGCTTGTAGCTGAAGTGGTGGAAGAAGATGAAGAAACTGGAAAAGAAAAAAGAGAGATA GTGGGAATGATTAGAGGATGTATTAAAACCGTTACATGTGATAAAAAATTTCCACGACCT TCCAAGTCTACCCTTTATGACGCTACTAAACCTCTCCCTATTTTCACCAAATTAGCCTAC M&C PC933520WOA 107 ATCTTAGGCCTTCGTGTATCTCCTTCACATAGGCGGATGGGGATAGCCTCGAAGTTGGTG AATAAACTTGAAGAATGGTTTCGTCAAAGTGGAGCTGAGTATTCTTACTTAGCCACTGAA AGTAAAAACGAAGCTTCTATATCTTTATTTACTCAAAAGAGTGGGTACTCAAAATTTCGG ACTCCTTCAATATTAGTCCAACCCGTTTTTGAACACCGGGCTAAGGTATCGGGTCGGGTG 5 ACAATCATCCGGGTTCGGCCTTCTGATGCCGAAATCCTCTACCGCAAGAAGTACTCCACT ACGGAATTCTTCCCACGAGACATCGATGCCGTGCTACACAATAGACTAACACTTGGTACG TTTGTTGCTGTGCCGCGTGAGTCATACGCGCCAGGATCATGGCCTGGGGCGGAGAAGTTT TTGTCGAACACGCCCGAATCGTGGGCTTTACTTAGTGTTTGGAACTGTAGTGAAGTGTGG AGGTTGGAAGTACGTGGTGCATCGCGCGTGGGTAAAGGATTGGCTAAAACAAGTAGAGTT 10 ATGGACCGGGCTTTACCTTTTCTTAAAATTCCTTCATTTCCAGAGGTTTTCAAGCCATTT GGGCTTCATTTTTTGTATGGGCTTGGAGGTGAAGGCCCACAATCTGTGAAGTATATAAAG GCTTTATGTGGATTTGCACATAACTTGGCCCAAGAAAAGGGGTGCAGTGTAGTGGCTACT GAGGTAGCTTCAAGAGATCCGCTTAAATTAGGAATCCCACACTGGAAGAAACTATCTTGC GCCGAAGATTTATGGTGTATTAAAAGACTTGGGGAAGATTACAGTGACGGGACTGTGGGT 15 GACTGGACTAAGTCACCGCCTGGTGTATCTATCTTTGTTGACCCTCGAGAGTTTTAA SEQ ID NO: 40 Beta vulgaris >KMT09828 promoterCTTTGTTTAAATAGTTTTAAAAAAAGCAGAGTCCCCTCAAAAAAAAACGGAGCCAAACTA CTTGACCCGCTGACCCACAAAAAGAAAAGCTAGTTTGCGCAAAACGGCGGTCGAGTCTTA 20 GAATGCGAAACTGGCGGCACTCTATCAGCCACGTCGTTAAATCCCGCAGTTACGGAATTT CGAAAACTACTTAAAACCAATTTTTACGTACAGCTAAAAACAAAAAATAAATAAAGATTA AATATGAATTAAATTAACGCCTTAAACATCATTAAACTCATGTATGATGTATGTATGTAT CATTGTTCCATACAACAGAGGAGCACATAACATGAATAGATGATAGTGATACCACAACAT GCATAGCATGATCGATCTAGTTGAGAAAGAAAAAAAAGCAAGAAAGGAACAAGACCCTAG 25 CCAGAGAAAAATAACGAGAGGATCGATATGATATTATTATGTTTAATGGTAATGAAATGT CTTGATCATCACATGATGGTGACGGTGATGTTGATGGTGATAGTGATGTTGATGTATGAT AATGGCAATGTATGGAAGTGGAGGGGAGGACTGATGAGAGATTTTGGAAAAGGTAATCAT TACAAATGGGGGTCATTATGGCATACGGTTCTTTAGCGCAGCATTAATGATTAAATTGAA AAGGAGAAAAATTTAATTACTTTCAGAATAGATATACTCACTTAGCTCAGCCGCCATTTC 30 ATTTTCACATCATTATTTATTTATTACTAGCACCTACAGATATGCCATAACAAAATTAAT GTAATTATGACTCTTGCCATTATAGAGTGATAATCTTACATATTGCAATTAATAATTCTT TTTACTATTTTACTCTTTTATAGTGCACTGTTCATCATTAGTAAATGACTTTTTTCATTG TTTTTTGGTTTTTCAATTTGATCAAACAGCCAAAAAAGTGTTTGGTTATTGGTTTTTGGT TGGCTGAAAAGCTGACTTTTAAACTAAAATGAAAACGCTGCCTGGAGCAGCTTTTTGGAT 35 TCGCTTTTGGCTTATTCGCTCACAAACTCTCTAATAAACAGCCAACAGCCCCAATTAGTT TTGCCAATTACTTATGCAAACAGTCGGCTTAAACAGACAGCTTAAAACGCCAATACATAT AGCTAGGGGTGTTCGCGGTTTGGGTTGGTGCGGTTTGGGATGTAACCCAACCCAAACCGC AAATATACGGTTTGAAAATTTTCCAACCCAAACCGCACCGCAATTTTTCCCAACCCAACC CAACCCAAACCGCGCAAAACGGTTTGGTGCGGTTTGGGTTGCGGTTTAACCATTTTTTCC 40 TCCGCGACATTACTTTTGTAATCTAAGATTAAAAAAAATTGTTACTAAGAGTCCAAATAG ATATGAAGTTGAATTCATAATTAAACATATATAATGATTTATCAAAAGTGTTAAAGTTTT TTACCTCAAACTTCTTTTTTCAATTTTATACTTGTATGTCACCTTCATAAGTGTATTTAA TTTTATTTTTTTTTAAATATGTAACTTTGCGGTTTGCGGTGCGGTTTTGGTTTACAAAAA M&C PC933520WOA 108 TTAAAACCCAACCCAAACCGCAAACCGCGGTTTGCTTGAAATCATAACCCAAACCGCAAT TTTGAAACCGCAAACCAACCCAAACCGCGAAAAAACGGTTTGGGTTGCGGTTTGGTGCGG TCCAGACCGCACTTTGAACAGCCCTACATATAGCCAATGCTTCCAGTTAGTCAAACAAGT CAACTAAACTAGCCAACAGGCACCAGCCGATTGTCAAACAACCACATATACTCAACAACA 5 ATAAAAGTTCATCTTTGATGTAAAGCACTACATTGGGCAAATTTTTCAAGTTTTTATCCG ATATAACATAATATGATACATATTGAAATCATAGTATCTTGATAATAGAGCTTTTTGTTC AAGTTTGATTTCATTACCATCAAAAAAGTGAGAATTACGTTCGATTTGACTAGAATGAGA GCATCTTGAAATAAGGTTGTTAACGTACGGTATAAATAATAGCGAGAAAACATACTAATC CAATTGAAATAGTCAAATTATTTAGAAGAAATTGTGGAAATTCTAAGAAACACTCTTTTT 10 GTCAAGCTACCAAAGAAAAGAAAAGTTAGTGTTTGGTGTTAGACTTTTTTTTCTCCAATC ATTTCCCTTCTTTTCCCCACCTTTTTTCCTTCTATTTTTTAGAACTTGATATGGTTCAAC TAATATATTAACAATCTTATTCAAAAAAACAGATACGAAAAAAAGAAAAGAGATTCTATA ATGTTGATACTTTTTTTTCCATTCAATTAAAAAAAAGCAGACAAAAAGGTGGGTATTTAA AGAATAGAAAACAAATTAATCAATCCATCCATCTTCCATTAAAAATCATTCTGAATTCCC 15 AAAAATAACAATTTATATTCTCCCTTTTAAATACAAACTTAACGGTTTTTCTTCTTCTCC TTTCCTATAAAAACAACCAAAAAAAACAAATCTTTATATTATATTCGAAATTCAACATCC TCTGTTCTACAACTATATCTATCTATTCTTCTCTCTATATATAAATATATTTATCATCCT TTGCTTATTATAGCCTTCTCTATACATACATAAAACATAAAAGCTTCCTTACAATTCACG TTATTATTTTCTTCTTACTCATCTTCTACCTCTGTCATATTAGCGGCGCCTCCTTGAAAA 20 AGCTGCAATTTTTACATTTTATATCTATCATATTTTTCTTCCCCCAAAAAACAGATTATA TTTGTTTGCTTCTATGATATGAATTTCGTGTCTATATATATAGAGGTACATGTAGAGTAG TTTTCCATAACAACACAGTGAGTTTGAGTTGTTAAGAGATAAAAGAAATAGAAGAAAAGA AAAG25 SEQ ID NO: 41 Gossypium raimondii >KJB48700 cds:protein_codingATGGGAGGTGATGACGAGTGCGTTATAGTGGTACGGGAGTTTGATCCTAGTAAAGATTTA GCAAGTGTAGAAGAAGTTGAAAAACGATGCGAAGTTGGTCCCAGCGGCAAACTCTCTCTC TTTACCGACCTCTTGGGTGACCCTATTTGCCGGGTCCGCCACTCCCCTGCTTTTCTCATG CTGGTGGCTGAATTAAGCTCCACCAAAGAAATAGTTGGGATGATAAGAGGTTGCATAAAA 30 ACCGTTACTTGTGGCAAAAAGCTCTCTCGGAATACCAAAACCAATGATCCCTCCAAACCT CTCCCTGTTTACACCAAAGTCGCTTACATTTTAGGCCTTCGGGTCTCCCCTTCCCACCGG AGAATGGGAATAGGGTTAAAGCTGGTACTAAGAATGGAAGAGTGGTTTGTCCAAAACGGC ACTGAATACTCTTACTTAGCCACGGAAAACGACAACCAAGCTTCCGTTAATCTCTTCACT GATAAATGCGGCTACTCCAAGTTCCGTACCCCTTCCATTTTGGTGAACCCCGTTTTCGCC 35 CATCGACTCCCTGTTTCCAACCGAGTCTCTCTGATTAAGCTGTCCCCATCCGACGCCGAG TCGCTATATCGGCGTCGTTTCTCCACCATCGAGTTCTTCCCTCGCGACATTGACTCGGTG CTTAATAACAGACTCAGCCTGGGGACCTTTTTGGCAGTGCCACGTGGATGTTGCTATACG CAAGAAACGTGGGCCGGGTGTGATAAGTTTTTATCCGACCCGCCAGAGTCATGGGCGGTT TTGAGTGTGTGGAATTGCAAGGACGTGTTTAGGTTGGAAGTCCGGGGCGCGTCGAGGATG 40 AGGAAAACGTTGGCTAAAACGACAAGGATAGTGGACAAGTTGTTGCCGTTCTTGAGGTTA CCGTCGATTCCGGAAGTGTTCAAGCCGTTCGGTTTGCACTTCCTGTACGGGGTGGGAGGG GAAGGGCCATCGGCGGCGAAGTTGGTTTATGCGCTGTGTGCACACGCGCATAACTTGGCC M&C PC933520WOA 109 AAAGAAGGAGGGTGCAGTGTGGTGGCGACTGAGGTGGCGAATGGTGAGCCGTTGAAAGCT GGGGTCCCACATTGGAAAAGGCTATCGTGCGATGCAGATTTATGGTGCATCAAACGGCTT GGGGAAGACTACAGTGACGGGTCCGTCGGTGACTGGACAAAATCACCCCCTGGACTTTCA ATTTTTGTAGACCCCAGAGAATTCTGA 5 SEQ ID NO: 42 Gossypium raimondii>KJB48700 promoter ACGCATTGTTGAGGTTGGTGAAATCAATACATATCCTCCATTTTTCGTTTACTTTTTTCA CCATGACCACATTCGAGACCCAATCTGGATATACAACTTCCCTGATAAAACTTGTCGAGA ATAATTTATCTACCTCCTGCCTCATAGCTTCCACCACATGCAGAGCAAACTTCTTCTTTT 10 TCTACTTAACAGGCTTCACTTTAGGGAGTACATTTAACCTTTGGAAAATTACCTGTGGGT TTATCCCTGGCATATCTACTGCTGACCAGGCAAAATCATCGGAGTTCGCCTTCAAACACT AAACCAAAGTGTTATTTTCTTCAGGTAATAATTTGATGAAATCTTTATTGTCATTTTTTC ATCGTCGTAGAATAGTTGTAAAGTTTCGGTGTGCTCTGCTGCTTCAAGTTTCTTACTTCT CTCTCATCCCTCGCATCTAGGCTATCCAAGTCTAAGGCTTGACTAGCTAACTCTGTCCTC 15 TGCGAACTTTGCCCCCTAGTTGCAGGCTCTCGTGCCTGTTTCACAAATAACATATAACAT TGCCTTACTGTTTACTGGTTGGATCACATGAACTCAATCCCTGTCTTTATCAAAAATTTA ATTTTCATACAAAATTTTAATACCACCATCTTCGCTGTTCTCATTATCGAACACCTAAAA ATGGAGTTATATGTCATAGGATGGTCCATAGCAAAAAATTATACATATTCTGTAGCCGTA TGCTCACCATCTCCCAAGGTGACTAGCAAAGTGATAGAACCCTTCACCTCTACTAGTGAT 20 TCGCAAAACCAAATAGAGGATTAGCCTTGGACAAAGCTTATTCTTTCAACCCCATCTTTT GATAAGCCTCCCAGGTTAAAACCTTCATAGCACTCCCACTGTCAACCAATATTATTCTAA CTTCAAAACCTACGATCGTTACTGATACAACCATTGGGTCATTTCTTTTTTTGTCGCACA CTAGCTCCTCATTGTCATTACCAAACTCAACACTCTATTTTCCCTACTGACACTGTCTCT TGGGTGCTTTAACTAACATGACACACCTCATGTGAGCTTTCCTCTTTGCATTAGAACTTT 25 CCCAATCCTCATCCATGCCTACAATCACATGGATAGTGCCCCTGACCTTCTGCTTACCTT TGGGATCACCACAAGAACTTTGACCCTGCTAAGTAGTAACCTAATCCACAAACTTGTGAG CTCACCATTTCTCACTTCCTCTTAAATAACATCCTTCAAAGTAAAACAATATTCAGTTTT ATGCTCCAGATCATCATGGAAAGCATATCTATCCTTAGAATTTGACCTCTTTGTACCATG CTTCATAGGAGACGGGTTAGGCAGGATGCTTAAACTCTTTACTTGATTCAAGATATAAGC 30 ACGAGTAATGTTCAATGAAGTGTTGGCTTTAAACTTGCCTTGAGGGATGTATCTTCTTAT CTGCTCGGGTTGAGATCCTTAAGGAACCTGAGGAAAATTACCCTATGTTTGCGTTTTCTG CTGCTCGTCTCCTCGGGCACGCAAAGCCTGATTAGGCGACCCCTATCTATTGCGAGCTCC TTCTCTTTTTCTGATTTGTCTACCATCTCTTTAAAATGTGTCATACGTCACCCTCTTAAT CTCTTCTACCTCCGCGAACTTGTGAGCCCGCTCATATAAGTTTGCCAAACTCTATGCCTG 35 TTATCAGTAAAGGAGTACTGTTCATATTCATTCTACACTCCGATAATGAAGGCATCGGTG GCCCAATGTCCTCCAGGTTCTTCGTATTCATTTGTTTGGTTGAATGTAATACTCATTTAA TGGATTGAGTATTGCTTTCCTACTACTTGTCTTCCCCTAATAACTATTCTTTAGTTTCAC CAAAGAATGACTATTCCATGCAACATTGAGTAAAAAAATATTAATTTCTAGCTTAGTCCT TAATTTAAAACTTGGACTTAGAAAATATAAAAATAAAATATTTAAAATATTTGAATATAA 40 TTTAATATAAAATATTTAATTTCAAAACATAATTTAAAATAAAATATTTTAATGGGCGAA TAAAAATATAGTAAAATAAAAAATGTTGTTAAAACATTAAATGTCTTTGAACATGTAATA GATAGAGAATTTTCATCTTTCATGAAACGACAAAATAAAGTCGTTATGTATGTGGAAACG M&C PC933520WOA 110 ATGGTCATGATATTATGATGATGATGATGATGATGATGATGATGATGATGGATGGATGGA TGGGTGGGTGGAATTGGAGGAGGGGGAGATTAGAGGGAAGGGAAATGATTGCCAAACTCC ATCCTTAGAAATTAAATGATTAAAATGAAAAAACAGAAAAAAAAGAAAAACAATAAAACT GCGCAGCCATAGCCGCCATTTCCAGTTCCACCCTTTTTGTGATGATTAATAATAAAAATT 5 GTTATTATTGTTTTTCCATTTATTTACAATCTCTCCCATTTTAACTTTCTTTCTATAAGA TAATATTTTCGTTACTAGTTTTAAAAAGAAAGAGAGAGAGAAAAAAGATAATATAGACAG ACAGAAAGACGAGAGAGGGAGAGATAGAGACTTTTAAAAGCACTGAATTAAATAAAAAAA GCCAGAAAAGAAAAGGAAAGGTTTAGTGATTCTAGAGAGATAGTAAAAAGAAGAAGCCAA CCAGTATTTAAAGGCAGCTTCAAACCAATTGTTGCTCTGGCCATCAACTTCTCCGTATTT 10 CCTCTCCTATGTCAAGGAAAAAAAACCCAAGAAGAAAATAAATACACAACAAGAAGCCCT TAGCTTAGCTTTTTTTCTTATGATCTCGAAGTCCTAAGCTTTCTACACCTTCTCTCTTTA CTGCTTTTGCTTGTTTCAGACGTTATTGACGGCTACTCTCTGACCTTGAAAAAGCTGCGT TTTCACCTATATAAACCCTCGTTTCCTTTTCCATCCATTTTCACTTTCACTATAACATAG AGTGTGAAAAAGAAGGAAAATT 15 SEQ ID NO: 43 Sorghum bicolor >EER93401 cds:protein_codingATGGTCGACGCGCCGGCCGTCGTGGTGGTGGTGGTGCGGGCGTACGACGACGCGCGCGAC CGCGTCGGCGTGGAGGAGGTGGAGCGCGCGTGCGAGGTCGGGTCCATGAGCGGCGGCGGC GGCAAGATGTGCCTCTTCACGGACCTCCTCGGCGACCCGCTCTGCCGCATTCGCCACTCG 20 CCGGACTCCCTCATGCTGGTCGCGGAGACAGCAACCGGCCCCAACAGCAGCACGGAGATC GCCGGACTCGTCCGCGGCTGCGTCAAGACCGTCGTCTCTGGCACTGCAGGCACCACCCAG GGCCAGCAGTCCAAGGACGACGACCCCATCTACACCAAGGTTGGCTACATCCTCGGCCTC CGCGTCTCGCCCAGCCACCGGAGGAAGGGGGTGGGGAAGAAGCTGGTGGATCGGATGGAG GAGTGGTTCCGTCAGAGGGGTGCCGAGTACTCGTACATGGCGACGGAGCAGGACAACGAG 25 CCGTCGGTGCGGCTCTTCACGGGCCGCTGCGGCTACGCCAAGTTCCGCACGCCGTCCGTG CTGGTGCACCCGGTGTTCCGCCACGCGCTCAAGCCCTCGCGGCGCGTTTCCATCGTGGAG CTCGATGCGCGGGAGGCCGAGCTGCTCTACCGCTGGCACTTCGCCAACGTCGAGTTCTTC CCCGCCGACATCGACGCCGTCCTGTCCAACGACCTGTCGCTGGGCACGTTCCTGGCCCTG CCGTCGGGGGCGCAGTGGGAGAGCGTGGAGGCGTTCCTGGCGTCGCCGCCGCCGTCGTGG 30 GCCGTGCTCAGCGTGTGGAACTGCATGGACGCGTTCCGCCTCGAGGTGCGCGGTGCGCCG CGCCTGATGCGCGCCGCGGCGGGCGCCACCCGGCTCGTGGACCGCGCCGCGCCGTGGCTC GGGATCCCCTCCATCCCGAACCTGTTCGCGCCGTTCGGGCTCTACTTCCTCTACGGCCTG GGCGGCGCGGGGCCCGACGCCCCCAGGCTCGCCCGCGCGCTGTGCCGCGAGGCGCACAAC ATGGCGCGGGACGGCGGGTGCGGCGTCGTGGCCACCGAGGTGGGCGCCTGCGAGCCCGTC 35 CGCGCCGGGGTGCCGCACTGGGCGCGCCTCGGCGCCGAGGACCTCTGGTGCATCAAGCGC CTAGCGGACGGCTACAGCACCGGCGGCCCGCTCGGGGACTGGACCAAGGCGCCGGCCAGG CACTCCATATTCATCGACCCGAGGGAGCTTTAA SEQ ID NO: 44 Sorghum bicolor >EER93401 promoter40 TTAGGTATTTCCAACATCAAAGAAAAAAAATGCAAGCTATATATGATGGTAAGATAGATT TGAATGCAAATGGTTCACACTATAAACATCACTATCAGTTAGAGAAAATAGTATGTTCTA GCTTATTTCAACTCCATTTCTCAAAAACTTATTAAAATGTACATGAAACTAAAGTCACCA M&C PC933520WOA 111 TACCTAAGTAGCTATTTGTGCACGAAAAAATGGGGAATGGGAATCAAACATTCTTAGTTA CTAGAGTGTTAGCTCATGGCTTACTAGGTAGAAGGAAATGTCACCTGCTAGGCCACCTTG ACTTGACCAAGCTGACAATCTAGATTTGAGAGACAAAGAGAGATTGAGTTTCTGAGGTGG CACTATGGGTTTAGGCTTCTGTAGAAATGAGATAGCCATCTAGAGGGAAGGTGAATAGTT 5 GTATTTGAAACTTAATCACAAAGGATAGCAAATATAAACTGTCTAGCCATGACTACACTA CTGTAATTCACAAGCACCCTAACAAGAGATTAAACAAGGGAGCAACTAGGGTGTCGGGCT AGAGATGGCTCACCTAACAATAGCAAGTCTAGTTGGCAACATCAATCAACCTAAGCAACT AGAGCACAAGCTAGCAAAAGGGATCTCCTACACATACTAGTAAGAAAAGGTAAACATTAG TGAGCAACTACACTAGAACAAGAGCAATAACACTAGCGTAAGATAACACAAGTAAATAGA 10 TAGTCTTAGAAAGTGCAAACCAATGGGAGACAATAGGTGGCATGATGATTTTTCTCCCAA GCTCAGACTACCTCACAAGGTAGGCTCCAACGTGAGCACTAGACTTGAGACGATCTATCA ACTTGTCACTGCATGTCAAGTTAAGCATGCATAATCACCCCTAGTGCCACTAGATGCTAT CTCCCTCAAGCAAATGCACTATATATATCACTCTAATTTCACCTTGAATGTGATCTCAAC TTGGCATATGTGAGTCGAGGGCTATGTTTGTCTCACCAACAAAGCATCAACCAATTGGCC 15 AACCCACCTATTTATAAGCCAACCAAATAAACTAGCCATTACAAGATTTACATGGCTTTT GGGGGAGCAACGAAAGATCTGGTGCTACCACTAGAGAACCCACTAGAGAAAGCCCATCAG CTAACTCAACAGTCATTAAAGATGTCAAATTTCCTTTATCGAAGAACTATATAAGGCAAC CAAAGAAACTAGAAAGAGTGCATAAGAGCTTGGGGTTGGGCAAACTAGTGGTTATTGGCA AAATCCAATAACCAGTGCGAAACAACCAGTGCGAAGTCTGATCTTTGGAGCCCACTCAGG 20 GCAACCCTGGAACCTTCTTTTGACAACACAAAAGGTATTCACCACCACTGGAAAAGTCTA TTCCTAGATCGTTAGACTTTTTTTTGACAATACAGATAGTGTTTGTCATAGAGCTACAGT GCACCCAGGCTTGTTGGCTCCATATCATCTAGTGTGTATACTCTTGTACAATTTGTTAGC CGATTGTTGCAACTAGCCATTGGAAGCCATTGTTGAACCCTTCGGTGCTGCCACTAGATC ATCTAGGTCTTAATGAACCACATACCAAATGACACTAACAAATCACTTTAAGTTATACTA 25 GTACTCTACGACAATAATTTTTGTGCTAGTTTAGATTAGATTACAACACATATAGTTCTT TGACAGAGGAAATTTGACATCCTCAGTGCTATGGTGCACACCAAAACAACTATGCATACC ATAAAAAAACTATGTGTGGAGTTTAGGACCAGGAATCTTTTTCTAGAAAAGATAAAAGAA AAATTTGAGGAGTGGATGAAAAAATTTCAACAACTAAATATAAGTGTAAAAGTTGAAATC CAAAATATAAATTCATGATATCACAAAAAGGAAGCAACAAGAACTAATATAATTTAAACT 30 AAGATTTTAACTAAATATTTAGTCCCCTGTCTCTATGCAAATGACATTTATTTTGGGGTA GAGCAGGGTATTAAAAACTCACAACTTAAAGGATAAAATACTATCTGGATGGGTATTTAT CATGTTAAAGAAAGATTGTACATATACTCCTTTCATGCCCGTAAAGAAAGTTGCTTAAGA CAAAGATACAATCTACAAAGTCTAACTTTGACTCTTTATTTTTATAAAAAATATTTATTA AAAAGTGGTATATGTATATTTATATGAAAGTATTTCATATGGGTATTTACCATGTTAAAA 35 AAAGAATATAGAAGAACAAGCTTAATTAGTTGCACGCACTAAACTTATTTAGAAAACTTT CAACACCTATATAGCGCAAGATTTTGGACCCTGCTGAAGCTCGAGTTGATATTTGCTAAA AGAAGATTAAGCCTACATAGTACTGAACAAAGTGAAAATGGAAAAGAGCTGGACCCAGTG TACCTTTTTCACCTGGTGGGGCCAACTCCGAGGTCCAAGGGTACACGAAGCCGGTATAAA TAGCAGCCGCAAGCGGCTCTCACGCCATCAACGGAGCTGAGCCTGATGAGCAGCAGTGGG 40 CAGCCGCCTTCCCCTGTGTGTTGCGGATCGCTACACCACCACTGTCACCACTCACCACTC AGTTACTCACCACATCCATCCGCGCCCACAGTGAGGCAGTGAGCACTGAGCAGCATCCCC CCTCTGCTATAGCTAGGGAGTAGCATAGTAGCAACCTGTGTGTTCTCTCCGATCCTATTC CTACCTGCCGTCTTCTTCCCTGCAAAGCCTGCGCGGCTCCGCTCCAGAGCTAGCTCCTAC M&C PC933520WOA 112 ACGCCGGCCTGCAGCATCTTTACCACCGCCGCCATCGCCGCCTCTCCTGCTAGCTAGCTG ACCCTGAGCCTGAGCTGCATCGATCGATCCCAACCGCGCGCGGTGTGTCCTGTCCTGTTC CCGCGCACGCAATCGCGCCCGCCTAAGCAAGCTAGACGGACGCCGCCAGTTATTGCACCG GCCGGCCGGGGATATCCGCGGGAGTGCGGAGTGGGCGGCGGATAGCTGCGCCGCTACTAC 5 TCTAGCTCCTTCGTGCGTGTGTGTGGGGAGGAGTAATTGTGGAGGAGGCAGGCACGCAAA AGGACGCCTGCGCCGGCCCGGCCCTCACGTGATTCATACTTGCGGCCTCCTCACTCCCTT CTTCTTCATCACTCCTCTCGCGTCGCGGCCTCACAACTACTATATCAGCAAGCACCAAAC ACACCCAGCTCCGGCTGGCCGGCCTCGATCTCGTCTTCTTCCCGGGAG10 SEQ ID NO: 45 – Intentionally skipped sequence (placeholder)GS2 / GRF4 Homologs SEQ ID NO: 46 Setaria italica>KQL31106 cds:protein_coding 15 ATGGCGATGCCGTATGCCTCTCTTTCCCCGGCAGGCGCCGACCACCGCTCGCCCACCGCC ACCGCTTCCCTCCTCCCCTTCTGCCGCTCCACCCCGCTCTCGGTCTCGGCGGCTAGCGGC GGCGGCGGTCTGGCAGAGGGCGCGCAGATGAGCGCGAGGTGGGCGGCGAGGCCGGTGCCG TTTACGCCGGCGCAGTACGAGGAGCTCGAGCAGCAGGCGCTCATATACAAGTACCTGGTC GCCGGCGTGCCCGTCCCGCCGGATCTCGTGCTCCCCATCCGCCGCGGCCTCGACTCCCTC 20 GCCACCCGCTTCTACGGCCACCCCACACTTGGGTACGGATCGTACTTCGGCAAGAAGCTG GATCCGGAGCCGGGCCGTTGCCGGCGTACAGACGGCAAGAAGTGGCGGTGCTCCAAGGAG GCCGCCCCGGACTCCAAGTACTGCGAGCGCCACATGCACCGCGGCCGCAACCGTTCAAGA AAGCCTGTGGAAACGCAGCTCGTGCCCCCGTCCCAGCCGCCGGCCACCGCGGCCGCGGCC GTCTCCGCCGCGGCGCCCCTTGCCGTCACCACCAACGGCAGCGGCTTCCAGAACCACTCT 25 CTTTACCCGGCCATCGCCGGCAGCACCGGTGGAGGCGGCGGGACCAGCAACATCTCCAGC ACGTTCTCCTCGCCGTTGGGGTCGTCGCCTCAGCTGCACATGGACAATGCTGCCAGCTAT GCATCTCTTGGCAGTGGAACGGCCAAAGATCTCAGGTACAATGCTTATGGGATAAGAACT TTGGCAGACGAGCACAATCAGCTTATTGCAGAAGCCATCGATTCATCAATGGAGAACCAG TGGCGCCTCCCGCCGTCCCAAAACTCCTCCTTCCCTCTCTCGAGCTACCCCCAGCTCGGG 30 GCGCTGAGCGACCTGGGTCAGAACACGGTCAGCTCGCTCTCGAAGATGGACAGGCAGCCA CTCTCCTTCCTAGGCAACGACTTCGGGGCGGTCGACTCCGGGAAGCAGGAGAACCAGACG CTGCGGCCCTTCTTCGATGAGTGGCCCAAGGCGAGGGACTCCTGGCCGGGCCTCTCCGAC GACAACACTAACCTCGCCTCGTTCCCGGCGACCCAGCTGTCGATCTCCATACCCATGGCG TCCTCGGACTTCTCCGTGGCAAGCTCTCAGTCGCCGAACGATGACTAA 35 SEQ ID NO: 47 Setaria italica>KQL31106 promoter GTATCAGGTATGACATTCAGTCGGACACTGTCGATCTTTGCAATATCTCATCCAGTTACG CGTGTAGTGCTCTGTTAAGAAACAATCTTGTTTCAGCGCAGTGGCGGAGCCATGAAAAAA TTATAGGTGGGGCGACTTCTAGCGAGAGAGTAATCAAAAGCAAAAATGCACTAATGCACA 40 CTACATCCACACATTGCAAGTGCTCGATTATAACGGAAAGGTGGATGAACAACACAATTG AATTTTAAGATGCTACCGCAAATTTGCAGCCTGTTCGCTTCAACTTATCAGCTACCGAAC AGTGTTTTTCTCTCACAACAAATCAGCCGTTTCAGCTTTTCAGCCGACTTATAAGTCCAG CCGAACAGGCCGTTGAATGATGAATTTATAGTTCTGGTACCAATGCGTACAAAAATAGAA M&C PC933520WOA 113 ATGGATGCACCAGTAATAAATTGTATAAAGAGATCGATTGACACGCTATACTCTAACCAA TGCTGATCTTTGTGCGCGCTGCTGTTGGTACCACGTCCTTTGCAACCGCAAGCCAGGCAG CGTTGGAAGCAAGGACTGGCGGGCAGCGTTGGTAGTGTTAGCGAAGGACGGTCCGCTGTT TTTCATTGACTTGGGCCCTCAATCTTTGGGCTGCTGTTGGGTTGGCCTAGGAAAAAACTA 5 ATAGACGCGATGGGTGCGTCATCTTCTTGCGCTCGCGGCTCTCTCGCCGACTCGTCCTCG CCATGACGGCCGATGCCTCCTCGACAGGAACACGGAGAGGTAGCTAAGGGGGCAGCCATC GGCCAAGGAGCAATGGAAGAAGGGGTGCTAGTGAGTTGAGCACGCCGGTAGCTGGGGCGG CCGCCCCACCCCACCGATAGGGTGGCTCCGCCCCTGGTTCAGCGGGTTAACTCAAATTAT GGTGATGCGTAATCATAACGACTATGTAATCCAGGTTTGATTTTCTTCTTAAGTGAAAAT 10 GCGGTGCTCCTCAAAAGGAAAATACTGACGAGTAAGTTAGGAGCTACTAACACCCGATGA AATTGAAGCATGAGAACTATATAAGGTTTTTCAAAAAGCATTGTGTGAGAGTTGGTGCTG AATTGAAAGAAAGTTGGATTACGCATGGCCATCTGATCTGTAATTTGGTTCGCCTTTTAA ATTTTTGTAAAATGATCTTGAAAAGAAATGGAATGAGTTGAAATGACCAAAGGGACATAT GATATGATCGCCTTTCATATAGTACCTTTGGATCAGGGGCGGAGCTTAATTCTGGCCGAG 15 GGTAGCCGTTGCCCCCATGTTGCTAGTAAGTTCCATTAGTGAACGCTGAATTTTACCATG TAATGAGTGATATCTGAGCCTCATCAAGCTTGCCACCGCTTCGGATGACCCTTATTCGAA CTTTGACGTTGTTACAATTCAATTGATGACTTTTCTAGAAATTCCATTCTAAAATGTGAA GATAACTTCATGTCAAGTGTCATGTCCAAAAACTTGGCTTTTGCTAGCGCTGATGTAAGT TTCTTACTAATTTTTTCCAAAATAAGTTTCTTACTAAACTTTTCTCGATCCACGGTCGCC 20 GCCGGTTGATCCTACGTAGTCTACTTACAAATGCTTTTGTAACAAATCCAGCGGTACCAT CATCATGCAAAAAATTGTAACGTGATTCTTTTCACAAAAGTTTCGATCTTGAAGATAGTA GTCTAGAAACAAAAGAAACTTGGCCAAGCTGCGTACTAGCTAGCGCACTAGCCACTAGGG GCCGATCAAGACGTCGCCTGACGGAAGACAGGAAGCAAGTAAAGAAATGGCAGGAGAGGA GGCAACGGCAACGGCGCCCCACACCCTCCTGTGCCGGCCTCGCATCGGTAACCGCGACGC 25 AGACACGCCCAAATCCACAGTCACCTCGCCTGGTCTCTCTCTCCTCCCATACGATGCTTG TCACCACCCTCGCTACAGTACTGCATCCAAATCCACTCCGCTCGCCCAAAAAATAGGCCC CGGAAGAAACAAAGGCGCATAGGAAACAGAAGAAAAAGCTGTACGTGAGGAGGGGAGAGC CGGAGGGGGGTAGGAGAGAGAAGCACGCTCACACGGTTCACACTACCTCACCAGTCACCA CGCCTGCTACGGACTACCACAACGCTAGAGGGACAGACGGCGAGGCAGGCGTCGCTTGTC 30 ATCTTGCAGCTGGATCGCCGGCCGGAGCGCAGCCACCCCGTCGCGCTCGTGTCATCCCAG TGTCGGGCTCGCAGCTCGCTCACACGAGACCACCAGCCTAGGCTCGCCGCTCGCCTCGCC TCACCCCGCCTCCTCGCGTGGAGACGACACGCAGAGCGCGGGGCAGACAGCGACGCAGAG AGAGACAAAGCGGGCAATAAAGGCGAGCGCGCGCGAGCGGGGGTGCTGAGCTGAGCGAGC GAAGCAAGCAAAGCACATCACGAGCCGCAGCCGAGCCAGCCAGCCAGCCCCCGGGGGATC 35 TCATTAAAGAGGGGGGTGCGGGTGCGGCCGGCCGCGGGGAGCAAGCAGCGCGCGAGAGAG ACAGCGAGAGG SEQ ID NO: 48 Zea mays>Zm00001eb251820_T002 cds:protein_coding ATGGCGATGCCGTATGCCTCTCTTTCCCCGGCAGGCGCCGCCGACCACCGCTCCTCCACA 40 GCCACGGCGTCCCTCGTCCCCTTCTGCCGCTCCACTCCGCTCTCCGCGGGCGGCGGGCTG GGCGAGGAGGACGCCCAGGCGAGCGCGAGGTGGCCGGCCGCGAGGCCGGTGGTGCCGTTC ACGCCGGCGCAGTACCAGGAGCTGGAGCAGCAGGCGCTCATATACAAGTACCTGGTGGCG GGCGTGCCCGTTCCGCCGGATCTCGTGGTTCCAATCCGCCGCGGCCTCGACTCCCTCGCA M&C PC933520WOA 114 ACCCGCTTCTACGGCCAACCCACACTCGGGTACGGACCGTACCTGGGGAGGAAACTGGAT CCGGAGCCCGGCCGGTGCCGGCGAACGGACGGCAAGAAGTGGCGGTGCTCCAAGGAGGCC GCCCCGGACTCCAAGTACTGCGAGCGCCACATGCACCGCGGCCGCAACCGTTCAAGAAAG CCTGTGGAAACGCAGCTCGCGCCCCAGTCCCAACCGCCCGCCGCCGCAGCCGTCTCCGCC 5 GCTCCGCCCCTGGCAGCCGCCGCCGCCGCCACCACCAACGGCAGCGGCTTCCAGAACCAC TCTCTCTACCCGGCCATCGCCGGCAGCACTGGTGGTGGAGGAGGAGTTGGCGGGTCCGGC AATATCTCCTCCCCGTTCTCCTCGTCGATGGGGGGATCGTCTCAGCTGCACATGGACAGT GCTGCCAGCTACTCCTACGCAGCTCTTGGTGGTGGAACTGCAAAGGATCTCAGGTACAAC GCTTACGGAATAAGATCTCTGGCGGACGAGCACAACCAGCTGATCGCAGAAGCCATCGAC 10 TCGTCGATAGAGAGCCAGTGGCGCCTCCCCAGCTCGTCGTTCCCGCTCTCGAGCTACCCA CATCTCGGGGCGCTGGGCGACCTGGGCGGCCAGAACAGCACGGTGAGCTCGCTGCCGAAG ATGGAGAAGCAGCAGCCGCCCTCGTCCTTCCTAGGGAACGACACCGGGGCCGGCATGGCC ATGGGCTCCGCCTCCGCGAAGCAGGAGGGCCAGACGCTGCGGCACTTCTTCGACGAGTGG CCCAAGGCGCGGGACTCCTGGCCGGGCCTCTCCGACGAGACCGCCAGCCTCGCCTCGTTC 15 CCCCCGGCGACCCAGCTGTCGATGTCCATACCCATGGCGTCCTCCGACTTCTCCGTGGCC AGCTCCCAGTCGCCCAACGGTGAGTCGCGTACGTTCCTGCTGGCCACGGACCGAAGGTGA SEQ ID NO: 49 Zea mays>Zm00001eb251820 promoter GCCAGGAATCGATCCACGTGAGAGGACACTTGGATTTGTTTTAGGTGGAGAAGTGAGAAT 20 GGGGAGTCAAAAATAGATGGATTTGTTCTCGATCCGGGTCAAAACGAACTACTAGATCGG ATTCCAAACGAGGCCAATTTGTTCAACCTGGAAGGAGGCTGGAAACATGCCAATCCCGCC GATACGAACACGTCCTATGGGTTGTTTGTAGATTTTCGACCAAAATTTGAAATTCAGTTT TTCCCTATTTTGGATCAGAATTTGAATTTTGGGTCCATGAAAATGAGCATGATACGTGTA AGGAGTAAATCGTCTTACCTGGGCCGAGGACTAATCCAATGGGCTACTTCCGCATAGCCC 25 ATGGGGCTATTGGCTTCATATTAAAATCGACTGTTTGGAATACTGCAGCTGCAATCTAGA GACAAAAAACACGGTAAAAGCAGTAGAAGCCAGTACGGATTAAGACCAAGCGAATAGGTT GATATTTGTTATGGCCTATGCCGACTCGAAATTAAACGTACTATACTTTTTTCGTGCTTG ACTATGCGAGACCGACGATCAAAAGTACAACTATACACTCAACTGCTCTTGCCTCAACCA GCGTCCTAGGATCAGAGGCCACTCTAAAGCCTACTTTGGATCAAAGGATAATTTTTCGAA 30 TAATGCATTTTTCCTGTGTTTTTAACTAATTCTTGTGAGGTTCCTATATATATCATTCTT GTGAAATTCTAGTGTTTCAACGGAAGCCTAAGTTTCGATGGGAAGAAAGGACATGTACTA GCAAGGAACCAAACTCCACGCATCATTCTTGCCTAGCCTTGCTTTATCGTGGCTACCTTG GACCAACAAAAGAACCAAGCAGCCCCAATGTATCTGATATGGAGCTAAAAATACAACCAA CTCATATTATACGTTGGATGTTTTGACTGCACTTGAGATGTTGTAAGACTTTCGGTACGC 35 TATACATATAGAGTTGAATATACAGTTGAAGACTGCTGCAGCGGTCAACTGTCTGATCTA CTGTAAACTCTATGAGGAAATCGGAAACGCTACTTCCAGAGTAGTGTAACTCCGACTGGA AAACTGTTGCAGAATACGGATAGCCTGATCAGTTAGACTGTCGGCTGCGGAGTTCAACTG TTGCAGAGTTAGAAAGAAATGATAAAATATATAGTAGTTAGTATAGAGTTGATATATAGA GTAAACATGACTGTAGAGGATTGTAGTATAGGGTAGATAGTTTTGCTGACCAGGACAAGA 40 TATTCCTTTTAGAGTATGAATTTAGAGTAGTATGAGTGCGGATAGCCTAACTTTGTAAGT ATTTTTAAAGCTTACTTTGCATACGGTCTTTGTGATCTACATCTTTACTATGGCTATTTC ATGATAATAACTAGATGAGATATATGACCAATCGAGTTGTACATATATGTTTGGGTTTTA ATTAAGGGCATAGTTAAAAGCACTGAGCTTTTAAGAAAACGATGTGGTTCTAAATATGGC M&C PC933520WOA 115 AGTTTATGCTTTGGTTTCTAGAAACTGAATTTCTAGCATATTTCCGTACTATTCTTAGTT GGTTTGGATAGAAACTACGACGATTATCACCGCTCTGAGGCCTAATGGCCTATGCACTTG ATTCTCTCCATGCCCACTCTGCCCTGTTCAAATGTTTAATTAATATTTAATTTAATAATT TTGAATTCAAGAATACGAGTTCAAGGTATATTTAAAATTGACATCAAAGAGAAATGAAAT 5 TAAAGCAATGATAGACTTGTCTTTGGGTGTGAAAAAAAGCTAGAAACTTATTTATAAAAA CCCAATTCTAAACATGTATACCTAATTTTTATTATAAATCGGTTTTTAGATAGAATCGTA AAGCCCTTGATCAGAGCATCCAACGAGCCATGAGGCCATGACGGAAGAGCGGAAGTGCAG ACGGCAACGGCGTTCCGCTTCATGCCGCACCCTCCAGTGTCCTGTGGCCTTTAAGTGCCG GCCTTGGGAACCGCGACGCAGACACAGCCCAAATCCGCAGTCACTCCTCCAACACGATGC 10 TTGTCACCACCCTTGCTACAGTGCCTGCATCCATATCCACTCCGCTCGCGCAAAAAATAT CCGAGTCGGAAACAAACAAAGCAGCATAGGAAACAGAAGAAAGCTGTACTAGTACGTGAG GACGAGGAGGGAGAGAGAGCAATACACAGAAGCCTGCTACCGTGCTACGGACTACCACAA CGCCAGAGGGACAACCGGACAGAGGGGGAGGCAGGCCTCGCTTGTCATCTAGCTAGGTCA GCCGGGGACGGGGTCGGAGCAGTAGAGCTAAAGCCAGAGGCCAGGCTCGTAGTAGTACGT 15 AGTAGTAGTGCCCTCCTCGTGTCATTTGGCCAGCCTTGTCCAGACGACCACACACACCAG ATTACGCTTAACATTCTGTTTGACATCTAAAACCAGCCGGCTTGATCCAAATGCCTCCCT AGGTAGTAGCTTAGTCTTGCTCGCCGCCTCTCCGGGAGACGACGACACGCCTGATGAGTG CCTGACGTTCCAGCGCGAGGCAGACAGCGACGCAGAGAGAGACAAAGCGGGCAATAAAGG CAGCCGCGCGCGAGCGAGGGAAGGGAGCGAAGCAAAGCACATCACGAGCCCAGCCTGCGC 20 CTGCGGAGGGAGGGGGCTCATTAAAGAGGGGGCGCGAGCGCGACCGGCCGCGGGGAGCAA GCAGCGCGCGAGAGAGACAGGTTGAG SEQ ID NO: 50 Secale cereale>SECCE2Rv1G0120160.1 cds:protein_coding ATGGCGATGCCCTTTGCCTCCCTGTCGCCGGCAGCCGACCACCACCGCTCCTCCCCCGTC 25 TTCCCCTTCTGCCGCTCCTCCCCTCTCTACCCGGCAGGGGAGGAGGCGGCGCACCAGCAC CAGCACCAGCAGCAGCAGCACGCGATGAGCGGCGGCGCGAGGTGGGCGGCGAGGCCGGCG CCCTTCACGGCGGCGCAGTACGAGGAGCTGGAGCAGCAGGCGCTCATCTACAAGTACCTC GTCGCCGGCGTCCCCGTCCCGCCGGATCTCCTCCTCCCCATCCGCCGCGGCTTCGACTCC CTCGCCTCGCGCTTCTACCACCACCACGCCCTTGGGTACGGGTCCTACTTCGGCAAGAAG 30 CTGGATCCGGAGCCGGGGCGGTGCCGGCGGACGGACGGCAAGAAGTGGCGGTGCTCCAAG GAGGCCGCCCAGGACTCCAAGTACTGCGAGCGCCACATGCACCGCGGCCGCAACCGTTCA AGAAAGCCTGTGGAAACGCAGCTCGTCGCAGCGCCCCACTCCCACTCCCACCCCCAGCAG CTGCAGCAGCAGGCCCCAGCCGCCGCGTTCCACGGCCACTCGCCGTATCCGGCGATCGCC ACTGGCGGCGGCGGCGGCGCGGCTGGCTCCTTCGCCCTGGGGTCTGCTCAGCTGCACATG 35 GACAATGCTGCTGCGCCTTACGCGACCGCTGGTGCTGCCGGGAACAAAGATTTCAGGTAC TCTGCCTATGGGTTTAGGACTTCGGCGATGGAGGAGCACAACCAGTTCATCACCGCGGCC ATGGACACGGCCATGGAGAACTACTCATGGCGCCTGATGCCGGCCCAGAACTCGGCGTTC TCACTCTCCAGCTACCCCATGCTGGGCACGCTGGGCGACCTGGACCAGAGCACGATCTGC TCGCTGGCCAAGACGGAGAGGGAGCCGCTGTCCTTCTTTGGCGGTGGCGGCGGCGGCTTC 40 GAGGACGACGAGTCGGTGGTGAAGCAGGAGAACCAGACGCTGCGGCCCTTCTTCGACGAG TGGCCCAAGGACAGGGACTCGTGGCCGGAGCTGCAGGAGCACGACGCCAACAGCAACGCC TTCTCGGCCACCAAGCTGTCCATCTCCATCCCGGTGACCAGCTCCGACTTCTCCACCACC M&C PC933520WOA 116 GCCGGCTCCCGCTCGCCCAACGGTATATACTCCCGGTGA SEQ ID NO: 51 Secale cereale>SECCE2Rv1G0120160 promoter GTGTGTGCTCCGTGCAGTACTTCACCCCAACTCAGCCCTAAGCCAAACTTGGGAGCTTTT 5 CTGATGGTCCACCTCTGGCGGCAGCACAAGCAGTCGTTCGACTCGATACTTGTTGGTGGC CTAGTGATACACCTAGCCGAGCACTTTGGCCTAAACTTGGAGGGTGTTGAGCATGTTGCA GGAGCCACTGAGTTTAACCATAAGGTTATGACCAGTTACAATTTCATATGGCCTATGTAG GGGGACCATTGCTACAAGTGCCCCTGCCTATACTAGCACTGTGTCCTTACGCTTGCCTCT ACCACAAGCTTTCTCCGTGCCTCAGCACCAGATGTGGGCCCTCTCTATTAGGGAACTCAA 10 TGAGTGTAGACACGCAGCACGGTTCGGGAATGATGTTCAGCCGGAACTGACACTTGAGTG GGACGCCGTTATTGAGGGTTGCAGTTGGAATGCTAGGGGAAATGCACCTAAGTATGTTCA TCGGTTTCACTATGGGCAATGTAGTGACTTTCTGCCAAAGGTAGAGTTTCAACAGCTGTC GTAGCAGGTTGGAGGGATCCATTCGACTTGTTCTCATACACCACTGACCCCATGATTTAG TACAATATTTAGAAATTCATGGCATCACATGTTTGACAATGAGGTTTGTATGATGTTCAT 15 ATGTGTTACAATGATCTTTGTTTTGTTTTATTTCTAACATTTCATGATGCTCCGAAATAG AGTGTATGTTCACATGTACAACAAAAGACTATGAAAGAAGTCAATACCAGTTAACCAAAT CATGATTACAGTGAGGGGTCATTTTAAGTTTTTCACGAGGTAAATCAAAAAATAAGCAAA AGGGACGTAAGCCGGAAAAAAATCGACTACAAATTTTTCAGGGGCTTATTCTCACTGCAG CTTCTCTAGGTAACTCTAAAGACGCTGCTTCTCCACAAGCCACGTCGCAAAAGAACTTGG 20 CCCTCTGGATTTTTTTTATGTTTTTCCATTCCTACAAAACTTAATTTGTAATCATGGACA TAGTCTACAATAATTTTCTTCTCTAACGGCTAGGTTTCTTTCAACCAATTTCTTTCTTCG CTATCATTTCTTTTGGTTGTGTAGTTGAGTTTGATGTGTTTTGTTATAAACGGTCGAGGA CTCATCCAACTATACACCGCCTATGCACATTTCTAGGTTATGATTGTCATTGTTCTGGAT TATGATAATACTATTTGCTAAGGAGATTTTGCGAATTTTAGTGTTTGACTAACCTTCACA 25 CCGACCACGCCGCCTCCTCTCCGGCTGCTCATTCTCACACAAGGTCTGACATGAGGTACT CTTTTGGATCCCATTCGGCTGCTCGTTGCACAGTCGAGAGGTGCAAACATCGTGGACCGG TGGCAAGAAACTTGCCTCACCCCGGCTTGGAAGGACACACCCTTTCTCATCAAGTTGGTG GTTAGGTGCATCTGGAAGCATCGTAGCACGGTCATCTAAAACAGCGCCCGGCTCGGCTTA GATTGACTCCTCGAGACAATTAATTCGAAAGCTAGACAGACGGATTTAATGGCACAAGTG 30 GGTTTGCGACTTTGCTTCTCAAGGAAACCACATGTGGCATTTAGGGGCATCATGTGCTGC TTGGGGTGTGCTCTTATAAGGGGTTGTGTACATCGGATGTAACAACTCTTCTCTCTATGA GTGCATCGAAACACAACCCAAAAAAGAAGACAGAGTAACACCTAGACCTAAACCCGATCC TAAAACTACAACCCCGGTCCCGACCCAAAGCCAAAAGGTAGCTCTTTTCTCAAGCCCGGC CACTAGTATGTGCAAAGCCCGACCCGAAGCCCGACCCAAAATCAAGAAATTGGAAGCCCA 35 AATGCTCTTTGGAGCTGTGAAATGAAAGGCCACCGATTCTCGATCGACCCACTCGTGGAG AGTAGACCTGGTCATGGGTTACCCGGCCCGACCGGCCCGACCCGGTGTTGCCCGCGGGCG GGACCTGGGCCTGAGCTTTGAGCCCGAAGCACGTGTCGGGCCGGGCCTGGGCTAGGCATT TTCATGTTTTAAGGAAGGGCCCGGCCCGTGGACCGGGCTTGGGCCAAAAAGCTAGGCCCG ACGACCGGGCTCGGGCCTAGGTTTTCTGCCTCGAGCTTTGGTAGGCCCGGTCCGGCCCGG 40 CCGATGGCCGGGTATAGTGGAGAGCAACCGATGCTACACAACGCAGGTGGTGGTGGTACC TTAAGCTGGAACCTACCCTGCCGTGTCCGCCTCACCCCCACACGGCGCTGCGCCTATCAC CACTTCACCAGGGCTCCTCCCTCCTATATATCCATCCCTTCTCGTCGTCGTCTCACCAGG CAGATCCACCCCTGCGCAGCGAGGGAAAGAGACACACAGCTCCACCAGGCAACTAGTAGG M&C PC933520WOA 117 AGTAAAAGGCAAAAGCACGGCACATTAAAAGAGAGTCCGGACCGGACCGGAGCCAAGCAG CAGCCGCAGCCGCAGCCGCAGCAGAGGAGAGAGAGCATAT SEQ ID NO: 52 Pisum sativum >Psat7g155280.1 cds:protein_coding 5 ATGAATAGCAGTGGTGGCAGTAACGGCGGCGGCGGAGGAGGGATGATGATGGGTTTCAGT AAATCATCACCTTTTACAGTTTCTCAGTGGCAGGAACTCGAACATCAAGCTTTGATCTTT AAGTATATGGTTGCTGGTCTTCCTGTTCCACCTGACCTCGTTCTTCCCTTTCACTCTCAC AACTTCTTTCACCATTCAATTCCTACCCTGAGTTATTGTTCCTTCTATGGAAAGAAGGTG GACCCGGAGCCAGGACGATGTAGGAGGACCGATGGAAAAAAATGGAGATGCTCTAAGGAA 10 GCATACCCGGACTCCAAATACTGTGAGCGCCACATGCACCGTGGCCGCAACCGTTCAAGA AAGCCTGTGGAATCACAAACAATGCCATCATCATCATCATCATCCACTCTTGTCGCTGCT CCCACCACTACACCTCCAAACTTCCACAACCTTCCCTCAACAAATGCCTTTGCTACTCAA GATTATCATCTCGATACCATTCCCTATGGGATTCCAAGTAAACAATACAGGTATCTTCAA GGTCTTAAGTCTGAGGGTGGAGAACATAGCTGCTTTGCAGAAGCTTCAGGAGGCGGCAAC 15 AAAGGTCTCCAAATGGAGTCTCAGCTAGAAAGCACATGGCCTTTGATGTCAACCAGAGTT TCCTCTTTTTCTGCATCAAAATCAAGTAATAATTCCATGTTGCAGAGTGATTATCCCCAG CATTCATTTTTATCAACTGAGTATGGATCTGGAGAAGCTGTGAAGGAAGAGGGCCAGCCT CTTCGATCCTTTTTCAACGAATGGCCTAAGAGCAGGGAATCATGGTCTGGTTTGGAAGAT GAAAGATCCAACCAAATAGCCTTCTCCACTACTCAACTCTCAATATCCATTCCCATGTCA 20 TCTTCTGACTTCTCTGCCACAAGCTCTCAGTCGCCACACGATAACTAG SEQ ID NO: 53 Pisum sativum >Psat7g155280.1 promoter TAAATATGATGTTGTTGATCTCCAGCCACCACCAATAAAGAGATAACCTGGAAGGCCTAA GAAAAAGAGGAATAGAGAAGTTGGAGAAATGGTGAGGGATGGGAGACACTTGAAAAGGGA 25 TAACCATGGAATCAAGTGCAGTTGTTGCCACATGGAAGGTCACAAGAAGGCCACGTGTAA ACTGCCACAACCAGTGGCACCTCCAACTCAGGTAGCTGAAGGAGCCTCAACTCAGGAACC ACAAGCAATAATTTCAAGTCATCCACCACAACCAACAGTTTCAAGTCATCCACCACAACC AACAATTTCAATTCAGCCACCACCCAAAAAGAAAAAAATCTCCAAAAGGGTGGGAAGCCT GTTTCAAGTCAACCTTGATGTTTATTGTGATGTTTGTTATGCCTTATAATGGCTTTTTGT 30 AATGTTATTTTAGTGCCTTATAATGGCTTATTGACAATGGCTTAGAATGCCTTTTTGTAA TGTTATTTTAGTGCCTTATAATGGCTTATTGACAAAGGATTATAATCGCTTTTTGTAATG TTATTTTAGTGCCTTATAATGGCTTATTGACAATGGCTTGTAATGTCTTTTTGTAATGTT ATTGATATTCCTTATAATGGCTTAATGACAATGACTTATAATGCATTTTTGCATCACTTT TTGGAATATTATAGATATGCATTATAACAAGTTATCTTTTGTTTCAACTTTTATACCAGT 35 TATATATCAGTGTTAAAACAAATGAATAATGTACAAACACAACACTTTGTTAAAACAAAC ACTAACATTCTTCATTAAATAAATAAAATATTACATTCATATTATGAAAACAAAATACTA AATTACCCTAATCCTAATTGAAATACTTTGTTATAAACACAAGGTTCAACCCTAAGCTCA ACATCCCAACCACAATGGACATTTTCAACCATCCTTTAGTATGGATAACCCCATTCTTCA ATTTGTTAATTTTCTTCTTTTGTCTTTCAAGCTTCAAATCCCTTTCATCAACAAATTCAT 40 CCCCTAACCACTTGAAGAAATTATACCCATTATCCATGTAATTTCTGTAATTTCTACAAC CCCAAAATTTTCTCCCATAGTTGACAATATTCACATAAGTAACAGTACTGAGAACACTAT CATTTTGGCAAAAACACTCAATTATATTTTTTCTACGAAGCACAGATCCAGATGTAGCAA M&C PC933520WOA 118 AGCCCGTGATTGAGGCACTGCTTGCATTCTTAGACATTGAAGCACTGATGGAAGATGGAT GAATACGAATGCACAATGGAAGATGGAAGATGGATGTATGAACATACAAATACATGATGG AAGATGGAAAATCAATGTATGAACATTGAAAAATGAAGGAGGAAGATGAAGATTGGGAAA ACACAAACCCTAAATAGTTTATACCACAACAGTTGAGTCTTAACTTGAATAGTTGCATAC 5 CTGGCATGCCACCTCAGCCATACTTAACATTTTCTCACCCTAACTAACGGAAGTGACCAA ATACTTTAACGAAAGAGAAGTTAAGATACCAATTATAACGATTTTTAGTTAGGGGACCAA CACGAAACTTCGCTTAAAGTTAAGGGACCTAAGAACTAATTATACCTAAAAAAATGTTAT TAAAAATAAAATAAATAACATATACAATCAAAATATTAGTACGCTCTATTTAAAACTATA AATTAATTATATAATATTTATTTATTTTTAAATTTTTAAATTTTTAAATATTAATAATTT 10 ATTCATATTAATATATAAATATTAATAATTAATTCGGTTAGGGCTTGACTTCTAAATTAG CCGTTTACACTCATTTAATTCAATTAGTAGGTAAATAATCACTTGTCCAATTATATAATA TTTAATCGATATTTTTATCTATTAAAATTATGTTTTACAATTTCTTACATAGACTATTTG GTAAGTCACATTTTCAAATTTATAATTTATAGTTTATAATTTATAAGCTCATATGATAAT TTAAAAATGTTTGCTAACAATCTTTTCATCATGAACTTATAACTTATTTTAATAATTTAT 15 AATTTATTTTATAGATGTTATTTTAAATAGCATTTTAACTTATAGTTTATAGTTTATCAT CTTTTCTTTCATTTTTATCTTTATAATTTATATTTTAACAAAAACTCTCTTTTATCCTTT ATAATTTATTTTGGTATAAAATAAAATAAATATATATTAAATGTCTTTTATGTTATTTTA TATTTATAAGTTAATTGAATCGTTAATTTTATCAAATATTTCAATTAACTTATCTGTTAT CGGTCACCAATCATCAGCTATAAATTATAAGTTATCAATCATTCGTCATCAACCATAACT 20 AAGCTAGTTTATCGGTCAATAATTATTTTACCAAATAAAATCATAAATAATTACAAAAAA AAATTCTAAAATTTCTTTTGTTATTTTTAATTTTTTATTCTACTTTATTTTGTTTATATT TCCTTCATCAATTAAACAATAAAAGTTAAAGTGTGTTTGTGTATTTTTATCAAGAATGAA CTATTTTTATTTTCATCAAGGAGTAATTTGTAAAGGGAGTTGAATGGAAATAGTAGTATT TGAGTTGGAGCATGAGTTGAGCATCCGGATACACAATCAACACGCGCCCGCGCAGTGTTT 25 GAGACAGGAAAAAAGTAACGGAAGAAAGCAAACCCCCGAAATCAAATCAAATGGTTCCTT CCGATTCCGACAAGCAAAGCTCGATCTGAGTCAACAACAACAGCTTACACACTATCGATA GTAGCAAGCACAGCAGTACAAGT SEQ ID NO: 54 Glycine max>KRH29333 cds:protein_coding 30 ATGAACAACAGCAGTGGCGGAGGAGGACGAGGAACTTTGATGGGTTTGAGTAATGGGTAT TGTGGGAGGTCGCCATTCACAGTGTCTCAGTGGCAGGAACTGGAGCACCAAGCTTTGATC TTCAAGTACATGCTTGCGGGTCTTCCTGTTCCTCTCGATCTCGTGTTCCCCATTCAGAAC AGCTTCCACTCTACTATCTCGCTCTCGCACGCTTTCTTTCACCATCCCACGTTGAGTTAC TGTTCCTTCTATGGGAAGAAGGTGGACCCTGAGCCAGGACGATGCAGGAGGACTGATGGA 35 AAAAAGTGGAGGTGCTCCAAGGAAGCATACCCAGACTCCAAGTACTGCGAGCGCCACATG CACCGTGGCCGCAACCGTTCAAGAAAGCCTGTGGAATCACAAACTATGACTCACTCATCT TCAACTGTCACATCACTCACTGTCACTGGGGGTAGTGGTGCCAGCAAAGGAACTGTAAAT TTCCAAAACCTTTCTACAAATACCTTTGGTAATCTCCAGGGTACCGATTCTGGAACTGAC CACACCAATTATCATCTAGATTCCATTCCCTATGCGATTCCAAGTAAAGAATACAGGTAT 40 GTTCAAGGACTTAAATCTGAGGGTGGTGAGCACTGCTTTTTTTCTGAAGCTTCTGGAAGC AACAAGGTTCTCCAAATGGAGTCACAGCTGGAAAACACATGGCCTTTGATGTCAACCAGA GTTGCCTCTTTTTCTACGTCAAAATCAAGTAATGATTCCCTGTTGCATAGTGATTATCCC M&C PC933520WOA 119 CGGCATTCGTTTTTATCTGGTGAATATGTGTCGGGAGAACACGTAAAGGAGGAGGGCCAG CCTCTTCGACCTTTTTTTAATGAATGGCCTAAAAGCAGGGAGTCATGGTCTGGTCTAGAA GATGAGAGATCCAACCAAACAGCCTTCTCCACAACTCAACTCTCAATATCCATTCCTATG TCTTCCAATTTCTCTGCAACGAGCTCTCAGTCCCCACATGGTGAAGATGAGATTCAATTT 5 AGGTAA SEQ ID NO: 55 Glycine max>KRH29333 promoter TGATTACTCGTCTCATATTCATTCTTGTACTCGTTCCGATAATAAAATTATAAAAATTTT ATTAATTACTTGTTAATTTTTTACAATTTTTACATAATTTAATATTCTTTTCTTGATGAT 10 TTTTGTAGACAAATGCGATGCAAGTACAAATAATTAAAAATAAATGTGTAATAATCAATT TTTTAAAATCAGATTTTCAATGTAATTTTTTACAATTTTTTTTTACAAATCTAAAGGAGA GGAGAATACCCGATACCCAACGAATATGAGAATGGGGTAACAAATTTTAATCCGTCGGAT ATCGGAGACGAGTACAGATAAATGTTGGGAAATCAGGGACTAGGGGCAATACCCATTTTT GTCGCGCCCCCATTGCCATGTCTAGTGGTGAGATTAAGGATAAATTGAGATTAATAAAAT 15 TTAAAGAGTACTAGTATTTGAACCAAATTAGATGATACTAATAAGTTAGTGCATATTTGA AAATTTGGTGTAAAATAACTTTTAAAGAAAAAAAAAATTTATCAAAAGATCATATTCTTA TTTTTATTTTCATTTTATTTTTATTTGTTAGATGTAAATTAGAACATGCACTACAATAAT TTTAATTATAAATCAAACATGTCACTAAAATTGGTACTACTGTATGTGTGGAGTATACTA TGTCACTATGCCACTGGTTGATGTTGAAAGGAAGATTTTTTGCCAAAACATACACTCTTT 20 CTCTCCAAAAATAAATAAATTAATATACACTAGTTTGGCTTTTAATTCCCAAATTACACC ATTTTTTTGTGACATTGAGATGTAGGGATTTGACAACCCGACTTCTCAGTGATTTTTATT TTTTTTTAATTTAAATTTTATTTTTATTCTAAATTTATGTTTTAGTTTAAATTATTATAC ACAAAAGTTAAGAAGTTAAAAAGTTGGGATTCATCCCTATTTTTTATCTATGGTTTTACT CCAATTTACTCTAATCAAGAATTAAGAGAATCTAACTTACTTGAATGTTATAAATCCTTC 25 ATACCTTATTTAATTCTTACCTATAAAAAATCCCAATCAAGAAAAAAATCCCAATTAAGA GAATCTAACTTACTTTAATTATAACCGAAACAAAGCTACGTAACTTGATTACAAAATGTA CGAGAAACCAAAATTAGTGATGGTGAAAAAAATCACCGACAAAAGTAAGAATCTACACGT GATCTGAGATCAGAGACATACTTTAAGAAGCAACAATCAACAGCCGAAAACCAAAATTAA AGGTATATATTCCTTAAATTGCTTTGTCCCTTTGACTTTTGCCATCGTGATGATTAATTA 30 AAGGTTTAGCAAACCCCTTCGAACTTCATACAATTGACTGAATTGAGAATTTTATTTTCA CATTCGAGGAAGCGATGCTACAACATCACTTTTTTTGTTCTGTATTGTGCTTTTTAACTG CCTTTTTTCTTCTTCTTTTTTTGCCTCCCTAACAAAGACATGTAAAAGTAATTGTAATAA TATTCGTTTCTTATGGAATGCAATCAGTTGATTGATGTAACTATAAACTATTATCTCCTT AATATCGAAAGACAAGTGAAGCCAAACACAAACAAGATAGGGCCTAGGGAGAGGTGTGGT 35 CCATGAATGATGAGGTATGGGTGACCAAACAATGAATGAATAATTGAAGCATCCTTGACC GTTGCTTGAGTTTGTGTCATCCTCAATAATATACTAGTCCCTTGGCTACAGAAACCGATA AGCCTAAAACTGGAATTGCACACATTTACGTTTTTGATTTTGATTTTGTTTTTGGCAATC TCGCCCCACATCAAATGTCACCCGCATTCCGGCAAGTAGTGGATGGTTTCCTCTAGCGGT GCTTTGCCTTTGGGCCACTGGGCCCGCAATTACTCCAGCCCATCATGCCTTGTTGCTGTC 40 CGTTAAAGGGTAGCATAATAAAATAAAAGTAGATCAACAAAATGAGAGCAAGTATTTCAA AAAAAAAAAAACATAGTAAAAAAACACTTCCTCTATTTATATTATCAAGATTTATTTATC TTAAAACATTCATTATCTCAAAAATACCTATATTACTTAATAGTATTTCATGAATTTAAA TCTAAGTTTACTATCAAACTCACCTTTTAAAACAATTATTACACAACAAGTTATAATTGA M&C PC933520WOA 120 ATGTCATAAAAAAAATTGATTATTGTGCTAACACGTGAAAAAAATTTATATTTAATTTTT TTATGTATAATTTGTTTGGACCAATGATAGAGATTAATTGTGATCTAATGAGTTATAAGA AATACGTGGCACATGATCCTAGACAAAAATAAATAAGAATTGTAAAATAATGTATTTTAT AGCTTTTCTGAAAGATTTTTTTTTTTAATTTCTTCTCATGCCCATACATGAATACATGAA 5 TGAGAATTTTTATTTTTATTTTTTTGTCTGAAATAAAGTTAAAAATTGGGAGCAGTGAAT GTTAAGGATGACTTTTGACTTGAATGCAACAAGAAGTAAAGTTCACTTTAAGTTGGAGGC TTGGAGCATCGCCATCCATAACACAACACAATCGACAATCCTAATGGTTCCGACAAAGCT CGACCTGAGTGTGATCTCATGATGTTTCTGCTCTAACTATGTTTGATTTGGATACCCAAC AACAAAAAGAGTGTTGTCGTGTTGTTGTAGTTAATAGTAATAGGACTAAGTAAGAGTAGT 10 GGAAAAC SEQ ID NO: 56 Helianthus annuus>mRNA:HanXRQr2_Chr13g0600651 cds:protein_coding ATGAGTGGCAGCACCACGGTGGCGGTTGCGGCGGCGAGCGGCGGCGGATTCTCCGGCTAC 15 CGGCCGCCGTTCACGCCGGTGCAGTGGCAGGAGCTGGAGCATCAAGCGTTGATATATAAA TACTTAGTTGCCGGTGTTCCGGTACCGTACGATCTTGTTGTTCCGATTAGAAGAAGCATG GAAGCGTTATCTGCTAGGTTTTTCAACAATCCGAATTCACTGGGCTATTGTTCGTATTAC GGGAAGAAGTTCGACCCAGAGCCAGGAAGGTGTAGAAGAACAGACGGCAAAAAATGGCGA TGTTCAAAAGATGCACATCCCGACTCCAAATACTGCGAGCGCCACATGCATCGTGGCCGA 20 AACCGTTCAAGAAAGCCTGTGGAATCCCAATCTCCGTCTCAGTCGTTGTCAACTGCGGTA TCGCGTGTCACCACCGTAAGCAGCACTGGTAGTGGTAATAATAGAAGTTACCAGAATCTC GGCAATAATAGCTTCCAGAACTCACATTTGTATCCCACCTCAAATTCCGGGAGTTTTGAT TTTGGGAGTAGTGCATCAAAGTTGCAGGTGCATGCTAGTGCCTATGGGATCAACAACAAC AATTATAGGTACGATCACGGGCTAACTGTCGACGTGGACGATAGAAATTACTCTTCCGGA 25 GCTTCAAGTTCAAGGGGTCTAGCGATAGAATCGAATGCAGACAGCACATGGCGTCTCGTG TCTAATCAAGTTCCCACAAGTTCGTTAATGGAATCAAGAAATGACTCTTATTTACAAACC AAATCCCCACGACTAACCACGGTTAATGCATTTGAGCCTGTGATTGACGCTACATCAGCG TCAAAGCCAAGTCAACAGCATTGTTTATTTGGTAGCAAGATTGAGTCGCCGGTTGAACCG AAACATGAGCAGCACTTGATGCGCCCGTTCTTTGATGAGTGGCCGGAGGCTAGGGAACCA 30 TGGTCCAGTCTTGATGCGGCTAGTGGCAAGAACTCGTTTTCCACGACTCAGTTGTCAATG TCGACGGCTCCTGAATATTCAACCGACTGCACCCCGAACGATGCTTAA SEQ ID NO: 57 Helianthus annuusHanXRQr2_Chr13g0600651 promoter GAGAGAGTGAACCACAGCCAATCAAATTTCTTCTTTTGTTTTTAATTTTTTTCTTTTTTA 35 TATATATAGTTAATGGGAGCTTTAAATGAGCTTTATTCCACGCTATGTAAGTTAACAATA ACGCCTCATGCTTATATGGAGAGAGTGAACCACAACCAATCAAATTTCTTCTTTTATTTT TAATTTTTTTCTTTTTTATATATATACTTAATAGTTAATGGGAGCTTTAAATGAGCTTTA TTCCACGCTATGTAAGTTAACAATAACGCCTCATGCTTATATGTCGCTTACGTGTCGCCT ACATGGCATGATAAAGTCCTAAGGGCTTTATTCCATACCACCCAGGCTTACTATGCTTTT 40 TTTATTGATATAAAATTTTGAAATTTCATGATTGTTGTTGTATGATTATTTGTTTCTTTA GTCTTCAATTTTCAATAATAAAGTTTCTTATTGATTTGTGATATCAATAAAAGTAGAGTA M&C PC933520WOA 121 AACTGCCATTTTGGTCATGTGGTTTGGGCACTTTTGCCATTTTAGTACAAATCTCAAATT TTTTACATTTGGGTCCCTGTGGTTTGCATTTTGTTGCCATTTTAGTCCAAAATCCAAAAA CCCCTATTTTTTGACTGTTGCAACCTCCTATTTTGTCTTTTAGTTCAGGGGATTTTGGTC ATTTTGCTTTTCATTTAACATTTTCTTAATTCAAAAATAATATATATATATATATATATA 5 TATATATTAATATACACACATACATATACACTCCAGTTCATCATCCCCAACATAAATCAG TCTTCTCTCTCTCTCTTTTCTCTCTTTATATAGCACCACCACCATCGCCACCTCCTCCCT CTCTCTCTCTTCCCTCTCTATAGCACCACCGCCACCTCCTCCCTCTTCTTACCACCGCCA CTCCGGCGTCGTCCCATCCAATAGTCACACCCACTCCGGCGTCGTCCCCCTTCTCACCAC CGGTGTCGATGATGAAGATGATTAAGGTGTCGATGATGAAGATGAATCGCTGAAGTTGAA 10 TCGCTGATTCTGCTGATGAAGGTGTCGATGATGAAGATGAACAGTTGATTGATTGATGCG GATGAAGGGGGTTCGAAGTTTCCCCAAATAAATCATCATCATCCCCAAATCTGAGTTTGA GTTGTATTTGTTGAGACATGAAAGTTTTAGATCTGAATTTGTGTTAATAAATTGCAAGAT TGTTGTGTTCATCAATCTGAAAGATTTGCAGGTTTTTGTGTTGTTGATCGACTGGTTTTG CAGATTGTTGTTTTTGATTAAGAAGGAGTATAGTTTACAGGTTCGACTATTGCAGATTGT 15 TGTTTGCTGTTGCAACGAAGAAGAACACAAGTTGCAGATTTTGTTTGCGGCGATCTGGGG GATGATGTTCATCAATGCCGCCGCATATCATCTTCCCCCAAATCAAAGGTCCAGTTTCAT GTGGGTTAATGGGTTTGATGCAAAATTAGGGTTTGGGTTTTGTGTGGCTCGGAGTAAAGA AAGAAGGTGAGATTGTCGGTTTGATTGTCTCCGTTTGATTGTCGGGATGGTGGTGTTGGC GGCAGTGGGATGGTGGTGGTGGGGATGATGCGGTGGTGGGGGTGGGGGTGGTGGTGGTGG 20 TGGAGGTGTGGTGGTGGTGGTGGTGCTTGTTGTTGTTCTGGCAGTGGTATTATTGTTGTT ATGGCAGTGACATTATTATTTATTAAATTTGAGTTATAATAAATCTAAATGACCAAATTA CCCCTGAGGATAAAGACAAAATAGAGGGGTTTAAGTTAAGGAATCTGGCCAAATTTAAAA TTTGGACTAAAATGACAACAAAATACAAACCATAGGGACCCAGATGTAAAAAGTTTAACA TTTGGACTAAAATGGCAAAAGTGCCCAAAGCAAAGGGACTAAAATGACAGTTTACTCATA 25 AAAGTAAGATTAACAATTTTTATTATATTCTAAAAACAAGTTGATCCTTATTCCACATGT TTTTATAGACATCTCTTTCACCAATATCCCATCAAATGGTTATGGTTATACACATTTGAT TTGAAATCTTTTATCGACTTGTTTTAAAACACTTTTAAAATATTTCATCACATCCAAAAG ATTCACACACAAGATTTATACAAACATTTAAGATCTTTTATCCACTTGCTTTAAAACACC TTTAAAATATTTCACCACATCCAAAATATTCATACACAAGATTTATACAAACATTTGAGA 30 TCTTTTATCCACTTGTTTTAAAACACTTTTAAAATATTTCACCACATCCAAAACATTCAC ACAAAAGATTTCTACAAACAATATAATAAAATATTATATTCTAATAACTTGATGTCATTG CATACATGGATGACAACAAATGACCAAATCTTGTCTCCATCCCAAAAACTTATCTTCAAT AAAAATAAAAATAATAATAAAAAAATAATAATAATGATTACAAAAATGTTAACATCCCCA CACCCCCAAATTCCCTTTCTTACCACGCAAACCCTAACCCTAGATCTGCCACTTCAAAGG 35 TAACTAAAATCATCTCAATCTCTCCATCAATTCACACACAAATTCACGCAACACATCTGT AATCCACTCCATTATCTCATCAATTTCAGTTTCATCTTCTCAAATCCACCGTAGAAATCG ACGGAG SEQ ID NO: 58 Hordeum vulgare>HORVU.MOREX.r3.2HG0193490.140 cds:protein_coding ATGGCGATGCCCTTTGCCTCCCTGTCGCCGGCAGCCGACCACCACCGCTCCTCCCCCATC TTCCCCTTCTGCCGCTCCTCCCCTCTCTACTCGGTAGGGGAGGAGGCGGCGCATCAGCAT CCTCATCCTCAGCAGCAGCAGCAGCAGCACGCGATGAGCGGCGCGCGGTGGGCGGCGAGG M&C PC933520WOA 122 CCGGCGCCCTTCACGGCGGCGCAGTACGAGGAGCTGGAGCAGCAGGCGCTCATCTACAAG TACCTCGTCGCCGGCGTCCCCGTCCCGCAGGACCTCCTCCTCCCCATCCGCCGCGGCTTC GAGACCCTCGCCTCGCGCTTCTACCACCACCACGCCCTTGGGTACGGGTCCTACTTCGGG AAGAAGCTGGATCCGGAGCCGGGGCGGTGCCGGCGGACGGACGGCAAGAAGTGGCGGTGC 5 TCCAAGGAGGCCGCTCAGGACTCCAAGTACTGCGAGCGCCACATGCACCGCGGCCGCAAC CGTTCAAGAAAGCCTGTGGAAACGCAGCTCGTCGCCAGCTCCCACTCCCAGTCCCAGCAG CACGCCACCGCCGCCTTCCACAACCACTCGCCGTATCCGGCGATCGCCACTGGCGGTGGC TCCTTCGCCCTGGGGTCTGCTCAGCTGCACATGGACACTGCTGCGCCTTACGCGACGACC GCCGGTGCTGCCGGAAACAAAGATTTCAGGTATTCTGCCTATGGAGTGAGGACGTCGGCG 10 ATCGAGGAGCACAACCAGTTCATCACCGCGGCCATGGACACCGCCATGGACAACTACTCG TGGCGCCTGATGCCGTCCCAGGCCTCGGCATTCTCGCTCTCCAGCTACCCCATGCTGGGC ACGCTGAGCGACCTGGACCAGAGCGCGATCTGCTCGCTGGCCAAGACTGAGAGGGAGCCA CTGTCCTTCTTCGGCGGCGGCGGCGACTTCGACGACGACTCGGCTGCGGTGAAGCAGGAG AACCAGACGCTGCGGCCCTTCTTCGACGAGTGGCCCAAGGACAGGGACTCGTGGCCGGAG 15 CTGCAAGACCACGACGCCAACAACAACAGCAACGCCTTCTCAGCCACCAAGCTGTCCATC TCCATGCCGGTCACCAGCTCCGACTTCTCTGGCACCACCGCCGGCTCCCGCTCGCCCAAC GGTATATACTCCCGGTGA SEQ ID NO: 59 Hordeum vulgare>HORVU.MOREX.r3.2HG0193490.1 promoter20 AAATTAACCAGAGGGTTGCCCCAAATAAGAGGGAGTATTTTTTTGAAAGAGGAGGAACAT TTTATTTGGGCAAAATTAGAAACTGGAATTGCCTCAAAAGAATTAGGAGAGGAGAGGGGA GTTTCCTATCGCACTAAAGACTCTGAGAGCATTTGAAGCACTCATAACTCAGAACACACA CACGCAACAGATGATGCAAGATGCCATGATGCAAACTAAATGACGATGCAACCAACAAAA TAAAATAGCCACACGACGGAAACGGAATAAAGGAGGGAATCTTCTCGAGCGTCGGTCTCG 25 GGATGTCACACATACACACTATGGGTGAGCCACTCTTCAAGATATCTTCACAAGTCCATT GCCACCACAATGGACGACAAGCTTCAAGCATGATATCTTCGTGATGATCCACTTAAACTT GCACATCGCAATCTTGATGACGATCACCACTTGACATAATTCTCCATGGGTTGTATGAGA TCTTCCTCTTGACGCAAGCCCATGGAAACACACCTAACCCCACATAGAACTCTCATAAAG ACCATGGGTTAGTACACAAACACGTAATGAAAAATGGTTACCATACCATGGGATCACTTG 30 ATCCCCGTCGGTACATCTTGTACGCTTTGTGTGTTGATCATCTTGATTTACTCTTTGTCT GACGATCTGGATCAACCTTGTGTCTCTATGACCATTCTTTGGATAATACCTTGAATACCA CCTTGGTCATCATATAACTCCTTGAACCCAACAGATGGACTTCAAGAAGTGCCTATGGAC AAATCATATAAATATAACTTAAGGCAACCATTAGTCCATAGGAATTGTCATCAATTACCA AACCACATATGGGGAGATATGCTGTAACACTATTTGTTATTCGTGTTGGCTATGATTCAG 35 AACTACTTAGTATTGTCGAACTAGTTGGTCTTTGCTGTGTTAAACTAGTTGCTGCCAATT TGTTATGTGTGACAATGTAAACCGGTTTGATATGTTCAGTGTGATGAACTATTTGTTCTT TGTGTGCAAATGAAGTATCTTTCGAGGTGTGATGCATACTAAAATGTGTGTGATCTGCAT ACTTTTCATTTGCCAACTATTGAATTGTTGTGCCAAATTTCTTGAATTGGCACATGGAGA GTATGGCGTCTTGGTGACCCAACTCCATGAAGCATATTATTTGAAAACAAATTCAAACTT 40 CAAACTTTCTAAACTGAAAAAAATCCAAAAAAAGATAAGCAACTATACAAGGATGAAATG TATATGTGTGTAAAAATTCAGGATAAAATACGTTTAAATGCAACCTCTACAAAAAAGTCA ATTTTTTGGATTTTAAGGATGAATAGTATCATTATTATTTTTGACTTGGAACTGTAAGGG TAACACCCCACGGCAAAAACTTTTATTCCCAAAACTTATACAATATATACGCATGAGGTA M&C PC933520WOA 123 AAAAGAGAAAAGGAACTAAAAAAGTACAGATGTAGAAAAGTACAAGACTCACCTCAAGGA GGTAAAAAGGAGATCCATCTAAGCATATCATCCTTAAACTTAACTTTAATCCTATGAGCT AAAGAGACATATTATGGATAAAAGTGGCTCTCCACCTCCTGAAAGAGGGCCTGATATTCC TGAAGACTCTATCATTTCTAACAGTCCAGATGTTCCAACAGGTCATGAACACAATCTCAG 5 CAAAGAAAGGTTTTTGGAAGTCCCGTGTGGCAGCAAAAGCAATCTCAGCCATATTGTCAC TAGCAGGCCATGAGATCTGAAGATAATTCCACACCCTCCAACTGAAGTTATAAAGTGCCT CAGATTTATCTTTTTCTACACAGCCCTCACTTTAACAGATTTGTCCTGAAAATTTACACA CATGTACATTATGACTCCATGTATGTGTGTACTTTTTTTTAGAATTATAAAAAATATGTA TTTTGGCAAATTCTGAATTTTGTAAAAAGGCCTCCATGAAGCTCGGTCTTCAAAAGCAAT 10 TTTCGCTTGAATTGAAACTATAAAGTTCAAATAAGTTTTTCAGACCCTACCGTCATACAC CTTGACGGTAGAATGTGAAACCCTACCATTATATAAACGAATTCCCGTTACAACAACTTT ACACACGAGGTCAGACTCCTACCGCCATAGTTCCTAATGGTAAGGTCTTGCATCCTATCG TCTTATACTTGGCGGTACGGCCGTTACGCCACGTGAGCCCTTCGGCTGGCAGTTGACGGC CGCTGTTGTTACTCGACTGTCAGATACCTATAAACCTATCGCCAACCTGTGTAACAATGA 15 AAAACGGTCAAATCCCGAAAAAATTTCGAAGCAGGATCGCATCCTGCTAAACTTTTGACA AATGGTCAAAACACGAAATTTTTGCCGCTCGTTGTGCCTCTGTAAGCTGGAAGCCTACGG TGTCGGCCTCACCCCCCACACGGTGCTGCCGCTGCTGCGCCCATCGCCAGCGCTTCACGC TATATATCCACCCCGTCGTCGTGTGAGTCTCACCAGGCAGATCGAGCCCTGCGCAGCGAG GGGAAAGAGACACACACAGCGCCACCAGGCAAGTAGTAGTAAAAGGCAAAAGCACGGCAC 20 ATTAAAAGAGAGGCCAGCCCAGCCCCGGACCGGACCGGAGCCAAGCAGCAGCCGCAGCCG CAGCCGCAGCAGAGGAGAGAGAGAGGGAGGGAGAAGCATAT SEQ ID NO: 60 Triticum aestivum>TraesCS6D02G245300.1 cds:protein_coding ATGGCGATGCCGTATGCCTCTCTTTCCCCGGCAGGCGACCGCCGCTCCTCCCCGGCCGCC 25 ACCGCCTCCCTCCTCCCCTTCTGCCGCTCCTCCCCCTTCTCCGCCGGCGGCGGCAATGGC GGCATGGGGGAGGAGGCGCGGATGGACGGGAGGTGGATGGCGAGGCCGGTGCCCTTCACG GCGGCGCAGTACGAGGAGCTGGAGCACCAGGCGCTGATATACAAGTACCTGGTGGCCGGC GTGCCCGTCCCGCCGGATCTCGTGCTCCCCATCCGCCGCGGCATCGAATCCCTCGCCGCC CGCTTCTACCACAACCCCCTCGCCATCGGGTACGGATCGTACCTAGGCAAGAAGGTGGAT 30 CCGGAGCCGGGCCGGTGCCGGCGCACGGACGGCAAGAAGTGGCGGTGCGCCAAGGAGGCC GCCTCCGATTCCAAGTATTGCGAGCGCCACATGCACCGCGGCCGCAACCGTTCAAGAAAG CCTGTGGAAACGCAGCTCGTCCCGCACACCCAGCCGCCGGCCGCCTCCGCCGTGCCGCCC CTCGCCACCGGCTTCCACAGCCACTCCCTCTACCCCGCCATCGGCGGCAGCACCAACGGT GGTGGAGGCGGGGGGAACAACGGCATGTCCATGCCCAGCACGTTCTCCTCCGCGCTGGGG 35 CCGCCTCAGCAGCACATGGGCAGCAATGCCGCCTCTCCCTACGCGGCTCTCGGTGGCGCC GGAACATGCAAAGATTTCAGGTATACCGCATATGGAATAAGATCTTTGGCAGACGAGCAC AGTCAGCTCATGACAGAAGCCATGAATACCTCCGTGGAGAACCCATGGCGCCTGCCGCCG TCGTCTCAAACGACCTCATTCCCGCTTTCAAGCTACGCTCCTCAGCTTGGAGCAACGAGT GACCTGGGTCAGAACAACAACCACAACAACAGCAGCAGCAACAGTGCCGTCAAGTCCGAG 40 CGGCAGCAGCCGCTCTCCTTCCCGGGGTGCGGCGACTTTGGCGGCGGCGGCATGGACTCC GCGAAGCAGGAGAACCAGACGCTGCGGCCGTTCTTCGACGAGTGGCCGAAGACGAGGGAC TCGTGGTCGGACCTGACGGACGACAACTCCAGCCTCGCCTCCTTCTCGGCCACCCAGCTG M&C PC933520WOA 124 TCGATCTCGATACCCATGACGTCCTCCGACTTCTCCGCCGCCAGCTCCCAGTCGCCCAAC GGTATGCTGTTCGCCGGCGAGATGTACTAG SEQ ID NO: 61 Triticum aestivum> TraesCS6D02G245300promoter 5 TAGCATTTTTTTGTTTTTCCCACCACCACCATTACAAATCTACTTTTCTTCACATGCTAT TCCCCACAAAAGCAAATGCAGTACTGTTTTCTGTTTTCACTTTTTTTACCTTGTTGTGCC GCATGTGCCGATCTTGTGTGGTCTGTCATTCCAATTTGATGTGCTCTATCTGCAATAGCG CTTTCCTGCAACTGTGAAGCTTGCCATGAGATGTTTGCCGACGTGGTTCCACCGTAACAA CCTGCGATCATGGTAGCAGTTGGTCATTGCATGGCAACTGTAGTTTGTTCAACATGGCAC 10 ATGTAGTGGTCCAACCATGCCACACAGATTTGTTCACACTTGTCATCCACGGTTGTTTAG ACATTGCAGTCACATTTGATTATACATGGCAACTGCAGTTGTCCAGGCATGGCAACTGCA GTTTGTCCAGCCATGGCAACCGAAGCTGTCCAGGCATGGCAACTGCAGTTGTCATATATG TTCCTTTCTCCTAACTCACTTTCCACAACTTCACTTTCATGCGCTCACCTGTGTACGATG TTGATGGCAGGTACATGGTTGCTGCCTGATGCCACGTGCTGCATGAAGGTCCTATATTTT 15 TTCGATATAACTCGTGAGTCTTTTGCCATCCTTGCAAGTTTAATTTCATGTCCACAACAA TAAGTTCCTTTTTCGTATGCGTTACCTTGATTTGCCACATTAGCTAGCTGAAGTTGGTTG CCCGTACATTTGTCAGCGTTAGCGCCCTGTGACGAAACTTGCCATGCTGCCCCCCTGATT GTGGTTTGGTCATAAGAACCTGCTAAAAAATAGATTGTAGGACATAAGTAGAATGACATC TTCATCGTCACGAACATGGCAACTGTTTCTTCTTTGTTTATCATGCATCGTACCGATGCC 20 CGCGTAGTGCTCTAGTAGTGGCAGCAATGGCCCTGAAGAAATGAGTTGATTGTACTCTGC TGCATCCCAAGGTGGCGTTTCCGGCCTTTGAGAAAGCCAAGGATCAGTGCCATCTTCGTG ATTCATTCTTCTGCTTTTTCTTTTCTGCTACTATGCTTTTAGTCACTGCATGAACAAGAA CGCATCAACAATCCACAAAAAGCGTTCTTGCTGTTTGCACGTAGAAGATAACACGGCAAT CTCATAATATTTTTTGCGTAGGCAACCAACACCTCATGGCAAGTAGGACATGCACATCCA 25 TTTTTCTTTTCTGAATTCTGGATGCCATCTATCATTTTGAAGCGATGGCAACAGAAAATA AAATAGGATGGCAAGCAATAATACATGGTGGCAACTATGGACAACGATAGATGGCAACTG ACGTTAGATACAAGTGGCAATTATTTTTCCTCCCTCCCCATGCCAAATTCCTCCTTTCTC TCCCTATTTTATAGTGATTACTACGCTACCAACTACTCGCATCAAAGCCAACCCAGAAGC TTGGCACAAGTCTAGCATAGTATATGGCAGATCTGGCGTATGTTGGTGGGAAAATGCAAA 30 GACACACAAATTCGTGGGGTGTTTGCCCTGATAGCGTGGATCCAGTCGCCATCTTCGTGG GCAAATTTTGCAAATTCAGATTTCTGGACAAAAGAAGATCGGGGATCCACCTGTTTTAGC TCGTCGTCTTGGGAGTGCGGGGAGGGGGGTAGGGTGGGGGTGGGGTGGGTGGTTAGCTGT GGGAAAGGCGCTAGGGATTTGCTCTGGTTGCCATGGCAACCAGAGAAGGAAGGCGACGGA GGTAGGGGATCGGGAGATGCGAGACAATGGCGGCAGGGCGGACCGGGGATCGGAAGGAGC 35 CCGGGACAGCTGGCGTGCTGAGTCGTGCGGGCAGCGCGGTCGTTTGGCCCGGACGTGTGG GCGGTTTTGCCACACACCGGACGTGCGGGTTGTGGCTGCGCGCGCCCGGATGCGGTTTTG CGGGCGAGTTCTTCTCCATGCCACACGAGGCGTGCGGCACAACCACCCGATACACCACAC GTGTGGCAGTTATCGGTGTTAAAAAAATGACGAGAGAAAAGTGGCGCAAACGGTTGCCCC GCACCCTCTCACGGACGGACTTTAAAAGTCGGCATTGGTAACCGCAACACAACACAGACA 40 GACGCACCCCAAGCCTCTCTCTATCTCTCTCTTCCCATGCAATAGTTGTCACCACTCGCT CGCTACAGTGCCCGCATTGCATCGCATCCACATCCATATGACCATATCCATTCCTCCCCA CGAGAAAAGGAGAGAGAGGGGAGAAATACTAGTCGTCGTCGTCGTAGTAGCTGGTACGTC TACGCTAGAGCGACAGGGAAAGAGGAGGGAGGGGGCGCTTGTCATCTACTCCTCCTCCTC M&C PC933520WOA 125 GCCCCTAGCTGGGATCCACAGCCTCCTCCTCCTCCTCGTGTCGGCCTCGTCCACATCCAC CGTCTCCTCCGAGCGAGGTGGACAGCGACGCGGCCACGGAGCGAGGGAGGGAGAGAGACA AAGCCGGTAATAAAGGCGGGGGCGCGCGCGCGCACAAGCCAAGCAAAGCACATTAACGAC GCCAGCCAGCCCGCGGGGAACCCCATTAAAGACGCTTCCGGGGGAGCGCCGTGGGCAAGC 5 ACAGGGGCTTAGCTTAGCTTGGCTTGTGTGTTGTGTGCGCGAGAGGGAGACAGCGGCCGA GAGAGAAAG SEQ ID NO: 62 TRITICUM AESTIVUM>TRAESCS6A02G269600.1 CDS:PROTEIN_CODING 10 ATGGCGATGCCGTATGCCTCTCTTTCCCCGGCAGGCGACCGCCGCTCCTCCCCGGCCGCC ACCGCCACCGCCTCCCTCCTCCCCTTCTGCCGCTCCTCCCCCTTCTCCGCCGGCGGCAAT GGCGGCATGGGGGAGGAGGCGCCGATGGACGGGAGGTGGATGGCGAGGCCGGTGCCCTTC ACGGCGGCGCAGTACGAGGAGCTGGAGCACCAGGCGCTCATATACAAGTACCTGGTGGCC GGCGTGCCCGTCCCGCCGGATCTCGTGCTCCCCATCCGCCGCGGCATCGAGTCCCTCGCC 15 GCCCGCTTCTACCACAACCCCCTCGCCATCGGGTACGGATCGTACCTGGGCAAGAAGGTG GATCCGGAGCCGGGCCGGTGCCGGCGCACGGACGGCAAGAAGTGGCGGTGCGCCAAGGAG GCCGCCTCCGACTCCAAGTACTGCGAGCGCCACATGCACCGCGGCCGCAACCGTTCAAGA AAGCCTGTGGAAACGCAGCTCGTGCCCCACTCCCAGCCGCCGGCCGCCTCCGCCGTGCCG CCCCTCGCCACCGGCTTCCACGGCCACTCCCTCTACCCCGCCGTCGGCGGCGGCACCAAC 20 GGTGGTGGAGGCGGGGGGAACAACGGCATGTCCATGCCCGGCACGTTCTCCTCCGCGCTG GGGCCGCCTCAGCAGCACATGGGCAACAATGCCGCCTCTCCCTACGCGGCTCTCGGCGGC GCCGGAACATGCAAAGATTTCAGGTATACCGCATATGGAATAAGATCTTTGGCAGATGAG CAGAGTCAGCTCATGACAGAAGCCATGAACACCTCCGTGGAGAACCCATGGCGCCTGCCG CCATCTTCTCAAACGACTACATTCCCGCTCTCAAGCTACTCTCCTCAGCTTGGAGCAACG 25 AGTGACCTGGGTCAGAACAACAGCAGCAACAACAACAGCGGCGTCAAGGCCGAGCGACAG CAGCAGCAGCAGCCGCTCTCCTTCCCGGGGTGCGGCGACTTCGGCGGCGGCGACTCCGCG AAGCAGGAGAACCAGACGCTGCGGCCGTTCTTCGACGAGTGGCCGAAGACGAGGGACTCG TGGTCGGACCTGACCGACGACAACTCGAACGTCGCCTCCTTCTCGGCCACCCAGCTGTCG ATCTCGATACCTATGACGTCCCCCGACTTCTCCGCCGCCAGCTCCCAGTCGCCCAACGGC 30 ATGCTGTTCGCCGGCGAGATGTACTAG SEQ ID NO: 63 Triticum aestivum>TRAESCS6A02G269600 promoter ATCTTTCACTAAAACCCTTCATACGCATTGCCTGCTGAAGAAAAGGCCATTTGACTTTAT CATAAGCCTTCTCAAAATCCACTTTTAAGATAACCCCATCTAGTTTCTTCGAGTGAATCT 35 CATGAAGCATTTCATGTAGGACAACCACTCCTTCGAGTATGTGTCTTCCAGGCATGAAAG TTGTCTGACTAGGCTGAACCACTGAGTGAGCAATTTGTGTCAATCTATTGGTTCCTACCT TTGTGAAGATTTTAAAACTCACATTCAGCAGGCAGATGGGTCGGAATTGTTCTATCCGGA CCGCCTCATTCTTCTTTGGTAATAACATTATCGTACCAAAGTTGAGGTGAAACAACTCCA AATGGCCATCGAAAAAATCGTGAAACATAGGCATAAGATCATCTTTAATGATATATGCCA 40 ACATTTCTTATAGAATTCAGCTAGAAACCCATCTGGTCCCGGTGCTTTATTCGGCTTCAT CTGAGTAATGGCCAGATGAACTTCTTTTTCAGTAAACGGTGCAGAGAGAACATCGTTCTC CGCTGCCTGTAATTGAGGTATATCATCAATCACAGACTCATCAAGAGACACCGTCGAAGT M&C PC933520WOA 126 ATTCGGAGGTCCAAAAAGGCCTTTATAATAGTTGGAAATATATGCTTTCAGATTCTCATG CCCCACTATTGTTCCCTCGTCCTGTTCAAGCTGTATGATCTTCTTCTTCCTATGTTTACC ATTTGCAACCATATGGAAAAACTGTGTGTTGTCGTCCCCTTGGACAACCTTAAGCGTTTT CGCACGCAACGCCCACTTGAGTTCCTCCTCTCTCAAGAGAGCATGTTGGCCTTGCTCAGC 5 CTCAGACTTGGTTCGATGCTCATTAACAGAAAGAAGTGTGGTTTCAGCCTTTACATCTAG TGTCTCAATGAGTTGAGTTAGACGTTCTTTTTTCTGCTTATAAATCCCAGTCTCATTCCT GGCCCATCCTCTCAGAAATTGTCGGAGGTTTCTAATCTTATTCTGCCATCTCTCGACATG TGTCCTTCCTGTAATTGGCTTAGCCCATTCGCATGCAATCATCTCCATAAATCCTTCTCG CTCAAACCAGCTTTACTCGAAAGAGAAGATGTTTTTGTTTGCAACATGGGTAGCCTCACC 10 CGAATCTAAAAAGAGTGGTGTATGATCTGAGATCCCTCTATGCATTGCATGGACCGACAC CAACGGATATTTTTGTTCCCACTCCACACTAGCAAGTACCCTATCCAGCTTTTCATAAGT CAGAACAGGTAACGAGTTGGCCCATGTAAACTGTCTACCGGTGAGCTCAATTTCTCTCAA ATTGAGGCTCTCGATAATCATGTTAAACATCATAGACCAACGTCCATCGAAATTGTCATT ATTCTTTTCTTCTCTTCTCCGAATGATATTAAAATCACCCCCGACTAGCAGTGGCAGATT 15 TTCATCTCCACAAATCCGCACTAGATGGGCAAGAAAATCGGGTTTAAATTGCTTGGAGGA GTGAGAGCATCTACAACCGGACTTAGCGAATCTGGGCTCTATAAGCCCGCGGGTGCCTCC GCGGACGGCCCTCCCTTGAGTTGCCGCACATTCACACATCTCAAATACGGATTCTTGAAT CCATGTATCCATGCACGTCCATCATACGATATAAATCATCCCAATTCAAATGTTTGAAAA CAAAATACGACAATGCAAAGCAAATCATAGTTCAATAATTCAGACATGCCAAATTAAAAT 20 CAATATCCGAGCATGATAGATCACTCGTTGGACGCCATCCATGCCCGCTTGCTCCGCGGC CATCCTTGCGGGCGGCGAGGATGGGGAGCAAGGGTGGCGGACGGCAAGGGCTTGGACACG AAAATAGGTGGATGAAGGCGGGAGAGAGGAGGGTTTAGTGAATTTTATGCAATTTATGTG GGGGGTTGGCCTGTCGGGTTCTACGTAATGGACGCGCCGAGGCATGAGGGATGCCGGTCA GCTTGGGTGTTTTAGATGCCCGTCCGGTCTTTTATTTTTAAGTCCGTAATTGGGCCGTTC 25 GCCGGACGTTCCATAGAGGTTTGGGGTGCCGGGAAGTAGATGCACAGTACTTCCGTTATC ACCACGACACAAGAAGCAAGCACATAGTACTGTTGTAAAAAAATGACGAGGGAAAAGTGG CGCAAACGGTTGCCCCGCACCCTCTCACGGACGGACTTTAAAAGTCGGCATTGGTAACCG CAACACAACACAGACAGACGCACCCCAAATCTCTCTCTCTCTCTCTTCCCATGCAATAGT TGTCGCCACTCGCTCGCTACAGTGACCGCATCGCATCGCATCCATGTCCATTCCTCCCCA 30 CGAGAAAAAGAGAGAGACAGCAGAAATACCAGTCGTCGTCGTCGTCGTCGTAGCCTGGTA CGTCTACGCTAGAGCGACAGGGAAAGAGGAGGGCGCTTGTCATCTACTCCTCCTCCTCGC CCGCTACTAGCTGGGATCCACAGCCTCCTCCTCCTCCTCGTGTCGGCCTCGTCCACATCC ACCATCTCCTCCGAGCGAGGTGGACAGCGACGCGGCCACGGAGCGAGTGAGAGAGACAAA GCCGGTAATAAAGGCGGGCGCGCGCGCGCGCACAAGCCAAGCAAAGCACATTAACGAGGC 35 CAGCCAGCCCGCAGGGAACCCCATTAAAGACGCTTCCGTGGGAGCGCCGTGGGGAAGCAA GCGAGCGAGCACAGGGGCTTGGCTTGCGCGTCGTGTGCTGTGTGCGCGAGAGGGAGACAG CGGCCGAGAGAGAAAG SEQ ID NO: 64 Triticum aestivum>TraesCS6B02G296900.1 cds:protein_coding40 ATGGCGATGCCGTATGCCTCTCTTTCCCCGGCAGGCGACCGCCGCTCCTCCCCGGCCGCC ACCGCCTCCCTCCTCCCCTTCTGCCGCTCCTCCCCGTTCTCCGCCGGCAATGGCGGCATG GGGGAGGAGGCGCGGATGGCCGGTAGGTGGATGGCGAGGCCGGCGCCCTTCACGGCGGCG CAGTACGAGGAGCTGGAGCACCAGGCGCTGATATACAAGTACCTGGTGGCCGGCGTGCCC M&C PC933520WOA 127 GTCCCGCCGGATCTCGTGCTCCCCATCCGCCGCGGCATCGAGACCCTCGCCGCCCGCTTC TACCACAACCCCCTCGCCATCGGGTATGGATCGTACCTGGGCAAGAAGGTGGATCCGGAG CCCGGCCGGTGCCGGCGCACGGACGGCAAGAAGTGGCGGTGCGCCAAGGAGGCCGCCTCC GACTCCAAGTATTGCGAGCGCCACATGCACCGCGGCCGCAACCGTTCAAGAAAGCCTGTG 5 GAAACGCAGCTCGTCTCGCACTCCCAGCCGCCGGCCGCCTCCGTCGTGCCGCCCCTCGCC ACCGGCTTCCACAACCACTCCCTCTACCCCGCCATCGGCGGCACCAACGGTGGTGGAGGC GGGGGGAACAACGGCATGCCCAACACGTTCTCCTCCGCGCTGGGGCCTCCTCAGCAGCAC ATGGGCAACAATGCCTCCTCACCCTACGCGGCTCTCGGTGGCGCCGGAACATGCAAAGAT TTCAGGTATACCGCATATGGAATAAGATCTTTGGCAGACGAGCACAGTCAGCTCATGACA 10 GAAGCCATGAATACCTCCGTGGAGAACCCATGGCGCCTGCCGCCATCGTCTCAAACGACC ACATTCCCGCTCTCAAGCTACGCTCCTCAGCTTGGAGCAACTAGTGACCTGGGTCAGAAC AACAACAGCAGCAGCAGCAACAGTGCCGTCAAGTCCGAACGGCAGCAGCAGCAGCAGCCC CTCTCCTTCCCGGGGTGCGGCGACTTCGGCGGCGGCGGCGCCATGGACTCCGCGAAGCAG GAGAACCAGACGCTGCGGCCGTTCTTCGACGAGTGGCCCAAGACGAGGGACTCGTGGTCG 15 GACCTGACCGACGACAACTCCAGCCTCGCCTCCTTCTCGGCCACCCAGCTGTCGATCTCG ATACCCATGACGTCCTCCGACTTCTCGGCCGCCAGCTCCCAGTCGCCCAACGGTATGCTG TTCGCCGGCGAAATGTACTAG SEQ ID NO: 65 Triticum aestivum>TraesCS6B02G296900 promoter20 ATCTTTCACTAAAACCCTTCATACGCATTGCCTGCTGAAGAAAAGGCCATTTGACTTTAT CATAAGCCTTCTCAAAATCCACTTTTAAGATAACCCCATCTAGTTTCTTCGAGTGAATCT CATGAAGCATTTCATGTAGGACAACCACTCCTTCGAGTATGTGTCTTCCAGGCATGAAAG TTGTCTGACTAGGCTGAACCACTGAGTGAGCAATTTGTGTCAATCTATTGGTTCCTACCT TTGTGAAGATTTTAAAACTCACATTCAGCAGGCAGATGGGTCGGAATTGTTCTATCCGGA 25 CCGCCTCATTCTTCTTTGGTAATAACATTATCGTACCAAAGTTGAGGTGAAACAACTCCA AATGGCCATCGAAAAAATCGTGAAACATAGGCATAAGATCATCTTTAATGATATATGCCA ACATTTCTTATAGAATTCAGCTAGAAACCCATCTGGTCCCGGTGCTTTATTCGGCTTCAT CTGAGTAATGGCCAGATGAACTTCTTTTTCAGTAAACGGTGCAGAGAGAACATCGTTCTC CGCTGCCTGTAATTGAGGTATATCATCAATCACAGACTCATCAAGAGACACCGTCGAAGT 30 ATTCGGAGGTCCAAAAAGGCCTTTATAATAGTTGGAAATATATGCTTTCAGATTCTCATG CCCCACTATTGTTCCCTCGTCCTGTTCAAGCTGTATGATCTTCTTCTTCCTATGTTTACC ATTTGCAACCATATGGAAAAACTGTGTGTTGTCGTCCCCTTGGACAACCTTAAGCGTTTT CGCACGCAACGCCCACTTGAGTTCCTCCTCTCTCAAGAGAGCATGTTGGCCTTGCTCAGC CTCAGACTTGGTTCGATGCTCATTAACAGAAAGAAGTGTGGTTTCAGCCTTTACATCTAG 35 TGTCTCAATGAGTTGAGTTAGACGTTCTTTTTTCTGCTTATAAATCCCAGTCTCATTCCT GGCCCATCCTCTCAGAAATTGTCGGAGGTTTCTAATCTTATTCTGCCATCTCTCGACATG TGTCCTTCCTGTAATTGGCTTAGCCCATTCGCATGCAATCATCTCCATAAATCCTTCTCG CTCAAACCAGCTTTACTCGAAAGAGAAGATGTTTTTGTTTGCAACATGGGTAGCCTCACC CGAATCTAAAAAGAGTGGTGTATGATCTGAGATCCCTCTATGCATTGCATGGACCGACAC 40 CAACGGATATTTTTGTTCCCACTCCACACTAGCAAGTACCCTATCCAGCTTTTCATAAGT CAGAACAGGTAACGAGTTGGCCCATGTAAACTGTCTACCGGTGAGCTCAATTTCTCTCAA ATTGAGGCTCTCGATAATCATGTTAAACATCATAGACCAACGTCCATCGAAATTGTCATT M&C PC933520WOA 128 ATTCTTTTCTTCTCTTCTCCGAATGATATTAAAATCACCCCCGACTAGCAGTGGCAGATT TTCATCTCCACAAATCCGCACTAGATGGGCAAGAAAATCGGGTTTAAATTGCTTGGAGGA GTGAGAGCATCTACAACCGGACTTAGCGAATCTGGGCT...
Claims
M&C PC933520WOA 176 CLAIMS:
1. A genetically altered plant, part thereof or plant cell characterised by reduced orabolished activity or expression of the GRAIN SIZE ON CHROMOSOME 3 5(GSE3) protein in said plant, part thereof or plant cell, wherein GSE3 comprisesa sequence as defined in SEQ ID NO: 1 or a functional variant or homolog thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1.10 2. The genetically altered plant, part thereof or plant cell of claim 1, wherein thehomolog is selected from a sequence comprising SEQ ID NO: 16, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41 and 43 or a functional variant thereof, whereinthe functional variant has at least 50% overall sequence identity to SEQ ID NO: 16, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41 or 43. 15 3. The genetically altered plant, part thereof or plant cell of claim 1 or 2, wherein theplant comprises at least one mutation in at least one GSE3 gene, wherein the mutation is preferably a loss or partial loss of function mutation, and wherein the GSE3 gene comprises a sequence as defined in SEQ ID NO: 3 to 4 or a functional20 variant or homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 3 to 4, and wherein the loss or partial loss of function mutation reduces or abolishes the activity or expression of GSE3.
4. The genetically altered plant, part thereof or plant cell of claim 1 or 2, wherein the25 plant, part thereof or plant cell comprises at least one RNAi that reduces orabolishes the expression of GSE3.
5. A genetically altered plant, part thereof or plant cell characterised by reduced orabolished activity or expression of a GRF4 (GROWTH REGULATING 4) protein30 and reduced or abolished expression of a GRF3 (GROWTH REGULATING 4) protein in said plant, part thereof or plant cell, wherein the GRF4 proteincomprises a sequence as defined in SEQ ID NO: 10 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a 35 sequence as defined in SEQ ID NO: 14 or a functional variant or fragment thereof,M&C PC933520WOA 177 and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO:
14.
6. The genetically altered plant, part thereof or plant cell of claim 5, wherein the5 homolog is selected from a sequence comprising SEQ ID NO: 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 87, 88, 89, 91, 93,95, 97, 99 and 101 or a functional variant thereof, wherein the functional varianthas at least 50% overall sequence identity to SEQ ID NO: 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70, 72, 74, 76, 78, 80, 82, 84, 86, 87, 88, 89, 91, 93, 95,10 97, 99 or 101.
7. The genetically altered plant, part thereof or plant cell of claim 5 or 6, wherein theplant comprises at least one mutation in at least one gene encoding GRF4 andat least one mutation in at least one gene encoding GRF3, wherein the mutation15 in GRF4 and the mutation in GRF3 is a loss or partial loss of function mutation,and wherein the gene encoding GRF4 comprises a sequence as defined in SEQID NO: 6, 7, 8 or 11 or a functional variant or homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 6, 7, 8or 11, and wherein the gene encoding GRF3 comprises a sequence as defined20 in SEQ ID NO: 12 or 13 or a functional variant or homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 12 or 13, and wherein the loss or partial loss of function mutations reduce orabolishes the activity or expression of GRF4 and GRF3.25 8. The genetically altered plant, part thereof or plant cell of claim 5 or 6, wherein theplant, part thereof or plant cell comprises at least one RNAi that reduces or abolishes the expression of a gene encoding GRF4 and at least one RNAi thatreduces or abolishes the expression of a gene encoding GRF3.30 9. The genetically altered plant, part thereof or plant cell of any of claims 1 to 8,wherein the plant is a male-sterile plant, preferably a cytoplasmic male-sterileplant or an environmental genic sterility male (EGMS) plant, preferably whereinthe EGMS plant is selected from a thermo-sensitive genic male-sterility (TGMS) plant or photoperiod-sensitive genic male-sterility (PGMS) plant. 35M&C PC933520WOA 178 10. The genetically altered plant, part thereof or plant cell of claim 1 to 8, wherein theplant is a maintainer plant line.
11. The genetically altered plant, part thereof or plant cell of any of claims 1 to 10,5 wherein the plant is a crop plant, preferably wherein the plant is selected from soybean, rice, wheat, maize, sorghum, soybean, rapeseed, cotton, sunflower, millet, barley, sugar-beet, rye, oats, ryegrass, beans, Brassica, sorghum and pea,and / or preferably wherein the plant part is selected from a grain or seed or pollen. 10 12. The genetically altered plant, part thereof or plant cell of claim 11, wherein theplant is rice and the plant is selected from the following rice lines: Xiaoligeng(XLG) or Y58S.15 13. The genetically altered plant, part thereof or plant cell of claim 11, wherein theplant is rice and the plant is selected from the following rice lines: Tianfeng B (TFB).
14. A grain or seed, wherein said grain or seed is characterised by reduced or20 abolished activity or expression of a GRAIN SIZE ON CHROMOSOME 3 (GSE3)protein in said grain or seed, wherein GSE3 comprises a sequence as defined inSEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1.25 15. A grain or seed, wherein said grain or seed is characterised by reduced orabolished activity or expression of GRF4 and GRF3 protein in said grain or seed,wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 30 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO:
14.
16. A method of hybrid breeding or mechanized hybrid seed production or increasing35 hybrid seed numbers, the method comprisingM&C PC933520WOA 179 a. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof orplant cell, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional 5 variant has at least 50% overall sequence identity to SEQ ID NO: 1; or b. reducing or abolishing the activity or expression of GRF4 and GRF3protein in a male-sterile plant, part thereof or plant cell, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functional10 variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functionalvariant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO:
14. 15 17. A method of increasing seed number in a plant, the method comprisinga. reducing or abolishing the activity or expression of a GRAIN SIZE ONCHROMOSOME 3 (GSE3) protein in a male-sterile plant, part thereof or20 plant cell, wherein GSE3 comprises a sequence as defined in SEQ ID NO: 1 or a functional variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 1; or b. reducing or abolishing the activity or expression of GRF4 and GRF325 protein in a male-sterile plant, part thereof or plant cell, wherein the GRF4 protein comprises a sequence as defined in SEQ ID NO: 10 or a functionalvariant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 10, and wherein the GRF3 protein comprises a sequence as defined in SEQ ID NO: 14 or a functional30 variant or fragment thereof, and wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 14; and optionally hybridising the male-sterile plant with a fertility restorer plant to obtain hybrid seed.M&C PC933520WOA 180 18. The method of claim 16 or 17, wherein the method further comprises increasingthe expression or activity of a protein encoded by the GS2 gene in a fertility restorer plant, plant part thereof or plant cell, wherein the GS2 gene comprises a sequence as defined in SEQ ID NO: 6, 7, 8 or 11 or a functional variant or5 homolog thereof, wherein the functional variant has at least 50% overall sequence identity to SEQ ID NO: 6, 7, 8 or 11, and optionally hybridising themale-sterile plant of claim 21 with the fertility restorer plant to obtain an F1 hybridplant or seed.10 19. A method of hybrid breeding, the method comprising hybridising the geneticallyaltered plants of any of claims 1 to 13 with a genetically altered fertility restorerplant, a part thereof or plant cell, wherein the fertility restorer plant (restorer line)is characterised by increased activity or expression of the protein encoded by the GS2 (GRAIN SIZE ON CHROMOSOME 2) gene, and obtaining F1 hybrid plants. 15 20. The method of any of claims 16 to 19, wherein the method comprises harvestingseeds obtained or obtainable from the male-sterile plant simultaneously with seeds obtained or obtainable from the fertility restorer plant, wherein preferably harvesting is mechanized harvesting. 20 21. An F1 hybrid plant or hybrid seed obtained or obtainable by the methods of anyof claims 16 to 20.
22. A method of hybrid breeding, the method comprising identifying and selecting25 plant that has a small-grain phenotype, the method comprising detecting in the plant or plant germplasm at least one polymorphism or mutation in the GRAIN SIZE ON CHROMOSOME 3 (GSE3) gene and / or promoter and selecting saidplant or progeny thereof, wherein the polymorphism or mutation reduces the expression and / or activity of GRAIN SIZE ON CHROMOSOME 3 (GSE3).30 35
Citation Information
Patent Citations
Application of CsHLS1 gene or protein coded by CsHLS1 gene in regulating and controlling organ size of cucumber plant
CN114990139A
Yield increase in plants overexpressing the HSRP genes
WO2007011681A2
Isolated novel nucleic acid and protein molecules from corn and methods of using those molecules to generate transgenic plant with enhanced agronomic traits
WO2009091518A2
GB194041200A