multiple spore-forming gene

By providing the nucleic acid and amino acid sequences of the Dip gene, apomixis was induced in sexual crops, solving the problem of apomixis gene identification and producing cloned seeds that are genetically identical to the parent plant, thus improving the efficiency and stability of cloned seed production.

CN108291234BActive Publication Date: 2026-04-21MASTER GENE LTD
View PDF 61 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MASTER GENE LTD
Filing Date
2016-09-05
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies make it difficult to effectively identify and utilize apomixis genes in sexual crops, which limits the application of apomixis in agriculture, especially in cloning seed production, fixing favorable genotypes, and solving problems such as loss of hybrid vigor and insufficient pollination.

Method used

Provides the nucleic acid and amino acid sequences of the Dip gene, as well as its homologs, fragments, and variants, and enables the preparation of apomixis seeds by inducing gametophyte apomixis through ploid sporophyte formation.

Benefits of technology

This technology enables the induction of apomixis in sexual crops, producing cloned seeds that are genetically identical to the parent plant. It solves the problem of identifying apomixis genes and improves the efficiency and stability of cloned seed production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0001650861580000391
    Figure BDA0001650861580000391
  • Figure BDA0001650861580000451
    Figure BDA0001650861580000451
  • Figure BDA0001650861580000461
    Figure BDA0001650861580000461
Patent Text Reader

Abstract

This invention provides the nucleotide and amino acid sequences of the Dip gene, as well as its (functional) homologs, fragments, and variants, which provide ploid sporophyte formation as part of apomixis. It also provides ploid sporophyte-forming plants, methods for their preparation, methods for using them, and methods for preparing apomixis seeds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and more specifically to plant biotechnology, including asexual plant breeding. Specifically, this invention relates to the identification of genes, variants, or fragments thereof, and the proteins and peptides they encode, associated with the process of apomixis (particularly gametophyte apomixis via diplospory formation). The invention also relates to methods for inducing gametophyte apomixis in plants and crops via diplospory formation using the genes, proteins, variants, and fragments of this invention, and methods for producing diplospory-forming plants and apomixis seeds. Background Technology

[0002] In botany, apomixis (also known as incomplete apomixis) refers to the formation of seeds through an asexual process. Apomixis occurs through a series of developmental processes that collectively transform the plant's sexual developmental program into an asexual one. Recurrent apomixis has been reported in over 400 species of flowering plants (Bicknell and Koltunow, 2004). Apomixis can occur in different forms, including at least two forms known as gametophytic apomixis and sporophytic apomixis (also known as adventive embryony). Examples of plants where gametophytic apomixis occurs include dandelion (Taraxacum sp.), mountain daisy (Hieracium sp.), Kentucky bluegrass (Poa pratensis), and orchardgrass (Tripsacum dactyloides). Examples of plants where sporophytic apomixis occurs include citrus (Citrus sp.) and mangosteen (Garcinia mangostana).

[0003] Interest in apomixis in general (especially in gametophyte apomixis) has increased in recent decades due to its potential uses in agriculture, particularly for the production of clonal seeds. Gametophyte apomixis is characterized by at least two developmental processes: (1) avoidance of meiosis (incomplete meiosis); and (2) development of the egg cell into an embryo without fertilization (parthenogenesis). Seeds produced by gametophyte apomixis are called apomixis seeds.

[0004] Because apomixis seeds are genetically identical to the parent plant, they are considered clones of the parent plant, and the process of producing such seeds is called clonal seed production. Apomixis has long been recognized as extremely useful in plant breeding (Asker 1979, Hermsen, JG Th. 1980. Breeding for apomixis in potato: Pursuing a utopian scheme. Euphaitica 29: 595-607, Asker and Jerling 1990, DeVielle Calzada et al., 1995). The advantage of apomixis is the ability to achieve true breeding (i.e., the unlimited proliferation of F1 hybrids with uniform genetic qualities) of hybrid vigor. In most crops, F1 hybrids are the best varieties because they tend to be associated with higher yields; this phenomenon is commonly referred to as "hybrid vigor." Since self-pollination of F1 hybrids leads to the loss of hybrid vigor in F2 sexual crops through recombination, each generation of F1 hybrids must be produced again through inbreeding with homozygous parents. Producing sexually occurring F1 seeds is a complex and expensive process that requires constant repetition. In contrast, apomixis F1 hybrids are true genetic organisms, meaning they are capable of pure reproduction.

[0005] In agriculture, apomixis are of great interest because of their ability to fix favorable genotypes (regardless of their genetic complexity) and to allow for the production of organisms capable of homogenization in a single step. This means that apomixis can be used to immediately fix multigene quantitative traits of interest. It should be noted that most yield traits are polygenetic. Asexual reproduction can be used for the accumulation (or stacking) of multiple traits (e.g., various resistances, several transgenes, or multiple quantitative trait loci). Without apomixis, to fix such a set of traits, each trait locus must be made homozygous individually and then combined into a hybrid. As the number of loci involved in a trait increases, generating homozygous trait loci through hybridization becomes laborious, time-consuming, and logically challenging. Similarly, selecting suitable parental lines for F1 hybrids requires a significant investment of time and effort. Furthermore, specific epistatic interactions between alleles are lost in the homozygous (parental line) stage and may not be recovered when combined in F1 hybrids. Apomixis makes it possible to fix this non-additive genetic variation.

[0006] Besides the immediate fixation of any genotype (regardless of its complexity), asexual reproduction has other important agricultural applications. Sexual interspecific hybrids and autoploids often suffer from sterility due to meiosis problems. Since apomixis skip meiosis, these problems in interspecific hybrids and autoploids are resolved. Because apomixis prevent female hybridization, it has been proposed to combine apomixis with male sterility for the control of transgenics, thereby preventing transgenic introgression hybridization in wild relatives of transgenic crops (Daniell, H. 2002. Molecular strategies for gene containment in transgenic crops. Nature biotechnology 20: 581-586). In insect-pollinated crops (e.g., Brassica), apomixis seed formation is not limited by insufficient pollinator services. This is becoming increasingly important given the growing health problems of pollinating bee populations (Varicella infestation, African killer bees, etc.). Since most viruses are not transmitted through seeds, apomixis can be used to maintain superior genotypes through cloning in tuber-propagated crops (such as potatoes), eliminating the risk of virus transmission through tubers. The storage cost of apomixis seeds is also lower than that of tubers. In ornamental plants, apomixis can replace labor-intensive and expensive tissue culture propagation. It is well known that apomixis typically significantly reduces the cost of variety development and propagation.

[0007] Major crops do not undergo apomixis, and most are sexual seed crops. Numerous attempts have been made to introduce apomixis into sexual crops. Specifically, because apomixis is genetically regulated, many researchers have attempted to identify the genes involved in the process. In naturally apomixis plants, apomixis has been identified as a source of apomixis genes (Ozias-Akins, P. and PJvan Dijk. 2007 in: Annu. Rev. Genet. 41: 509-537). However, little is known about the genetic and molecular background of apomixis, and attempts to identify apomixis genes have so far failed to yield agriculturally applicable genes. This is primarily because identifying and isolating apomixis genes has proven to be a difficult task. Naturally apomixis plants are often polyploid, and localized cloning is challenging in polyploid plants. Other complicating factors include recombination in apomixis-specific chromosomal regions, repetitive sequences, and segregational aberrations during hybridization. Furthermore, the genomes of apomixis plants have not yet been sequenced, complicating the search for the entire apomixis gene pool. Therefore, apomixis genes have not yet been cloned and / or isolated. Attempts to introduce apomixis into sexual crops can be summarized as follows:

[0008] a) The introgression of apomixis genes from wild apomictic plants into crop species through extensive hybridization has not yet been successful. For example, attempts to transfer apomixis from *Tripsacum dactyloides* to maize and millet [Savidan, Y. (2001). Transfer of apomixis through wide crosses. In Flowering of Apomixis: From Mechanisms to Genetic Engineering, Y.; apomixis from *Pennisetum squamulatum* into pearl millet. Savidan, J.G. Carman, and T. Dresselhaus, eds (Mexico: CIMMYT, IRD, European Commission DG VI), pp. 153–167; Morgan, R., Ozias-Akins, P., and Hanna, WW (1998). Seed set in an apomictic BC3 pearl millet. Int. J. Plant Sci. 159,89–97.; WO97 / 10704.].

[0009] b) Mutants in sexually type species, particularly in Arabidopsis thaliana. For example, WO2007066214 describes the use of an incomplete meiotic mutant called Dyad in Arabidopsis thaliana. However, Dyad is a recessive mutation with very low penetrance. Its practical use in crop species is therefore very limited.

[0010] c) De novo generation of apomixis through hybridization between two sexual ecotypes did not result in apomixis of agricultural interest (US20040168216 and US20050155111).

[0011] d) Cloning candidate apomixis genes via transposon tags in maize. US20040148667 discloses homologs of the elongate gene, which are presumed to induce apomixis. However, according to Barrell and Grossniklaus (2005) in Plant Journal Vol:34, pp 309-320, the elongate gene skips the second meiotic division and therefore cannot maintain the maternal genotype.

[0012] Furthermore, US20060179498 describes so-called "reverse breeding" as an alternative to apomixis. However, compared to apomixis, which does not require any laboratory methods (because it is an in vivo method performed by the plant itself without any external (human) intervention), reverse breeding represents a complex and laborious in vitro laboratory method. Moreover, in reverse breeding, hybridization must be performed once the parental lines are reconstructed (homozygous diploid gametes).

[0013] Therefore, there is a need for alternative methods for inducing apomixis in sexual crops, which at least do not have some of the limitations of existing technologies. In particular, there is a need for methods for producing ploid sporophyte-forming plants and apomixis seeds. There is also a need to identify alternative genes and proteins involved in apomixis processes (especially ploid sporophyte formation) that are suitable for the aforementioned methods and can substantially mimic the apomixis pathway in sexual crops. Invention Overview

[0015] This invention provides the nucleic acid and amino acid sequences of the Dip gene, as well as its (functional) homologs, fragments, and variants, which provide ploid sporophyte formation as part of apomixis. It also provides ploid sporophyte-forming plants, methods for their preparation, methods for using them, and methods for preparing apomixis seeds. Invention Details

[0017] definition

[0018] As used herein, the term "sexual plant reproduction" refers to the developmental pathway in which a diploid somatic cell, called a "megaspore mother cell," undergoes meiosis to produce four reduced megaspores. One of these megaspores undergoes mitosis to form a female gametophyte (also called an embryo sac) containing a meiotic egg cell (i.e., a cell with a reduced number of chromosomes compared to the mother cell) and two meiotic polar nuclei. A pollen grain sperm cell fertilizes the egg cell to produce a diploid embryo, while a second sperm cell fertilizes the two polar nuclei to produce a triploid endosperm (a process known as double fertilization).

[0019] As used herein, the term "megasporocyte mother cell" refers to a diploid cell that produces megaspores through meiosis (usually meiosis) to generate four haploid megaspores that will develop into female gametophytes. In angiosperms (also known as flowering plants), megasporocyte mother cells produce megaspores that develop into female gametophytes through two distinct processes: megasporogenesis (formation of megaspores in the nucellus or megasporangium) and gynogeny (development of megaspores into female gametophytes).

[0020] As used herein, the term "asexual reproduction" refers to the process by which plants reproduce without fertilization and without gamete fusion. Asexual reproduction produces new individuals that, unless a mutation occurs, are genetically identical to and similar to the parent plant. There are two main types of asexual reproduction in plants: vegetative propagation (i.e., the vegetative sheet of the initial plant involving budding, tillering, etc.) and apomixis.

[0021] As used in this article, the term "apomixis" refers to the formation of seeds through an asexual process.

[0022] As used herein, the term "plophysophyte formation" refers to the derivation of unmeiotic embryo sacs directly from megasporocytes via mitosis or via an aborted meiotic event. Three main types of plophysophyte formation have been reported, named after the plants in which they occur: the dandelion, *Ixeris*, and *Antennaria* types. In the dandelion type, promeiosis begins but is subsequently aborted, resulting in two unmeiotic didiglia, one of which produces an embryo sac via mitosis. In the *Ixeris* type, after promeiosis, the nucleus undergoes two further mitotic divisions at equal intervals to produce an octenocyte embryo sac. The dandelion and *Ixeris* types are referred to as meiotic plophysophyte formation because they involve modifications of meiosis. Conversely, in the *Antennaria* type, called mitotic plophysophyte formation, the megasporocyte does not initiate meiosis and divides directly three times to produce an unmeiotic embryo sac. In gametophyte apomixis via ploid sporophyte formation, unmeiotic megaspores produce unmeiotic gametophytes. These unmeiotic megaspores arise from either mitotic-like division (mitotic ploid sporophyte formation) or modified meiosis (meiotic ploid sporophyte formation). In both sporophytic and ploid sporophytic gametophyte apomixis, unmeiotic ovocytes develop parthenogenetically into embryos. The apomixis of dandelion are diploid, meaning that the first female meiosis (meiosis I) is skipped, resulting in two unmeiotic megaspores with the same genotype as the parent plant. One of these megaspores degenerates, and the other surviving unmeiotic megaspore produces an unmeiotic female gametophyte (or embryo sac) containing an unmeiotic ovocyte. This unmeiotic ovocyte develops into an embryo with the same genotype as the parent plant without fertilization. Seeds produced by apomixis in gametophytes are called apomixis seeds.

[0023] The term "ploid sporophyte formation function" refers to the ability of a plant, preferably in the female ovary, and more particularly in the megasporocyte and / or female gamete, to induce ploid sporophyte formation. Therefore, a plant incorporating ploid sporophyte formation function is capable of undergoing the ploid sporophyte formation process, i.e., resuming the production of unmeiotic gametes through the first meiotic division.

[0024] The term "ploid sporophyte formation as part of gametophyte apomixis" refers to the ploid sporophyte formation portion of the apomixis process, that is, the role of ploid sporophyte formation in seed formation during asexual reproduction. In particular, in addition to the function of ploid sporophyte formation, apomixis function is also required to establish the apomixis process. Therefore, the combination of ploid sporophyte formation and parthenogenesis may lead to apomixis.

[0025] Apomixis is known to occur in different forms, including at least two forms known as gametophytic apomixis and sporophytic apomixis (also called adventitious embryo form). Examples of plants that undergo gametophytic apomixis include dandelion (Taraxacum sp.), mountain daisy (Hieracium sp.), Kentucky bluegrass (Poa pratensis), and orchardgrass (Tripsacum dactyloides). Examples of plants that undergo sporophytic apomixis include citrus (Citrus sp.) and mangosteen (Garcinia mangostana).

[0026] As used herein, the term "plosporophytic plant" refers to a plant that reproduces apomixis via gametophyte formation, or a plant that has been induced (e.g., through genetic modification) to reproduce apomixis via gametophyte formation. In both cases, plosporophytic plants produce apomixis seeds when combined with apomixis factors.

[0027] As used herein, the term "apomixis seed" refers to a seed obtained from an apomixis plant species or from a plant or crop that has undergone apomixis (particularly gametophyte apomixis formed from ploid sporophytes). Asexually reproduced seeds are characterized by being clones and genetically identical to their parent plants, and by their ability to germinate into plants with true genetic capacity.

[0028] A "clone" of a cell, plant, part of a plant, or seed is characterized by its genetic similarity to its siblings and the parent plant from which it originated. Individual clones have nearly identical genomic DNA sequences; however, mutations can lead to subtle differences.

[0029] As used herein, the term "true inheritance" or "true-inherited organism" (also known as a purebred organism) refers to an organism that always passes on a certain phenotypic trait to its offspring with no change or almost no change. An organism is said to be truly inherited for each trait (where each trait is truly inherited), and the term "true inheritance" is also used to describe a single inherited trait.

[0030] As used herein, the term "F1 hybrid" (or F1 hybrid) refers to the first generation of offspring that is significantly different from the parental type. The F1 hybrid is used in genetics and may appear as an F1 cross in selective breeding. Offspring with significantly different parental types produce new, uniform phenotypes with combinations of parental characteristics. "F1 hybrid" is associated with unique advantages, such as heterosis, and is therefore highly desirable in agricultural practice. In one embodiment of the invention, the methods, genes, proteins, variants, or fragments taught herein can be used to fix the genotype of F1 hybrids, regardless of their genetic complexity, and allow for the production of purebred organisms in a single step.

[0031] As used herein, the term "allele" refers to any one or more variant forms of a gene at a specific locus. In the diploid cells of an organism, the alleles of a given gene are located at a specific location or locus on a chromosome. One allele exists on each chromosome of a pair of homologous chromosomes. Diploid or polyploid plant species may contain a large number of different alleles at a specific location. In one embodiment, the Dip locus of the wild dandelion species (accession) as taught herein may contain various Dip or dip alleles, which may differ slightly in the nucleic acid and / or the encoded amino acid sequence.

[0032] As used herein, the term “locus” refers to one or more specific locations or sites on a chromosome where, for example, one or more genes or genetic markers are located. For example, a “Dip locus” as taught herein refers to a location in the genome where the Dip gene (and two (or more) Dip alleles) as taught herein are found.

[0033] As used herein, the term "dominant allele" refers to the relationship between alleles of a gene, where the effect of the phenotype of one allele (i.e., the dominant allele) masks the contribution of the second allele (the recessive allele) at the same locus. The first allele is dominant, and the second allele is recessive. For genes on autosomes (any chromosome other than sex chromosomes), alleles and their associated traits are autosomal dominant or autosomal recessive. Dominance is a key concept in Mendelian and classical genetics. For example, a dominant allele may encode a functional protein, while a recessive allele may not. In one implementation, the gene and its segments or variants, as taught herein, refer to the dominant allele of the Dip gene.

[0034] As used herein, the term "female ovary" refers to the outer shell within which spores are formed. It can be composed of a single cell or can be multicellular. All plants, fungi, and many other lineages form an ovary at some point in their life cycle. The ovary produces spores through mitosis or meiosis. Typically, within each ovary, the megasporocyte undergoes meiosis to produce four haploid megaspores. In gymnosperms and angiosperms, only one of these four megaspores is functional at maturity, and the other three degenerate. The remaining megaspore undergoes mitosis and develops into a female gametophyte, which eventually produces an egg cell.

[0035] As used herein, the term "female gamete" refers to the cell that fuses with another type of cell ("male") during fertilization (conception) in a sexually reproducing organism. In species that produce two morphologically different types of gametes, and in which each individual produces only one type, a female is any individual that produces the larger type of gamete (called an ovule (egg) or ovum). In plants, female ovules are produced by the ovary of a flower. When mature, haploid ovules produce female gametes ready for fertilization. Male cells are (primarily haploid) pollen and are produced by the anthers.

[0036] As used herein, the term "pollination" refers to the process by which pollen is transferred from the anther (male part) to the stigma (female part) of the plant, thereby enabling fertilization and reproduction. It is unique to angiosperms, which are flowering plants. Each pollen grain is a male haploid gametophyte adapted to be transported to the female gametophyte, and in the process of double fertilization, it influences fertilization by producing male gametes (or gametophytes). A successful angiosperm pollen grain (gametophyte) containing a male gamete is transported to the stigma, where it germinates and its pollen tube descends along the style to the ovary. Its two gametes descend along the tube to the gametophyte containing the female gamete, which is held within the carpel. One nucleus fuses with the polar nuclei to produce endosperm tissue, and the other nucleus fuses with the ovule to produce the embryo. Even most natural apomixis require pollination for the sexual development of the endosperm. However, in a small number of apomixis, such as in dandelions and sagebrush, endosperm development does not occur through a polar nucleus fertilization process known as spontaneous endosperm development. In Arabidopsis thaliana, some mutations are known to cause spontaneous endosperm development.

[0037] As used herein, the term "apomixis" refers to a form of incomplete meiosis in which embryonic growth and development do not occur through fertilization. The genes and proteins of this invention can be combined with apomixis factors (e.g., genes or chemical factors) to produce apomixis offspring.

[0038] As used herein, the term “vacuole protein sorting associated protein type 13” (abbreviated as VPS13) refers to the protein encoded by the Vps13 gene, which participates in the control of protein cycling steps through the trans-Golgi network to the vacuole and cell membrane.

[0039] As used herein, the term “genetic marker” or “polymorphic marker” refers to a region on genomic DNA that can be used to “mark” a specific location on a chromosome. If a genetic marker is closely associated with a gene, or if the genetic marker is “located” within a gene (in gene markers), it “marks” the DNA in which the gene is found, and can therefore be used for (molecular) marker analysis, as taught herein, to select or not select the presence of the gene, for example in marker-assisted breeding / selection (MAS) methods. Non-restrictive examples of genetic markers are AFLP (Amplified Fragment Length Polymorphism, EP534858), microsatellites, RFLP (Restriction Fragment Length Polymorphism), STS (Sequence Marker Site), SNP (Single Nucleotide Polymorphism), SFP (Single Feature Polymorphism; see Borevitz et al. (2003) In: Genome Research Vol: 13, pp 513-523), SCAR (Sequence Characterization Amplified Region), CAPS marker (Enzyme Digestion Amplified Polymorphic Sequence), etc. The farther a marker is from a gene, the greater the likelihood of recombination (exchange) between the marker and the gene, thus losing the association (and co-segregation of the marker and gene). The distance between genetic loci is measured by recombination frequency and given in cM (centiMorgans; 1 cM is the meiotic recombination frequency between 1% of two markers). Due to the significant differences in genome size between species, the actual physical distance represented by 1 cM (i.e., the kilobases, kb between two markers) varies considerably between species. It should be understood that when referring to “connecting” markers in this article, this also includes markers “located” within the gene itself.

[0040] As used herein, the term “marker-assisted selection” (abbreviated as “MAS”) refers to the process of screening plants to screen for the presence and / or absence of one or more genetic and / or phenotypic markers in order to accelerate the transfer of DNA regions containing the markers (and optionally lacking flanking regions) into (elite) breeding lines.

[0041] As used herein, the term "molecular marker assay" (or test) refers (directly or indirectly) to a DNA-based assay that determines the presence or absence of a specific allele (e.g., the Dip allele) in a plant or plant part. Preferably, it allows for the determination of whether a specific allele at the Dip locus is homozygous or heterozygous in any single plant. For example, in one embodiment, nucleic acids linked to the Dip locus are amplified using PCR primers, the amplification products are enzymatically digested, and based on the electrophoretic analysis pattern of the amplification products, it can be determined which Dip allele is present in any single plant, and the conjugation of the allele at the Dip site (i.e., the genotype of each locus). Non-limiting examples of molecular marker assays include sequence-characterized amplified region (SCAR) marker assays, cleaved amplified polymorphic sequence (CAPS) marker assays, etc.

[0042] As used herein, the term “heterozygous” refers to a genetic condition in which two (or more in the case of polyploids) distinct alleles are present at a particular locus such as the Dip locus (e.g., dominant Dip allele / recessive dip allele), but are located separately on the corresponding homologous chromosome pair in the cell.

[0043] As used herein, the term “homozygous” refers to a genetic condition in which two (or more, in the case of polyploidy) identical alleles are located at a specific locus (e.g., the Dip locus, for example, for a dominant Dip allele / recessive dip allele), but are located on corresponding pairs of homologous chromosomes in the cell.

[0044] As used herein, the term “variety” refers to a group of plants that conforms to UPOV conventions and is considered as a unit within a single botanical classification at the lowest known level, defined by the expression of a trait produced by a given genotype or combination of genotypes, which can be distinguished from any other group of plants by the expression of at least one of the said trait, and is considered as a unit in view of its fitness to remain unchanged (stable) during reproduction.

[0045] As used herein, the terms “polypeptide” and “protein” are used interchangeably and refer to molecules composed of chains of amino acids without regard to a specific mode of action, size, three-dimensional structure, or origin.

[0046] As used herein, the terms “isolated polypeptide” or “isolated protein” are used interchangeably and refer to proteins that are no longer present in their natural environment, such as proteins present in test tubes (in vitro) or in recombinant bacterial or plant host cells.

[0047] As used herein, the term "nucleic acid" refers to any polymer or oligomer of pyrimidine and purine bases, preferably cytosine, thymine, and uracil, and adenine and guanine, respectively (see Albert L. Lehninger, Principles of Biochemistry, at 793-800 (Worth Pub. 1982), which is incorporated herein by reference in its entirety for all purposes). This invention contemplates any deoxyribonucleic acid, ribonucleic acid, or peptide nucleic acid component and any chemical variant thereof, such as methylated, hydroxymethylated, or glycosylated forms of these bases. Polymers or oligomers may be heterogeneous or homogeneous in composition and may be isolated from naturally occurring sources or produced artificially or synthetically.

[0048] According to the present invention, the terms "polynucleotide", "nucleic acid molecule", "nucleotide sequence" or "nucleic acid sequence" refer to polymeric DNA or RNA molecules in single-stranded or double-stranded form, particularly DNA encoding proteins, its variants or fragments.

[0049] As used herein, the terms “isolated polynucleotide,” “isolated nucleic acid molecule,” “isolated nucleotide sequence,” or “isolated nucleic acid sequence” refer to a polynucleotide that is no longer in its natural environment, i.e., substantially isolated from other cellular components that naturally accompany its native sequences or proteins (e.g., ribosomes, polymerases, and many other sequences and proteins). The term includes polynucleotides that have been isolated from their natural environment and includes recombinant or cloned nucleic acid isolates and chemically synthesized analogs or analogs biosynthesized through heterologous systems, such as nucleic acid sequences in bacterial host cells or plant nuclei or plastid genomes.

[0050] As used herein, the terms "functional DIP gene or protein" or "functional DIP gene or protein variant or fragment" (e.g., orthologs or mutants, and portions of a gene) refer to the ability of a gene and / or its encoded protein to modify or induce apomixis (particularly gametophyte apomixis through ploid sporophyte formation) in a plant by altering the expression levels of one or more genes in said plant (e.g., by overexpression or silencing), both quantitatively and / or qualitatively. For example, the functionality of a putative DIP protein obtained from plant species X can be tested by various methods. Preferably, if the protein is functional, silencing the Dip gene encoding the protein in plant species X using, for example, a gene silencing vector will cause a reduction (i.e., a reduction in chromosome number) or inhibition of ploid sporophyte formation, while overexpression in susceptible plants will result in enhanced ploid sporophyte formation. Similarly, supplementation with a functional DIP protein will be able to restore or confer ploid sporophyte formation. Those skilled in the art will have no difficulty in testing functionality.

[0051] As used herein, the term "gene" refers to a DNA sequence containing a region (transcriptional region) operatively linked to a suitable regulatory region (e.g., a promoter), which is transcribed into an RNA molecule (e.g., mRNA) in a cell. Thus, a gene may contain several operatively linked sequences, such as a promoter, a 5' leader sequence containing, for example, a translation initiation-related sequence, a (protein) coding region (cDNA or genomic DNA), and a 3' untranslated sequence containing, for example, a transcription termination site.

[0052] As used herein, the term "chimeric gene" or "recombinant gene" refers to any gene in nature that is not normally present in a species, particularly a portion of one or more nucleotide sequences of a nucleic acid sequence present in a gene, which exists in nature in an unrelated form. For example, a promoter in nature is unrelated to part or all of the transcriptional region or to another regulatory region. The term "chimeric gene" is understood to include an expression construct in which a promoter or transcriptional regulatory sequence is operatively linked to one or more coding sequences or antisense sequences (inverse complementation of sense strands) or inverted repeat sequences (sense and antisense, whereby the RNA transcript forms double-stranded RNA upon transcription).

[0053] The term "3'UTR" or "3' untranslated sequence" (also known as the "3' untranslated region" or "3' end") refers to a nucleic acid sequence found downstream of the coding sequence of a gene that contains, for example, a transcription termination site and (in most, but not all, eukaryotic mRNAs) a polyadenylation signal (e.g., AAUAAA or its variants). After transcription termination, the mRNA transcript can be cleaved downstream of the polyadenylation signal and a poly(A) tail can be added to participate in the transport of mRNA to the cytoplasm (the site of translation).

[0054] As used herein, the term "5'UTR," "leader sequence," or "5' untranslated region" refers to the region of mRNA transcription and the corresponding DNA, between the +1 position of the mRNA transcription start and the translation start codon of the coding region (usually AUG on mRNA or ATG on DNA). The 5'UTR typically contains sites important for translation, mRNA stability and / or degradation (turnover), and other regulatory elements.

[0055] As used herein, the term "expression of the gene or variant thereof" refers to the process of transcribing a DNA region operatively linked to a suitable regulatory region (particularly a promoter) into biologically active RNA (i.e., capable of being translated into a biologically active protein or peptide (or an active peptide fragment) or being active itself (e.g., in post-transcriptional gene silencing or RNAi)). In some embodiments, an active protein refers to a protein with sustained activity. The coding sequence is preferably in a sense orientation and encodes the desired, biologically active protein or peptide or an active peptide fragment. In gene silencing methods, the DNA sequence is preferably in the form of antisense DNA or inverted repeat DNA containing a short sequence of the target gene in an antisense orientation or both sense and antisense orientations. "Ectopic expression" refers to the expression of the gene in a tissue where it is not normally expressed.

[0056] As used herein, the term "transcriptional regulatory sequence" refers to a nucleic acid sequence capable of regulating the transcription rate of a (coding) sequence operatively linked to a transcriptional regulatory sequence. Therefore, the transcriptional regulatory sequence as defined herein will include all sequence elements (promoter elements) necessary to initiate transcription, maintain and regulate transcription, including, for example, attenuators or enhancers. While primarily referring to upstream (5') transcriptional regulatory sequences of the coding sequence, this definition also includes regulatory sequences found downstream (3') of the coding sequence.

[0057] As used herein, the term "promoter" refers to a nucleic acid segment that controls the transcription of one or more genes, located upstream of the transcription start site of that gene in the transcriptional direction, and structurally characterized by the presence of a DNA-dependent RNA polymerase binding site, a transcription start site, and any other DNA sequence (including, but not limited to, transcription factor binding sites, repressor and activator protein binding sites, and any other nucleic acid sequence known to those skilled in the art that directly or indirectly functions to regulate the amount of transcription from the promoter). Optionally, the term "promoter" may also include a 5' UTR region (e.g., a promoter may include one or more portions upstream (5') of the translation start codon of a gene, as this region may play a role in regulating transcription and / or translation).

[0058] As used in this article, the term “constitutive promoter” refers to a promoter that is active in most tissues under most physiological and developmental conditions.

[0059] As used herein, the term "inducible promoter" refers to a promoter that is physiologically (e.g., regulated by external application of certain compounds) or developmentally.

[0060] As used herein, the term "tissue-specific promoter" refers to a promoter that is active only in a specific type of tissue or cell. "Promoter active in plants or plant cells" refers to the promoter's general ability to drive transcription within plants or plant cells. It does not imply any spatiotemporal activity of the promoter.

[0061] As used herein, the term "operably linked" refers to the connection of polynucleotide elements in a functional relationship. Nucleic acids are "operably linked" when they are placed in a functional relationship with another nucleic acid sequence. For example, if a promoter or transcriptional regulatory sequence affects the transcription of a coding sequence, it is operably linked to the coding sequence. Operable linking means that the linked DNA sequences are typically contiguous and, where necessary, the coding regions of two proteins are linked contiguously and within the reading frame to produce a chimeric protein.

[0062] As used herein, the term "chimeric protein" or "hybrid protein" refers to a protein composed of various protein domains or motifs that are not found in nature but whose linkages form a functional protein that exhibits the functionality of the linked domains. Chimeric proteins can also be fusion proteins of two or more naturally occurring proteins.

[0063] As used herein, the term "domain" refers to any part or region of a protein that has a specific structure or function, which can be transferred to another protein to provide a new hybrid protein that has at least the functional characteristics of the domain.

[0064] As used herein, the term "target peptide" refers to an amino acid sequence that targets a protein or protein fragment to an intracellular organelle, such as a plastid, preferably a chloroplast, mitochondria, or the extracellular space or a non-plastoid (secretory signal peptide). The nucleic acid sequence encoding the target peptide may be fused with a nucleic acid sequence encoding the amino-terminal (N-terminus) of the protein or protein fragment (in the box), or may be used to replace the native target peptide.

[0065] As used herein, the terms "nucleic acid construct" or "vector" refer to an artificial nucleic acid molecule produced using recombinant DNA technology and used to deliver exogenous DNA or RNA into a host cell. The vector backbone can be, for example, a binary or superbinary vector (see, for example, US5591616, US2002138879, and WO9506722), such as co-integration vectors or T-DNA vectors known in the art and described elsewhere herein, in which a chimeric gene is integrated, or, if a suitable transcriptional regulatory sequence is already present, only the desired nucleic acid sequence (e.g., a coding sequence, antisense, or inverted repeat sequence) is integrated downstream of the transcriptional regulatory sequence. Vectors typically contain other genetic elements to facilitate their use in molecular cloning, such as selection markers, multiple cloning sites, etc. (see below).

[0066] As used herein, the terms "host cell," "recombinant host cell," "transformed cell," or "transgenic cell" refer to a new individual cell (or organism) resulting from at least one nucleic acid molecule, particularly containing a chimeric gene encoding a desired protein or nucleic acid sequence, which, after transcription, produces antisense RNA or inverted repeat RNA (or hairpin RNA), or siRNA or miRNA for silencing a target gene / gene family, which has been introduced into the cell. The host cell is preferably a plant cell or a bacterial cell. The host cell may contain a nucleic acid construct as an extrachromosomal (attachment) replication molecule, or more preferably, a chimeric gene integrated into the nuclear or plasmonic genome of the host cell.

[0067] As used herein, the terms “recombinant plant” or “recombinant plant part” or “transgenic plant” refer to a plant or plant part (e.g., seed, fruit, or leaf) that contains a chimeric gene as taught herein at the same locus in all cells and plant parts, although the gene may not be expressed in all cells.

[0068] As used herein, the term "elite event" refers to a recombinant plant in which a recombinant gene has been selected to contain a recombinant gene at a location in the genome that imparts favorable or desired phenotypic and / or agronomical characteristics. The flanking DNA of the integration site can be sequenced to characterize the integration site and distinguish it from other transgenic plants containing the same chimeric gene at other locations in the genome.

[0069] As used herein, the term "selectable marker" is a term well-known in the art and is used herein to describe any genetic entity that, when expressed, can be used to select cells containing a selectable marker. Selectable marker gene products confer, for example, antibiotic resistance or, more preferably, herbicide resistance, or another selectable trait such as a phenotypic trait (e.g., a change in pigmentation) or nutritional requirements. The term "reporter gene" is primarily used to refer to visible markers such as green fluorescent protein (GFP), eGFP, luciferase, GUS, etc.

[0070] As used herein, the terms "gene ortholog" or "protein ortholog" refer to a homologous gene or protein found in another species that has the same function as the gene or protein, but (usually) has undergone sequence differentiation since the time of gene differentiation in this species (i.e., a gene that evolved from the common ancestor of the species). In one embodiment, orthologs of the dandelion Dip gene can be identified in other plant species based on sequence comparisons (e.g., based on the percentage of sequence identity across the entire sequence or specific domains) and functional analysis.

[0071] The term "same linear region" as used in this article refers to a term used in comparative genomics and refers to the same region on the chromosomes of two related species.

[0072] As used herein, the term “strict hybridization conditions” refers to conditions that can be used to identify nucleic acid sequences substantially identical to a given nucleic acid sequence. Strict conditions are sequence-dependent and vary under different conditions. Typically, strict conditions are chosen to be approximately 5°C lower than the thermal dissolution temperature (Tm) of a specific sequence at defined ionic strengths and pH. Tm is the temperature at which 50% of the target sequence hybridizes with a perfectly matched probe (at defined ionic strengths and pH). Strict conditions are typically chosen where the salt concentration is approximately 0.02 mol at pH 7 and the temperature is at least 60°C. Decreasing the salt concentration and / or increasing the temperature increases strictness. Strict conditions for RNA-DNA hybridization (using, for example, a 100 nt probe for Northern blotting) are those that include, for example, at least one wash for 20 minutes in 0.2X SSC at 63°C, or equivalent conditions. The stringent conditions for DNA-DNA hybridization (using Southern blotting with, for example, a 100 nt probe) include, for example, washing at least once for 20 minutes (usually twice) in 0.2X SSC at a temperature of at least 50°C, typically about 55°C, or equivalent conditions. See also Sambrook et al. (1989) and Sambrook and Russell (2001).

[0073] As used herein, the term "highly stringent conditions" refers to conditions achievable, for example, by hybridization at 65°C in an aqueous solution containing 6×SSC (20×SSC containing 3.0M NaCl, 0.3M sodium citrate, pH 7.0), 5×Denhardt's (100X Denhardt's containing 2% Ficoll, 2% polyvinylpyrrolidone, 2% bovine serum albumin), 0.5% sodium dodecyl sulfate (SDS), and 20 μS alkylsulfonate denatured vector DNA (single-stranded fish sperm DNA, with an average length of 120–3000 nucleotides) as non-specific competitive agents. Following hybridization, highly stringent washing can be performed in several steps, culminating in a final wash (approximately 30 minutes) in 0.2–0.1×SSC, 0.1% SDS at the hybridization temperature.

[0074] As used herein, the term "moderate stringency" refers to conditions equivalent to hybridization in the solutions described above, but performed at approximately 60–62 °C. In this case, the final wash is performed at the hybridization temperature in 1×SSC, 0.1% SDS.

[0075] As used herein, the term “low stringency” refers to hybridization in solutions equivalent to those described above at approximately 50–52 °C. In this case, the final wash is performed at the hybridization temperature in 2 × SSC, 0.1% SDS. See also Sambrook et al. (1989) and Sambrook and Russell (2001).

[0076] When used in the context of amino acid or nucleic acid sequences, the terms “substantially identical”, “substantially similar”, “substantially similar”, “variant”, or “sequence identity” as used herein refer to two amino acid sequences or two nucleic acid sequences that share at least a certain percentage of sequence identity when optimally aligned, for example, using the GAP or BESTFIT programs with default parameters. GAP uses the Needleman and Wunsch global alignment algorithms to align two sequences across their full length, maximizing the number of matches and minimizing the number of gaps. Typically, using GAP default parameters, the gap generation penalty is 50 (nucleic acid) / 8 (protein) and the gap extension penalty is 3 (nucleic acid) / 2 (protein). For nucleic acids, the default scoring matrix used is nwsgapdna, and for proteins, the default scoring matrix is ​​Blosum62 (Henikoff & Henikoff, 1992, PNAS 89, 915-919). It is evident that when an RNA sequence is considered substantially similar to or has a certain degree of sequence identity with a DNA sequence, thymine (T) in the DNA sequence is considered equivalent to uracil (U) in the RNA sequence. Sequence alignment and percentage sequence identity scores can be determined using computer programs such as the GCG Wisconsin Package, Version 10.3, obtained from Accelrys Inc., 9685 Scranton Road, San Diego, CA 92121-3752 USA. Alternatively, the program “needle” in EmbossWIN (version 2.10.0) can be used with the same GAP parameters as described above, or with a gap opening penalty of 10.0 and a gap extension penalty of 0.5, using DNAFULL as the matrix. To compare sequence identity between sequences of different lengths, local alignment algorithms are preferred, such as the Smith-Waterman algorithm (Smith TF, Waterman MS (1981) J. Mol. Biol 147(1); 195-7), using, for example, the program “water” in EmbossWIN. The default parameters are a gap opening penalty of 10.0 and a gap extension penalty of 0.5. For proteins, Blosum62 is used and for nuclei, the DNAFULL matrix is ​​used.

[0077] The terms “comprising” and “including” as used herein, and their synonyms, refer to a situation where the terms are used in their non-restrictive sense to indicate the inclusion of the item following the word, but not the exclusion of items not specifically mentioned. It also includes the more restrictive verb “consisting of.” Furthermore, the indefinite article “a” or “an” in reference to an element does not preclude the possibility of more than one of that element, unless the context explicitly requires the presence of one and only one of that element. Therefore, the indefinite article “a” or “an” generally means “at least one.” Further understanding is that when “sequence” is referred to herein, it generally refers to an entity molecule having a specific subunit sequence (e.g., amino acids).

[0078] As used herein, the term "plant" includes plant cells, plant tissues or organs, plant protoplasts, plant cell tissue cultures of regenerable plants, plant callus, plant cell masses and intact plant cells in a plant, or parts of a plant such as embryo, pollen, ovules, sporangia, fruit, flower, leaf (e.g., harvested lettuce), seed, root, root tip, etc.

[0079] As used herein, the term “gene silencing” refers to the downregulation or complete suppression of gene expression of one or more target genes (e.g., endogenous Dip genes). The use of repressive RNA to reduce or eliminate gene expression is well-established in the field and is a topic of several reviews (e.g., Baulcombe 1996, Stam et al. 1997, Depicker and VanMontagu, 1997). Numerous techniques are available for achieving gene silencing in plants, such as chimeric genes that generate antisense RNA of all or part of the target gene (see, for example, EP 0140308B1, EP 0240208B1, and EP 0223399B1), or sense RNA (also known as co-suppression), see EP 0465572B1. However, the most successful method to date, however, is the generation of both sense and antisense RNA of the target gene (“inverted repeat sequences”), which forms double-stranded RNA (dsRNA) in the cell and silences the target gene. Methods and vectors for dsRNA production and gene silencing have been described in EP 1068311, EP 983370 A1, EP 1042462A1, EP 1071762 A1, and EP 1080208 A1. Therefore, the vectors of this invention can contain an active transcriptional regulatory region in plant cells, which is operatively linked to a sense and / or antisense DNA fragment of the DIP gene of this invention. Typically, short (sense and antisense) segments of the target gene sequence, such as 17, 18, 19, 20, 21, 22, or 23 nucleotides of coding or non-coding sequences, are sufficient. Longer sequences, such as 50, 100, 200, or 250 nucleotides or more, can also be used. Preferably, the short sense and antisense fragments are separated by spacer sequences (e.g., introns) that form loops (or hairpins) upon dsRNA formation. Any short fragment or a fragment or variant thereof of SEQ ID NO: 4 and / or SEQ ID NO: 5 may be used to prepare a silencing vector derived from the DIP gene and a transgenic plant in which one or more target genes are silenced in all or some tissues or organs (depending on the promoter used).

[0080] A convenient way to generate hairpin constructs is to use a general-purpose carrier such as... The technology carriers pHANNIBAL and pHELLSGATE (see Wesley et al. 2004, Methods Mol Biol. 265:117-30; Wesley et al. 2003, Methods Mol Biol. 236:273-86 and Helliwell & Waterhouse 2003, Methods 30(4):289-95.), all of which are incorporated herein by reference. Attached Figure Description

[0081] Figure 1A Seed head of the triploid plant A68 (wild type) in the absence of cross-pollination, exhibiting complete apomixis. Note the dark center of the fully developed seed.

[0082] Figure 1B The typical seed head of the A68 ploidypophyte formation deletion (LoD) mutant in the absence of cross-pollination. Typical LoD mutants under these conditions have seed heads smaller than the A68 wild type. Note the spotted center, numerous white undeveloped seeds lacking the parthenogenetic gene, and a few developing seeds with the parthenogenetic gene. Parthenogenesis is a gene expressed in the gametophyte and therefore segregates and is replaced by meiosis when ploidypophyte formation is deleted.

[0083] Figure 2 Associated sequence polymorphisms and ploid sporophyte formation phenotypes within the large dandelion germplasm resource group. Differences between sexual (dip) and diploid alleles (Dip) are shown in gray. Invention Details

[0085] In a first aspect, the present invention relates to isolated polynucleotides comprising the nucleic acid sequence of SEQ ID NO: 1 or a nucleic acid sequence having at least 50% or 70%, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 96% or 97%, and most preferably at least 98% or 99% sequence identity with the nucleic acid sequence of SEQ ID NO: 1.

[0086] In a second aspect, the present invention relates to isolated polynucleotides comprising the nucleic acid sequence of SEQ ID NO: 2 or a nucleic acid sequence having at least 50% or 70%, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 96% or 97%, and most preferably at least 98% or 99% sequence identity with the nucleic acid sequence of SEQ ID NO: 2.

[0087] The isolated polynucleotide containing the nucleic acid sequence SEQ ID NO: 1 or SEQ ID NO: 2 was identified as part of the presumed vacuole protein sorting-related protein gene (Vps13) of the broad genus *Taraxacum*. The Vps13 gene is a large gene. Therefore, the nucleic acid sequences of SEQ ID NO: 1 and SEQ ID NO: 2 can be contained in a single isolated nucleic acid sequence, i.e., as part of the same nucleic acid sequence. Therefore, the isolated nucleic acid sequence can contain SEQ ID NO: 1 and SEQ ID NO: 2 or variants thereof. It is understood that the Vps13 gene can contain numerous exons and introns, as well as other gene-related sequences, such as the promoter and terminator sequences contained in SEQ ID NO: 1, extending to the 5' and 3' ends (open reading frame; ORF) of the protein-coding sequence shown (SEQ ID NO: 2), and thus can be larger than SEQ ID NO: 2. Therefore, the percentage of sequence identity may not therefore involve the complete sequence of the isolated nucleic acid sequence. Rather, it is only the percentage of sequence identity that the nucleic acid sequence contained in the isolated nucleic acid sequence has with SEQ ID NO: 1 or SEQ ID NO: 2. Therefore, it should be understood that the percentage of sequence identity should be calculated relative to the nucleic acid sequences contained in the isolated nucleic acid sequences, wherein the first and last nucleotides of the nucleic acid sequence match the nucleic acid sequences of SEQ ID NO: 1 and / or SEQ ID NO: 2. Therefore, when calculating the percentage of sequence identity, it is preferable to calculate only relative to the sequences corresponding to SEQ ID NO: 1 and / or SEQ ID NO: 2. It should also be understood that SEQ ID NO: 1 and SEQ ID NO: 2, or variants thereof, are coding sequences, i.e., sequences encoding amino acids. Therefore, intron sequences may be scattered throughout such coding sequences in DNA. Therefore, in the case of calculating sequence identity from DNA sequences, portions of the sequence (e.g., introns) that do not show alignment with SEQ ID NO: 1 and / or SEQ ID NO: 2 are not considered.

[0088] In one embodiment, the isolated polynucleotide as taught herein has a nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2 as taught herein, or a variant thereof.

[0089] In one embodiment, an isolated polynucleotide comprising the nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2 or a variant or fragment thereof, as taught herein, may be referred to as “Dip or DIP polynucleotide” or “Dip or DIP gene” or “apomixis polynucleotide or apomixis gene” or “ploid sporophyte formation polynucleotide or ploid sporophyte formation gene”.

[0090] In one embodiment, the isolated polynucleotide as taught herein comprises a nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2 as taught herein, or a variant thereof, and / or the expression product of the polynucleotide and / or the protein encoded by the polynucleotide, the polynucleotide being capable of providing plosporophyte formation function to the plant or plant cells, or capable of inducing plosporophyte formation or using plosporophyte formation as part of gametophyte apomixis, preferably of the type carried out by plosporophyte formation, preferably in crops currently considered sexual crops. Gametophyte apomixis through plosporophyte formation produces offspring that are genetically identical to the parent plant. Thus, in one embodiment, the isolated polynucleotide as taught herein, or a variant thereof, can be used to produce offspring that are genetically identical to the parent plant without the need for fertilization and hybridization.

[0091] In a preferred embodiment, the Dip polynucleotide or gene and its variants and / or the expression product of the polynucleotide and / or the protein encoded by the polynucleotide, as taught above, are capable of providing plosporophyte formation function to plants or plant cells, preferably of the type that occurs in sexual crops via plosporophyte formation upon introduction into plants or plant cells.

[0092] It should be understood that the term "isolated polynucleotide" or variants thereof (e.g., genomic DNA, cDNA, or mRNA) includes naturally occurring, artificial, or synthetic nucleic acid molecules. Nucleic acid molecules can encode any polypeptide or variant thereof taught herein. These nucleic acid molecules can be used to produce polypeptides or proteins or variants thereof as taught herein. Due to the degeneracy of the genetic code, multiple nucleic acid molecules can encode the same polypeptide (e.g., polypeptides or proteins or variants thereof containing the amino acid sequence SEQ ID NO: 3 and / or SEQ ID NO: 7 or 12, as taught herein).

[0093] In one embodiment, the isolated polynucleotide as taught herein includes any variant nucleic acid molecule that encompasses any nucleic acid molecule containing a nucleotide sequence having more than 50%, preferably more than 55%, preferably more than 60%, preferably more than 65%, preferably more than 70%, preferably more than 75%, preferably more than 80%, preferably more than 85%, preferably more than 90%, preferably more than 95%, preferably more than 96%, preferably more than 97%, preferably more than 98%, and preferably more than 99% sequence identity with the nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2. Variants also include nucleic acid molecules derived from nucleic acid molecules having the nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2 by one or more nucleic acid substitutions, deletions, or insertions. Preferably, compared to SEQ ID NO: 1 or SEQ ID NO: 2, such nucleic acid molecules contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more nucleic acid substitutions, deletions, or insertions up to 100, 90, 80, 70, 60, 50, 45, 40, 35, 30, 25, 20, 15. Sequence identity can be determined by any suitable means available in the art. For example, bioinformatics can be used to perform pairwise alignment between nucleic acid sequences to identify regions that may be similar due to functional, structural, or evolutionary relationships between sequences. It should also be understood that many methods can be used to identify, synthesize, or isolate variants of the polynucleotides taught herein, such as nucleic acid hybridization, PCR techniques, computer simulation analysis, and nucleic acid synthesis.

[0094] In one implementation, the term "variant" also includes naturally occurring variants found in nature (e.g., in other dandelion species or other plants). The nucleotide sequences of said variants isolated from other dandelion species or other plants may include a dominant Dip allele and a recessive Dip allele from different plant species (e.g., including different dandelion varieties, lines, germplasm, or breeding lines). For example, without being bound by theory, the EMS mutation identified in the examples can be considered a variant of recessive Dip because of the loss of ploidysporophyte formation function, while the wild-type sequence can be considered a dominant Dip because the wild-type sequence provides ploidysporophyte formation function.

[0095] In one embodiment, the polynucleotides isolated from the variants of the present invention, such as homologs or orthologs, may also be found and / or isolated from plants other than those belonging to the genus *Taraxacum*. The isolated polynucleotides may be isolated from other wild or cultivated apomixis or non-apomixis plants and / or other plants using known methods such as PCR, strict hybridization, etc. Therefore, variants of SEQ ID NO: 1 and / or SEQ ID NO: 2 also include nucleotide sequences naturally found in other *Taraxacum* plants, strains, or varieties, and / or naturally present in other plants of other species. Such nucleotides may be identified, for example, in a BLAST search, or by re-identifying the corresponding sequences in the plant.

[0096] In one embodiment, isolated polynucleotide variants as taught herein include, for example, isolated polynucleotides of the present invention derived from a different “source” than SEQ ID NO: 1 and / or SEQ ID NO: 2, which are of dandelion origin. Thus, specifically, the present invention includes genes or alleles from a plant (e.g., wild or cultivated plants and / or from other plants) in which ploid sporophyte formation is present (as part of apomixis via gametophyte formation). These homologues can be readily isolated using the provided nucleotide sequences and / or their complementary sequences or portions thereof as primers or probes. For example, moderately stringent, stringent, or highly stringent nucleic acid hybridization methods can be used. For example, fragments of the sequences of SEQ ID NO: 1 and / or SEQ ID NO: 2, or their complementary sequences, can be used. The fragment used in this hybridization method may contain at least 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, 3000 or more consecutive nucleic acids from SEQ ID NO:1 and / or SEQ ID NO:2.

[0097] It should be understood that, due to the degeneracy of the genetic code, multiple nucleic acid sequences can encode the same amino acid sequence. For optimal expression in the host, the isolated nucleic acid sequences according to the present invention can be codon optimized by adjusting codon usage to the most preferred codons in the plant gene, particularly by optimizing the natural gene for the target plant genus or species using an available codon usage table (e.g., more adapted to expression in the target plant) (Bennetzen & Hall, 1982, J. Biol. Chem. 257, 3026-3031; Itakura et al., 1977 Science 198, 1056-1063). Codon usage tables for various plant species have been published, for example, by Ikemura (1993, in "Plant Molecular Biology Labfax", Croy, ed., Bios Scientific Publishers Ltd.) and Nakamura et al. (2000, Nucl. Acids Res. 28, 292.), and are available in major DNA sequence databases (e.g., EMBL in Heidelberg, Germany). Therefore, synthetic DNA sequences can be constructed to produce identical or substantially identical proteins. Various techniques for modifying codon usage to be preferred by host cells can be found in patents and scientific literature. The exact method of modifying codon usage is not critical to this invention.

[0098] The DNA sequence described above can be routinely modified with minor modifications as described above, i.e., through PCR-mediated mutagenesis (Ho et al., 1989, Gene 77, 51-59., White et al., 1989, Trends in Genet. 5, 185-189). Modifications to the DNA sequence can also be routinely introduced by resynthesizing the desired coding region using available techniques.

[0099] In one embodiment, the isolated polynucleotide or its variants in this invention can be modified to provide an optimal translation initiation environment for the N-terminus of the DIP protein by adding or deleting one or more amino acids at the N-terminus of the protein. Generally, for optimal translation initiation, it is preferred that the protein of this invention to be expressed in plant cells begin with a Met-Asp or Met-Ala dipeptide. Thus, an Asp or Ala codon can be inserted after the existing Met, or the second codon Val can be replaced by a codon of Asp (GAT or GAC) or Ala (GCT, GCC, GCA, or GCG). The DNA sequence can also be modified to remove illegal splicing sites.

[0100] The polynucleotides or variants thereof isolated according to the present invention are preferably "functional," i.e., they are preferably capable of providing ploid sporophyte formation function to plants, preferably as part of gametophyte apomixis, and preferably of the type occurring in plants or plant cells or sexual crops via ploid sporophyte formation. In one embodiment, isolated polynucleotides or variants thereof are provided that are homologous to polynucleotides comprising nucleic acid sequences SEQ ID NO: 1 and / or SEQ ID NO: 2 derived from the genus Taraxacum, said isolated polynucleotides being isolated from apomixis plants. Thus, in this embodiment, the polynucleotides or variants thereof isolated according to the present invention are isolated from apomixis plants. Such isolated polynucleotides or variants thereof may be particularly capable of providing ploid sporophyte formation function to plants in plants or plant cells or (sexual) crops.

[0101] It should be understood that variants of the polynucleotides taught herein perform the same function as polynucleotides containing the nucleic acid sequences of SEQ ID NO: 1 or SEQ ID NO: 2 taught herein, i.e., the ability to provide ploid sporophyte formation to plants or plant cells, preferably as part of inducing ploid sporophyte formation or gametophyte apomixis in plants or plant cells or sexual crops, particularly when introduced into plants or plant cells or sexual crops. It is further understood that any isolated polynucleotide and its variants as taught herein can encode any polypeptide and its variants as taught herein.

[0102] In one embodiment, the expression product of the polynucleotide and its variants as taught herein is an RNA molecule, preferably an mRNA molecule, siRNA, or miRNA molecule.

[0103] In one embodiment, fragments of polynucleotides and their variants as taught herein and / or expression products of said fragments and / or proteins encoded by said fragments are capable of providing ploid sporophyte formation function to plants or plant cells, preferably as part of inducing gametophyte apomixis.

[0104] In a preferred embodiment, the fragments as taught herein and / or the proteins encoded by said fragments are capable of providing ploid sporophyte formation function, preferably inducing ploid sporophyte formation or as part of inducing gametophyte apomixis.

[0105] In one implementation, the expression product of the fragment as taught herein is an RNA molecule, preferably an mRNA molecule, siRNA, or miRNA molecule.

[0106] In one embodiment, the fragment as taught herein may be a separate polynucleotide having a length of at least 20, 30, 40, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000, 1500, 2000, or 3000 nucleotides containing the nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2 and its variants as taught herein.

[0107] In a preferred embodiment, the fragment taught herein has a nucleic acid sequence of SEQ ID NO: 4, 6 or 11.

[0108] In a more preferred embodiment, the expression product of the fragment taught herein has the nucleic acid sequence of SEQ ID NO: 5.

[0109] In one embodiment, the expression product of the fragment as taught herein encodes a polypeptide comprising the amino acid sequence described in SEQ ID NO: 7 and / or 12.

[0110] Chimeric genes and vectors

[0111] In one implementation, a chimeric gene may comprise any polynucleotide, its fragments, and variants taught herein.

[0112] In one implementation, any polynucleotide, fragment thereof, and variant thereof taught herein can be operatively linked to a promoter when included in a vector taught herein. Any promoter known in the art may be used, and is suitable for linking to polynucleotides, fragments thereof, and variants taught herein. Non-limiting examples of suitable promoters include promoters that allow constitutive or regulatory expression, weak expression, and strong expression, etc. Any methods known in the art may be used to integrate polynucleotides, variants thereof, or fragments taught herein into chimeric genes.

[0113] In some implementations, it may be advantageous to operatively link the polynucleotides, fragments thereof, and variants taught herein to a so-called “constitutive promoter.” Alternatively, it may be advantageous to operatively link the polynucleotides, fragments thereof, and variants taught herein to a so-called “inducible promoter.” An inducible promoter may be a physiologically regulated promoter (e.g., through external administration of certain compounds).

[0114] In one implementation, the promoter operatively linked to the isolated polynucleotide, its variant, or fragment taught herein can be, for example, a constitutively active promoter, such as: a strong constitutive or enhanced 35S promoter (“35S promoter”) of cauliflower mosaic virus (CaMV) isolate CM 1841 (Gardner et al., 1981, Nucleic Acids Research 9, 2871-2887), CabbbB-S (Franck et al., 1980, Cell 21, 285-294) and CabbbB-JI (Hull and Howell, 1987, Virology 86, 482-493); a 35S promoter as described in Odell et al. (1985, Nature 313, 810-812) or US5164316, or a promoter from the ubiquitin family (e.g., Christensen et al., 1992, Plant Mol.Biol.18,675-689,EP 0 342 926, see also Cornejo et al.1993, Plant Mol.Biol.23,567-581 for the maize ubiquitin promoter, gos2 promoter (de Pater et al.,1992 Plant J.2,834-844), emu promoter (Last et al.,1990, Theor.Appl.Genet.81,581-588), Arabidopsis actin promoter (e.g., the promoter described by An et al. (1996, Plant J.10,107.)), rice actin promoter (e.g., the promoter described by Zhang et al. (1991, The Plant Cell 3,1155-1165)), and US Promoters described in 5,641,876 or the rice actin 2 promoter as described in WO070067; cassava leaf vein mosaic virus promoter (WO 97 / 48819, Verdaguer et al. 1998, Plant Mol. Biol. 37, 1055-1067); pPLEX series promoters from ground clover dwarf virus (WO 96 / 06932, especially the S7 promoter); alcohol dehydrogenase promoters (e.g., pAdh1S (GenBank accession numbers X04049, X00581)); and TR1' and TR2' promoters (respectively "TR1' promoter" and "TR2' promoter") that drive the expression of 1' and 2' genes of T-DNA, respectively (Velten et al.).The promoters of Scrophularia mosaic virus, histone gene promoters (such as the Ph4a748 promoter from Arabidopsis thaliana (PMB 8:179-191)) described in EMBO J 3, 1984, US6051753, and EP426641, or others.

[0115] Because constitutive expression of chimeric genes, gene constructs, or vectors in plants can be highly detrimental to plant health (high cost), it is preferable in one implementation to use promoters whose activity can be induced. Examples of inducible promoters are trauma-induced promoters, such as the MPI promoter described by Cordera et al. (1994, The Plant Journal 6, 141), which is induced by trauma (e.g., caused by insects or physical injury), or the COMPTII promoter (WO0056897) or the PR1 promoter described in US6031151. Alternatively, the promoter can be induced by chemicals such as dexamethasone as described in Aoyama and Chua (1997, Plant Journal 11: 605-612) and US6063985, or by tetracycline (TOPFREE or TOP10 promoter, see Gatz, 1997, Annu Rev Plant Physiol Plant Mol Biol. 48: 89-108 and Love et al. 2000, Plant J. 21: 579-88).

[0116] Non-constitutive promoters, but specific to one or more tissues or organs of the plant, can be used. Preferred promoters are tissue-specific. The promoter may preferably be developmentally regulated, for example, leaf-preferred or epidermal-preferred, whereby the nucleic acid sequence is expressed only or preferentially in cells of a specific tissue or organ, and / or only during a specific developmental stage, preferably in the female ovary, megasporocyte, and / or female gamete. For example, the Dip gene can be selectively expressed in plant leaves by placing the coding sequence under the control of a light-inducible promoter, such as the promoter of the ribulose-1,5-bisphosphate carboxylase small subunit gene of the plant itself or another plant (such as peas (as disclosed in U.S. Patent 5,254,799) or Arabidopsis species such as those disclosed in US 5034322 and other patents).

[0117] The term "inducible" does not necessarily require the promoter to be completely inactive in the absence of an inducer. Low levels of nonspecific activity may exist, as long as this does not lead to significant losses in plant yield or quality. Therefore, inducible is preferred when it increases promoter activity, resulting in increased transcription of downstream coding regions upon contact with the inducer.

[0118] In a preferred embodiment, the promoter of an endogenous gene is used to express a protein comprising the amino acid sequence of SEQ ID NO: 3 or a variant or fragment thereof (e.g., SEQ ID NO: 7 and / or 12) taught herein. For example, according to the invention, the promoter of the dandelion Dip allele or a corresponding promoter from another plant species can be isolated and operatively linked to a nucleic acid sequence encoding the protein of the invention. The protein is preferably capable of providing ploidysporophyte formation function, preferably as part of ploidysporophyte formation or gametophyte apomixis. The promoter, typically the upstream transcriptional regulatory region within approximately 2000 base pairs (bp) upstream of the transcription start site and / or translation start codon of a polynucleotide encoding an amino acid sequence containing the amino acid sequence of SEQ ID NO: 3 or a fragment or variant thereof (e.g., SEQ ID NO: 7 and / or 12) (e.g., homologs from other Taraxacum and / or other plants), can be isolated from apomixis and / or other plants using known methods such as TAIL-PCR (Liu et al. 1995, Genomics 25(3):674-81; Liu et al. 2005, Methods Mol. Biol. 286:341-8), Linker-PCR, or inverse PCR (IPCR). It should be understood that, since the gene sequence is part of the putative vacuole protein sorting-related protein gene Vps13 (SEQ ID 1) of the genus *Taraxacum*, the promoter comprises a sequence of SEQ ID 1 located at the 5' end (SEQ ID 2) of the gene coding region, or a sequence of another region of SEQ ID 1 located at the 5' end of a subgenomic region expressed as mRNA, miRNA, or siRNA. The expressed mRNA, siRNA, or miRNA will include the female gametophyte stage, i.e., its expression activity can be traced back to the location and time of expression of the sporophyte formation phenotype or the developmental stage leading to that stage.

[0119] In one embodiment of the invention, an endogenous promoter derived from a polynucleotide encoding a protein comprising the amino acid sequence of SEQ ID NO: 3 or a fragment or variant thereof, as taught herein, may be used, such as homologs from other dandelion-derived and / or other plants. Longer sequences than these may also be used. For any of the said nucleic acid sequences, a region approximately 2000 bp upstream of the translation start codon in the coding region may contain a transcriptional regulatory element. Thus, in one embodiment, a nucleotide sequence 2000 bp, 1500 bp, 1000 bp, 800 bp, 500 bp, 300 bp, or less upstream of the translation or transcription start site of the said polynucleotide may be isolated, and its promoter activity may be tested, and if functional, the sequence may be operatively linked to a polynucleotide encoding a protein comprising the amino acid sequence of SEQ ID NO: 3 or a fragment or variant thereof (e.g., SEQ ID NO: 7 and / or 12), as taught herein. Promoter activity of the entire sequence and its fragments can be tested, for example, by deletion analysis, by deleting the 5' and / or 3' ends of the transcription start site region, and by testing promoter activity using known methods (e.g., operatively linking the promoter to the deletion or multiple deletions of the reporter gene).

[0120] In another embodiment, the promoter drives the expression of the miRNA and siRNA molecules of the present invention.

[0121] In the plants, plant cells, or sexual crops described in this invention, whether the Dip allele, derived from a ploidypophyte formation function, has the ability to provide or induce ploidypophyte formation (preferably as part of gametophyte asexual reproduction) may depend on the molecular function of the polypeptide or the molecular function of a protein encoded by an isolated polynucleotide as taught herein. In one embodiment, a protein encoded by an isolated polynucleotide, fragment thereof, or variant thereof, taught herein, may have a dominant function provided by expressing or overexpressing a protein comprising the amino acid sequence of SEQ ID NO: 3 or a variant or fragment thereof (e.g., SEQ ID NO: 7 and / or 12). When expressed in a plant, the isolated polynucleotide encoding the protein is capable of providing or enhancing ploidypophyte formation function in the plant, or of inducing or enhancing ploidypophyte formation in the plant, plant cell, or crop.

[0122] For example, when a polynucleotide containing the nucleic acid sequence of SEQ ID NO: 1 and / or SEQ ID NO: 2 or a fragment or variant thereof (e.g., SEQ ID NO: 4, 5, 6, or 11) is synthesized in a plant by means of a suitable plant promoter and having a functional number of the protein encoding it, the function of ploid sporophyte formation, preferably as part of gametophyte apomixis, or the occurrence of ploid sporophyte formation, can be induced or significantly enhanced compared to plants lacking said protein. Function (i.e., the ability of the polynucleotides, their variants, or fragments, taught herein, to induce or cause ploid sporophyte formation in plants) can be tested by introducing such nucleic acid sequences into a suitable host plant (e.g., a non-diploid dandelion line) and allowing them to be expressed in the host plant, and by analyzing the effect of the transformant on ploid sporophyte formation function in bioassays (as described in the examples taught herein).

[0123] In one embodiment, silencing the expression of a polynucleotide, variant, or fragment thereof that encodes a protein containing the amino acid sequence of SEQ ID NO: 3 or a variant or fragment thereof (e.g., SEQ ID NO: 7 and / or 12) as taught herein may result in loss of function, i.e., reduced or absent plosporophyte formation or the absence of apomixis through plosporophyte formation. Therefore, those skilled in the art can readily identify polynucleotides, variants, or fragments thereof that encode a protein containing the amino acid sequence of SEQ ID NO: 3 or a fragment or variant thereof (e.g., SEQ ID NO: 7 and / or 12) and / or a fragment thereof (as described herein, preferably capable of providing plosporophyte formation as part of apomixis in a plant or plant cell or crop).

[0124] In one embodiment, a chimeric gene as taught herein is provided, comprising any isolated polynucleotide (SEQ ID NO: 1 or SEQ ID NO: 2), a variant or fragment thereof as taught herein (e.g., SEQ ID NO: 4, 5, 6, or 11). The chimeric gene is preferably capable of providing ploid sporophyte formation function to the plant, plant cells, or crop described in this invention.

[0125] In one embodiment, a polynucleotide as taught herein (e.g., SEQ ID No: 1 or SEQ ID NO: 2), a variant or fragment thereof as taught herein (e.g., SEQ ID NO: 4, 5, 6 or 11), or a chimeric gene as taught herein may be included in the genetic construct.

[0126] In a preferred embodiment, the genetic constructs taught herein may comprise open reading frames of isolated polynucleotides of the present invention (e.g., SEQ ID NO: 2), variants thereof, or fragments thereof (e.g., SEQ ID NO: 4, 5, 6, or 11) as taught herein.

[0127] In one embodiment, the isolated polynucleotides taught herein (e.g., SEQ ID NO: 1 or SEQ ID NO: 2), variants or fragments thereof as taught herein (e.g., SEQ ID NO: 4, 5, 6 or 11) may be contained within a nucleic acid vector.

[0128] The construction of the chimeric genes, genetic constructs, and vectors described in this invention is generally known in the art. Preferably, the chimeric genes, genetic constructs, and vectors are capable of providing ploid sporophyte formation function to plants, or of inducing ploid sporophyte formation or inducing gametophyte apomixis through ploid sporophytes in plants, plant cells, or crops. Chimeric genes can be generated by modifying endogenous gene sequences. For example, a recessive allele (i.e., dip) can be modified to become a dominant allele (i.e., Dip), thereby enabling the dominant allele to provide ploid sporophyte formation function or to induce ploid sporophyte formation or gametophyte apomixis through ploid sporophytes in plants, plant cells, or crops. Alternatively, an endogenous gene that is capable of providing ploid sporophyte formation function or inducing ploid sporophyte formation or gametophyte apomixis through ploid sporophytes but is not expressed can be modified, i.e., by modifying the endogenous promoter sequence to express the endogenous gene. Such modifications may include (targeted) mutagenesis, thereby causing mutations in at least 1, 5, 10, 20, 50, 100, 200, 500, or 1000 nucleotides of an endogenous gene. An example of such modification can be found in Example 5, where the four mutations found in EMS mutations confer a loss of the ploidypophyte formation phenotype; therefore, reversing these mutations can provide the acquisition of the ploidypophyte formation phenotype.

[0129] In one embodiment, the chimeric gene taught herein can be generated by operably linking a nucleic acid sequence encoding the protein (or variant or fragment) described herein to a promoter sequence suitable for expression in a host cell using standard molecular biology techniques. The promoter sequence may already be present in the vector, and thus the nucleic acid sequence is simply inserted downstream of the promoter sequence in the vector. In one embodiment, the chimeric gene comprises a suitable promoter for expression in plant or microbial cells (e.g., bacteria), operably linked to the nucleic acid sequence of the present invention, optionally followed by a 3' untranslated nucleic acid sequence. The nucleic acid sequence described herein is optionally located after a 5' untranslated region (UTR). As described below, the promoter, 3' UTR, and / or 5' UTR may be derived, for example, from an endogenous Dip gene, or from other sources. Furthermore, the nucleic acid sequence described herein may also contain intron sequences, which may be contained within the 3' UTR or 5' UTR sequence, but may also be incorporated into the coding sequence of the nucleic acid sequence of the present invention.

[0130] In one embodiment, the chimeric gene, genetic construct, and vector as taught herein are preferably capable of expressing a nucleic acid sequence encoding the amino acid sequence of the present invention, wherein the amino acid sequence of the present invention is preferably capable of providing ploid sporophyte formation function to the plant, preferably as part of gametophyte apomixis in the plant, plant cell, or crop. Therefore, the chimeric gene, genetic construct, and vector preferably contain the dominant Dip allele of the present invention.

[0131] In one embodiment, the nucleic acid vector as taught herein may contain an active promoter sequence in a plant cell, said promoter sequence being operatively linked to an isolated polynucleotide (e.g., SEQ ID NO: 1 or SEQ ID NO: 2), a variant or fragment thereof as taught herein (e.g., SEQ ID NO: 4, 5, 6 or 11), or a chimeric gene or genetic construct as taught herein.

[0132] In a preferred embodiment, the promoter sequence of a nucleic acid vector, as taught herein, may include:

[0133] a) The natural promoter sequence of the nucleic acid sequence of SEQ ID NO: 1 and / or SEQ ID NO: 2;

[0134] b) A functional fragment of the promoter sequence of a); or

[0135] c) A nucleic acid sequence comprising a natural promoter sequence having at least 70%, preferably at least 80%, more preferably at least 90%, and most preferably at least 95% sequence identity with the nucleic acid sequence of SEQ ID NO: 1 and / or SEQ ID NO: 2;

[0136] d) The natural promoter sequence of the nucleic acid sequence of SEQ ID NO: 6;

[0137] e)d) functional fragments of the promoter sequence; or

[0138] f) A nucleic acid sequence comprising a natural promoter sequence having at least 70%, preferably at least 80%, more preferably at least 90%, and most preferably at least 95% sequence identity with the nucleic acid sequence of SEQ ID NO: 6.

[0139] In a preferred embodiment, the promoter of the nucleic acid vector, as taught herein, is a female ovary-specific promoter, preferably a megasporocyte-specific promoter and / or a female gamete-specific promoter.

[0140] isolated peptides

[0141] In a third aspect, the present invention relates to isolated polypeptides comprising the amino acid sequence of SEQ ID NO: 3 or an amino acid sequence having at least 50% or 70%, preferably at least 80%, more preferably at least 90%, even more preferably at least 95%, even more preferably at least 96% or 97%, most preferably at least 98% or 99% sequence identity with the amino acid sequence of SEQ ID NO: 3.

[0142] In a preferred embodiment, the polypeptide taught herein has the amino acid sequence of SEQ ID NO: 3 or a variant or fragment thereof.

[0143] In one embodiment, the isolated polypeptide as taught herein comprises the amino acid sequence of SEQ ID NO: 3 as taught above, and variants or fragments thereof, which may be referred to as a DIP polypeptide or protein or an apomixis-related polypeptide or protein.

[0144] In one embodiment, the DIP polypeptide or protein, and its variants or fragments, as taught above, are capable of providing plosporophyte formation function to plants or plant cells, preferably as part of inducing plosporophyte formation or gametophyte apomixis in crops. Thus, in one embodiment, isolated polypeptides or proteins as taught herein can be used to produce offspring that are genetically identical to the parent plant without the need for fertilization and hybridization.

[0145] In a preferred embodiment, DIP polypeptides or proteins, and variants or fragments thereof, preferably as part of apomixis, as taught above, are capable of providing plosporophyte formation function to plants or plant cells or can induce plosporophyte formation in crops (particularly when introduced into plants or plant cells).

[0146] A polypeptide or protein having the amino acid sequence of SEQ ID NO: 3 or a variant thereof, as taught herein, is identified as the presumed vacuole protein sorting-related protein gene Vps13 or a portion thereof from the genus *Taraxacum*. The Vps13 gene is a large gene. Therefore, the amino acid sequence of SEQ ID NO: 3 may be contained in a single isolated protein, i.e., as part of the same amino acid sequence, or as multiple parts of that same amino acid sequence. Thus, the isolated protein may contain both SEQ ID NO: 3 or a variant thereof. It should be understood that, since the Vps13 gene can constitute a large protein, the percentage of sequence identity may not be relative to the complete sequence of the isolated protein when compared to the size of the amino acid sequence of SEQ ID NO: 3 or a variant thereof. Instead, only the amino acid sequence contained in the isolated protein may have the said percentage of sequence identity with SEQ ID NO: 3. Therefore, it should be understood that the percentage of sequence identity should be calculated relative to the amino acid sequence contained in the isolated protein, wherein the first and last amino acids of the amino acid sequence are aligned with the amino acid sequence of SEQ ID NO: 3. Therefore, when calculating the percentage of sequence identity, it is preferable to calculate only relative to the sequence corresponding to SEQ ID NO: 3.

[0147] It should be understood that the polypeptides taught herein also include variant polypeptides having the amino acid sequence of SEQ ID NO: 3, said variant having more than 50%, preferably more than 55%, more than 60%, more than 65%, more than 70%, preferably more than 75%, more than 80%, more than 85%, more than 90%, more than 95%, preferably more than 96%, preferably more than 97%, preferably more than 98%, preferably more than 99% sequence identity with the amino acid sequence of SEQ ID NO: 3. Variant polypeptides having the amino acid sequence of SEQ ID NO: 3 also include polypeptides derived from polypeptides having the amino acid sequence of SEQ ID NO: 3 by substitution, deletion, or insertion of one or more amino acids. Preferably, compared with a polypeptide having the amino acid sequence of SEQ ID NO: 3, such a polypeptide contains 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more up to 100, 90, 80, 70, 60, 50, 45, 40, 35, 30, 25, 20, 15 amino acid substitutions, deletions or insertions.

[0148] In one embodiment, the variant polypeptides taught herein may differ from the provided amino acid sequence by one or more amino acid deletions, insertions, and / or substitutions, and include natural and / or synthetic / artificial variants.

[0149] In one embodiment, the term "variant polypeptide" also includes naturally occurring variant polypeptides found in nature (e.g., in cultivated or wild lettuce plants and / or other plants). The isolated protein also includes fragments of the isolated protein, i.e., non-full-length peptides. Fragments comprise peptides or variants thereof consisting of at least 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000 or more consecutive amino acids encoded by SEQ ID NO: 3, particularly consisting of or composed of 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 400, 500, 600, 700, 800, 900, 1000 or more consecutive amino acids or variants thereof from SEQ ID NO: 3.

[0150] The isolated polypeptides or variants thereof, as taught herein, are preferably capable of providing ploid sporophyte formation function to plants, and preferably capable of inducing ploid sporophyte formation or gametophyte apomixis in plants, plant cells, or crops. ploid sporophyte formation is... This means that the isolated polypeptides, fragments, and variants according to the invention are capable of inducing ploid sporophyte formation. The ploid sporophyte formation function of the invention involves skipping the first female meiosis (meiosis I), resulting in two unmeiotic megaspores with the same genotype as the maternal plant. One megaspore degenerates, and the other surviving unmeiotic megaspore produces an unmeiotic female gametophyte (or embryo sac) containing an unmeiotic egg cell. This unmeiotic egg cell develops into an embryo without fertilization and has the same genotype as the maternal plant, i.e., a clone of the maternal plant.

[0151] In one implementation, the isolated polypeptide or variant thereof, as taught herein, may be isolated from a natural source and resynthesized by chemical synthesis (using, for example, a peptide synthesizer, provided by Applied Biosystems) or by recombinant host cells through expression of nucleic acid sequences encoding the isolated polypeptide, fragments thereof, and variants thereof, as taught herein.

[0152] In one embodiment, the isolated polypeptide or variant thereof, as taught herein, may contain conserved amino acid substitutions within the following categories:

[0153] Basic (e.g., Arg, His, Lys);

[0154] Acidic (e.g., Asp, Glu);

[0155] Nonpolar (e.g., Ala, Val, Trp, Leu, Ile, Pro, Met, Phe, Trp); or

[0156] Polar (e.g., Gly, Ser, Thr, Tyr, Cys, Asn, Gln).

[0157] In addition, non-conservative amino acid substitutions may also fall within the scope of this invention.

[0158] In one embodiment, the isolated polypeptide or variant thereof, as taught herein, may also be a chimeric polypeptide, such as a polypeptide composed of at least two distinct domains. Since SEQ ID NO: 3 is derived from or partially derived from the Vps13 gene, SEQ ID NO: 3 or a variant thereof may be exchanged with a corresponding sequence in a Vps13 protein that does not provide or poorly provides plosporophyte formation function or cannot induce gametophyte apomixis through plosporophyte formation in plants or plant cells or crops. Thus, chimeric polypeptides or proteins capable of providing or improving plosporophyte formation function, or capable of plosporophyte formation or improved plosporophyte formation in plants or plant cells or crops, can be obtained. The chimeric polypeptide as taught herein may also have a portion or more of the amino acid sequence of SEQ ID NO: 3. Furthermore, the chimeric polypeptides taught herein may comprise an N-terminal domain of one protein (e.g., derived from the genus *Taraxacum* or another plant species) and an intermediate and / or C-terminal domain of another protein (e.g., derived from the genus *Taraxacum* or another plant species). Such chimeric proteins may have improved plosporophyte formation function compared to the native protein, or may contribute to improving or potentially improving plosporophyte formation in plants, plant cells, or crops.

[0159] Amino acid sequence identity can be determined by any suitable means available in the art. For example, amino acid sequence identity can be determined by pairwise comparison using the Needleman and Wunsch algorithms as defined above and the GAP default parameters. It should also be understood that many methods can be used to identify, synthesize, or isolate variants of the peptides taught herein, such as immunoblotting, immunohistochemistry, ELISA, amino acid synthesis, etc.

[0160] It should also be understood that any variant or fragment of the DIP peptide taught herein may perform the same function and / or have the same activity as the DIP peptide taught herein. The functionality or activity of any DIP peptide or variant thereof may be determined by any method known in the art and deemed suitable for these purposes by a person skilled in the art.

[0161] In one embodiment, fragments of the polypeptide (SEQ ID NO: 3) or variants thereof, as taught herein, are capable of providing plosporogenesis function to plants or plant cells capable of inducing plosporogenesis or gametophyte apomixis.

[0162] In one embodiment, fragments of the polypeptide and its variants as taught herein may have at least 10, 20, 30, 40, 50, 100, 150, 200, 250, 300, 400 or 500 consecutive amino acids of the polypeptide.

[0163] In one embodiment, fragments of polypeptides and their variants as taught herein have the amino acid sequence of SEQ ID NO: 7 and / or 12.

[0164] method

[0165] On the other hand, the present invention relates to a method for producing apomixis seeds, comprising the following steps:

[0166] a) Transform plants, plant parts or plant cells with any polynucleotide taught herein (e.g. SEQ ID NO: 1 or SEQ ID NO: 2) or its variants or fragments (e.g. SEQ ID NO: 4, 5, 6 or 11), or chimeric genes taught herein, or genetic constructs taught herein and / or nucleic acid vectors taught herein to produce primary transformants;

[0167] b) Growing flowering plants and / or flowers from the primary transformant, wherein the polynucleotides, variants or fragments, chimeric genes, constructs and / or vectors as described above are present and / or expressed at least in the female ovary, preferably in the megaspore mother cell and / or female gamete; and

[0168] c) Pollinating the primary transformant to induce seed production, preferably using pollen from a tetraploid plant or self-pollination pollen from the primary transformant.

[0169] It should be understood that step c) can be omitted when the primary transformant develops autonomous endosperm.

[0170] In one implementation, the apomixis seed obtained by the methods taught herein is a clone of a primary transformant as taught herein.

[0171] In one embodiment of step (a), the plant or plant part may be transformed with a chimeric gene containing any polynucleotide as taught herein (e.g., SEQ ID NO: 1 or SEQ ID NO: 2) or its variants or fragments (e.g., SEQ ID NO: 4 or SEQ ID NO: 5).

[0172] In a preferred embodiment, the chimeric gene comprises SEQ ID NO: 2.

[0173] In one embodiment, the chimeric gene may be included in the genetic construct or vector of the present invention.

[0174] In another embodiment, the chimeric gene may also comprise a modified endogenous gene. This modification may include, but is not limited to, modification via targeted mutagenesis or the use of nucleases such as CRISPR / Cas. When introduced into the plant, plant part, or plant cell, the chimeric gene preferably provides ploid sporophyte formation or induces ploid sporophyte formation or gametophyte apomixis through ploid sporophyte formation in the plant, plant part, or plant cell. The vector can be used to transform host cells into which the chimeric gene is inserted in the nuclear genome or in plastid, mitochondrial, or chloroplast DNA, and the vector can be expressed using a suitable promoter (e.g., McBride et al., 1995 Bio / Technology 13,362; US 5,693,507). One advantage of plastid genome transformation is the reduced risk of transgene (one or more) diffusion. Plastosome genome transformation can be performed in a manner known in the art, see, for example, Sidorov VA et al. 1999, Plant J. 19: 209-216 or Lutz KA et al. 2004, Plant J. 37(6): 906-13.

[0175] In one implementation, the polynucleotide or variant or fragment taught herein contained in the chimeric gene, as taught above, is operatively linked to a promoter sequence, wherein the promoter sequence comprises:

[0176] (a) The endogenous promoter sequence of the nucleic acid sequence of SEQ ID NO: 1 and / or SEQ ID NO: 2;

[0177] (b) A functional fragment of the natural promoter sequence;

[0178] (c) A nucleic acid sequence whose endogenous promoter sequence contains at least 70% sequence identity with the nucleic acid sequence of SEQ ID NO: 1 or SEQ ID NO: 2; or

[0179] Functional fragments of the nucleic acid sequences in (d) and (c);

[0180] e) The natural promoter sequence of the nucleic acid sequence of SEQ ID NO: 6;

[0181] f)d) functional fragments of the promoter sequence;

[0182] g) A nucleic acid sequence comprising a natural promoter sequence having at least 70%, preferably at least 80%, more preferably at least 90%, and most preferably at least 95% sequence identity with the nucleic acid sequence of SEQ ID NO: 6; or

[0183] Functional fragments of the nucleic acid sequence h)g).

[0184] It should be understood that, as described above, the chimeric gene of the present invention can represent a dominant allele. Therefore, transformation of a plant, plant part, or plant cell with such a dominant chimeric gene will be sufficient to provide the plant, plant part, or plant cell with ploid sporophyte formation function or to induce ploid sporophyte formation or gametophyte apomixis through ploid sporophyte formation in the plant, plant part, or plant cell.

[0185] In one embodiment, a polynucleotide is provided capable of encoding a protein (SEQ ID NO: 3) or a variant or fragment thereof (e.g., SEQ ID NO: 7 and / or 12) as taught herein, and capable of providing ploid sporophyte formation function to a plant, plant part, or plant cell, or inducing gametophyte apomixis through ploid sporophyte formation as described above in the plant, plant part, or plant cell. Such polynucleotides can be used to prepare chimeric genes, and vectors containing these are used to transfer the chimeric gene into a host cell and produce the protein in the host cell (e.g., a cell, tissue, organ, or organism derived from a transformed cell). Vectors used to produce said protein (or protein fragment or variant) in plant cells are referred to herein as “expression vectors.” The host cell is preferably a plant cell.

[0186] Any plant may be a suitable host, but the most preferred host is a plant species that can benefit from enhanced or reduced plosporophyte formation. Varieties or breeding lines with other favorable agronomic characteristics are particularly preferred. By producing transgenic plants and inducing plosporophyte formation, along with suitable control plants, it is readily possible to test whether the genes and / or proteins (or variants or fragments thereof) provided herein confer the desired increase in plosporophyte formation on the host plant.

[0187] In one implementation scheme, suitable host plants may be selected from maize / corn (Corn), wheat (Wheat), barley (e.g., malted barley), oats (e.g., oats), sorghum (Sorghum bicolor), rye (Secalecereale), soybean (Glycine, e.g., G. max), cotton (Cotton, e.g., upland cotton, sea island cotton), brassica (e.g., rapeseed, mustard greens, cabbage, Chinese cabbage, etc.), sunflower (oilseed sunflower), safflower, yam, cassava, alfalfa (alfalfa), rice (Oryza, e.g., indica or japonica rice varieties), forage grasses, pearl barnyard grass (P. glaucum), tree species (pine, poplar, fir, plantain, etc.), tea, coffee, oil palm, coconut, vegetable species (e.g., peas, zucchini, legumes (e.g., green beans), peppers, cucumbers, artichokes, asparagus, eggplant, broccoli, garlic, leeks. Lettuce, onion, radish, turnip, tomato, potato, cabbage, carrot, cauliflower, chicory, celery, spinach, lettuce, fennel, beet), fleshy fruit plants (grape, peach, plum, strawberry, mango, apple, plum, cherry, apricot, banana, blackberry, blueberry, citrus, kiwi, fig, lemon, lime, nectarine, raspberry, watermelon, orange, grapefruit, etc.), ornamental plants (such as rose, petunia, chrysanthemum, lily, gerbera), herbs (mint, parsley, basil, thyme, etc.), woody plants (such as poplar, willow, oak, eucalyptus), fiber plants such as flax (Linumusitatissimum) and hemp (Cannabis sativa).

[0188] In a preferred embodiment, the host plant may be a plant species selected from the genera *Taraxacum*, *Lactuca*, *Vaccinium*, *Capsicum*, *Solanum*, *Cucumis*, *Zea*, *Cotton*, *Soybean*, *Castor*, *Oryza*, and *Sorghum*.

[0189] In one embodiment, a polynucleotide (SEQ ID NO: 1 or SEQ ID NO: 2), a variant or fragment thereof (e.g., SEQ ID NO: 4, 5, 6 or 11) is preferably included in the chimeric gene of the present invention and is capable of encoding a protein (SEQ ID NO: 3) or a variant or fragment thereof (e.g., SEQ ID NO: 7 and / or 12) and is capable of providing ploid sporophyte formation function to a plant, plant part or plant cell, or inducing ploid sporophyte formation or gametophyte apomixis by ploid sporophyte formation in a plant, plant part or plant cell, which is capable of being stably inserted into the nuclear genome of a single plant cell in a conventional manner, and the plant cell thus transformed can be used in a conventional manner to produce transformed plants that have an altered phenotype due to the presence of the protein in a specific cell at a specific time. In this regard, T-DNA vectors containing polynucleotides, variants thereof, or fragments as taught herein, capable of encoding proteins, variants thereof, or fragments as taught herein, capable of providing plosporophyte formation function or inducing plosporophyte formation or inducing gametophyte apomixis through plosporophyte formation, can be used in *Agrobacterium tumefaciens* to transform plant cells, and the subsequently transformed plants can be regenerated from the transformed plant cells using, for example, the steps described in EP 0 116 718, EP 0 270 822, PCT publication WO84 / 02913, and published European patent application EP 0 242246, and in Gould et al. (1991, *Plant Physiol* 95, 426-434). The construction of T-DNA vectors for *Agrobacterium*-mediated plant transformation is well known in the art. The T-DNA vector can be a binary vector as described in EP 0 120 561 and EP 0 120 515, or a co-integration vector that can be integrated into the Agrobacterium Ti-plasmid via homologous recombination, as described in EP 0 116 718. Lettuce transformation protocols have been described, for example, in Michelmore et al., 1987 and Chupeau et al., 1989.

[0190] Preferred T-DNA vectors each contain a promoter operatively linked to a nucleic acid sequence that functionally encodes a protein capable of providing ploidy sporophyte formation (e.g., encoding SEQ ID NO: 3 or a variant or fragment thereof (e.g., SEQ ID NO: 7 and / or 12)). The promoter is operatively linked to a sequence between the nucleotide sequence or the T-DNA boundary sequence, or at least located to the left of the right boundary sequence. The boundary sequence is described in Gielen et al. (1984, EMBO J3, 835-845). Of course, other types of vectors can be used to transform plant cells, such as direct gene transfer (as described in EP 0 223 247), pollen-mediated transformation (as described in EP 0 270 356 and WO85 / 01856), protoplast transformation (as described in US 4,684,611), plant RNA virus-mediated transformation (as described in EP 0 067 553 and US 4,407,956), liposome-mediated transformation (as described in US 4,536,475), and other methods. Known methods such as electroporation or triparental hybridization can be used to introduce T-DNA vectors into Agrobacterium.

[0191] Similarly, the selection and regeneration of transformed plants from transformed cells is well known in the art. Clearly, procedures can be specifically tailored to different species, and even to different varieties or cultivars of a single species, in order to regenerate transformants at a high frequency.

[0192] The methods taught herein can be used to obtain plants, plant parts, or plant cells with altered levels of plosporophyte formation, particularly transgenic plants containing significantly enhanced levels of plosporophyte formation. Such plants can be prepared using various methods, as further described below.

[0193] Plants obtained by or achievable through the methods of this invention can be used in conventional plant breeding programs to produce more transgenic plants as taught herein. Single-copy transgenic plants can be obtained using, for example, Southern blot analysis or PCR-based methods. Third Wave Technologies, Inc. selects technologies for differentiation. The presence of chimeric genes can distinguish transformed cells and plants from untransformed cells and plants. Flanking sequences of plant DNA at transgene insertion sites can also be sequenced, enabling the development of "event-specific" detection methods for routine use. See, for example, WO0141558, which describes elite event detection kits (e.g., PCR kits) based on, for example, integrated sequences and flanking (genomic) sequences.

[0194] In one embodiment, the polynucleotides, variants thereof, or fragments taught herein are capable of providing ploid sporophyte formation function, or inducing ploid sporophyte formation, or inducing gametophyte apomixis through ploid sporophyte formation, in plants, plant parts, or plant cells, for example, by expressing the protein, variants thereof, or fragments thereof of the present invention capable of providing ploid sporophyte formation function, or inducing gametophyte apomixis through ploid sporophyte formation, in plants, plant parts, or plant cells, inserted into the plant cell genome, such that the inserted coding sequence is located downstream (i.e., 3') and under the control of a promoter that can direct expression in the plant cell. This can preferably be accomplished by inserting a chimeric gene into the plant cell genome, particularly the nucleoid or plastid (e.g., chloroplast) genome.

[0195] The nucleic acid sequence of the present invention, or its corresponding sequence, capable of providing ploid sporophyte formation function to plants, is preferably inserted into the plant genome such that the coding sequence is located upstream (i.e., 5') of a suitable 3' untranslated region (“3' end” or 3'UTR). Suitable 3' ends include the CaMV 35S gene (“3'35S”), the carmine synthase gene (“3'nos”) (Depicker et al., 1982 J. Mol. Appl. Genetics 1, 561-573.), the octopus alkaloid synthase gene (“3'ocs”) (Gielen et al., 1984, EMBO J 3, 835-845), and the T-DNA gene 7 (“3'gene 7”) (Velten and Schell, 1985, Nucleic Acids Research 13, 6981-6998), which function as a 3' untranslated DNA sequence in transformed plant cells, etc. In one embodiment, the 3'UTR and / or 5'UTR of the dandelion allele are capable of providing ploid sporophyte formation function, i.e., using SEQ ID NO: 1 and / or SED ID NO: 2 (or variants or fragments thereof). The 3'UTR and / or 5'UTR may also be used in another embodiment, as it can also be used in combination with other coding regions or other nucleic acid constructs.

[0196] DIP-encoding nucleic acid sequences can be optionally inserted as heterozygous gene sequences into the plant genome, thereby enabling the sequence that provides the plant with ploid sporophyte formation function to be linked within the reading frame with genes encoding optional or scoreable markers (US 5,254,799; Vaeck et al., 1987, Nature 328, 33-37), such as the neo (or nptII) gene encoding kanamycin resistance (EP 0 242 236), thereby enabling the plant to express an easily detectable fusion protein.

[0197] Preferably, for the purpose of selection, but also for the control of weed selection, the transgenic plants of the present invention can also be transformed with DNA encoding proteins that confer resistance to herbicides, such as broad-spectrum herbicides, for example herbicides based on glufosinate as the active ingredient (e.g., glufosinate-ammonium). Or BASTA; resistance is conferred by the PAT or bar gene; see EP 0242 236 and EP 0 242 246) or glyphosate (e.g. Resistance is conferred by EPSPS genes (see, for example, EP0508 909 and EP0 507 698). Using herbicide resistance genes (or other genes that confer the desired phenotype) as selective markers also has the advantage of avoiding the introduction of antibiotic resistance genes.

[0198] Alternatively, other selectable marker genes, such as antibiotic resistance genes, can be used. Since retaining antibiotic resistance genes in the transformed host plant may be unacceptable, these genes can be removed again after selection of the transformants. Different techniques exist for removing transgenes. One method to achieve removal is by side-linking the chimeric gene to a lox site and, after selection, crossing the transformed plant with a plant expressing CRE recombinase (see, for example, EP506763B1). Site-specific recombination results in the excision of the marker gene. Another site-specific recombination system is the FLP / FRT system described in EP686191 and US5527695. Site-specific recombination systems such as CRE / LOX and FLP / FRT can also be used for gene stacking purposes. Furthermore, single-component excision systems have been described, see, for example, WO9737012 or WO9500555.

[0199] All or part of the nucleic acid sequences of the present invention can provide plants with ploid sporophyte formation capabilities, for example, because they encode the proteins of the present invention. They can also be used to transform microorganisms, such as bacteria (e.g., *Escherichia coli*, *Pseudomonas*, *Agrobacterium*, *Bacillus*, etc.), fungi or algae, or insects, or for the preparation of recombinant viruses. Transformation of bacteria with all or part of the nucleic acid sequences of the present invention integrated into a suitable cloning vector is performed in a conventional manner, preferably using the conventional electroporation techniques described in Maillon et al. (1989, FEMS Microbiol. Letters 60, 205-210) and WO 90 / 06999. For expression in prokaryotic host cells, the codon usage of the nucleic acid sequences can be optimized accordingly. Intron sequences should be removed, and other adjustments can be made in known manners to obtain optimal expression.

[0200] The DNA sequence of the nucleic acid sequence of the present invention can be further modified in a translation-neutral manner, that is, for the amino acid sequence, by modifying the potentially repressive DNA sequence present in the gene part and / or by introducing changes in codon usage, for example, as described above, adjusting codon usage to be optimized for plants, preferably specific related plant genera.

[0201] As described above, according to one embodiment of the invention, the protein or chimeric protein of the invention, capable of providing plants with ploid sporophyte formation function, targets intracellular organelles such as plastids, preferably chloroplasts and mitochondria, and may also be secreted from the cell, potentially optimizing protein stability and / or expression. Similarly, the protein may be localized to vacuoles. For this purpose, in one embodiment of the invention, the chimeric gene of the invention comprises a coding region encoding a signal peptide or target peptide linked to the protein coding region of the invention. Particularly preferred peptides included in the proteins of this invention are transport peptides for targeting chloroplasts or other plastids, particularly those from repetitive transport peptide regions of plant genes whose gene products are targeted to plastids, such as the optimized transport peptide of Capellades et al. (US 5,635,618), the transport peptide of ferredoxin-NADP+ oxidoreductase from spinach (Oelmuller et al., 1993, Mol. Gen. Genet. 237, 261-272), the transport peptide described by Wong et al. (1992, Plant Molec. Biol. 20, 81-93), and the targeting peptide in published PCT patent application WO 00 / 26371. Also preferred are signal secretion peptides of proteins linked to such extracellular peptides, such as the secretion signal of potato glycoside inhibitor II (Keil et al., 1986, Nucl. Acids Res. 14, 5641-5650), the secretion signal of the rice alpha-amylase 3 gene (Sutliff et al., 1991, Plant Molec. Biol. 16, 579-591), and the secretion signal of tobacco PR1 protein (Cornelissen et al., 1986, EMBO J. 5, 37-40). Particularly useful signal peptides of the present invention include chloroplast transport peptides (e.g., Van Den Broeck et al., 1985, Nature 313, 358), or the optimized chloroplast transport peptides that induce protein transport to chloroplasts as described in US5,510, 471 and US5,635, 618. Secretion signal peptides or peptides that target proteins to other plastids, mitochondria, endoplasmic reticulum, or other organelles may also be used. Signal sequences used to target intracellular organelles or for extracellular or secreted signals into the cell wall in plants are found in naturally occurring targeted or secreted proteins, preferably... et al. (1989, Mol. Gen. Genet. 217, 155-161), And Weil (1991, Mol. Gen. Genet. 225, 297-304), Neuhaus and Rogers (1998, Plant Mol. Biol. 38, 127-144), Bih et al. (1999, J. Biol. Chem. 274, 2284-22894), Morris et al. (1999, Biochem. Biophys. Res. Commun. 255, 328-333), Hesse et al. (1989, EMBO) The proteins described in J. 8, 2453-2461, Tavaldoraki et al. (1998, FEBS Lett. 426, 62-66.), Terashima et al. (1999, Appl. Microbiol. Biotechnol. 52, 516-523), Park et al. (1997, J. Biol. Chem. 272, 6876-6881), and Shcherban et al. (1995, Proc. Natl. Acad. Sci. USA 92, 9245-9249).

[0202] In one embodiment, several nucleic acid sequences encoding proteins capable of providing ploid sporophyte formation function to plants are selectively co-expressed in a single host under the control of different promoters. Co-expression host plants can be readily obtained by transforming plants already expressing the proteins of the present invention, or by hybridizing plants transformed with different proteins of the present invention. Therefore, the present invention also provides plants or plant parts having multiple nucleic acid sequences having the same or different isolated nucleic acid sequences of the present invention, wherein each nucleic acid sequence is capable of providing ploid sporophyte formation function to the plant. It is understood that the term "multiple" in this context refers to per cell. Alternatively, several nucleic acid sequences of the present invention (each of which is capable of providing ploid sporophyte formation function to the plant) may be present on a single transformation vector, or separate vectors may be used for simultaneous co-transformation and selection of transformants containing multiple chimeric genes. Similarly, one or more genes encoding proteins capable of providing ploid sporophyte formation function of the present invention may be expressed in a single plant along with other chimeric genes, such as encoding proteins that enhance or inhibit ploid sporophyte formation or are involved in apomixis. It is understandable that different proteins can be expressed in the same plant, or each protein can be expressed in a single plant and then combined in the same plant by hybridizing the single plants. For example, in hybrid seed production, each parent plant can express a single protein. After hybridizing the parent plants to produce hybrids, the two proteins are combined in the hybrid plant.

[0203] Plants generated from multiple chimeric genes of the present invention under the control of different promoters are also an embodiment. Thus, the enhancement or inhibition of the ploidophyte formation phenotype can be fine-tuned by expressing appropriate amounts of the protein of the present invention capable of providing ploidophyte formation function to the plant at appropriate times and locations. This fine-tuning can be accomplished by determining the most suitable promoter and / or by selecting a transformation “event” that displays the desired expression level.

[0204] Transformants expressing the desired levels of proteins capable of providing ploid sporophyte formation are selected, for example, by analyzing copy number (Southern blot analysis), mRNA transcription levels (e.g., RT-PCR using primer pairs or flanking primers), or by analyzing the presence and levels of the ploid sporophyte-forming protein in various tissues (e.g., SDS-PAGE; ELISA assays, etc.). For regulatory reasons, single-copy transformants are preferred, and flanking sequences at chimeric gene insertion sites are analyzed, preferably sequenced, to characterize the transformation outcome. Transgenic events expressing high or moderate DIP expression are selected for further development until high-performance elite events with stable Dip transgenes are obtained.

[0205] Furthermore, it is conceivable that a plant possessing several chimeric genes may have a first chimeric gene encoding a protein capable of providing plosporophyte formation function, and a second chimeric gene capable of suppressing or silencing the first chimeric gene. The second chimeric gene is preferably under the control of an inducible promoter. Such a plant may be particularly advantageous because it allows control over plosporophyte formation function. By inducing expression of said promoter, plosporophyte formation function in the plant may be lost. Furthermore, this control may also be obtained, or can be achieved, by introducing the chimeric gene according to the invention into a plosporophyte-forming plant, which is also capable of suppressing or silencing the endogenous gene (i.e., its naturally encoded amino acid of the invention) that provides plosporophyte formation function to the plant.

[0206] By selecting conserved nucleic acid sequence portions of the nucleic acid sequence of the present invention, alleles in a host plant or plant part can be silenced. As mentioned above, such silencing may result in suppression of ploid sporophyte formation in the plant. Therefore, this document also covers plants containing chimeric genes that include transcriptional regulatory elements operatively linked to sense and / or antisense DNA fragments of the nucleic acid sequence of the present invention, and that the chimeric genes are capable of exhibiting repressive or enhanced ploid sporophyte formation. The transcriptional regulatory element may be a suitable promoter, which may be an inducible promoter.

[0207] Transformed plants expressing one or more proteins capable of providing ploid sporophyte formation function to the plants of the present invention may also contain other transgenes, such as genes conferring disease resistance or tolerance to other biotic and / or abiotic stresses. To obtain such plants with “superimposed” transgenes, other transgenes may be introduced into the transformed plant, or the transformed plant may be subsequently transformed with one or more other genes, or several chimeric genes may be used to transform plant lines or varieties. For example, several chimeric genes may be present on a single vector or on different vectors co-transformed.

[0208] In one embodiment, the following genes are combined with one or more chimeric genes described in this invention: known disease resistance genes (especially genes conferring enhanced resistance to necrotic pathogens), virus resistance genes, insect resistance genes, abiotic stress resistance genes (e.g., drought tolerance, salt tolerance, heat or cold tolerance, etc.), herbicide resistance genes, etc. Therefore, the superimposed transformants can possess a wider range of tolerances to biotic and / or abiotic stresses, such as pathogen resistance, insect resistance, nematode resistance, salinity, cold stress, heat stress, and water stress. Additionally, as described above, in this embodiment, silencing or inhibiting ploidypophyte formation can be combined with gene expression methods in a single plant.

[0209] It should be understood that plants or plant parts containing the chimeric gene of the present invention preferably do not exhibit undesirable phenotypes, such as reduced yield, increased susceptibility to diseases (especially to necrophores), or unwanted structural changes (dwarfing, deformity), etc., and if these phenotypes are observed in primary transformed plants, they can be removed by conventional methods. Any plant described herein may be homozygous or hemizygous for the chimeric gene of the present invention.

[0210] On the other hand, the present invention relates to a method for producing clones of hybrid plants, comprising the following steps:

[0211] a) Plants that reproduce sexually by cross-fertilization with pollen from the plants taught in this paper to produce F1 hybrid seeds;

[0212] b) At least in the female ovary, preferably in the megasporocyte and / or in the female gamete, select F1 plants that contain and / or express polynucleotides or variants or fragments thereof as taught herein or polypeptides or variants or fragments thereof as taught herein;

[0213] c) Optionally, the selected F1 plants are pollinated to induce seed production, preferably with pollen from tetraploid plants; and

[0214] d) Harvesting seeds; and

[0215] e) Optionally, a hybrid clone plant is grown from the seed.

[0216] When the selected F1 plant develops autonomous endosperm, step c can be omitted.

[0217] In one implementation, the cloning of step (e) of the method taught herein is an apomixis clone.

[0218] In one implementation, the method taught herein includes obtaining the hybrid plant.

[0219] On the other hand, the present invention relates to a method for conferring ploid sporophyte formation in plants, plant parts, or plant cells, or for inducing gametophyte apomixis through ploid sporophyte formation in plants, plant parts, or plant cells, comprising the following steps:

[0220] a) Transform the plant, plant part, or plant cell with any polynucleotide, variant, or fragment thereof, chimeric gene, genetic construct, and / or nucleic acid vector taught herein; and

[0221] b) Optionally regenerating plants, wherein the polynucleotide, its variants or fragments, genes, constructs and / or vectors are present and / or expressed at least in the female ovary, preferably in megasporocytes and / or in female gametes.

[0222] In one implementation, the polynucleotides, variants or fragments taught herein are integrated into the genome of the plant, plant part or plant cell.

[0223] In one implementation, the method as taught herein includes obtaining ploid sporophyte-forming plants.

[0224] On the other hand, the present invention relates to a method for conferring ploid sporophyte formation on a plant, plant part, or plant cell, or for inducing ploid sporophyte formation on a plant, plant part, or plant cell, or for inducing gametophyte apomixis through ploid sporophyte formation, comprising the steps of:

[0225] a) Modifying an endogenous polynucleotide, a variant of a polynucleotide, or a fragment thereof, preferably a gene for a vacuole protein sorting-related protein, in a plant, plant part, or plant cell, such that the modified plant, plant part, or plant cell contains any polynucleotide, a variant thereof, or a fragment thereof, as taught herein; and

[0226] b)r can be used to regenerate plants.

[0227] In one implementation, the modified polynucleotide, polynucleotide variant, or fragment of the method taught herein in step (a) is expressed and / or encodes a polypeptide.

[0228] In one embodiment, the modified polynucleotide or polynucleotide fragment of step (a) of the method taught herein is present at least in the female ovary, preferably in the megasporocyte and / or female gamete.

[0229] In one implementation, step (a) of the method taught herein is modified by the following:

[0230] a) Introducing or expressing at least one site-specific nuclease in the plant, plant part, or plant cell, preferably said nuclease being selected from: Cas9 / RNA CRISPR nucleases, zinc finger nucleases, broad-spectrum nucleases, and TAL effector nucleases; and / or through...

[0231] b) Oligonucleotide-directed mutagenesis using oligonucleotides, preferably said oligonucleotides being single-stranded oligonucleotides; and / or by...

[0232] c) Chemical mutagenesis, preferably using ethyl methanesulfonate.

[0233] In one implementation, the method taught herein includes obtaining ploid sporophyte-forming plants.

[0234] In one embodiment, the modification, particularly in the genus *Taraxacum*, comprises the deletion of nucleotides encoding amino acid residues GGGGW at positions 96-100 of the endogenous dip amino acid sequence shown in SEQ ID NO: 10 and / or the deletion of nucleotides encoding amino acid residues PPT at positions 108-110 of the endogenous dip amino acid sequence shown in SEQ ID NO: 10. In other organisms, nucleic acids encoding amino acid residues corresponding to the amino acid residues GGGGW or PPT found in *Taraxacum* may be deleted. Those skilled in the art will be able to identify the correct amino acid residues to be deleted and the corresponding nucleotide sequences encoding these amino acid residues.

[0235] In one implementation, the modifications include one or more, for example, such as Figure 2 All differences between the nucleotide sequences of dip (sexual allele; SEQ ID NO: 13) and Dip (ploid sporophyte formation allele) are shown.

[0236] In one embodiment, the whole plant, seed, cell, tissue, and progeny of any transformed plant obtainable by the methods taught herein are included herein and the presence of the chimeric gene, genetic construct, or vector taught herein can be identified in DNA by detection, for example, by using whole-genome DNA as a template and PCR analysis using specific PCR primer pairs, such as specific primer pairs designed for sequences SEQ ID NO: 1 and / or SEQ ID NO: 2 and / or SEQ ID NO: 4, 6, or 11 or variants thereof as described above. “Event-specific” PCR diagnostic methods can also be developed, wherein the PCR primers are based on flanking plant DNA of the inserted chimeric gene, see US6563026. Similarly, event-specific AFLP fingerprints or RFLP fingerprints can be developed that identify transformed or modified plants or any plant, seed, tissue, or cell derived therefrom.

[0237] Plants and seeds

[0238] On the other hand, the present invention relates to plants, plant parts or plant cells comprising chimeric genes, genetic constructs and / or nucleic acid vectors taught herein, wherein the genes, constructs and / or vectors are present and / or expressed at least in the female ovary, preferably in the megasporocyte and / or in the female gamete.

[0239] In one implementation, the plant seeds taught herein are apomixis seeds.

[0240] In one implementation, the seed taught herein is a clone of the plant from which it develops.

[0241] In a preferred embodiment, the plants, plant parts, plant cells, or seeds taught herein are derived from the genera: dandelion, lettuce, pea, capsicum, solanum, cucumber, maize, cotton, soybean, wheat, rice, allium, brassica, sunflower, beet, chicory, chrysanthemum, foxtail grass, rye, barley, alfalfa, common bean, rose, lily, coffee, flax, hemp, cassava, carrot, squash, watermelon, and sorghum.

[0242] use

[0243] On the other hand, the present invention relates to the use of any isolated polynucleotide, variant or fragment thereof as taught herein, for inducing ploid sporophyte formation in plants.

[0244] On the other hand, the present invention relates to the use of any isolated polynucleotide or fragment or variant thereof as taught herein for the prevention of segregation of multiple genes, QTLs or transgenes.

[0245] On the other hand, the present invention relates to the use of any isolated polynucleotide or fragment or variant thereof as taught herein for gene stacking.

[0246] On the other hand, the present invention relates to the use of any isolated polynucleotide or fragment or variant thereof, as taught herein, as a marker for development and / or identification of ploid sporophyte formation traits.

[0247] In one embodiment, the polynucleotides (SEQ ID NO: 1 or SEQ ID NO: 2), variants or fragments thereof (e.g., SEQ ID NO: 4, 5, 6 or 11), proteins that they can encode (SEQ ID NO: 3) or variants or fragments thereof (e.g., SEQ ID NO: 7 and / or 12), and polynucleotide sequences encoding any protein capable of providing ploidypophyte formation function or inducing gametophyte apomixis through ploidypophyte formation, and variants thereof, may be used as genetic markers for marker-assisted selection of the following: alleles capable of providing ploidypophyte formation function in dandelion species (and / or other plant species), and for transferring different or identical ploidypophyte formation alleles to and / or combining them in plants of interest, and / or transferring them to and / or combining them in plants that can be used to produce intraspecific or interspecific hybrids with plants in which ploidypophyte formation alleles (or variants) are present.

[0248] Based on these sequences, a wide variety of marker assays can be developed. The development of marker assays typically involves identifying polymorphisms among alleles so that the polymorphism is a genetic marker “marking” a specific allele. The polymorphism is then used in the marker assay. For example, nucleic acid sequences of SEQ ID NO: 1 and / or SEQ ID NO: 2 and / or SEQ ID NO: 4, 5, 6, or 11 or variants thereof according to the invention can be associated with the presence, absence, reduction, inhibition, or enhancement of ploidypophyte formation. This is accomplished, for example, by screening one or more such sequences in ploidypophyte-forming plant material and / or non-ploidypophyte-forming plant material to associate a specific allele with the absence or presence of ploidypophyte formation function. Therefore, PCR primers or probes can be synthesized that detect the presence or absence of SEQ ID NO: 1 and / or SEQ ID NO: 2 and / or SEQ ID NO: 4, 5, 6, or 11 or variants or fragments thereof in samples obtained from plant material (e.g., RNA, cDNA, or genomic DNA samples). By comparing the sequences or portions thereof, polymorphic markers that may be associated with ploidypophyte formation can be identified. Polymorphic markers (e.g., SNP markers linked to the Dip or dip allele) can then be developed into rapid molecular assays for screening plant material for the presence or absence of ploid sporophyte formation alleles. Therefore, the presence or absence of these “genetic markers” indicates the presence of the Dip allele linked to them, and the detection of the genetic markers can be used instead of the detection of the Dip allele. Examples of these markers are disclosed in the Examples section.

[0249] Preferably, a simple and rapid marker assay is used, which can rapidly detect specific Dip or dip alleles (e.g., alleles conferring ploidy sporophyte formation, such as Dip, in contrast to alleles without this function, such as dip) or combinations of alleles in a sample (e.g., a DNA sample). Therefore, in one embodiment, the use of the nucleic acid sequences of SEQ ID NO: 1 and / or SEQ ID NO: 2, or variants or fragments thereof (SEQ ID NO: 4 or SEQ ID NO: 5, 6, or 11) having at least 70%, 80%, 90%, 95%, 98%, 99% or more nucleic acid identity with them, or one or more fragments thereof, in a molecular assay is provided for determining the presence or absence of the Dip allele and / or the dip allele in a sample, and / or whether the sample is homozygous or heterozygous relative to said alleles.

[0250] This determination may include, for example, the following steps:

[0251] (a) Provide ploid and nonploid sporophyte forming plant material and / or its nucleic acid samples;

[0252] (b) Identify nucleotide sequences derived from the Vps13 gene, for example, including sequences from the material of (a) corresponding to SEQ ID NO: 1 and / or SEQ ID NO: 2 or variants and / or fragments thereof (SEQ ID NO: 4, 5, 6 or 11), thereby identifying polymorphisms between nucleotide sequences;

[0253] (c) Associate polymorphism with ploid sporophyte formation characteristics in plants, thereby associating polymorphism with ploid sporophyte formation and nonploid sporophyte formation alleles at the Dip locus;

[0254] The identified relevant polymorphisms may optionally be further used in step (d).

[0255] (d) Use the polymorphic markers to develop marker assays for germplasm screening or characterization and MAS.

[0256] Therefore, in one embodiment of the invention, PCR primers and / or probes, molecular markers, and kits for detecting DNA or RNA sequences of alleles derived from ploid sporophyte formation genes (i.e., Dip and / or dip alleles) are provided. Degenerate or specific PCR primers capable of amplifying Dip and / or dip DNA (e.g., nucleic acid sequences or variants or fragments thereof from SEQ ID NO: 1 and / or SEQ ID NO: 2, as well as SEQ ID NO: 4, 5, 6, or 11) from a sample can be synthesized based on said sequences (or variants thereof) that are well known in the art (see Dieffenbach and Dveksler (1995) PCR Primer: A Laboratory Manual, Cold Spring Harbor Laboratory Press and McPherson at al. (2000) PCR-Basics: From Background to Bench, First Edition, Springer Verlag, Germany). For example, any consecutive nucleotides of length 9, 10, 11, 12, 13, 14, 15, 16, 18 or longer in those sequences (or complement chains) can be used as primers or probes. The polynucleotide sequences of the present invention can also be used as hybridization probes. The Dip gene / allelic assay kit may contain Dip and / or dip allele-specific primers and / or Dip and / or dip allele-specific probes. The primers and / or probes can be used to detect Dip and / or dip DNA in a sample using relevant protocols. For example, such assay kits can be used to determine whether a plant has been transformed with the Dip gene (or a portion or variant thereof) of the present invention, or to screen for the presence of the Dip allele (or Dip homologs or orthologs) and optionally, the determination of complexes in dandelion germplasm and / or other plant species germplasm.

[0257] Therefore, in one embodiment, a method is provided for detecting the presence of a nucleotide sequence encoding a DIP protein in plant tissues (e.g., in dandelion tissue or nucleic acid samples thereof). The method includes:

[0258] a) Obtain plant tissue samples, such as dandelion tissue samples or nucleic acid samples thereof.

[0259] b) Analyze the presence or absence of one or more markers associated with the Dip allele in a nucleic acid sample using molecular marker assays, wherein the marker assays detect SEQ ID NO: 1 and / or SEQ ID NO: 2 and / or SEQ ID NO: 4, 5, 6 or 11, or any one of the sequences containing at least 70% nucleotide identity in the sample, and optionally...

[0260] c) Select plants that contain one or more of the markers (e.g., dandelion plants).

[0261] Further applications of ploid sporophyte formation

[0262] Plomesporophyte formation is an element of apomixis, and genes involved in plomesporophyte formation can be combined with parthenogenetic genes to produce apomixis and used in the applications listed above. These genes can be introduced into sexual crops through transformation. Knowledge of the structure and function of apomixis genes can also be used to modify endogenous sexual reproduction genes to make them apomixis genes. A preferred use is to place apomixis genes under an inducible promoter so that apomixis can be shut down when sexual reproduction produces new genotypes, and activated when apomixis is needed to propagate elite genotypes.

[0263] However, the ploid sporophyte-forming polynucleotide or gene of the present invention can also be used in entirely new ways, rather than directly as an element of apomixis. The ploid sporophyte-forming gene can be used for sexual polyploidization, producing polyploid offspring from diploid plants. Polyploid plants generally exhibit heterosis and produce higher yields than diploid plants (Bingham, ET, RWGroose, DR Woodfield & K.K. Kidwell, 1994. Complementary gene interactions in alfalfa are greater in autopolyploids than diploids. Crop Sci 34:823–829.; Mendiburu, AO & S.J. Peloquin, 1971. High yielding tetraploids from 4x-2x and 2x-2x matings. Amer Potato J 48:300–301). The Dip gene, which provides the plant of this invention with the function of ploid sporophyte formation (or a chimeric gene, vector, or genetic construct), avoids female meiosis I and thus produces first division recovery (FDR) egg cells, which transmit the entire maternal genome, including the interactions of all heterozygous and epistatic genes (Mok, DWS and SJPeloquin. 1972. Three mechanisms of 2npollen formation in diploid potatoes. Am. Potato J. 49:362-363.; Ramana, MS. 1979. A re-examination of the mechanisms of 2n gamete formation in potato and its implications for breeding. Euphaitica 28:537-561). Offspring produced by FDR gametes are superior to those produced by second division recovery (SDR) gametes, which only transmit a portion of the parental heterozygosity and epistaticity to their offspring. Unreduced gametes of the FDR and SDR types produce hybrid offspring after hybridization, exhibiting significantly increased heterozygosity compared to somatic polyploidization achieved through chemical treatments (such as colchicine). Therefore, FDR gametes, like those induced by the Dip gene, represent the optimal gamete type for sexual polyploidization.FDR gametes have been shown to be useful for improving polyploid crops such as potatoes, alfalfa, blueberries, and some forage crops (Ramanna, MS and Jacobsen E. 2003. Relevance of sexual polyploidization for crop improvement – ​​a review. Euphaitica 133:3-8; Mariani, A. & S. Tavoletti, 1992. Games with Somatic Chromosome Number in the Evolution and Breeding of Polyploid Polysomic Species. Proc Workshop, Perugia, Tipolithographia Porziuncola-Assisi (PG) Italy, pp. 1–103; Veilleux, R., 1985. Diploid and polyploid gametes in crop plants: Mechanisms of formation and utilization in plant breeding. Plant Breed Rev 3:252–288). In these applications, the Dip gene is expressed only during female megasporogenesis, and males undergo meiosis, which is highly beneficial. This allows the Dip gene to infiltrate into the diploid gene pool by reducing pollen grains, creating new beneficial gene combinations through hybridization. Another very useful characteristic of the Dip gene for plant breeding is its dominance, thus heterozygotes express the ploid sporophyte formation phenotype. This significantly simplifies the use of the Dip gene in breeding programs.

[0264] One specific application of sexual polyploidization is the production of triploids, which can be used to produce seedless fruits. Triploids can also serve as a source of trisomics, which is very useful for localization studies.

[0265] However, in apomixis, the combination of ploid sporophyte formation and parthenogenesis in a single plant, using ploid sporophyte formation in one generation and parthenogenesis in the next, will link the crop's sexual gene pool at both diploid and polyploid levels through the ploidy-induced increase in ploidy and the ploidy-induced decrease in ploidy at the partloid and polyploid levels. This is very practical because polyploid populations may be more suitable for mutation induction, as they are more tolerant of more mutations. Polyploid plants may also be more robust. However, diploid populations are more suitable for selection, and diploid hybridization is more conducive to the construction of genetic mapping, BAC libraries, etc. Parthenogenesis in polyploids produces double haploids that can hybridize with diploids. ploid sporophyte formation in diploids produces unreduced FDR oocytes, which can be fertilized by pollen from polyploids to produce polyploid offspring. Therefore, the alternation of ploid sporophyte formation and parthenogenesis in different reproductive generations connects the diploid and polyploid gene pools.

[0266] The following non-limiting examples illustrate different embodiments of the invention. Unless otherwise stated in the examples, all recombinant DNA techniques are performed according to the standard protocols described below, for example, in Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual, Second Edition, Cold Spring Harbor Laboratory Press, and Sambrook and Russell (2001) Molecular Cloning: A Laboratory Manual, Third Edition, Cold Spring Harbor Laboratory Press, NY; and in Volumes 1 and 2 of Ausubel et al. (1994) Current Protocols in Molecular Biology, Current Protocols, USA. Standard materials and methods for plant molecular work are described in Plant Molecular Biology Labfax (1993) by RDDCroy, jointly published by BIOS Scientific Publications Ltd (UK) and Blackwell Scientific Publications, UK. Example

[0267] Example 1: Genetic localization of the DIP locus

[0268] 1.1 Apomixis Recombinant Groups

[0269] Genetic mapping of the ploid sporophyte formation (Dip) locus was performed by hybridization between the diploid sexual dandelion plant TJX3-20 and the triploid parthenogenetic plant A68. TJX3-20 was selected as the male-sterile (non-pollen-producing) seed parent to prevent the production of a high proportion of self-pollinated offspring, a result of pollen-guided effects common in diploid × triploid dandelion hybrids (Tas en Van Dijk 1999). Average seed formation in the TJX3-20 × A68 hybrid was low, ranging from 1% to 3%. Numerous hybridizations produced a total of 190 offspring. Only viable euploid offspring were produced: 97 diploids, 92 triploids, and 1 tetraploid (ploidy levels determined by PARTEC flow cytometry, Van Dijk al. 2003). None of the diploids were apomixis, unlike the triploids which were classified as either apomixis or non-apomixis.

[0270] Phenotyping of 1.2-fold sporophyte formation

[0271] To genetically locate the DIP locus, triploid progeny plants were phenotypiced based on ploid sporophyte formation versus aploid sporophyte formation (meiosis). Triploid progeny plants that produce triploid seeds without cross-pollination are apomixis and therefore also ploid sporophytic. For ploid sporophyte formation phenotyping in apomixis plants, a so-called testcross was performed (Ozias Akins and Van Dijk 2007). Triploid progeny from the TJX3-20×A68 cross were crossed with diploid sexual pollen donors. Seeds were harvested and germinated, and the ploidy level of the progeny was determined by flow cytometry (Partec ploidy analyzer, van Dijk et al. 2003). If the progeny consisted only of tetraploid plants, the triploid maternal plant was subtractively identified as ploid sporophytic because the diploid pollen donor produced haploid pollen grains. If the offspring consist of plants with triploid or lower ploidy levels, it can be concluded that the oocytes of the maternal plant have a reduced number of chromosomes and that the maternal plant itself is formed from nonploid sporophytes.

[0272] 1.3 Genetic map of the DIP chromosomal region

[0273] According to the method described by Wu et al. (1992), single-agent dominant markers (simple, e.g., 001) can be located in autopolyploid plants. Seven AFLP (Vos et al. 1995) markers (from Vijverberg et al. 2004) closely associated with the Dip locus were located in 76 triploid progeny plants from the TJX3-20xA68 hybrid: (see Table 1 for AFLP primer codes) E40M60-505 (505 indicates the size of the base pair fragment; short code: S4), E38M48-215 (S8), E42M50-440 (S7), E35M52-235 (S10), E38M48-215 (S9), E45M53-090 (A4), and E37M59-135 (A5). To locate the Dip locus, the triploid progeny plants were subjected to ploid sporophyte formation phenotypes using the pseudo-testcross method described above. Table 2 shows the genotypes of four triploid progeny plants (AS99, AS112, AS193, and AS196) that had recombination events in the DIP chromosomal region.

[0274] EcoRI

[0275] EcoRI Selective nucleotides E35 ACT E37 ACG E38 ACT E40 AGC E42 AGT E43 ATA E45 ATG E49 CAG E60 CTC MseI M40 AGC M42 AGT M48 CAC M50 CAT M52 CCC M53 CCG M59 CTA M60 CTC

[0276] Table 1. Selective nucleotides of the AFLP primers used.

[0277]

[0278] Table 2. Marker maps of recombination (TJX320×A68) and deletion (A68_i124) in the Dip region. (+) indicates the presence of the marker; (-) indicates the absence of the marker.

[0279] Example 2. Deletion localization of the DIP locus

[0280] 2.1 Apomixis-deficient population

[0281] Because the number of seeds in the TJX3-20×A68 hybrid is too low to produce the thousands required for fine genetic mapping, an alternative method is needed. Therefore, deletion mapping is used to precisely locate this chromosomal region. Gamma radiation causes random deletions of varying sizes throughout the genome, regardless of hot or cold spots of recombination. Gamma radiation deletion has been successfully used to map apomixis genes in grass species (Catanach AS, Erasmuson SK, Podivinsky E, Jordan BR, Bicknell RA (2006) Deletion mapping of genetic regions associated with apomixis in Hieracium. Proc Natl Acad Sci USA 103(49):18650-18655). First, the optimal dose of gamma radiation (50% seedling survival) for clone A68 was determined by irradiating dried dandelion seeds with a series of test doses in the range of 100-800 g / L produced from a 60Co source (Isotron BV, Ede, The Netherlands). For the final experiment, 3 × 2000 seeds were irradiated with three different doses: one-third at 250 Gy, one-third at 300 Gy, and one-third at 400 Gy. The seeds were allowed to germinate on moist filter paper in petri dishes at room temperature. A total of 3075 plants were grown in pots in a greenhouse (350 plants treated with 200 Gy, 1600 plants treated with 300 Gy, and 1125 plants treated with 400 Gy). The plants were grown in a heated greenhouse for two months (21°C during the day, 16 hours of light, and 18°C ​​at night). Next, the plants were kept at 2–10°C for two months to induce flowering. After this vernalization period, the plants were again grown in a heated greenhouse under the same conditions. Over 90% of the plants flowered and produced seeds. The plants were classified according to whether they exhibited the apomixis deletion phenotype (LoA). Apomixis A68 plants spontaneously produce seeds and form large, white seed heads with a dark brown center, in which the seeds (botanically termed achenes: solitary seed fruits) are attached to the receptacle (see [link to plant section]). Figure 1A ).

[0282] In cases of apomixis deletion, the center of the seed head is lighter, and the diameter of the seed spike is typically reduced due to the seed's inability to develop normally. The apomixis deletion phenotype was screened from over 13,000 seed spikes. Finally, 102 plants were identified as having the apomixis deletion phenotype. Most of these plants produced both apomixis deletion and apomixis seed heads, indicating they were chimeras. This is due to the fact that the shoot meristem of irradiated seeds is multicellular (M1 generation).

[0283] 2.2 ploidy-sporophyte formation deletion phenotype

[0284] The absence of apomixis in irradiated plants may be due to the absence of plophypophyte formation, parthenogenesis, or other causes. The absence of plophypophyte formation in apomixis was detected by testcross (see above). Because these non-Dip plants retained the parthenogenetic phenotype, plants with absent plophypophyte formation also arose spontaneously (therefore without any cross-pollination; see [link to article]). Figure 1B A smaller number of triploid and hypoploid triploid offspring. Since parthenogenesis is expressed in the gametophyte, it separates from the egg cells of plants that form nonploid sporophytes.

[0285] 2.3 Low-resolution missing location

[0286] When a portion of one of the three homologous chromosomes is deleted, a single-dose AFLP / SCAR marker located in the deletion region is lost. To determine which Dip loci were partially lost in 102 plants with apomixis, the presence / absence of the following Dip-associated single-dose markers was investigated: S8, S7, S9, S10, A4, and A5.

[0287] A total of 23 plants with apomixis were found to have lost two or more of these markers. Most of these plants were also phenotypically classified as ploid sporophytogenesis deletion, confirming that the Dip-gene was lost through deletion. The number of missing markers is an indicator of deletion size (Catanach et al 2006). Plant i124 retained all of these markers except S9 and A4, indicating that this plant has the smallest deletion at the Dip-locus. Five plants with the smallest deletion (including i124) were used to create non-chimeras via tissue culture. Leaves were sterilized and explants were grown in vitro to regenerate whole plants. AFLP analysis confirmed that these plants were homozygous and still carried the DIP deletion.

[0288] Example 3. DNA sequencing of the DIP locus

[0289] 3.1 Precise localization of DIP loci using AFLP markers in deletion populations

[0290] To identify novel AFLP markers within the detected minimum Dip deletion (i124), a novel marker screening strategy, Bullked Deletion Analysis (BDA), similar to Bullked Separation Analysis (Michelmore et al. 1991), was developed. The presence of AFLP fragments was compared in three DNA samples: Sample A: DNA from a plant with the minimum Dip deletion (i124); Sample B: a pool of DNA from three plants with larger deletions in the Dip region; and Sample C: unirradiated DNA from an A68 clone. Only AFLPs lacking in both Sample A and Sample B were found within the minimum deletion. Considering the pooled sample B prevented selection of deletions outside the Dip locus. Candidate AFLPs from BDA were validated on deleted plants with deletions in each ploidytophthoriform region. Screening 966 different AFLP primer combinations yielded three novel Dip deletion markers (DD1: E43M40-68, DD2: E49M42-215, and DD3: E60M42-76) located within a Dip deletion in plant i124. Based on the number of AFLP markers screened using 966 AFLP primer combinations and the three marker deletions, the estimated size of the Dip deletion in plant i124 is less than 450 kb. The DD2 marker was successfully cloned and sequenced (SEQ ID NO: 12).

[0291] 3.2 Gene segregation via BAC landing and walking

[0292] To construct the complete physical BAC contiguous group at the Dip locus of the apomixis clone A68, a BAC library was screened. The A68 BAC library was constructed by the Arizona Genome Institute and was available via the AGI website (http: / / www.genome.arizona.edu / orders / ) as TO_Ba. The BAC library had an average insert size of 113 kb, covering 10 genomic equivalents (Taraxacum genome size: 835 Mb / 1C). It was constructed at the HindIII site of the pAGIBAC1 vector and contained 73,728 clones. The BAC insert library was double-sampled onto four sheets of nylon filter paper. DNA from clones in the BAC library was pooled (192 superpools: 384 BACs per plate; each plate also pooled 96 BAC DNAs in 4 pools). AFLP analysis of the pooled BAC DNA was performed to screen for BACs containing the S10, A4, DD1, DD2, and DD3 markers in the BAC insert library. The BAC library was also screened using the DD2 sequence (SEQ ID NO: 12) (Ross et al. 1999) via overgo hybridization through a nylon filter. For each marker, a BAC insert was selected and fully sequenced using GS-FLX sequencing technology. New overgo oligonucleotide probes were developed by using the ends of seed BACs to extend BAC contigs (BAC walks).

[0293] In addition to BAC walking, a physical map of the A68BAC library was constructed using sequence-based tagging (Whole Genome Profiling - Van Oeverenet al2011). BAC walking and WGP localization provided consistent BAC contigs for the DIP region. A minimal BAC tiling path was constructed based on the shared WGP tag using Finger Printed Contig (FPC) software. BACs with minimal tiling paths were sequenced using GS FLX technology. Individual 454 reads were assembled using Newbler software. In most cases, two BAC variants were found, with sequence identity varying between 95-99%. These variants were interpreted as distinct alleles or haplotypes. The presence of the DD2 marker (SEQ ID NO: 12) distinguished between the Dip and dip BAC minimal tiling paths.

[0294] 3.3 Location of missing break points on the minimum overlapping path of BAC

[0295] To locate deletion breakpoints and determine the minimal overlay path covering the smallest Dip deletion in i124, PCR primers for one gene were designed for each BAC sequence. Genes were amplified by PCR, and the DNA was directly Sanger sequenced on an ABI 3730XL. This produced a complex raw sequencing dataset in the ABI trace file for A68, with many bimodal peaks. However, in i124, the pattern was frequently simplified and was a subset of the A68 pattern, as is expected when one of the least distinct alleles is missing. When a gene's sequence pattern showed bimodal peaks in both A68 and i124, it was inferred that the gene was not missing in i124. BACs in the middle of the minimal overlay path typically indicated the missing gene, while BACs at the ends showed no indication of a deletion. Therefore, it was concluded that the minimal overlay path spanned the deletion in i124.

[0296] Example 4. Unbiased identification of ploid sporophyte formation genes within the Dip locus

[0297] 4.1 The generation of EMS apomixis knockout

[0298] We infer that when apomixis in dandelions is genetically controlled, knockout mutations could likely be induced by mutagens such as ethyl methanesulfonate (EMS). Since we can predict genes within the Dip locus, the Dip gene should be able to be identified by resequencing genes at the Dip locus in ploidysis formation deletion mutants. When we find several ploidysis formation mutants, they should all have mutations in the same gene (the Dip gene), while mutations in genes at the Dip locus unrelated to the ploidysis formation phenotype would not be enriched. This would thus identify the functional Dip gene.

[0299] To generate an EMS-induced apomixis knockout, 1800 plants grown from A68 seeds were treated with 0.35% EMS for 16 hours at room temperature. After seed formation, the plants were screened for a ploid sporophyte formation deletion phenotype (see description above). Six putative DIP deletion mutants (LoD1 to LoD6) were detected, although two of them, LoD3 and LoD5, did not produce seeds in the test cross. Since the LoD plants were isolated for parthenogenesis, some viable M2 seeds were produced, and M2 plants were grown from them. To our knowledge, this is the first successful generation of an apomixis deletion mutant via EMS treatment. Attempts to generate apomixis deletions via EMS treatment in other species have been unsuccessful (Asker and Jerling 1990, Praekelt and Scott 2001).

[0300] 4.2 High-throughput resequencing of genes predicted in the physical spacer map of ploidy sporophyte formation deletion in apomixis-eluting EMS mutants

[0301] Genes in the Dip and dip BAC minimally covered pathways (see above) were predicted using the Augustus gene prediction software (Stanke M., R Steinkamp, ​​S Waack and B. Morgenstern (2004) "AUGUSTUS: a web server for gene finding in eukaryotes" Nucleic Acids Research, Vol. 32, W309-W312) using Arabidopsis gene models. Gene annotation was performed by BLASTing the predicted protein sequences against a non-redundant database from NCBI, with a 40% protein identity threshold. A total of 129 dandelion genes were predicted in the Dip and dip BAC minimally covered pathways.

[0302] Leaf materials were collected from dandelion A68, the A68LoD1-6EMS mutant, and the LoD deletion line (A68i124). Genomic DNA was extracted using the CTAB program (Rogstad 1992). Genomic DNA was extracted using the standard program on a FLUOstar Omega (BGM LABTHEC) with Quant-iTTM. DNA samples were quantified using dsDNA reagent (Invitrogen). The DNA samples were diluted to a concentration of 20 ng / μl and then LoD samples were combined to produce two pools (pool A = LoD1 + LoD2 + LoD3; pool B = LoD4 + LoD5 + LoD6).

[0303] Specific primers were designed for PCR amplification of 129 predicted genes, primarily targeting their coding sequences. A total of 295 primer pairs were designed. The dandelion apomixis A68 clone, the A68_i124 deletion line (LoD phenotype), and A68LoD EMS mutant pools A and B were selected as targets for amplicon screening. The aim was to associate the EMS mutant phenotype with the EMS mutation and thereby identify the DIP gene.

[0304] 295 amplicones were generated from each selected target via PCR. 50 μl PCR reactions were performed per sample, containing 80 ng DNA, 50 ng forward primer, 50 ng reverse primer, 0.2 mM dNTP, 1 U Herculase H II fusion DNA polymerase (Stratagene), and 1X Herculase H II reaction buffer. PCR was performed using the following thermogram: 95 °C for 2 min; then 95 °C for 30 s, 30 s at 55 °C, and 30 s at 72 °C, for 35 cycles; followed by cooling to 4 °C. Equal volumes of PCR products from the samples were used for GS FLX fragment library samples.

[0305] Amplicon screening was performed using the GS FLX+PLATFORM (Roche Applied Science), which allows for massively parallel picoliter-level amplification and pyrosequencing of single DNA molecules. Amplicon sample libraries were constructed using the standard Roche protocol. Barcodes (Multiplex Identifiers, MIDs) were added during experimental preparation. Samples with MID tags were combined for simultaneous amplification and sequencing (multiplexing). The amplicon libraries (A68, A68_i124, A68_EMS pools A and B) were sequenced using a complete picoliter microtiter plate (PTP) (70x75 mm) with two regions. Sequencing was performed according to the manufacturer's instructions (Roche Applied Science).

[0306] The bioinformatics analysis for mutation screening consists of five parts:

[0307] (1) GS FLX+ data processing was performed using Roche GS FLX+ software. The base-called reads were trimmed and filtered for quality, and then converted to FASTA format.

[0308] (2) Sample processing. Sequence reading begins based on specific barcode identification. The barcode sequences are trimmed and the sequence readings for each sample are saved to the database separately.

[0309] (3) Amplicon processing. Amplicon initiation is based on target-specific primer sequence recognition. The sequence reads of each amplicon are clustered using CAP3 (95% homology, 40 nucleotide overlap).

[0310] (4) Polymorphism detection. Identify all potential SNPs and INDLES in each clustered amplicon.

[0311] (5) Detection of EMS SNPs. Identification of EMS-induced SNPs. Such SNPs are expected to be present only in EMS mutant plants (EMS pool A or B). Consider six independent EMS mutants pooled together (3 in pool A, 3 in pool B), and an EMS-induced SNP is detected in either pool A or B, but not in both. A SNP is considered a true EMS-SNP if it matches the following parameters: (a) not present in A68 and A68_i124; (b) detected in either pool A or pool B.

[0312] A total of 6 putative EMS mutations (C→T or G→A) were identified, 4 of which were found in a single gene and showed very high protein BLAST homology with Arabidopsis protein vacuolar protein sorting (VPS) 13-like proteins (gi|10129653|emb|CAC08248.1|) (Table 3).

[0313] amino acid initiation Amino acid termination Blast score E value 643 785 204.91 1.9e-054 1020 1393 237.65 2.6e-064 1645 2097 256.91 4.1e-070 2142 2618 303.91 3.0e-084 2608 3384 728.78 3.7e-212 3390 3737 327.79 1.9e-091 3621 3931 273.09 5.6e-075

[0314] Table 3. Protein homology between SEQ ID NO: 3 and Arabidopsis VPS13-like protein (gi|10129653|emb|CAC08248.1|). Tera-BLASTP search for protein queries (DeCypher, TimeLogic™ standard settings).

[0315] This is a large gene, representing 34 of the 295 exons sequenced, equivalent to 11% of the total re-sequencing nucleotides. Enrichment of mutations in the Dip gene was expected through selection for the ploidysporogenesis deletion phenotype. All four ToVps13 EMS mutations were found in the Dip haplotype, and none were found in the dip haplotype. We calculated the probability of this mutation distribution across the gene sequence based on the following probabilities: The predicted size of ToVps13 is 11% of the total re-sequencing region. Due to the three haplotypes, the size of a single ToVps13 haplotype is 3.7% of the total re-sequencing region. The probability that the first EMS mutant is located in the Dip haplotype is 0.33. The probabilities that the second, third, and fourth EMS mutations are located in the same gene within the same haplotype are 0.037 × 0.037 × 0.037 = 5.1 × 10E-5. The probability of the first EMS mutation being located in the correct haplotype and the second, third, and fourth mutations being located in the same DNA region within the same haplotype is 0.33 × 5.1 × 10E⁻⁵ = 1.67E⁻⁵. Since this could also occur in other DNA regions, the probability of the entire resequencing region is 100 / 11 × 1.67 × E⁻⁵ = 1.54 × 10E⁻⁴. Therefore, the probability of this random assignment occurring is 1.54 / 10,000. Thus, the Vps13 sequence is very likely included in ploidysporogenesis.

[0316] In both LoD plants, a second EMS mutation was found in the re-sequencing regions, one in an oligopeptide transporter and the other in a putative transporter gene. In both cases, the mutation was not in the Dip haplotype, but in the dip haplotype. Therefore, we conclude that these two EMS mutations are not related to the ploid sporophyte formation phenotype. In the putative LoD3 and LoD5 plants, no EMS mutations were detected in the re-sequencing regions. These plants do not produce offspring in testcrosses (see above) and are likely female-sterile mutations rather than apomixis deletion mutations.

[0317] Example 5. Association mapping of DIP loci in various unrelated sexual and apomixis dandelions

[0318] To provide further evidence regarding the ploidy-sporophytogenesis phenotype involved in SEQ ID NO: 1, the association between sequence SEQ ID NO: 4 and ploidy-sporophytogenesis was investigated in a group of apomixis (=ploidy-sporophytogenesis) plants and a group of sexually (=diploid) plants. Both groups consisted of 13 unrelated plants, as diverse as possible in terms of geographical origin and taxonomic groups (different sections and species within the genus *Taraxacum*). Plloidity levels were determined by flow cytometry according to the methods described by Tas and Van Dijk (1999, Heredity 83:707-714). Breeding systems were identified by seed formation isolated from pollinators: apomixis isolation produced complete seed formation, while sexually isolated groups did not. Parts of SEQ ID NO: 4 were resequencing, 1–300 nt or 7–586 nt, first by Illumia paired-end sequencing and second by sequencing on a genome sequencer (GS) FLX+PLATFORM (Roche Applied Science). Using Decypher (TimeLogic) with standard settings, the nucleotide BLAST sequence for SEQ ID NO: 4 was analyzed. Table 4 shows the highest nucleotide sequence identity and the lowest E value for each plant. It is clear from this table that all apomixis carried the sequencing region of SEQ ID NO: 4, while none of the sexual organisms carried this DNA fragment. Therefore, there is the largest association imbalance between this sequence and ploidypophyte formation. If the nucleotide region is not functionally involved in ploidypophyte formation, recombination and mutagenesis will erode the association imbalance between the nucleotide region and ploidypophyte formation over time. Therefore, the perfect association between apomixis and SEQ ID NO: 4 confirms that this sequence is essential for ploidypophyte formation.

[0319]

[0320]

[0321] B. Apomixis (ploid sporophyte formation)

[0322]

[0323]

[0324] B. Apomixis (ploid sporophyte formation)

[0325]

[0326]

[0327] Table 4. Association plot between apomixis and SEQ ID NO: 4. The sequence was analyzed using nucleotide BLAST against SEQ ID NO: 4 using Decypher (TimeLogic) with standard settings. The highest nucleotide identity and lowest E value for each plant are given. Indet. indicates indeterminate. IBOT refers to the Institute of Botany, Czech Republic, geographic origin unknown.

[0328] Example 6. Expression of the DIP gene in megasporocytes of apomixis and near-isogenic ploidyte formation deletion mutants

[0329] To investigate the expression of the DIP candidate gene, RNA-seq was performed from isolated megasporocytes (MMCs) and female gametophytes (FGs) from apomixis (A68) and their syngenesis deletion line (i124). Preliminary studies clearly indicate that megasporogenesis in dandelion occurs in the buds of very young inflorescences (approximately 0.5 cm in diameter), prior to stem elongation, while the buds are still within the rosette of the plant. For later stages (female gametophytes; FGs), buds were collected when the stem was 1 cm long.

[0330] The fresh ovary was cut open and soaked in a mannitol mixture of pectinase, pectin lyase, hemicellulase, and cellulase. The ovule was then separated from the surrounding tissue by manual microdissection using a needle. The oil unit (Eppendorf) collects the separated ovules from 20 batches and immediately freezes them at -80°C until further processing. RNA was extracted from a pool of 20 ovules using an RNA isolation kit. AmbionMessageAmp was used. TM The II aRNA amplification kit linearly amplifies RNA via in vitro reverse transcription. Different pools of 20 ovules from the same genotype and tissue are considered biological replicates.

[0331] A total of 10 samples were sequenced in 6 Illumina HiSeq lanes (3 bioreplicas of A68MMC, 3 bioreplicas of FG, and 4 bioreplicas of MMC i124). Based on the samples, overlapping read pairs were merged using FLASH software. http: / / ccb.jhu.edu / software / FLASH / Merged (unfiltered) reads were performed using Trinity software. http: / / trinityrnaseq.sourceforge.net / Assembly was performed according to Trinity's "Abundance Estimation Using RSEM" protocol. http: / / trinityrnaseq.sourceforge.net / analysis / abundance_ estimation.html Transcript abundance was estimated. Then, the "Trinity Transcript Identification" protocol was used. http: / / trinityrnaseq.sourceforge.net / analysis / diff_expression_analysis.htmlIdentify differentially expressed isotypes.

[0332] More than 40 meiotic genes (e.g., Dmc1, Spo11, Rad50) were detected in the reassembled expressed genes, indicating that the correct developmental ovule stages of MMC and FG were studied. SEQ ID NO: 4 was reassembled and showed moderate expression in the apomixis A68 at both MMC and FG stages. In Table 5, expression was quantified as FPKM values ​​(fractions per kilobase per million localized reads). In the deletion mutant i124, SEQ ID NO: 4 was not expressed, but it was expressed in its ploid sporophyte-forming homolog A68. Therefore, the expression data confirm Vps13 gene deletion and expression at both MMC and FG developmental stages.

[0333] Expression and association mapping analyses conducted to date indicate that the nucleic acid molecule currently annotated as the 3 major ends of the Vps13 gene, as shown in SEQ ID NO: 4, is independently transcribed, either as a novel gene or as a differentially spliced ​​variant of the Vps13 gene, and is similar to the sporulation gene Spo2 of *Saccharomyces pombe*. The Spo2 gene encodes a 15-kDa protein consisting of 133 amino acid residues, which has been incorrectly annotated as the last exon of the *Saccharomyces pombe* Vps13 gene. In fact, the Spo2 gene is directly downstream of the Vps13 gene and is transcribed independently (Nakase et al 2008, *Molecular Biology of the Cell*. Vol. 19, 2476–2487).

[0334] Notably, the mRNA sequence of SEQ ID NO: 5 does not contain the ATG start codon and has a short open reading frame that may be translated. However, when using ribosomal profiling in budding yeast (Saccharomyces cerevisiae), Brar's laboratory (University of California, Berkeley) found that... http: / / www.unal-and-brar-labs.org / brar-sorfs Non-canonical translations of thousands of novel short peptides during meiosis have been identified. These meiosis-specific expressed short open reading frames (sORFs) lack the ATG start codon and translate into peptides shorter than 80 amino acids, thus failing to be predicted by standard gene software. sORFs reside in regions previously unknown to contain expressed sequences. sORFs can also be short alternative isoforms of proteins with known functions. The presence of these short peptides during meiosis has been confirmed by canonical methods. However, the function of these thousands of short meiosis-specific peptides remains a mystery.

[0335]

[0336] Table 5. Expression of SEQ ID NO: 4 in megasporocytes and in female gametophytes of apomixis A68 and Dip-deficient line i124. Absolute expression was measured as fragment per characteristic kilobase per million reads (FPKM). Mean and standard error were calculated. The percentage of allele-specific expression is indicated.

[0337] Example 7. Overexpression of ToDIP and Todip in Arabidopsis thaliana

[0338] Two ToDIP sequence fragments (SEQ ID NO: 11) and (SEQ ID NO: 9) with an artificial ATG start codon were cloned into vectors with a 35S promoter. Three independent Arabidopsis flower dip transformation experiments were performed using these constitutive overexpression vectors. In each experiment, 15 to 30 T0 plants were obtained per allele.

[0339] 35S::Todip overexpression transformants were indistinguishable from wild-type plants and were fully fertile. In contrast, some plants in all three experiments with 35S::ToDIP overexpression transformants were partially sterile (20% in the first experiment, and 10% in the second and third experiments).

[0340] Ovules were removed using the method described in Yadegari, R., et al. (1994) (Cell differentiation and morphogenesis are uncoupled in Arabidopsis raspberry embryos. Plant Cell, 6, 1713–1729), and megasporocyte (MMC) and female gametophyte (FG) development were studied using a Nomarski microscope. MMC and FG development appeared normal in all studied 35S::dip transformants, as in wild-type Arabidopsis. However, 35S::ToDIP plants frequently exhibited megasporocyte abnormalities, extra micronuclei near the megaspore, and FG development arrest, such as arrest at the FG1 stage, absence of vacuoles, and ruptured embryo sacs. Similar disturbances in FG development were observed in the Arabidopsis dyad mutant affecting both female and male meiosis (Ravi M et al. (2008) Gamete formation without meiosis in Arabidopsis. Nature 451: 1121–1124). Therefore, the observed abnormal MMC and FG phenotypes of 35S::ToDIP may indicate the presence of disrupted female meiosis.

[0341] These ToDIP phenotypes are dominant because they were observed in hemizygotes T0. This is consistent with the dominance of the dandelion DIP allele. In the first experiment, pollen development was also affected (extra nuclei) in some plants, but in the second and third experiments, pollen development appeared normal. At least in the second and third experiments, the phenotypic effects of the DIP constructs were female-meiosis specific, consistent with the function of DIP in dandelion.

[0342] In summary, the dandelion DIP allele was found to produce female-specific meiotic dominance in heterologous plant species. This effect was not observed with the dandelion dip allele. The Arabidopsis overexpression phenotype provides strong supporting evidence that the DIP sequence induces a ploidypophyte formation phenotype in dandelion.

[0343] Example 8. Function of the DIP gene in dandelion

[0344] To further confirm the ploid sporophyte formation function of SEQ ID NO: 4, dandelion i124 plants lacking the DIP allele were transformed with a plasmid containing SEQ ID NO: 4, which was then fused with different promoters and regulatory elements in an appropriate vector. The following promoter sequences were used:

[0345] 1. The natural dandelion promoter of SEQ ID NO: 4 (approximately 1500 bp of SEQ ID NO: 1, upstream of SEQ ID NO: 4)

[0346] 2. The promoter of the dandelion ortholog of Arabidopsis DMC1 (At3g22880) (Klimyuk VI and Jones JD 1997. AtDMC1, the Arabidopsis homologue of the yeast DMC1 gene: characterization, transposon-induced allelic variation and meiosis-associated expression. Plant J.:11:1-14). This gene has a meiosis-specific promoter.

[0347] 3. 35S promoter. This promoter leads to overexpression of SEQ ID NO: 4.

[0348] Wahler et al. (2009) published a protocol for transforming dandelion plants. Since i124 carries all other elements of apomixis, complementation of ploid sporophyte formation restores apomixis, which can be readily determined by high seed formation in the triploid plant and genetic markers in the T1 progeny. The offspring plants contain the complete maternal genome, without the segregation of maternal markers.

[0349] Example 9. Introduction of ploid sporophyte formation in sexual crops through transformation

[0350] Following the protocols of Dreni, L et al. 2011 (Plant Cell 23:2850-2863) and Dias, BBA et al. 2006 (Plant Pathology 55:187-193), diploid sexual plants of rice and lettuce were used for transformation, respectively. The same construct with the promoter and SEQ ID NO: 4, as disclosed in Example 8, was used. Triploid offspring were produced by crossing T0 ploid sporophyte-forming plants with diploid pollen donors. Triploidy was determined by root tip chromosome counting or by flow cytometry. Both are standard methods (Tas and Van Dijk 1999, Heredity 83:707-714). It was further demonstrated that ploid sporophyte formation could be identified in the offspring plant analysis by analysis of genetic markers. In addition to paternal markers, the offspring would carry the complete maternal genotype.

[0351] Example 10. Introducing ploid sporophyte formation into sexual crops via genome editing

[0352] Targeted genome editing technologies, such as CRISPR-CAS9, TALENS, and ZFN (zinc finger nucleases), are commonly used in the field to generate mutations in existing genes. This is done not only by generating knockout alleles, but also by introducing mutations encoded by so-called “repair DNA” (e.g., Doudna JA and Gersbach CA2015 Genome editing: the end of the beginning Genome Biology (2015) 201516:292, and the references cited therein).

[0353] This DNA fragment typically encodes a segment of a (target) gene sequence, in which a change is introduced that leads to alteration of gene function. Typically, such a sequence replaces the targeted gene sequence in a genome editing event through homologous recombination, thereby introducing selective mutations in the genome of a host cell (e.g., a plant cell) in a targeted manner.

[0354] This embodiment includes introducing an alteration of the dip homologue in a given plant species that results in a functional change to DIP, i.e., altering the function of the naturally occurring recessive non-ploid sporophyte formation allele through the dominant ploid sporophyte formation (DIP) allele.

[0355] Dip homologues are easily identified in many plant species. CRISPR CAS-mediated genome editing using dandelion-based "repair" plasmids can convert natural dip homologues into their DIP sisters by simply modifying SNPs and insertion / deletion mutations (indels) that conform to the differences between dandelion DIP and dip alleles. sequence list <110> MasterGene Co., Ltd. <120> ploid sporophyte formation gene <130> P29610PC00 <150> NL2015398 <151> 2015-09-04 <160> 13 <170> PatentIn version 3.5 <210> 1 <211> 23361 <212> DNA <213> Taraxacum officinale <400> 1 ttttcaccta gaagagcaca accctcttca agtgaagcac acactcgaag tgtatttccg 60 tatttcagtc cagctggaaa cagaagacca ttaatgatat tttagcaaga aagctttcat 120 tcatttgttc atcatagcag catcctactt cctacaacat tgtattttga aaaatctacc 180 aattaaatag caatatcatc caaatacaaa gcaagtttca ctgctatata ttgctaaacc 240 tctctttaat ccctagatat aaagggca aaaaagtgga acatagtac ctttactctg 300 tcatttgatg ataaatgttg tttttagtt agtatcaaac tgaaattgtg ctaatcacaa 360 tggtgttgaa tgttggcaa attctattta tttaataaat cataggtaaa acatgacat 420 aaactatt ttactctacc atttgacac cactcttgat gtatcaact gattatatga 480 caccccaccc catataatca gataatcact tccaaaaagc attackcacc accccttat 540 attttcacta actatcccaa ttgcctgtct ataattcag aattaataaa tctactcagc 600 tgattgattt tgataaata ataatgacac attgcatact gatagataaa gttctgcagc 660 acccaatca atgtcctaa aaccttaagg tgagacac agcccaagg attgagat 720 cgacatatac tctttacag ttccttaata tcaatgacaa aactgaatcg aatttaacgg 780 ggggaaaaat aaatacacca atttcataac ttaaaacacc tggttaatcc tcgcgactac 840 aaatcaaac agctgaat tttaccaga gaaaaata aaaccctagt tgatcgcatt 900 gatactgag ctacacacat ccaattacat acgaaaaat acgaatcaga agtttacgaa 960 acgaaacaag atctagtcaa taaagaaaac tgaaactaag acagatcaga aaagtggttt 1020 ctcgaacgta cctgatcgaa gtgaagatgt cagatctgat aaactggtat cttgtagtaa 1080 aagatgaatc gagagataaa aatgaagaac gaatgcagaa ttgtaatata aaatacacaa 1140 aactcaagat ttatttctgt tttatgttga aaatttgaat gaaaatgatt ctaacagaga 1200 aagggtttgg tctttccttt ctgtttctct gtcaacaaca atttagggaa ggcgtagaga 1260 gatgagagca gacagcgcac ctttgcgcgg agaatttctg ggatatatct atttatttcc 1320 cttttacctt tcttttcatt atgttttttt ttttttcttc agaaaattca aactttcttt 1380 ttcaatattc tgagtacgct aatttctagc tggcatgtag caagtcttaa tttctagctg 1440 gcatgtagca agtctttgtt gggtggtata gcatttccct ctttttaata gtaagtgaaa 1500 atgacttaat agggtaaaaa cgtttcgaac ttgtccatat ttggtcatta tcctattttt 1560 tgtacacaat acaccattaa gatttatatg accggttaag tttcgtccat tgccgatttt 1620 cataggtaat catcgtgaaa tcagtaagca acctcatgaa tcacccactc ctgccgccac 1680 aaccacagcc gaagcttcat ttgagcccct ggatgaaaat ctcttctcgg atctaacccc 1740 actctgttat ttactccaat ccacctcccc caaatcaccc ccaacattaa tctttcctcc 1800 gttctaccac caccaatcct ctcaagattc atcctttctc atcgccgtga cgttttccag 1860 caaaacccat gtccaccacc gccgccgtca aaccagtccg gaaattcccg ccatcgtgcc 1920 cggatgtaac cccgccgtca aaccagtccg gaaatcaaca acacaacccc ctgttttgca 1980 tcaccagagc ttgaatccgg gtcgtggagt tcgatttaaa ctatcttcaa ctcaaaatct 2040 ggtaacatta gcagctagtt ccaccaatca tacccctgat ttcgattcta ggtttccaaa 2100 tcacgattcc tcatacatga acaactggtc taatcaagaa gaagacaacg acgatgggtg 2160 tcttgatgac tctacaatta atgaaacctc gattcatccg ggtaccaaaa tcacaagagt 2220 gtgaaaggag gaattcggat aaaatcctca attccatccc acctcgggcc gagaaagttg 2280 agcaaggctg ctcatgaaca tgaagtcgat gaggatggtg aaatgtggat taaagtgcct 2340 agaaacacaa atttgtttta cgggaaccat caaaacggca atagttttaa tcctaattgg 2400 gatcaaggga agaaaagaag aaatgaaggg atatagggagg tggttcaatc gattaagcta 2460 ttaggagatg ggttcatgaa gatgctgact gatttcagga ggttgctgac tgatttcacg 2520 atgattacct attaaaaccg gtaaatggac gaaacttaac cggtcatata aaccttaatg 2580 gtgtattgtg tacaaaaaaa taggataatg accaaatgtt gacaagttcg aaacgttttt 2640 accctattaa gtcattttcc ctaatagtta actaattat cttttcggct ctaaattat 2700 ttttgttata attcttcta aatagactgg tttatctttc ttaattactg tatgaaacat 2760 ttcacaaact tattactttt taaaaatata tcatatttct caagataaaa agtaagataa 2820 aagaattaat atatctatgt gttacaaatc aaaaaatgat aaattgttgg atgccgtagt 2880 attttgccac gactaaaagt taacatatta tttctagggc attttaata ataaagaaac 2940 tgacttatat atatgtaaca tattatttgt tcatatttag tcattatata tcttttgtca 3000 tattattaaa tttgttacaa ttttttatt ttagccacta caatttgcaa ttgcaggtca 3060 tagtgaccaa atgtaaacaa ttgcaacaag ttcgatggta taatgtgaca aaaaaatata 3120 atgaataat atgttttcaa acattctttt ttatatttag ccactacaat ttgtcatctg agttggaata aatatgtttt caaacattct ttttttaaat tcaaaatgaa tgttctcgtg 3300. 3300. 3300. 3300. 3300. 3300. 3300. 3300 gacgaaatac fathers fathers ttacttccgg tagtcgtgtt tttaaaatat attcaatctt attatgta attachment ggttatta taaacata ttattatt aaatatttt attgatag acaaaaaata attachment tagctaatg aaaaaacc ataaaaaat gattaatagg ttattggtgt catatttatt ctaaacgcta aacatatatg gttctaaaaa aatgtctgat gaatgttcaa aatatggtac aaaatatgt ctcaagaact gtaacatcct taatttttag aagttatgtt ttagttaa attack tagctaa tttgttaagt tttaggataa taatgaaaga aataagaga ctcaaggtta ttttgcattt ttgacaagta gagggaccac tatgaaaaa tgtggaatt aaaagcctgg aactctcttc cagcttctaa gatcgcgaga gaaagagagc agtgaaaggg aggggcgatt ttgagagcaa accaagaaca agagctaaat tctgatcatt gaagagatta ggagtgatac taaagaggat 3900 tcaagcacaa aaaggtaaga aatcatcttc tatttacaag aaattcgttt ttgatgttga 3960 gggtagaaac ccaaaatcga aatcagatga aattgaaggt ttaagggctg gattctaacc 4020 tctttggtc attagaagtg atttctagct tttgctcacc tttgaattga ttatttgaag 4080 gttttgggga agatcaccat gaacttagaa tttgctattg gggtttgatt atatgctttg 4140 attccatccc ttagacatgt cataaagcct gtaaatgagt gataagaagt attgctaaaa 4200 ggttaggact tgtagaccaa agtttggaag ttgcaagaaa ttggaaaaga gaagctgcct 4260 agactgcgga cacgggtcgt gtcccacagg ctggaatcat tttgagtttc gatttcttta 4320 tttttttatc caattccaat aattccaagt gcaataagct tagatatgat gataaatgag 4380 ttgcacaaaa gttgaatata tttggaaata cccaagggta ttttggtcat tttaatgacc 4440 tttagactat tgttttgtag aaaacaaacg aagggagcaa taggaaggag gtctagctgg 4500 gggttgcaag tatctgaaaa cctctacaac ctaaggtgag tacgtgtgat tcatttcccc 4560 tctttgaggg tatgtttatg tgataggaat ctatgtatgc aaagtaggtt gttattgtaa 4620 agggaaatgt tagtctttct agttgcaggt actatgttgc tatgatgatg tatagaatga 4680 tatgaagata tgaatgctag aaagaatgtg tacgaaatga atgagaatgt tgtccggact 4740 atgagagcct cgcaaggggg cctagaggat tctattggga cggtacctcc cttcgcgaaa 4800 tagaactcct aatgtgtatg agagccttgg tggactgatt tgttcttttc accttggcct 4860 agacttgggg tggttcctaa taggacatag accctaagtg taaaggaatg atcattgaag 4920 gggatcccga tatattgtgg ataatggcac aagaattaaa tagaattgat gatgttataa 4980 attgaattta aatgaatgtt ttagattata tccgcgtatc tcaccagacc ttgtctgata 5040 tgttgttttg tgccatgtat tccaggtagt agttctcatg cataggagat ttggatattt 5100 tgaagttata ttagcaagtg gaatgaagac ccggttgtta tccattctcc aattcagttt 5160 atactatggt tgttttgat acaagttgtt aaatacccca aacttaatta catcttttga 5220 aagtaaatgt ttatgatatg taatagttt taaacttgat atgttttctt ggttaaacca 5280 tagtcataaa aaaatatttt aaaaagttat gaaaggttgg ggtgtttcaa gaacagtccg 5340 acagcggccc tcctccgacg gtgtttccg gcgagttcct ttctactccg gctagcttcc 5400 ccattattgg attaacctat tactgcacct ctccctccgt ctcgaacatg tcacattgtt 5460 ctcggttctc tctgccatga cataacaac aaacatatgt atagttgtgt gtaggcattg 5520 agtttaaacg gaggagaaaa tatgccatgg ccagctcacc cgctgtcgcc atcccagcca 5580 gcattagcc agccgtcgct cacatccacc gaaaagctga tttccacatc gtcgctcccg 5640 gccgtcacct acacgtaggt tcgtaagtttt ccaatatctt agttgcgttt tcggatggct 5700 gataaaaatg catccgttga caaggaagat attgtgattt gtgagtaaac tattttgctg 5760 gaggaaggtg agttagatca catggcgagc atggtaagat gtgaggtggg taaatttcca 5820 ttcaattact tggggcttcc cataggggca aacatgaaat tatctaaaca ttggaacatt 5880 atagttgata agtttgagaa aagattatca aactggaaag ctgagaattt gtcctttagg 5940 gggcggctga ctttgataaa atcggtgatg gggagtcttc cgttgtttta tttctcgatg 6000 tttagggctc cgaagaaagt ggtcgataag ctagaaggga taagaagaag gttcttatgg 6060 ggcggtaaga agtcggaaaa aaagattcac tgggtttcgt ggggaaggt gataaaatca 6120 6180 aaatggtttt ggcggctcaa aactgaaagg gacagcctgt gggttagatg tgtcacggct 6240 tgtcataata tcaaacttat tgatgggaaa cgggtggcta aagcttcctt gaagggagata 6300 tggtggaaca tcatgagctg tgttgaagag ttaaaaacga aaagaatttc tgtggagtca 6360 aagttagtaa ggcaactggg aaacggcaaa cacacgcatt tttggaagga tagatggtta 6420 cacaacaaag ttttaaaaga tgaccttccg gagttgtaca aaatagaagg ggacaaaaat 6480 tgtatggtaa accaaagact ggtttgggac aacaacgaga aaatattcaa gcaagcttgg 6540 gactggaaaa gaccgatcag gaggaaga gaaccaaag aactcgaaac tttaataatt 6600 ttgacaaatg ggatacaatt aaaggaaata gaagataatt ggagatggaa ggagggatcg 6660 gacgggaaat tttcggtggg aaattgagg aaactctttg cttatcagga gcaggccgag 6720 gttgatggtg gattcgattg gatcaattgg gtccccttga aggtgaattg cttgcctgg 6780 aggttgaaac aagagcgagt tcctgtgatg tgtaaattag caaagagagg ggtatatgtg 6840 gaatccaaaa tatgtaaaat ctgtcaacag gaagaagagg aaaccgaaca tgcttttttc 6900 aggtgtgcac atgcgcatca ggtgtgggac tggttcaaga tgtggtcggg tctgatgcgg 6960 gaaatccctc taaacttcag atccatggag gcggagatca aggctggtgc tggtgacaaa 7020 aaatcggtga aactaggaat ggctttggct tatgtgatgc tgtggactat ctggaaaatc 7080 aggaatggtg cagtcttcaa caacagaaaa gcgagggcaa tgaacacgac ggatgaggtt 7140 caattaatct cctttaattg gataaaaaat agaagtagat gcaattggat caaatggtgt 7200 gattggggtg ttaagccttg catgaactgt aatatgtagc tgtactcttt gttctctctag 7260 catcttgcta gaagttttgt tttgctttta tataccaac gccgttcaaa aaaaaaaaaa 7320 ctattttgct gagtttttga cgaaatgagt gttattttta aaacattttt cgctttgtta 7380 ttttaaaaca tttaccaaaaa gggtagtaaa catatttaca aaaagagtgt ttttaaaat 7440 taacaaccag tgcgtaaaac acccccggtt ccttcttcct tgcgtcgact tcagaaatcc 7500 atctcctacc ggccgacttc accagcggcg gctccctctc catctccggc gccctctcca 7560 gcgaccgctc ccttccacc ggcgacgtct tctccagagc cggcgaccca caggtagggt 7620 tcttgtttct ttaggggacg tctttgaatt tgcctgaaat tgcctgaaac aacagtgaga 7680 tttcaagtca aatagttgtt gacctgaaat tgctttttgt gtttcaaatg gaaaggaata 7740 ggtgttgatt taaatcaata gcttgcatag cctacagtgg gattcaagtc acaggtaggg 7800 ttctttcttt agaaattgcc tgaaacacct atttgaccta ttcctttcca ttcgaaattg 7860 cttgcatagc ctacagtggg attcaagtca aagctattgt gatttaaata atataaaggt 7920 cgtacagcag acaacttgta aaggaacatg ctctgatcat gatccatccc ctgtcagtca 7980 tttgatcctg taggcccc atgctaattg gcactgaggt gactgtcttc aggtaacaca 8040 ttattgtgta ttacttcaat taatcatgtc cattttatgt tttcatctat gttaatcatt 8100 tcccgctatt tattattccg tgcttactgt actaacattg cagttttata attggtatag 8160 gaaatgtttg aaggtttagt acggcagctg atattaggtt atcttggcca atatattaaa 8220 gatatacaca gagaacaact caagatcaca ctgtggaatg gtaagtcacc actctctatc 8280 atattcatga caaatagttc acgtagtttg ttattcatgg cttgctactt ttcattaata 8340 tcagcacaat atatttactt ttagtcattt acatttatgt ggtttttaag ggtataacag 8400 cataagctga atactcatat cactcacatt tatggtgaac accaatatgc catttcttct 8460 atgccccctt cctctttatt tatttatttg taacttggtc tgaggcatgt tttctcgtct 8520 ttgatctctt tatgactcct tcaatatttc gccatattat atattcgtat caatatccag 8580 tgaatttctt atatctgatg tcatttttat gtgaagaaca tactgatcac aatctatgct 8640 attgaaattt tgctctcatt ttaatgcaga ggaagtgttt ttggaaaatg tggagttaat 8700 tctggaagct tttgattatc ttgaacttcc gtttgctcta aagcaaggtg atcaataata 8760 tccgtaaact ttcaagtttc aatactttgt ttatattgtt atttttgatg tggtttctca 8820 tttttgggtt ggttgtcatt atttgacaca tgtgagtggt tgataggacg ggttgggagg 8880 ctaagcatta gaattccttg gaaaaagctt ggttgggatc ctattaat aatcttagag 8940 gatatattag tttgtgcctc tcaacgtgag gatgaagagg taagtagtat ctcatttcaa 9000 ggtataataa gtattgtgca tctccattat attatataac ttttcactaa tgtgaacagt 9060 tatcctttt gtagtggagt gttgatgatg ttgaaagacg agaatttct ggaaaaaagg 9120 ccaaacttgc tgcagcagaa ttggcaaagt tatcgcaacg tgtatgtggt gagtgactat 9180 tttccattc tatgtttcat tgtataact cctgttattc attgtggact atgataat 9240 gcattaattt aaagaagttg aagactacca ttgctggtag gctttgtatt ttagtgaaat 9300 agtaactatg ttttgccagt cctatatgta aattaaacag cttttcagc tttcttcact 9360 gttggcagtc ttttaatgc tcagatcttc cagagtttta taagatcacc ttttactat 9420 actttttat tgttgtaggg aaaagtcagt gttagttaat catagctata tacttatatt 9480 cactgttaat gcatttaatt ctttcatccc aaaagatgca ttttttgac acaatacata 9540 ataaaaatgt tggttattgg tcaaatataa aaaatattca tttatgaggt caaaataaag 9600 atttttattt tttgaaccag aatcaaaatt tcggatatta aataacaata gttaatggga 9660 aaacacatta agaaatgttg cttgtcgtat atgtatatgg atatatggga atatgggatc 9720 agtgacattc ttggatacac cctttttttt ttttgcataa aaccagtaaa atctttttca 9780 cattatttgt gctaatatgt tgtctttttc tgcagataat cagactggga aatcatttat 9840 gtcatacatt actgccaagg tagtgagatt tcatgttgct tgaacttgtt aaaaatgtca 9900 aaaagtcata tgcctgtttt tccgttttac ttatcatgta atcttagttt tcttaatctt 9960 catttgaaat caggtttgat tttacgtgta aatggtctga tagttgtgtc actagtattt 10020 atttatctat ttattcccat ccagattatt gatggcattc aagtcaccat caggaatgtc 10080 catatcgtat atagagatat ttcaaatgaa aaatcccaaa ctgtatttgg tgtgaagttg 10140 gctagtttga ctgcaatgaa gcaaaactat gctgggtata cctctttttc ttactctaca 10200 gtcaatcaca tacttaaatt ttcacaaaac tcttaatgac tttgatatat gtatatatat 10260 cttcagggta ttaagtggaa aggtgagagt tggccaagta aacaaaattg ttgagataca 10320 aggtttggaa atatactgta aaacctttca tggatcttca acagacatcc atactgaaaa 10380 tggtgaagac tccatggcaa tggtggctgc aagttatgat aatgatgaac atgctcactt 10440 gttggcccca gtcaacgtat ctgcttctct ttcggtatgc atcatgcact tgtaccatct 10500 cataacaaat tttctagatt tttattgtta tgttgtactt tgcaggtgaa taggtctgga 10560 aggctggaga ataatgcagc acaatactcg gttgatattg agttgtctgg cttggttagt 10620 tagttcctta acttttctaa tcataattaa ttattaagtt attatgtcag attatctaac 10680 atgttagtat gtcaccattt ttaggtattg tccctagatg aagatcagtt gcagcaaata 10740 ctgtatctat atgaatatct atgcacatgt cggctaagag agaagtgagt ggattgcttt 10800 ttcaccacct tatcctttag aattacattt tcattatcct tttattgttt ttttaagata 10860 tggacgatat cgtccttggg ggaaacctat atcagagaga caattgggat ggcagataca 10920 gtggtggcat tatgctcaac actctgtgtt atctgatgtt cgtaaaagac tgaagaaaac 10980 ttcatggaaa taccttggag aacgtctgta agtttaattc tttatttaat tttaacatat 11040 acctgcaagt tttttatgat gaaatcttga aatctgttac atgtacagag gcagaaaacg acggtatgta aatctgtaca aattgaaact cgaatgtctt cgaaaagaac aggtgagttt tcatgattta aacatcaagg tgtatatgat ttaaacatta ttttaatta tgaattatga tcagcctttg gatgatgaaa ttgtaatgga gttggaccaa atggagaaag tgtctgatat 11340. ttgagttaca gatctgctgc tgagatga cttcaggtat gattagatc attgctattg gtatctataa cacatgcatt tatttaactt atacctttta ttaaagacta tgaaaatgta ggagttcttg gtggattcac cttctggtat tggaggtagt gaagtgaata ctaccattga caagtcaatg gatgatgacc aaacatctgg caaaccgcaa ggatggttga aatggctgtc ccgtggtatg ctaggtgctg gaggtacaga cgattccagc cagttttctg gtgttgtttc agatgaagta atcaaggtaa ggtgaaatga tttcaattcc aattgaatta attachment attachment gatacaggat attachment caacaaagtt tcatcctgct ccttcccctg tcttggatgc ttctggaact gatagggttc ttttgacctc catcaaatgc 11760 tctatacatc aaatttgc aacacttcgc aataagtatg acgtattaat taattaatta 11820 aatatctaat tatctaatta tctttatagt tttgtaagac ttacttcttt tttctaggaa 11880 gttggatcga gctattggtg aagtggtttt tgaggggaat gttgtggagt gcatgatttg 11940 ggaggaatct gctgttgtta ctgcatcaat caattctgta gagatgatta atccattaaa 12000 caatcaagcc atttactta ttaaaagggt gtgtgttctt tttgttttg ttacatggaa 12060 agtcttgagt tatttatttt tgaatcttt tctatttctt tcaggtcatc tctgaggaga 12120 gttttcttga agaggagaaa ccgtctttaa atatccaagc ttacattcca caagcaaatc 12180 gtgagggtga cttgacattg aaggtacccc accttgcatt ttttgaaaaa tattttcttg 12240 tgctataaat ttctgataat tacagtgact cttaaacttc atttaggtt ttgcttgagc 12300 cgattgaagt gacgtgtgat ccaacatatc ttgttaattt catggagcta tatactgtgt 12360 tgggttccta tacctctcat gaagaaaggg tcagttttta catttctcaa agtagatttc 12420 ttttaagaac cttctggaag atgcagattg gaagttatat ctagatagtg aattcaatgt 12480 ttctcatata aaaacactga acatacttta ccttatgaga ttttgaataa caacttccat 12540 tagtttaatg ttacaggtcc taaactcgct taatgggata aatgatgtaa agtcacgtct 12600 aatatccaaa gccaagtatg tcatctacta tacctttagg gtctagtaat gggattgata 12660 gcaacttcat gtttaaagct atatctatct cttctcttct tttgttttgc aggtatattt 12720 tgtcaggccg gaagagaatg atgtgggata ttagtttgat aaacatcaag ataaatattc 12780 catgggagaa tgggaactca gagatgcata aattggtaat ttaatttctt attacatggt 12840 ctacaaactc acatatttg cattaacttt aaacaaaaac atccaggtac ttgaattaac 12900 agctgtcacc tttgcatcca agcgcgatat cggctctttt gcaccagata tcaatgtacc 12960 atctcaattc atgaggaatc tgattgatga caattcttca aacgagcttc tagaaggaac 13020 tcacattcaa gatctgtacg atctcttgga aatcaaaata atcgacttcc aggtcagcat 13080 atcccatcat ctctcatctt tcacatcact gatgtttgac tttgactttg actttgactt 13140 cctgcagata aacttgtttg tgcctttcta tccatatact tttcctatct tggagaaact 13200 caatgcctct tctgctttat catggtgcat tgttcaggac gaatctttac taaaagcact 13260 agaggtaggt aaacggtaaa tctcacattt actaattact gttgattttt ttggcacact 13320 ttttaccatc aataattggt ttcaggttta tgtattggtg gctacccttt tggcacatgt 13380 gtcaccatca ataattggct cattttaga actagttgaa agcatgaaca tgctgcatca 13440 tacttcacaa ttgggcgcca catcagcaac ctcgtcaatt gaaccaagga actctagcag 13500 tatctctgtt attgctaatt tggagtctgc tagcatcatt gttgaccttg aaaatggctt 13560 agaagctagc tgcacactaa ctgtgtctct tcaggacttg gatatgagga tgggtagtat 13620 gaaatccaca caatctttct ggatatgtac aagggattta aaagtaactt ctcggttgtt 13680 ggaaagtggt gatgacctgg accttataat atgtctgcct caaagtactt cacctaatga 13740 tgggtgtctt gtgctacatt atgatggtaa tttgagcata tgtttgagtg atttggatct 13800 tcattgctat ccacatattg ttggattgct ggttgagttt tctggaaagc tatccacata 13860 tagtccttca aatgccaaaa atcaagattt tgtggacagt aacagtaaca ctatactctc 13920 agattcttat atagacttcc agaggtttgg ttgttccaaa atttcagtta atcattaccc 13980 ttttgttaca atatacaatg acagatctct tcttaacctc gatacttcac ttattaacat 14040 caagaaggtt cataagacaa atagttcaaa gttgagggca aaaaaggata atcatcaagt 14100 aggtgctcta gttgtaatga atcttgatct caacagcatt agactacatc ttcatgactc 14160 ttcatccatt gttgcatctg ttacacttcc tgtttccaaa tcctcttttg ctatccatga 14220 gaactttctg gatgtgttat tttcaactga gggattgagt ctttcatcac agtggtatcc 14280 tcagactcta caagactctc tgtggggccc tgcttcactg aatctttctc cagtgatcaa 14340 cattcgtgtg agaaaaggga accatggaat cgaactggat tttagtgttc aaaatgtttc 14400 ctgcatattg ccaactgagt tcctggctgc actcattggt tacttctcat tgcctgattg 14460 gagctactca aatccaaatg agtcatcacc tactactacc aacaccaata ccaataacaa 14520 cagcatcagt ttcacttaca agtttgagat attggactcg gttttattta cacctgtggc 14580 taatcctgat cacgagttta taaagcttaa tattccacag atgtactgca cattcattga 14640 tagcattgat tcagacactc tgttgaaaga aatcccttta gagtgctctg ttccagttgg 14700 atttactgga aatcagaatt attgtctgaa tgtatttggg agggatttat ctctacatca 14760 tattatttgt cgaaaagaca atgcttctga agtgacaagt gtcagcttga tcgcgcctt 14820 tagtggcgat atatggatca caataccata tgaatctaat tcttcttatg caacatgtat 14880 catgtcaagg gttagtaaat gtcagtttac tgttgaaggt aaatattttc taccaaactc 14940 tgagttttgt tcatgtttga tacttgatta atgcaatggt gtttcatttt tccaggaaga 15000 gaaatacttg gctgcattgg agcattacag gatgttgtcg accaattctc atctgttggc 15060 aatctatcca cttgtttcac ttctgatgtc tcagagtttc tcaatttaaa agaaaattat 15120 gtggttccag ttcccattga atcttcaact gtcagcttta cagaaatcag atgctctgtt 15180 caatccatgt cagtagaact ttactctgac aagatgaatg gtaacagact tatcgccaaa 15240 tccgacatga aattcgcatg tggaatatca atgaaaacag acaagcctct ttctcttgat 15300 atatcattca cttgtttcac actttcctca cttcttacct ctgttgtctt actggaatgc 15360 acatcatgta ccaaaaatgt accggttctc aacatgcagt tcttgatgtc agatgatggt 15420 aaaaaccacc tgcgattctc ccttccttgt gttaacattt ggttgttctt gtctgaatgg 15480 agtcaagttg ttgacctggt caattcttgc tgtgaacctg caatccagaa tgaggaaccg 15540 gaaaaatcca catcagctcc ggtttctcgt gttgatactg cagaaaactc acctcaatcc 15600 accactgtct ccagctatcc atctttagaa gatcggtttt cattaaccgt aaagtcggat 15660 cttattggtg taaaaatccg tattcctgtt caagtttctg gagaagtagt taaatacttt 15720 ggggccccac aagttcgaga gcagagttta gttacaggaa gagatcacgg aagttttctg 15780 tttatttatc tacaaagtag gtgcactgag gtgaacatga agggcgaaac ggtaaatttg 15840 aagtcgaatc tggggaaagc aatgggaaca gttgaactgt ttcagaacaa gagtgtccat tcttggcccc tttttcagtt actggaaatc ccgaatctgg 15960. tcttggcccc atggaccgca tgcatttaaa gacagagatt cactgtgata atcttgatgt ttggctctca catcatgcat tttacttctg gcaaactatg ctgtttatgt ttccggaagg ttccggatcc 16140. gaatctcctc aacttccagt tggcagtgtc aatttcagat tccacttacg aaaactctcc attcttttaa cagatgaaaa ggtatgtaaa ctgtaaactg taagtaaac tgttttctca 16260. gtttttaatt agtggagctc caatggtcct ttattggaga ttctcatggg aagtttgtta tttcacgga tcataactgc aaacatgatg gaagggtcaa tcgacagtga ccttcaagtc aactacaaca acatccataa agtcctatgg gagccttttc tcgaaccatg gagttcca atacctta gaagacaaca aggaaaaagc accctcgaga acccccccccccccccccccccccccccccccc aaacctccc atcaatgtca ccgaatcctt catcgaggtc gctttcagaa cattcgacat gatcaaagac gcctcagatc tcataagtct taacattctt cctgaaaaca attcaagatt gacaaaaccc 16620 catacaaatg aaaacacact cgcgaataga tacgcccctt acacactcga aaacttgaca 16680 tcacttcctc tggtatttta catctcaaaa ggtgaggggt tcaatatgac atcattgaaa 16740 gatggaaaac acgtgcagcc aggctgttca tatcccgtga atattgataa caacccggaa 16800 gaacaaacgt ttgggtttag gcctagtcat tctactgaca atctaggtgg cgacatgcaa 16860 ttcgctgatg ctcaacatca ttatatagtt gttcaacttg aaggaacttc tacactttcg 16920 gctccagttt ctattgatct tgttggtgtc agcttctttg aggtggattt ttctactaat 16980 ataggacttg atgtttctaa aggtggctat gttgttcctg tagtaatcga tgtttcagtc 17040 caacggtaca caaagctagt ccgcttgtac tcaacggtac gtttatttct ttccactcct 17100 aatatatata tacatatata tatatatata tattgatgtg atgacgtggc aggtcatact 17160 gacgaatgca acatcaatgc catttgaagt acggttcgat atcccatttg gtgtatctcc 17220 caagattcta gaccctgtat accccggcca tgagtttcct cttcctctac atttagcaga 17280 atcagggcga ataagatggc gtccattagg aagcacttac ctatggagtg aagcttacag 17340 catttcaaat attctctcaa atgaaagcaa gattggacat ttgcggtcgt ttgtctgtta 17400 tccttctctt ccaagtagcg acccctttcg gtgttgtgtg tccgttcatg atgtgtgttt 17460 gccatctgct ggaagggtaa taaacaagag gggatcatct tcatctctgt ataatattaa 17520 tactcatgat catggtgaca agatagagaa ccaggatctc tcgaataaga ggtgtattca 17580 cttgattact ttgagtaatc ctttgatagt gaaaaactat ttgccagttg aagtgtctgt 17640 ggtgattgag agtggtggag tttcacgtag catgctgctt tcagaggttg ttatcagtta 17700 tcttattata aactttgaca ttttcttaat gtattgtttt ataacatatg caggttgaaa 17760 ctttttttta tcatattgat tcttcacatg atctttcgtt aacttttgag atacacgggt 17820 ttagaccttc tcttttgaag tttccacgtg ctgaaaagtt cagtggaatt gctaaattca 17880 gtggaacaaa gttttcttct tcagaatcca tcaactttgc tgctgataat tctaaaggtg 17940 ctgcacactt acttttctca tctccctttt tctctcttac attaaaagaa gctgatgtgt 18000 ctttgtttag gtccattata tgtgacaatg gaaaaggtga tggatgcatt ctctggggct 18060 cgtgaaatct gcatatttgt gcccttctta ttgtacaact gctgtggttt tccccttacg 18120 attgcaaatt cgactaatga tctcacaatg cgtgacactt tgccttcatg ttatgatttg 18180 gatgaagagg acccgttttt gggcaaaaaa gatggtctaa gccttttgtt ttccaaccaa 18240 gtttccaata atgatcctat gagtaatctt gtttccacta gaaaggaatc tttcacctca 18300 tctggatcaa ccaaaaaaaa catcggcaca caaaaaccgt ctttacacga tcaggaaaaa 18360 agtcaacttt cgcatagtca acaacttgac tttgacgaaa caactcgcaa aaaagtcaac 18420 ttccgcatgt attctcccga ccctaacatc gcttcgagtg agatcatggt gagagttagt 18480 cgatgccggt ctgatgctga catggcaagt acctcagact acacgtggtc cagtcaattc 18540 ttcctggtcc caccagccgg ttcaaccaca gtccttgtcc cccgatcatc aaccaatgct 18600 tcatatgtat tatcagttgc ctctagtgct atttccgggc catattctgg aaggacaagg attatcaatt tccagcctag attackcgttatc gcagtcagga cttgtgctat aggcagaaag gttccgactt tatataccat ttgaaagcag gacaacactc ccatatccat tggacagaca taacaaggta catcaatcat tttcccctaa ctttctgtta ccaaattcac atgcttatgt ggcagttgca taaaataaaa cagggagtta ctggtgtctg ttcgtttcga tgagccaggg tggcagtggt caggctgctt ttttccagaa catctaggtg atacacagct gaagatgaga aactatgtta gtggtgcagt tagtatggtt cgtgtagagg tccaaaatgc tgacgatgca atcagagatg acaagattgt tgggaaccct catggtgaat ctgggacaaa tttgattctt ttgtctgatg atgatactgg ctttatgcca tacagggttg ataatttctc aaaggaggtt agttacttac tccattcttc atttctaatt ctttccattt gatctatgat ataagtttgg tgaaatgttg ccagaggtta cgtatttacc aacaaaaatg tgaagcattt gagactgtta tacattcata tacgtcttgc ccttacgcat gggatgaacc ctcctaccca catcgtttaa ctgttgaggt atgcatttga tttaaaaaaca gttggctatt aaattattca 19380 tatattcatt ctgttgttaa atctccaagg tgtttgctga gagggtagta gggtcttaca 19440 ctctagatga tgctaaagag tacaaacctg tggttttacc ttcaacctcc gaggtaaatt 19500 tgtaattttt acaaacatgg atgggttaaa tctgtacta actactcttt catgatttct 19560 tttggtagaa acctgaaaga aggttgctaa tatctgtcca tgctgaagga gcattgaagg 19620 tcctgagcat cattgattca agctatcata tatttgatga tgtaaaaatc ccacgttctc 19680 ctcggttaac tgagaaaaga gaatacgacc aaaaacaaga aagttcactt ctttatcaag 19740 aaagagtatc gatttccatt ccattcattg ggatttctgt tatgagctct caaccacagg 19800 taccattact ttatacaaat atcttgtcta tatgtggtga atttcaccat tcttataaaa 19860 tgtttttgt aggagttgct ttttgcatgt gcaaggaaca caatgattga cctggttcaa 19920 agtctggacc agcaaaatct ttccttgaaa atctcctctt tacagattga taaccaactg 19980 ccaccacac cctaccccgt tatatct ttcgaccatg agtataaaca cccccaacc 20040 tctcagataa aaaaaaaya tgagctgtg ttctcattgg ctgcagcaaa atggaggaat 20100 aaagatagag ctttgctttc attcgagcat aaacttaa ggtacatgat tataacaata 20160 tgtcaataaa tagacagtca gaaattagtc aattactgag aattttggtt gttgcagaat 20220 ggcagatttc catcttgagc ttgaacagga tgtgatttta agtctgttg atttctccaa 20280 ggcagtatcc tcaggttcc atagcagagg aatgccacat atggattcag tcgtgcatcc 20340 tctttcctca aacttgagtg gaataaagc aactaaattg gctgaaaga ctgaaattga 20400 gggtgaaagt ttcccttt taccatcaat agtgccaatt ggtgcaccat ggcagaaaat 20460 atatctcctg gcaagaaaac agaagaaaat attgtggaa cttcttgaag tggccccat 20520 cactttaacc ctaagtaag acgatcgaca ctcagagggg catatgtgta aatgtgttga 20580 attggtcttt ttttcctgca gcttttcgag cagtccatgg atgctagga atggaatact 20640 tacatcagga gatatctta tccatgtaag tgtgaccatt ggaataggt tactgatagt 20700 agtaaagagt aatttgtatg ggattgcaga gaggtctgat ggctcttgct gacgtggagg 20760 gagcacggat ccatctaagg cggttaacaa tctctcatca gttggccagc ttggaatcca 20820 tacgagagat cttaatcata cattataccc gccaacttct ccacgaaatg tacaaggtta 20880 gattacatct ctacatatat aaatcttcaa ttgtgattat ttttcttttt cttcttcagg 20940 tatttggttc agctggggta ataggcaacc ccatgggttt cgcaagaagt gtaggacttg 21000 gcatccgaga cttcctctca gttccagcca gaagcttcat gcaggtaata aataaataat 21060 taatattaa atccaagtta attagtaatg tcaattttaa ttatatcgca gagccccgca 21120 ggacttatca cggggatggc acagggaact acaagtcttc taagcaatac ggtttacgcc 21180 ataagtgatg ctgccaccca agtcagtaga gccgcacaca aggtatatat atatacat 21240 tatattatac ttttctcaat tctaagattt catcatttac ttgaagggca ttgttgcatt 21300 tacaatggac gaccccccat ctgcagcaga aatgggcaaa ggtgtaataa atgaagtttt 21360 ggaggggctg actggtcttc tccaatcacc aataagagga gctgaaaaac acgggcttcc 21420 aggagtcctt tcaggtatag cactaggagt aacgggtcta gtggcaaggc cagccgccag 21480 catactggaa gtaacagaaa aaaccgcccg cagcataaga aaccgaagca aactctacca 21540 catccgcctc cgggtccgcc tcccgagacc gctaaccccc aaccaccacc cgttaaaacc 21600 ctactcgtgg gaccaagcag tcggcctctc cgtcctcacc aataccaatt ccaattccga 21660 ttccgacctc aaagacgaaa ccctcgtcct ctccaaatcc ctcaaacaaa agggcaaatt 21720 cgtcattatc acccaacggt tactcctcat tgttacctcc tcgagcctaa cgaatttagg 21780 tcaacccaat ttcaaaggcg tccctgcgga ccccgattgg gtggttgaag ccgagataac 21840 gttggatagt gtgatacacg tggatgttga tggagaggtg gtgcatattg tcgggagtag 21900 ttctgatgtg gtggttagac agaatgttgg tgggaagcag cggtggtata atccgttgcc 21960 gctgtttcag acgaatttgg agtgtttagg gaaggaggag gcgggggagt tgttgaaggt 22020 gttgttggtg acgattgaga gagggaagga gagagggtgg ggccgggggt gtgtgtaccg 22080 tctgcatcag agtaatgtta ggtgatgtat attttttc tacatataaa gttactatag 22140 gagaaaaagg actgatatt atattaca tacctgaaac aaggaaacgt ttctttcaa 22200 aattttggct gtattat ttgtcgacc atgttgggct aaatggcca atttattact 22260 tatgacatgg ttaaaaaata ttggtgtctt gttttgtata attacaattt atattagtat 22320 cgatgcaatg windowtgta gaagcgcta ccgtataaaaacatagt catgaggtta 22380 cacctagtg gggtcaaggg aaaaaaaca ttttaaacgt tttcagat tgcattcca 22440 gccgcctata actaccattc taatcttat ccgacatctt aacacgatat agggatttt 22500 actttacttg tcaaacttg aagttgacag ttgatata tccagtcctc tttccaata 22560 gtaacttttc cagccgccca taactatcat tctttctttg ttataactat gttctatcta 22620 atttagttgt ttttattat ttttgataaa atatgtttt agttgttttct gatttttgttt 22680 atatttaatt attattgttt tttgctaa aacatatcaa atttgttttg aattttgtct 22740 atgtatagtt tttatgataa attatttgta aatttagccg gttaaaatga atttttcggc 22800 caagcatgaa cattcgatg aacgaacgca catgcttttt ataggttatc tacttgagca 22860 cgtgcttcaa ttcatattt gcattcaata atccattgaa ctacttgtaa tttgctcta 22920 gttccttgcc agagtggttt agggatattt tcatttatat ggtgactttt 22980 tctaaaaag gtaaacattt ttttagatt gcaataatc gtcataagg actggaaaaa 23040 tcaaaaaaa caatgatgca ttttactaaaaaaaacaata gtgacaaaat atgaaaaaa 23100 tcagaaaat tgacgtttta cataaaaaat aacgactaaa tatagacaaa acgagaacta 23160 agttatcctt gtaagttatt atcccctgat tatattatca aacaaatttc gtatgacaca 23220 atacaggtct aaaaacctac ctcacctact tacatccgcc aaaatttac tcttatctat 23280 tcatattaaa cctatacttt gaaattttga actccatgaa ttaatgcta aaatttctaa 23340 atgatgaaa accctaattc a 23361 <210> 2 <211> 11799 <212> DNA <213> Tarrhacum officina <400> 2 atgtcagatc tgataaactg taagcaacct catgaatcac ccactcctgc cgccacaacc 60 acagccgaag cttcatttga gcccctggat gaaaatctct tctcggatct aaccccactc 120 tgttatttac tccaatccac ctcccccaaa tcacccccaa cattaatctt tcctccgttc 180 taccaccacc aatcctctca agattcatcc tttctcatcg ccgtgacgtt ttccagcaaa 240 acccatgtcc accaccgccg ccgtcaaacc agtccggaaa ttcccgccat cgtgcccgga 300 tgtaaccccg ccgtcaaacc agtccggaaa tcaacaacac aaccccctgt tttgcatcac 360 cagagcttga atccgggtcg tggagttcga tttaaactat cttcaactca aaatctggta 420 acattagcag ctagttccac caatcatacc cctgatttcg attctaggtt tccaaatcac 480 gattcctcat acatgaacaa ctggtctaat caagaagaag acaacgacga tgggtgtctt 540 gatgactcta caattaatga aacctcgatt catccgggag tgatactaaa gaggattcaa 600 gcacaaaaag gcattgagtt taaacggagg agaaaatatg ccatggccag ctcacccgct 660 gtcgccatcc cagccagcat tagcccagcc gtcgctcaca tccaccgaaa agctgatttc 720 cacatcgtcg ctcccggccg tcacctac gtaggttcgg ctccgaagaa agtggtcgat 780 aagctagaag ggataagaag aaggttctta tggggcggta agagtcgga aaaaagatt 840 cactgggttt cgtgggagaa ggtgataaaa tcaaagata aagaaggtttt aggggtgaat 900 ggattgagca gcatgaatat ggccttgcta gtgaaatggt tttggcggct caaaactgaa 960 agggacagcc tgtgggttag atgtgtcacg gcttgtcata atcaact tattgatggg 1020 aaacgggtgg ctaaagcttc cttgaaagga gtatggtgga acatcatgag ctgtgttgaa 1080 gagttaaaaa cgaaggaat ttctgtggag tcaagttag taaggcact gggaaacggc 1140 aaacacacgc attttggaa ggatagatgg ttacacaca aagttttaaa agatgacctt 1200 ccggagttgt acaaataga agggacaa aattgtatgg taaaccaag actggtttgg 1260 gaacaacg agaaatatt caagcagct tgggactga aaagaccgat caggagagga 1320 aggaaacca aagaactcga aactttaata attttgacaa atgggataca atttaaggaa 1380 atagagata attggagatg gaaggaggga tcggacggga aattttcggt gggaaaattg 1440 aggaaactct ttgcttca ggagcaggcc gaggttgatg gtggattcga ttgatcaat 1500 tgggtcccct tgaggaaga agaggaacc gaacatgctt tttcaggtg tgcacatgcg 1560 catcaggtgt gggactggtt caatgtgg tcggtctga tgcggat ccctctaaac 1620 ttcagatcca tggaggcgga gatcaggct ggtgctggtg acaaaaatc ggtgaaacta 1680 ggaatggctt tggcttgt gatgctgtgg actatctgga aaatcaggaa tggtgcagtc 1740 ttcaacaca gaaaagcgag ggcaatgaac acgacggatg agaatccat ctcctaccgg 1800 ccgactcac cagcggcggc tccctcca tctccggcgc ccttccagc gaccgctccc 1860 tctccaccgg cgacgtctc tccagaccg gcgacccaca gcccatgct aattggcact 1920 gaggaaatgt ttgaaggttt agtacggcag ctgatattag gttatctttgg ccaatatatt 1980 aagatatac acagagaa actcagatc acactgtgga atgaggagt gttttggaa 2040 aatgtggagt taattctgga agctttgat tatcttgaac tccgtttgc tctaaagcaa 2100 ggacggttg ggaggctaag cattagaatt ccttggaaaa agctggttg ggatcctatt 2160 ataataatct tagaggatat attagttgtgt gcctctcaac gtgaggatga agagtggagt 2220 gttgatgatg ttgaaagacg agaatttct ggaaaaaagg ccaaacttgc tgcagcagaa 2280 ttggcaaagt tatcgcaacg tgtatgtgat aatcagactg ggaaatcatt tatgtcatac 2340 attactgcca agattattga tggcattcaa gtcaccatca ggaatgtcca tatcgtatat 2400 agagatattt caaatgaaaa atcccaaact gtatttggtg tgaagttggc tagtttgact 2460 gcaatgaagc aaaactatgc tggggtatta agtggaaagg tgagagttgg ccaagtaaac 2520 aaaattgttg agatacaagg tttggaaata tactgtaaaa cctttcatgg atcttcaaca 2580 gacatccata ctgaaaatgg tgaagactcc atggcaatgg tggctgcaag ttatgataat 2640 gatgaacatg ctcacttgtt ggccccagtc aacgtatctg cttctctttc ggtgaatagg 2700 tctggaaggc tggagaataa tgcagcacaa tactcggttg atattgagtt gtctggcttg 2760 gtattgtccc tagatgaaga tcagttgcag caaatactgt atctatatga atatctatgc 2820 acatgtcggc taagagagaa atatggacga tatcgtcctt gggggaaacc tatatcagag 2880 agacaattgg gatggcagat acagtggtgg cattatgctc aacactctgt gttatctgat 2940 gttcgtaaaa gactgaagaa aacttcatgg aaataccttg gagaacgtct aggcagaaaa 3000 cgacggtatg taaatctgta caaattgaaa ctcgaatgtc ttcgaaaaga acagcctttg 3060 gatgatgaaa ttgtaatgga gttggaccaa atggagaaag tgtctgatat agaagatata 3120 ttgagttaca gatctgctgc tgagaatgaa cttcaggagt tcttggtgga ttcaccttct 3180 ggtattggag gtagtgaagt gaatactacc attgacaagt caatggatga tgaccaaaca 3240 tctggcaaac cgcaaggatg gttgaaatgg ctgtcccgtg gtatgctagg tgctggagaggt 3300 acagacgatt ccagccagtt ttctggtgtt gtttcagatg aagtaatcaa ggatatttat 3360 gaggcaacaa agtttcatcc tgctccttcc cctgtcttgg atgcttctgg aactgatagg 3420 gttcttttga cctccatcaa atgctctata catcaaattt ctgcaacact tcgcaataag 3480 aagttggatc gagctattgg tgaagtggtt tttgaggggga atgttgtgga gtgcatgatt 3540 tgggaggaat ctgctgttgt tactgcatca atcaattctg tagagatgat taatccatta 3600 aacaatcaag ccattttact tattaaaagg gtcatctctg aggagagttt tcttgaagag 3660 gagaaaccgt cttaaatat ccaagcttac attccacaag caaatcgtga gggtgacttg 3720 acattgaagg ttttgcttga gccgattgaa gtgacgtgtg atccaacata tcttgttaat 3780 ttcatggagc tatatactgt gttgggttcc tatacctctc atgaagaaag ggtcctaaac 3840 tcgcttaatg ggataaatga tgtaaagtca cgtctaatat ccaaagccaa gtatattttg 3900 tcaggccgga agagaatgat gtgggatatt agtttgataa acatcaagat aaatattcca 3960 tgggagaatg ggaactcaga gatgcataaa ttggtacttg aattaacagc tgtcaccttt 4020 gcatccaagc gcgatatcgg ctcttttgca ccagatatca atgtaccatc tcaattcatg 4080 aggaatctga ttgatgacaa ttcttcaaac gagcttctag aaggaactca cattcaagat 4140 ctgtacgatc tcttggaaat caaaataatc gacttccagg acgaatcttt actaaaagca 4200 ctagaggttt atgtattggt ggctaccctt ttggcacatg tgtcaccatc aataattggc 4260 tcattttag aactagttga aagcatgaac atgctgcatc atacttcaca attgggcgcc 4320 acatcagcaa cctcgtcaat tgaaccaagg aactctagca gtatctctgt tattgctaat 4380 ttggagtctg ctagcatcat tgttgacctt gaaaatggct tagaagctag ctgcacacta 4440 actgtgtctc ttcaggactt ggatatgagg atgggtagta tgaaatccac acaatctttc 4500 tggatatgta caagggattt aaaagtaact tctcggttgt tggaaagtgg tgatgacctg 4560 gaccttaataa tatgtctgcc tcaaagtact tcacctaatg atgggtgtct tgtgctacat 4620 tatgatggta atttgagcat atgtttgagt gatttggatc ttcattgcta tccacatatt 4680 gttggattgc tggttgagtt ttctggaaag ctatccacat atagtccttc aaatgccaaa 4740 aatcaagatt ttgtggacag taacagtaac actatactct cagattctta tatagacttc 4800 cagaggtttg gttgttccaa aatttcagtt aatcattacc cttttgttac aatatacaat 4860 gacagatctc ttcttaacct cgatacttca cttattaaca tcaagaaggt tcataagaca 4920 aatagttcaa agttgagggc aaaaaaggat aatcatcaag taggtgctct agttgtaatg 4980 aatcttgatc tcaacagcat tagactacat cttcatgact cttcatccat tgttgcatct 5040 gttacacttc ctgtttccaa atcctctttt gctatccatg agaactttct ggatgtgtta 5100 ttttcaactg agggattgag tctttcatca cagtggtatc ctcagactct acaagactct 5160 ctgtggggcc ctgcttcact gaatctttct ccagtgatca acattcgtgt gagaaaaggg 5220 aaccatggaa tcgaactgga ttttagtgtt caaaatgttt cctgcatatt gccaactgag 5280 ttcctggctg cactcattgg ttacttctca ttgcctgatt ggagctactc aaatccaaat 5340 gagtcatcac ctactactac caacaccaat accaataaca acagcatcag tttcacttac 5400 aagtttgaga tattggactc ggttttattt acacctgtgg ctaatcctga tcacgagttt 5460 ataaagctta atattccaca gatgtactgc acattcattg atagcattga ttcagacact 5520 ctgttgaaag aaatcccttt agagtgctct gttccagttg gatttactgg aaatcagaat 5580 tattgtctga atgtatttgg gagggattta tctctacatc atattatttg tcgaaaagac 5640 aatgcttctg aagtgacaag tgtcagcttg atcgcgcctt ttagtggcga tatatggatc 5700 acaataccat atgaatctaa ttcttcttat gcaacatgta tcatgtcaag ggttagtaaa 5760 tgtcagttta ctgttgaagg aagagaaata cttggctgca ttggagcatt acaggatgtt 5820 gtcgaccaat tctcatctgt tggcaatcta tccacttgtt tcacttctga tgtctcagag 5880 tttctcaatt taaaagaaaa tttgtggtt ccagttccca ttgaatcttc aactgtcagc 5940 tttacagaaa tcagatgctc tgttcaatcc atgtcagtag aactttactc tgacaagatg 6000 aatggtaaca gacttatcgc caaatccgac atgaaattcg catgtggaat atcaatgaaa 6060 acagacaagc ctctttctt tgatatatca ttcacttgtt tcacactttc ctcacttctt 6120 acctctgttg tcttactgga atgcacatca tgtaccaaaa atgtaccggt tctcaacatg 6180 cagttcttga tgtcagatga tggtaaaaac cacctgcgat tctcccttcc ttgtgttaac 6240 atttggttgt tcttgtctga atggagtcaa gttgttgacc tggtcaattc ttgctgtgaa 6300 cctgcaatcc agaatgagga accggaaaaa tccacatcag ctccggtttc tcgtgttgat 6360 actgcagaaa actcacctca atccaccact gtctccagct atccatcttt agaagatcgg 6420 ttttcattaa ccgtaaagtc ggatcttatt ggtgtaaaaa tccgtattcc tgttcaagtt 6480 tctggagaag tagttaaata ctttggggcc ccacaagttc gagagcagag tttagttaca 6540 ggaagagatc acggaagttt tctgtttatt tatctacaaa gtaggtgcac tgaggtgaac 6600 atgaagggcg aaacggtaaa tttgaagtcg aatctggga aagcaatggg aacagttgaa 6660 6720 gaagccgaat ctggtaatga tgatatggac cgcatgcatt taagagaga gattcactgt 6780 gataatcttg atgtttggct ctcacatcat gcattttact tctggcaaac tatgctgttt 6840 atgtttccgg aaggttccgg atccgaatct cctcaacttc cagttggcag tgtcaatttc 6900 agattccact tacgaaaact ctccattctt ttaacagatg aaaagtggag ctccaatggt 6960 cctttattgg agattctcat gggaagtttg ttattcacg gaatcataac tgcaaacatg 7020 atggaagggt caatcgacag tgaccttcaa gtcaactaca acaacatcca taaagtccta 7080 tgggagcctt ttctcgaacc atggaagttc caataacct taagaaca acaaggaaaa 7140 agcaccctcg agaacagtcc agtcatgacc gatatccgcc tcgagtcatc aacaaacctc 7200 aacatcaatg tcaccgaatc cttcatcgag gtcgctttca gaacattcga catgatcaaa 7260 gacgcctcag atctcataag tcttaacatt cttcctgaaa acaattcaag attgacaaaa 7320 ccccatacaa atgaaaacac actcgcgaat agatacgccc cttacacact cgaaaacttg 7380 acatcacttc ctctggtatt ttacatctca aaaggtgagg ggttcaatat gacatcattg 7440 aaagatggaa aacacgtgca gccaggctgt tcatatcccg tgaatattga taacaacccc 7500 gaacaaa cgtttgggtt taggcctagt cattctactg acaatctagg tggcgacatg 7560 caattcgctg atgctcaaca tcattatata gttgttcaac ttgaaggaac ttctacactt 7620 tcggctccag tttctattga tcttgttggt gtcagcttct ttgaggtgga tttttctact 7680 aataggac ttgatgtttc taaaggtggc tatgttgttc ctgtagtaat cgatgtttca 7740 gtccaacggt acaaaagct agtccgcttg tactcaacgg tcatactgac gaatgcaaca 7800 tcaatgccat ttgaagtacg gttcgatatc ccatttggtg tatctcccaa gattcttagac 7860 cctgtatacc ccggccatga gtttcctctt cctctacatt tagcagaatc agggcgaata 7920 agatggcgtc cattaggag cacttaccta tggagtgag cttacagcat ttcaatatt 7980 ctctcaaatg aaagcaagat tggacattg cggtcgttg tctgttatcc ttctcttcca 8040 agtagcgacc cctttcggtg tgtgtgtcc gttcatgatg tgtgtttgcc atctgctgga 8100 agggtaataa acagagggg atcatctca tctctgtata atattaac tcatgatcat 8160 ggtgacaaga tagagaacca ggatctctcg aaagaggt gtattcactt gattactttg 8220 agtaatcctt tgagtgaa aaactatttg ccagttgaag tgtctgtggt gattgagagt 8280 ggtggagttt cacgtagcat gctgctttca gaggttgaaa ctttttta tcatattgat 8340 tcttcacatg atctttcgtt aactttgag atacacgggt ttagaccttc tcttttgag 8400 tttccacgtg ctgaaagtt cagtggaatt gctaaattca gtggaacaa gttttctct 8460 tcagaatcca tcaacttgc tgctgaat tctaaaggtc cattatatgt gatatggaa 8520 aaggtgatgg atgcattctc tggggctcgt gaatctgca tatttgtgcc cttcttattg 8580 tacaactgct gtggttttcc ccttacgatt gcaattcga ctaatgatct cacaatgcgt 8640 gacactttgc cttcatgtta tgatttggat gaagaggacc cgtttttggg caaaaaagat 8700 ggtctaagcc ttttgttttc caaccaagtt tccaataatg atcctatgag taatcttgtt 8760 tccactagaa aggaatcttt cacctcatct ggatcaacca aaaaaaacat cggcacacaa 8820 aaaccgtctt tacacgatca ggaaaaaagt caactttcgc atagtcaaca acttgacttt 8880 gacgaaacaa ctcgcaaaaa agtcaacttc cgcatgtatt ctcccgaccc taacatcgct 8940 tcgagtgaga tcatggtgag agttagtcga tgccggtctg atgctgacat ggcaagtacc 9000 tcagactaca cgtggtccag tcaattcttc ctggtcccac cagccggttc aaccacagtc 9060 cttgtccccc gatcatcaac caatgcttca tatgtattat cagttgcctc tagtgctatt 9120 tccgggccat attctggaag gacaaggatt atcaatttcc agcctagata cgttatcagt 9180 aatgcatgca gtcaggactt gtgctatagg cagaaaggtt ccgactttat ataccatttg 9240 aaagcaggac aacactccca tatccattgg acagacataa caagggagtt actggtgtct 9300 gttcgtttcg atgagccagg gtggcagtgg tcaggctgct tttttccaga acatctaggt 9360 gatacacagc tgagatgag aaactatgtt agtggtgcag ttagtatggt tcgtgtagag 9420 gtccaaaatg ctgacgatgc aatcagagat gatagattg ttgggaaccc tcatggtga 9480 tctgggacaa atttgattct ttgtctgat gatgatactg gctttatgcc atacaggggtt 9540 gataatttct caaggagag gttacgtatt taccacaaa atgtgaagc atttgagact 9600 gttatacatt catatacgtc ttgccttac gcatgggatg aaccctccta cccacatcgt 9660 ttaactgttg aggtgtttgc tgagaggta gtaggttctt acacctga tgatgctaaa 9720 gagtacaaac ctgtggtttt accttcaacc tccgagaac ctgaagaag gttgctata 9780 tctgtccatg ctgaggagc attgaaggtc ctgagcatca ttgatcaag ctatcatata 9840 tttgatgatg taaaaatccc acgttctcct cggttaactg agaaagaga atacgaccaa 9900 aaacaagaaa gttcacttct ttatcagaa agagtatcga ttcattcc attcattggg 9960 atttctgtta tgagctctca accacaggag tgctttg catgtgcaag gaacacaatg 10020 attgacctgg ttcaagtct ggaccagcaa atctttct tgaaaatctc ctctttacag 10080 attgataacc aactgccaac cacaccctac cccgttatat tatctttcga ccatgagtat 10140 aaacaactcc caacctctca gataaaaaac aaagatgagc ctgtgttctc attggctgca 10200 caaaatgga ggataaaga tagagctttg ctttcattcg agcatataaa cttaagaatg 10260 gcagatttcc atcttgagct tgaacaggat gtgattttaa gtctgtttga tttctccaag 10320 gcagtatcct caaggttcca taggagga atgccacata tggattcagt cgtgcatcct 10380 ctttcctcaa acttgagtgg aaataagca actaaattgg ctgaaaagac tgaaattgag 10440 ggtgaaagtt tccctctttt accatcaata gtgccaattg gtgcaccatg gcagaaata 10500 tatctcctgg caagaaaaca gaaaaata tatgtggaac ttcttgaagt ggcccccatc 10560 actttaaccc taagcttttc gagcagtcca tggatgctaa ggaatggaat acttacatca 10620 ggagaatatc ttatccatag aggtctgatg gctcttgctg acgtggaggg agcacggatc 10680 catctaaggc ggttaacaat ctctcatcag ttggccagct tggaatccat acgagagatc 10740 ttaatcatac attatacccg ccaacttctc cacgaaatgt acaaggtatt tggttcagct 10800 ggggtaatag gcaaccccat gggtttcgca agaagtgtag gacttggcat ccgagacttc 10860 ctctcagttc cagccagaag cttcatgcag agccccgcag gacttatcac ggggatggca 10920 cagggaacta caagtcttct aagcaatacg gtttacgcca taagtgatgc tgccacccaa 10980 ggcattgttg catttacaat ggacgacccc ccatctgcag cagaaatggg caaaggtgta 11040 ataaatgaag ttttggaggg gctgactggt cttctccaat caccaataag aggagctgaa 11100 aaacacgggc ttccaggagt cctttcaggt atagcactag gagtaacggg tctagtggca 11160 aggccagccg ccagcatact ggaagtaaca gaaaaaaccg cccgcagcat aagaaaccga 11220 agcaaactct accacatccg cctccgggtc cgcctcccga gaccgctaac ccccaaccac 11280 cacccgttaa aaccctactc gtgggaccaa gcagtcggcc tctccgtcct caccaatacc 11340 aattccaatt ccgattccga cctcaaagac gaaaccctcg tcctctccaa atccctcaaa 11400 caaaagggca aattcgtcat tatcacccaa cggttactcc tcattgttac ctcctcgagc 11460 ctaacgaatt taggtcaacc caatttcaaa ggcgtccctg cggaccccga ttgggtggtt 11520 gaagccgaga taacgttgga tagtgtgata cacgtggatg ttgatggaga ggtggtgcat 11580 attgtcggga gtagttctga tgtggtggtt agacagaatg ttggtgggaa gcagcggtgg 11640 tataatccgt tgccgctgtt tcagacgaat ttggagtgtt tagggaagga ggaggcgggg 11700 gagttgttga aggtgttgtt ggtgacgatt gagagaggga aggagagagg gtggggccgg 11760 gggtgtgtgt accgtctgca tcagagtaat gttaggtga 11799 <210> 3 <211> 3932 <212> PRT <213> Taraxacum officinale <400> 3 Met Ser Asp Leu Ile Asn Cys Lys Gln Pro His Glu Ser Pro Thr Pro 1 5 10 15 Ala Ala Thr Thr Thr Ala Glu Ala Ser Phe Glu Pro Leu Asp Glu Asn 20 25 30 Leu Phe Ser Asp Leu Thr Pro Leu Cys Tyr Leu Leu Gln Ser Thr Ser 35 40 45 Pro Lys Ser Pro Pro Thr Leu Ile Phe Pro Pro Phe Tyr His His Gln 50 55 60 Ser Ser Gln Asp Ser Ser Phe Leu Ile Ala Val Thr Phe Ser Ser Lys 65 70 75 80 Thr His Val His His Arg Arg Arg Gln Thr Ser Pro Glu Ile Pro Ala 85 90 95 Ile Val Pro Gly Cys Asn Pro Ala Val Lys Pro Val Arg Lys Ser Thr 100 105 110 Thr Gln Pro Pro Val Leu His His Gln Ser Leu Asn Pro Gly Arg Gly 115 120 125 Val Arg Phe Lys Leu Ser Ser Thr Gln Asn Leu Val Thr Leu Ala Ala 130 135 140 Ser Ser Thr Asn His Thr Pro Asp Phe Asp Ser Arg Phe Pro Asn His 145 150 155 160 Asp Ser Ser Tyr Met Asn Asn Trp Ser Asn Gln Glu Glu Asp Asn Asp 165 170 175 Asp Gly Cys Leu Asp Asp Ser Thr Ile Asn Glu Thr Ser Ile His Pro 180 185 190 Gly Val Ile Leu Lys Arg Ile Gln Ala Gln Lys Gly Ile Glu Phe Lys 195 200 205 Arg Arg Arg Lys Tyr Ala Met Ala Ser Ser Pro Ala Val Ala Ile Pro 210 215 220 Ala Ser Ile Ser Pro Ala Val Ala His Ile His Arg Lys Ala Asp Phe 225 230 235 240 His Ile Val Ala Pro Gly Arg His Leu His Val Gly Ser Ala Pro Lys 245 250 255 Lys Val Val Asp Lys Leu Glu Gly Ile Arg Arg Arg Phe Leu Trp Gly 260 265 270 Gly Lys Lys Ser Glu Lys Lys Ile His Trp Val Ser Trp Glu Lys Val 275 280 285 Ile Lys Ser Lys Asp Lys Glu Gly Leu Gly Val Asn Gly Leu Ser Ser 290 295 300 Met Asn Met Ala Leu Leu Val Lys Trp Phe Trp Arg Leu Lys Thr Glu 305 310 315 320 Arg Asp Ser Leu Trp Val Arg Cys Val Thr Ala Cys His Asn Ile Lys 325 330 335 Leu Ile Asp Gly Lys Arg Val Ala Lys Ala Ser Leu Lys Gly Val Trp 340 345 350 Trp Asn Ile Met Ser Cys Val Glu Glu Leu Lys Thr Lys Gly Ile Ser 355 360 365 Val Glu Ser Lys Leu Val Arg Gln Leu Gly Asn Gly Lys His Thr His 370 375 380 Phe Trp Lys Asp Arg Trp Leu His Asn Lys Val Leu Lys Asp Asp Leu 385 390 395 400 Pro Glu Leu Tyr Lys Ile Glu Gly Asp Lys Asn Cys Met Val Asn Gln 405 410 415 Arg Leu Val Trp Asp Asn Asn Glu Lys Ile Phe Lys Gln Ala Trp Asp 420 425 430 Trp Lys Arg Pro Ile Arg Arg Gly Arg Glu Thr Lys Glu Leu Glu Thr 435 440 445 Leu Ile Ile Leu Thr Asn Gly Ile Gln Leu Lys Glu Ile Glu Asp Asn 450 455 460 Trp Arg Trp Lys Glu Gly Ser Asp Gly Lys Phe Ser Val Gly Lys Leu 465 470 475 480 Arg Lys Leu Phe Ala Tyr Gln Glu Gln Ala Glu Val Asp Gly Gly Phe 485 490 495 Asp Trp Ile Asn Trp Val Pro Leu Lys Glu Glu Glu Glu Thr Glu His 500 505 510 Ala Phe Phe Arg Cys Ala His Ala His Gln Val Trp Asp Trp Phe Lys 515 520 525 Met Trp Ser Gly Leu Met Arg Glu Ile Pro Leu Asn Phe Arg Ser Met 530 535 540 Glu Ala Glu Ile Lys Ala Gly Ala Gly Asp Lys Lys Ser Val Lys Leu 545 550 555 560 Gly Met Ala Leu Ala Tyr Val Met Leu Trp Thr Ile Trp Lys Ile Arg 565 570 575 Asn Gly Ala Val Phe Asn Asn Arg Lys Ala Arg Ala Met Asn Thr Thr 580 585 590 Asp Glu Lys Ser Ile Ser Tyr Arg Pro Thr Ser Pro Ala Ala Ala Pro 595 600 605 Ser Pro Ser Pro Ala Pro Ser Pro Ala Thr Ala Pro Ser Pro Pro Ala 610 615 620 Thr Ser Ser Pro Glu Pro Ala Thr His Ser Pro Met Leu Ile Gly Thr 625 630 635 640 Glu Glu Met Phe Glu Gly Leu Val Arg Gln Leu Ile Leu Gly Tyr Leu 645 650 655 Gly Gln Tyr Ile Lys Asp Ile His Arg Glu Gln Leu Lys Ile Thr Leu 660 665 670 Trp Asn Glu Glu Val Phe Leu Glu Asn Val Glu Leu Ile Leu Glu Ala 675 680 685 Phe Asp Tyr Leu Glu Leu Pro Phe Ala Leu Lys Gln Gly Arg Val Gly 690 695 700 Arg Leu Ser Ile Arg Ile Pro Trp Lys Lys Leu Gly Trp Asp Pro Ile 705 710 715 720 Ile Ile Ile Leu Glu Asp Ile Leu Val Cys Ala Ser Gln Arg Glu Asp 725 730 735 Glu Glu Trp Ser Val Asp Asp Val Glu Arg Arg Glu Phe Ser Gly Lys 740 745 750 Lys Ala Lys Leu Ala Ala Ala Glu Leu Ala Lys Leu Ser Gln Arg Val 755 760 765 Cys Asp Asn Gln Thr Gly Lys Ser Phe Met Ser Tyr Ile Thr Ala Lys 770 775 780 Ile Ile Asp Gly Ile Gln Val Thr Ile Arg Asn Val His Ile Val Tyr 785 790 795 800 Arg Asp Ile Ser Asn Glu Lys Ser Gln Thr Val Phe Gly Val Lys Leu 805 810 815 Ala Ser Leu Thr Ala Met Lys Gln Asn Tyr Ala Gly Val Leu Ser Gly 820 825 830 Lys Val Arg Val Gly Gln Val Asn Lys Ile Val Glu Ile Gln Gly Leu 835 840 845 Glu Ile Tyr Cys Lys Thr Phe His Gly Ser Ser Thr Asp Ile His Thr 850 855 860 Glu Asn Gly Glu Asp Ser Met Ala Met Val Ala Ala Ser Tyr Asp Asn 865 870 875 880 Asp Glu His Ala His Leu Leu Ala Pro Val Asn Val Ser Ala Ser Leu 885 890 895 Ser Val Asn Arg Ser Gly Arg Leu Glu Asn Asn Ala Ala Gln Tyr Ser 900 905 910 Val Asp Ile Glu Leu Ser Gly Leu Val Leu Ser Leu Asp Glu Asp Gln 915 920 925 Leu Gln Gln Ile Leu Tyr Leu Tyr Glu Tyr Leu Cys Thr Cys Arg Leu 930 935 940 Arg Glu Lys Tyr Gly Arg Tyr Arg Pro Trp Gly Lys Pro Ile Ser Glu 945 950 955 960 Arg Gln Leu Gly Trp Gln Ile Gln Trp Trp His Tyr Ala Gln His Ser 965 970 975 Val Leu Ser Asp Val Arg Lys Arg Leu Lys Lys Thr Ser Trp Lys Tyr 980 985 990 Leu Gly Glu Arg Leu Gly Arg Lys Arg Arg Tyr Val Asn Leu Tyr Lys 995 1000 1005 Leu Lys Leu Glu Cys Leu Arg Lys Glu Gln Pro Leu Asp Asp Glu 1010 1015 1020 Ile Val Met Glu Leu Asp Gln Met Glu Lys Val Ser Asp Ile Glu 1025 1030 1035 Asp Ile Leu Ser Tyr Arg Ser Ala Ala Glu Asn Glu Leu Gln Glu 1040 1045 1050 Phe Leu Val Asp Ser Pro Ser Gly Ile Gly Gly Ser Glu Val Asn 1055 1060 1065 Thr Thr Ile Asp Lys Ser Met Asp Asp Asp Gln Thr Ser Gly Lys 1070 1075 1080 Pro Gln Gly Trp Leu Lys Trp Leu Ser Arg Gly Met Leu Gly Ala 1085 1090 1095 Gly Gly Thr Asp Asp Ser Ser Gln Phe Ser Gly Val Val Ser Asp 1100 1105 1110 Glu Val Ile Lys Asp Ile Tyr Glu Ala Thr Lys Phe His Pro Ala 1115 1120 1125 Pro Ser Pro Val Leu Asp Ala Ser Gly Thr Asp Arg Val Leu Leu 1130 1135 1140 Thr Ser Ile Lys Cys Ser Ile His Gln Ile Ser Ala Thr Leu Arg 1145 1150 1155 Asn Lys Lys Leu Asp Arg Ala Ile Gly Glu Val Val Phe Glu Gly 1160 1165 1170 Asn Val Val Glu Cys Met Ile Trp Glu Glu Ser Ala Val Val Thr 1175 1180 1185 Ala Ser Ile Asn Ser Val Glu Met Ile Asn Pro Leu Asn Asn Gln 1190 1195 1200 Ala Ile Leu Leu Ile Lys Arg Val Ile Ser Glu Glu Ser Phe Leu 1205 1210 1215 Glu Glu Glu Lys Pro Ser Leu Asn Ile Gln Ala Tyr Ile Pro Gln 1220 1225 1230 Ala Asn Arg Glu Gly Asp Leu Thr Leu Lys Val Leu Leu Glu Pro 1235 1240 1245 Ile Glu Val Thr Cys Asp Pro Thr Tyr Leu Val Asn Phe Met Glu 1250 1255 1260 Leu Tyr Thr Val Leu Gly Ser Tyr Thr Ser His Glu Glu Arg Val 1265 1270 1275 Leu Asn Ser Leu Asn Gly Ile Asn Asp Val Lys Ser Arg Leu Ile 1280 1285 1290 Ser Lys Ala Lys Tyr Ile Leu Ser Gly Arg Lys Arg Met Met Trp 1295 1300 1305 Asp Ile Ser Leu Ile Asn Ile Lys Ile Asn Ile Pro Trp Glu Asn 1310 1315 1320 Gly Asn Ser Glu Met His Lys Leu Val Leu Glu Leu Thr Ala Val 1325 1330 1335 Thr Phe Ala Ser Lys Arg Asp Ile Gly Ser Phe Ala Pro Asp Ile 1340 1345 1350 Asn Val Pro Ser Gln Phe Met Arg Asn Leu Ile Asp Asp Asn Ser 1355 1360 1365 Ser Asn Glu Leu Leu Glu Gly Thr His Ile Gln Asp Leu Tyr Asp 1370 1375 1380 Leu Leu Glu Ile Lys Ile Ile Asp Phe Gln Asp Glu Ser Leu Leu 1385 1390 1395 Lys Ala Leu Glu Val Tyr Val Leu Val Ala Thr Leu Leu Ala His 1400 1405 1410 Val Ser Pro Ser Ile Ile Gly Ser Phe Leu Glu Leu Val Glu Ser 1415 1420 1425 Met Asn Met Leu His His Thr Ser Gln Leu Gly Ala Thr Ser Ala 1430 1435 1440 Thr Ser Ser Ile Glu Pro Arg Asn Ser Ser Ser Ile Ser Val Ile 1445 1450 1455 Ala Asn Leu Glu Ser Ala Ser Ile Ile Val Asp Leu Glu Asn Gly 1460 1465 1470 Leu Glu Ala Ser Cys Thr Leu Thr Val Ser Leu Gln Asp Leu Asp 1475 1480 1485 Met Arg Met Gly Ser Met Lys Ser Thr Gln Ser Phe Trp Ile Cys 1490 1495 1500 Thr Arg Asp Leu Lys Val Thr Ser Arg Leu Leu Glu Ser Gly Asp 1505 1510 1515 Asp Leu Asp Leu Ile Ile Cys Leu Pro Gln Ser Thr Ser Pro Asn 1520 1525 1530 Asp Gly Cys Leu Val Leu His Tyr Asp Gly Asn Leu Ser Ile Cys 1535 1540 1545 Leu Ser Asp Leu Asp Leu His Cys Tyr Pro His Ile Val Gly Leu 1550 1555 1560 Leu Val Glu Phe Ser Gly Lys Leu Ser Thr Tyr Ser Pro Ser Asn 1565 1570 1575 Ala Lys Asn Gln Asp Phe Val Asp Ser Asn Ser Asn Thr Ile Leu 1580 1585 1590 Ser Asp Ser Tyr Ile Asp Phe Gln Arg Phe Gly Cys Ser Lys Ile 1595 1600 1605 Ser Val Asn His Tyr Pro Phe Val Thr Ile Tyr Asn Asp Arg Ser 1610 1615 1620 Leu Leu Asn Leu Asp Thr Ser Leu Ile Asn Ile Lys Lys Val His 1625 1630 1635 Lys Thr Asn Ser Ser Lys Leu Arg Ala Lys Lys Asp Asn His Gln 1640 1645 1650 Val Gly Ala Leu Val Val Met Asn Leu Asp Leu Asn Ser Ile Arg 1655 1660 1665 Leu His Leu His Asp Ser Ser Ser Ile Val Ala Ser Val Thr Leu 1670 1675 1680 Pro Val Ser Lys Ser Ser Phe Ala Ile His Glu Asn Phe Leu Asp 1685 1690 1695 Val Leu Phe Ser Thr Glu Gly Leu Ser Leu Ser Ser Gln Trp Tyr 1700 1705 1710 Pro Gln Thr Leu Gln Asp Ser Leu Trp Gly Pro Ala Ser Leu Asn 1715 1720 1725 Leu Ser Pro Val Ile Asn Ile Arg Val Arg Lys Gly Asn His Gly 1730 1735 1740 Ile Glu Leu Asp Phe Ser Val Gln Asn Val Ser Cys Ile Leu Pro 1745 1750 1755 Thr Glu Phe Leu Ala Ala Leu Ile Gly Tyr Phe Ser Leu Pro Asp 1760 1765 1770 Trp Ser Tyr Ser Asn Pro Asn Glu Ser Ser Pro Thr Thr Thr Asn 1775 1780 1785 Thr Asn Thr Asn Asn Asn Ser Ile Ser Phe Thr Tyr Lys Phe Glu 1790 1795 1800 Ile Leu Asp Ser Val Leu Phe Thr Pro Val Ala Asn Pro Asp His 1805 1810 1815 Glu Phe Ile Lys Leu Asn Ile Pro Gln Met Tyr Cys Thr Phe Ile 1820 1825 1830 Asp Ser Ile Asp Ser Asp Thr Leu Leu Lys Glu Ile Pro Leu Glu 1835 1840 1845 Cys Ser Val Pro Val Gly Phe Thr Gly Asn Gln Asn Tyr Cys Leu 1850 1855 1860 Asn Val Phe Gly Arg Asp Leu Ser Leu His His Ile Ile Cys Arg 1865 1870 1875 Lys Asp Asn Ala Ser Glu Val Thr Ser Val Ser Leu Ile Ala Pro 1880 1885 1890 Phe Ser Gly Asp Ile Trp Ile Thr Ile Pro Tyr Glu Ser Asn Ser 1895 1900 1905 Ser Tyr Ala Thr Cys Ile Met Ser Arg Val Ser Lys Cys Gln Phe 1910 1915 1920 Thr Val Glu Gly Arg Glu Ile Leu Gly Cys Ile Gly Ala Leu Gln 1925 1930 1935 Asp Val Val Asp Gln Phe Ser Ser Val Gly Asn Leu Ser Thr Cys 1940 1945 1950 Phe Thr Ser Asp Val Ser Glu Phe Leu Asn Leu Lys Glu Asn Tyr 1955 1960 1965 Val Val Pro Val Pro Ile Glu Ser Ser Thr Val Ser Phe Thr Glu 1970 1975 1980 Ile Arg Cys Ser Val Gln Ser Met Ser Val Glu Leu Tyr Ser Asp 1985 1990 1995 Lys Met Asn Gly Asn Arg Leu Ile Ala Lys Ser Asp Met Lys Phe 2000 2005 2010 Ala Cys Gly Ile Ser Met Lys Thr Asp Lys Pro Leu Ser Leu Asp 2015 2020 2025 Island Ser Phe Thr Cys Phe Thr Leu Ser Ser Leu Leu Thr Ser Val 2030 2035 2040 Val Leu Leu Glu Cys Thr Ser Cys Thr Lys Asn Val Pro Val Leu 2045 2050 2055 Asn Met Gln Phe Leu Met Ser Asp Asp Gly Lys Asn His Leu Arg 2060 2065 2070 Phe Ser Leu Pro Cys Val Asn Ile Trp Leu Phe Leu Ser Glu Trp 2075 2080 2085 Ser Gln Val Val Asp Leu Val Asn Ser Cys Cys Glu Pro Ala Ile 2090 2095 2100 Gln Asn Glu Glu Pro Glu Lys Ser Thr Ser Ala Pro Val Ser Arg 2105 2110 2115 Val Asp Thr Ala Glu Asn Ser Pro Gln Ser Thr Thr Val Ser Ser 2120 2125 2130 Tyr Pro Ser Leu Glu Asp Arg Phe Ser Leu Thr Val Lys Ser Asp 2135 2140 2145 Leu Ile Gly Val Lys Ile Arg Ile Pro Val Gln Val Ser Gly Glu 2150 2155 2160 Val Val Lys Tyr Phe Gly Ala Pro Gln Val Arg Glu Gln Ser Leu 2165 2170 2175 Val Thr Gly Arg Asp His Gly Ser Phe Leu Phe Ile Tyr Leu Gln 2180 2185 2190 Ser Arg Cys Thr Glu Val Asn Met Lys Gly Glu Thr Val Asn Leu 2195 2200 2205 Lys Ser Asn Leu Gly Lys Ala Met Gly Thr Val Glu Leu Phe Gln 2210 2215 2220 Asn Lys Ser Val His Ser Trp Pro Leu Phe Gln Leu Leu Glu Ile 2225 2230 2235 Asp Ile Glu Ala Glu Ser Gly Asn Asp Asp Met Asp Arg Met His 2240 2245 2250 Leu Lys Thr Glu Ile His Cys Asp Asn Leu Asp Val Trp Leu Ser 2255 2260 2265 His His Ala Phe Tyr Phe Trp Gln Thr Met Leu Phe Met Phe Pro 2270 2275 2280 Glu Gly Ser Gly Ser Glu Ser Pro Gln Leu Pro Val Gly Ser Val 2285 2290 2295 Asn Phe Arg Phe His Leu Arg Lys Leu Ser Ile Leu Leu Thr Asp 2300 2305 2310 Glu Lys Trp Ser Ser Asn Gly Pro Leu Leu Glu Ile Leu Met Gly 2315 2320 2325 Ser Leu Leu Phe His Gly Ile Ile Thr Ala Asn Met Met Glu Gly 2330 2335 2340 Ser Ile Asp Ser Asp Leu Gln Val Asn Tyr Asn Asn Ile His Lys 2345 2350 2355 Val Leu Trp Glu Pro Phe Leu Glu Pro Trp Lys Phe Gln Ile Thr 2360 2365 2370 Leu Arg Arg Gln Gln Gly Lys Ser Thr Leu Glu Asn Ser Pro Val 2375 2380 2385 Met Thr Asp Ile Arg Leu Glu Ser Ser Thr Asn Leu Asn Ile Asn 2390 2395 2400 Val Thr Glu Ser Phe Ile Glu Val Ala Phe Arg Thr Phe Asp Met 2405 2410 2415 Ile Lys Asp Ala Ser Asp Leu Ile Ser Leu Asn Ile Leu Pro Glu 2420 2425 2430 Asn Asn Ser Arg Leu Thr Lys Pro His Thr Asn Glu Asn Thr Leu 2435 2440 2445 Ala Asn Arg Tyr Ala Pro Tyr Thr Leu Glu Asn Leu Thr Ser Leu 2450 2455 2460 Pro Leu Val Phe Tyr Ile Ser Lys Gly Glu Gly Phe Asn Met Thr 2465 2470 2475 Ser Leu Lys Asp Gly Lys His Val Gln Pro Gly Cys Ser Tyr Pro 2480 2485 2490 Val Asn Ile Asp Asn Asn Pro Glu Glu Gln Thr Phe Gly Phe Arg 2495 2500 2505 Pro Ser His Ser Thr Asp Asn Leu Gly Gly Asp Met Gln Phe Ala 2510 2515 2520 Asp Ala Gln His His Tyr Ile Val Val Gln Leu Glu Gly Thr Ser 2525 2530 2535 Thr Leu Ser Ala Pro Val Ser Ile Asp Leu Val Gly Val Ser Phe 2540 2545 2550 Phe Glu Val Asp Phe Ser Thr Asn Ile Gly Leu Asp Val Ser Lys 2555 2560 2565 Gly Gly Tyr Val Val Pro Val Val Ile Asp Val Ser Val Gln Arg 2570 2575 2580 Tyr Thr Lys Leu Val Arg Leu Tyr Ser Thr Val Ile Leu Thr Asn 2585 2590 2595 Ala Thr Ser Met Pro Phe Glu Val Arg Phe Asp Ile Pro Phe Gly 2600 2605 2610 Val Ser Pro Lys Ile Leu Asp Pro Val Tyr Pro Gly His Glu Phe 2615 2620 2625 Pro Leu Pro Leu His Leu Ala Glu Ser Gly Arg Ile Arg Trp Arg 2630 2635 2640 Pro Leu Gly Ser Thr Tyr Leu Trp Ser Glu Ala Tyr Ser Ile Ser 2645 2650 2655 Asn Ile Leu Ser Asn Glu Ser Lys Ile Gly His Leu Arg Ser Phe 2660 2665 2670 Val Cys Tyr Pro Ser Leu Pro Ser Ser Asp Pro Phe Arg Cys Cys 2675 2680 2685 Val Ser Val His Asp Val Cys Leu Pro Ser Ala Gly Arg Val Ile 2690 2695 2700 Asn Lys Arg Gly Ser Ser Ser Ser Leu Tyr Asn Ile Asn Thr His 2705 2710 2715 Asp His Gly Asp Lys Ile Glu Asn Gln Asp Leu Ser Asn Lys Arg 2720 2725 2730 Cys Ile His Leu Ile Thr Leu Ser Asn Pro Leu Ile Val Lys Asn 2735 2740 2745 Tyr Leu Pro Val Glu Val Ser Val Val Ile Glu Ser Gly Gly Val 2750 2755 2760 Ser Arg Ser Met Leu Leu Ser Glu Val Glu Thr Phe Phe Tyr His 2765 2770 2775 Ile Asp Ser Ser His Asp Leu Ser Leu Thr Phe Glu Ile His Gly 2780 2785 2790 Phe Arg Pro Ser Leu Leu Lys Phe Pro Arg Ala Glu Lys Phe Ser 2795 2800 2805 Gly Ile Ala Lys Phe Ser Gly Thr Lys Phe Ser Ser Ser Glu Ser 2810 2815 2820 Ile Asn Phe Ala Ala Asp Asn Ser Lys Gly Pro Leu Tyr Val Thr 2825 2830 2835 Met Glu Lys Val Met Asp Ala Phe Ser Gly Ala Arg Glu Ile Cys 2840 2845 2850 Ile Phe Val Pro Phe Leu Leu Tyr Asn Cys Cys Gly Phe Pro Leu 2855 2860 2865 Thr Ile Ala Asn Ser Thr Asn Asp Leu Thr Met Arg Asp Thr Leu 2870 2875 2880 Pro Ser Cys Tyr Asp Leu Asp Glu Glu Asp Pro Phe Leu Gly Lys 2885 2890 2895 Lys Asp Gly Leu Ser Leu Leu Phe Ser Asn Gln Val Ser Asn Asn 2900 2905 2910 Asp Pro Met Ser Asn Leu Val Ser Thr Arg Lys Glu Ser Phe Thr 2915 2920 2925 Ser Ser Gly Ser Thr Lys Lys Asn Ile Gly Thr Gln Lys Pro Ser 2930 2935 2940 Leu His Asp Gln Glu Lys Ser Gln Leu Ser His Ser Gln Gln Leu 2945 2950 2955 Asp Phe Asp Glu Thr Thr Arg Lys Lys Val Asn Phe Arg Met Tyr 2960 2965 2970 Ser Pro Asp Pro Asn Ile Ala Ser Ser Glu Ile Met Val Arg Val 2975 2980 2985 Ser Arg Cys Arg Ser Asp Ala Asp Met Ala Ser Thr Ser Asp Tyr 2990 2995 3000 Thr Trp Ser Ser Gln Phe Phe Leu Val Pro Pro Ala Gly Ser Thr 3005 3010 3015 Thr Val Leu Val Pro Arg Ser Ser Thr Asn Ala Ser Tyr Val Leu 3020 3025 3030 Ser Val Ala Ser Ser Ala Ile Ser Gly Pro Tyr Ser Gly Arg Thr 3035 3040 3045 Arg Ile Ile Asn Phe Gln Pro Arg Tyr Val Ile Ser Asn Ala Cys 3050 3055 3060 Ser Gln Asp Leu Cys Tyr Arg Gln Lys Gly Ser Asp Phe Ile Tyr 3065 3070 3075 His Leu Lys Ala Gly Gln His Ser His Ile His Trp Thr Asp Ile 3080 3085 3090 Thr Arg Glu Leu Leu Val Ser Val Arg Phe Asp Glu Pro Gly Trp 3095 3100 3105 Gln Trp Ser Gly Cys Phe Phe Pro Glu His Leu Gly Asp Thr Gln 3110 3115 3120 Leu Lys Met Arg Asn Tyr Val Ser Gly Ala Val Ser Met Val Arg 3125 3130 3135 Val Glu Val Gln Asn Ala Asp Asp Ala Ile Arg Asp Asp Lys Ile 3140 3145 3150 Val Gly Asn Pro His Gly Glu Ser Gly Thr Asn Leu Ile Leu Leu 3155 3160 3165 Ser Asp Asp Asp Thr Gly Phe Met Pro Tyr Arg Val Asp Asn Phe 3170 3175 3180 Ser Lys Glu Arg Leu Arg Ile Tyr Gln Gln Lys Cys Glu Ala Phe 3185 3190 3195 Glu Thr Val Ile His Ser Tyr Thr Ser Cys Pro Tyr Ala Trp Asp 3200 3205 3210 Glu Pro Ser Tyr Pro His Arg Leu Thr Val Glu Val Phe Ala Glu 3215 3220 3225 Arg Val Val Gly Ser Tyr Thr Leu Asp Asp Ala Lys Glu Tyr Lys 3230 3235 3240 Pro Val Val Leu Pro Ser Thr Ser Glu Lys Pro Glu Arg Arg Leu 3245 3250 3255 Leu Ile Ser Val His Ala Glu Gly Ala Leu Lys Val Leu Ser Ile 3260 3265 3270 Ile Asp Ser Ser Tyr His Ile Phe Asp Asp Val Lys Ile Pro Arg 3275 3280 3285 Ser Pro Arg Leu Thr Glu Lys Arg Glu Tyr Asp Gln Lys Gln Glu 3290 3295 3300 Ser Ser Leu Leu Tyr Gln Glu Arg Val Ser Ile Ser Ile Pro Phe 3305 3310 3315 Ile Gly Ile Ser Val Met Ser Ser Gln Pro Gln Glu Leu Leu Phe 3320 3325 3330 Ala Cys Ala Arg Asn Thr Met Ile Asp Leu Val Gln Ser Leu Asp 3335 3340 3345 Gln Gln Asn Leu Ser Leu Lys Ile Ser Ser Leu Gln Ile Asp Asn 3350 3355 3360 Gln Leu Pro Thr Thr Pro Tyr Pro Val Ile Leu Ser Phe Asp His 3365 3370 3375 Glu Tyr Lys Gln Leu Pro Thr Ser Gln Ile Lys Asn Lys Asp Glu 3380 3385 3390 Pro Val Phe Ser Leu Ala Ala Ala Lys Trp Arg Asn Lys Asp Arg 3395 3400 3405 Ala Leu Leu Ser Phe Glu His Ile Asn Leu Arg Met Ala Asp Phe 3410 3415 3420 His Leu Glu Leu Glu Gln Asp Val Ile Leu Ser Leu Phe Asp Phe 3425 3430 3435 Ser Lys Ala Val Ser Ser Arg Phe His Ser Arg Gly Met Pro His 3440 3445 3450 Met Asp Ser Val Val His Pro Leu Ser Ser Asn Leu Ser Gly Asn 3455 3460 3465 Lys Ala Thr Lys Leu Ala Glu Lys Thr Glu Ile Glu Gly Glu Ser 3470 3475 3480 Phe Pro Leu Leu Pro Ser Ile Val Pro Ile Gly Ala Pro Trp Gln 3485 3490 3495 Lys Ile Tyr Leu Leu Ala Arg Lys Gln Lys Lys Ile Tyr Val Glu 3500 3505 3510 Leu Leu Glu Val Ala Pro Ile Thr Leu Thr Leu Ser Phe Ser Ser 3515 3520 3525 Ser Pro Trp Met Leu Arg Asn Gly Ile Leu Thr Ser Gly Glu Tyr 3530 3535 3540 Leu Ile His Arg Gly Leu Met Ala Leu Ala Asp Val Glu Gly Ala 3545 3550 3555 Arg Ile His Leu Arg Arg Leu Thr Ile Ser His Gln Leu Ala Ser 3560 3565 3570 Leu Glu Ser Ile Arg Glu Ile Leu Ile Ile His Tyr Thr Arg Gln 3575 3580 3585 Leu Leu His Glu Met Tyr Lys Val Phe Gly Ser Ala Gly Val Ile 3590 3595 3600 Gly Asn Pro Met Gly Phe Ala Arg Ser Val Gly Leu Gly Ile Arg 3605 3610 3615 Asp Phe Leu Ser Val Pro Ala Arg Ser Phe Met Gln Ser Pro Ala 3620 3625 3630 Gly Leu Ile Thr Gly Met Ala Gln Gly Thr Thr Ser Leu Leu Ser 3635 3640 3645 Asn Thr Val Tyr Ala Ile Ser Asp Ala Ala Thr Gln Gly Ile Val 3650 3655 3660 Ala Phe Thr Met Asp Asp Pro Pro Ser Ala Ala Glu Met Gly Lys 3665 3670 3675 Gly Val Ile Asn Glu Val Leu Glu Gly Leu Thr Gly Leu Leu Gln 3680 3685 3690 Ser Pro Ile Arg Gly Ala Glu Lys His Gly Leu Pro Gly Val Leu 3695 3700 3705 Ser Gly Ile Ala Leu Gly Val Thr Gly Leu Val Ala Arg Pro Ala 3710 3715 3720 Ala Ser Ile Leu Glu Val Thr Glu Lys Thr Ala Arg Ser Ile Arg 3725 3730 3735 Asn Arg Ser Lys Leu Tyr His Ile Arg Leu Arg Val Arg Leu Pro 3740 3745 3750 Arg Pro Leu Thr Pro Asn His His Pro Leu Lys Pro Tyr Ser Trp 3755 3760 3765 Asp Gln Ala Val Gly Leu Ser Val Leu Thr Asn Thr Asn Ser Asn 3770 3775 3780 Ser Asp Ser Asp Leu Lys Asp Glu Thr Leu Val Leu Ser Lys Ser 3785 3790 3795 Leu Lys Gln Lys Gly Lys Phe Val Ile Ile Thr Gln Arg Leu Leu 3800 3805 3810 Leu Ile Val Thr Ser Ser Ser Leu Thr Asn Leu Gly Gln Pro Asn 3815 3820 3825 Phe Lys Gly Val Pro Ala Asp Pro Asp Trp Val Val Glu Ala Glu 3830 3835 3840 Ile Thr Leu Asp Ser Val Ile His Val Asp Val Asp Gly Glu Val 3845 3850 3855 Val His Ile Val Gly Ser Ser Ser Asp Val Val Val Arg Gln Asn 3860 3865 3870 Val Gly Gly Lys Gln Arg Trp Tyr Asn Pro Leu Pro Leu Phe Gln 3875 3880 3885 Thr Asn Leu Glu Cys Leu Gly Lys Glu Glu Ala Gly Glu Leu Leu 3890 3895 3900 Lys Val Leu Leu Val Thr Ile Glu Arg Gly Lys Glu Arg Gly Trp 3905 3910 3915 Gly Arg Gly How To Get Arg With His Gln Ser Asn Val Arg 3920 3925 3930 <210> 4 <211> 909 <212> DNA <213> Tarrhacum officina <400> 4 gaaaccgaag siaacctac cacatccgcc tccgggtccg cctcccgaga ccgctaaccc 60 ccaaccacca cccgttaaaa ccctactcgt gggaccaagc agtcggcctc tccgtcctca 120 ccaataccaa ttccaattcc gattccgacc tcaagacga aaccctcgtc ctctccaaat 180 ccctcaaca aaagggcaaa ttcgtcatta tcaccacg gttactccctc attgttacct 240 cctcgagcct aacgaattta ggtcaaccca atttcaaagg cgtccctgcg gaccccgatt gggtggttga agccgagata acgttggata gtgtgataca cgtggatgtt gatggagagg 360 tggtgcatat tgtcgggagt agttctgatg tggtggttag acagaatgtt ggtgggaagc 420 agcggtggta taatccgttg ccgctgtttc agacgaattt ggagtgttta gggaaggagg 480 aggcggggga gttgttgaag gtgttgttgg tgacgattga gagagggag gagagagggt 540 ggggccgggg gtgtgtgtac cgtctgcatc agagtaatgt tagtgatgt atattttttt 600 tctacatata aagttactat aggagaaaaa ggactggata ttatattata catacctgaa acaaggaaac gttttctttc aaaatttttgg ctgtattatt attttgtcga ccatgttggg ctaaaatggc caattattta cttatgacat ggttaaaaaa tattggtgtc ttgttttgta tattacaat ttatattagt atcgatgcaa tgtaagattg tagaaagcgc taccgtata aacaacataa gtcatgaggt tacaccctag tggggtcaag ggaacaaaaa cattttaaac gtttttcag 909 <210> 5 <211> 909 <212> RNA <213> Tarrhacum officina <400> 5 gaaaccgaag caaccuac cacauccgcc uccggguccg cucccgaga ccgcuaaccc 60 squirrel cccguuaaaa squickcucku gggaccaagc agucggccuc squirrel 120 ccaauaccaa uuccaauucc gauuccgacc ucaagacga aacccucguc cucuccaau 180 240 ccucgagccu aacgaauuua ggucaaccca auuucaagg cgucccugcg gaccccgauu 300 gggugguga agccgagaua accuggaua gugugauaca cguggauguu gauggagagg 360 uggugcauau ugucgggagu agucugaug uggugguag acagauuguu gggggaagc 420 agcgguggua uaauccguug ccgcuguuuc agacgaauuu ggaguguuua gggaaggagg 480 aggcggggga goooooooooooooooooooough 540 ggggccgggg guguguguac cgucugcauc agaguaugu uagguguac auauuuuuuu 600 ucuaaaaaaaaaaaccuau aggaaaaa ggacuggaua uuaaauaaaaaaccugaa 660 acaaggaaac guuuucuuuc aaaauuuugg cuguauuauu auuuugucga ccauguuggg 720 cuaaaauggc caauuauuua cuuaugacau gguuaaaaaa uauugguguc uuguuuugua 780 uaauuacaau uuauauuagu aucgaugcaa uguaagauug uagaaagcgc uaccguauaa 840 aacaacauaa gucaugaggu uacacccuag uggggucaag ggaacaaaa cauuuuaaac 900 guuuuucag 909 <210> 6 <211> 582 <212> DNA <213> Taraxacum officinale <400> 6 aaccgaagca aactctacca catccgcctc cgggtccgcc tcccgagacc gctaaccccc 60 aaccaccacc cgttaaaacc ctactcgtgg gaccaagcag tcggcctctc cgtcctcacc 120 aataccaatt ccaattccga ttccgacctc aaagacgaaa ccctcgtcct ctccaaatcc 180 ctcaaacaaa agggcaaatt cgtcattatc acccaacggt tactcctcat tgttacctcc 240 tcgagcctaa cgaatttagg tcaacccaat ttcaaaggcg tccctgcgga ccccgattgg 300 gtggttgaag ccgagataac gttggatagt gtgatacacg tggatgttga tggagaggtg 360 gtgcatattg tcgggagtag ttctgatgtg gtggttagac agaatgttgg tgggaagcag 420 cggtggtata atccgttgcc gctgtttcag acgaatttgg agtgtttagg gaaggaggag 480 gcgggggagt tgttgaaggt gttgttggtg acgattgaga gagggaaggga gagagggtgg 540 ggccgggggt gtgtgtaccg tctgcatcag agtaatgtta gg 582 <210> 7 <211> 194 <212> PRT <213> Taraxacum officinale <400> 7 Asn Arg Ser Lys Leu Tyr His Ile Arg Leu Arg Val Arg Leu Pro Arg 1 5 10 15 Pro Leu Thr Pro Asn His His Pro Leu Lys Pro Tyr Ser Trp Asp Gln 20 25 30 Ala Val Gly Leu Ser Val Leu Thr Asn Thr Asn Ser Asn Ser Asp Ser 35 40 45 Asp Leu Lys Asp Glu Thr Leu Val Leu Ser Lys Ser Leu Lys Gln Lys 50 55 60 Gly Lys Phe Val Ile Ile Thr Gln Arg Leu Leu Leu Ile Val Thr Ser 65 70 75 80 Ser Ser Leu Thr Asn Leu Gly Gln Pro Asp Phe Lys Gly Val Pro Ala 85 90 95 Asp Pro Asp Trp Val Val Glu Ala Glu Ile Thr Leu Asp Ser Val Ile 100 105 110 His Val Asp Val Asp Gly Glu Val Val His Ile Val Gly Ser Ser Ser 115 120 125 Asp Val Val Val Arg Gln Asn Val Gly Gly Lys Gln Arg Trp Tyr Asn 130 135 140 Pro Leu Pro Leu Phe Gln Thr Asn Leu Glu Cys Leu Gly Lys Glu Glu 145 150 155 160 Ala Gly Glu Leu Leu Lys Val Leu Leu Val Thr Ile Gln Arg Gly Lys 165 170 175 Glu Arg Gly Trp Gly Arg Gly Cys Val Tyr Arg Leu His Gln Ser Asn 180 185 190 Val Arg <210> 8 <211> 51 <212> DNA <213> Taraxacum officinale <400> 8 gaattcctct atgttttgga accttcttgt tgttattccg caccacttta a 51 <210> 9 <211> 479 <212> DNA <213> Taraxacum officinale <400> 9 aattccaatt ccgattccga cctcaaagac gaaaccctcg tcctctccaa atccctcaaa 60 caaaagggca aattcgtcat tatcacccaa cggttactcc tcattgttac ctcctcgagc 120 ctaacgaatt taggtcaacc cgatttcaaa ggcgtccctg cggaccccga ttgggtggtt 180 gaagccgaga taacgttgga tagtgtgata cacgtggatg ttgatggaga ggtggtgcat 240 attgtcggga gtagttctga tgtggtggtt agacagaatg ttggtggtgg tggtgggtgg 300 gggaagcagc ggtggtataa tccgccgacg ccgttgccgc tgtttcagac gaatttggag 360 tgttaggga aggaggaggc gggggagttg ttgaaggtgt tgttggtgac gattcagaga 420 gggaaggaga gagggtgggg ccgggggtgt gtgtaccgtc tgcatcagag taatgttag 479 <210> 10 <211> 159 <212> PRT <213> Taraxacum officinale <400> 10 Asn Ser Asn Ser Asp Ser Asp Leu Lys Asp Glu Thr Leu Val Leu Ser 1 5 10 15 Lys Ser Leu Lys Gln Lys Gly Lys Phe Val Ile Ile Thr Gln Arg Leu 20 25 30 Leu Leu Ile Val Thr Ser Ser Ser Leu Thr Asn Leu Gly Gln Pro Asp 35 40 45 Phe Lys Gly Val Pro Ala Asp Pro Asp Trp Val Val Glu Ala Glu Ile 50 55 60 Thr Leu Asp Ser Val Ile His Val Asp Val Asp Gly Glu Val Val His 65 70 75 80 Ile Val Gly Ser Ser Ser Asp Val Val Val Arg Gln Asn Val Gly Gly 85 90 95 Gly Gly Gly Trp Gly Lys Gln Arg Trp Tyr Asn Pro Pro Thr Pro Leu 100 105 110 Pro Leu Phe Gln Thr Asn Leu Glu Cys Leu Gly Lys Glu Glu Ala Gly 115 120 125 Glu Leu Leu Lys Val Leu Leu Val Thr Ile Gln Arg Gly Lys Glu Arg 130 135 140 Gly Trp Gly Arg Gly Cys Val Tyr Arg Leu His Gln Ser Asn Val 145 150 155 <210> 11 <211> 455 <212> DNA <213> Taraxacum officinale <400> 11 aattccaatt ccgattccga cctcaaagac gaaaccctcg tcctctccaa atccctcaaa 60 caaaagggca aattcgtcat tatcacccaa cggttactcc tcattgttac ctcctcgagc 120 ctaacgaatt taggtcaacc caatttcaaa ggcgtccctg cggaccccga ttgggtggtt 180 gaagccgaga taacgttgga tagtgtgata cacgtggatg ttgatggaga ggtggtgcat 240 attgtcggga gtagttctga tgtggtggtt agacagaatg ttggtgggaa gcagcggtgg 300 tataatccgt tgccgctgtt tcagacgaat ttggagtgtt tagggaagga ggaggcgggg 360 gagttgttga aggtgttgtt ggtgacgatt gagagaggga aggagagagg gtggggccgg 420 gggtgtgtgt accgtctgca tcagagtaat gttag 455 <210> 12 <211> 151 <212> PRT <213> Taraxacum officinale <400> 12 Asn Ser Asn Ser Asp Ser Asp Leu Lys Asp Glu Thr Leu Val Leu Ser 1 5 10 15 Lys Ser Leu Lys Gln Lys Gly Lys Phe Val Ile Ile Thr Gln Arg Leu 20 25 30 Leu Leu Ile Val Thr Ser Ser Ser Leu Thr Asn Leu Gly Gln Pro Asn 35 40 45 Phe Lys Gly Val Pro Ala Asp Pro Asp Trp Val Val Glu Ala Glu Ile 50 55 60 Thr Leu Asp Ser Val Ile His Val Asp Val Asp Gly Glu Val Val His 65 70 75 80 Ile Val Gly Ser Ser Ser Asp Val Val Val Arg Gln Asn Val Gly Gly 85 90 95 Lys Gln Arg Trp Tyr Asn Pro Leu Pro Leu Phe Gln Thr Asn Leu Glu 100 105 110 Cys Leu Gly Lys Glu Glu Ala Gly Glu Leu Leu Lys Val Leu Leu Val 115 120 125 Thr Ile Glu Arg Gly Lys Glu Arg Gly Trp Gly Arg Gly Cys Val Tyr 130 135 140 Arg Leu His Gln Ser Asn Val 145,150 <210> 13 <211> 938 <212> DNA <213> Tarrhacum officina <400> 13 gaaaccgaag caacctac cacatccgcc tccgggtccg cctcccaga ccgctaaccc 60 ccaaccaccc gttaaaaccc tactcgtggg accaagcagt cggccctcc gtcctcacca 120 attccaattc cgattccgac ctcaaagacg aaaccctcgt ccttccaaa tccctcaaac 180 aaaagggcaa attcgtcatt atcacccaac ggttactcct cattgttacc tcctcgagcc 240 taacgaattt aggtcaaccc gatttcaaag gcgtccctgc ggaccccgat tgggtggttg 300 aagccgagat aacgttggat agtgtgatac acgtggatgt tgatggagag gtggtgcata 360 ttgtcgggag tagttctgat gtggtggtta gacagaatgt tggtggtggt ggtgggtggg 420 ggaagcagcg gtggtataat ccgccgacgc cgttgccgct gtttcagacg aatttggagt 480 gtttagggaa ggaggaggcg ggggagttgt tgaaggtgtt gttggtgacg attcagagag 540 ggaaggagag agggtggggc cgggggtgtg tgtaccgtct gcatcagagt aatgttaggt 600 gatgtatatt ttttttgtac atataaagtt tactatagga gaaaaaggac tggatattat 660 attatacata cctgaaaaca aggaaacgtt ttctttcaaa atttggctg tattattatt 720 ttgtcgacca tgttgggcta aaatggccaa ttatttactt atgacatggt taaaaaatat 780 tggtgtcttg ttttgtataa ttacaattta tattcctttt tagactataa cttatgcaat 840 gtaagattgt gtataaaaca acataagtca tgaggttaca ccctagtggg atcaaggggg 900 ctacaccccg gaacaaaaac attttaaacg tttttcag 938

Claims

1. An isolated polynucleotide fragment consisting of a nucleic acid sequence SEQ ID NO: 11 and encoding a protein consisting of SEQ ID NO: 12, wherein the polynucleotide, and / or the expression product of the polynucleotide fragment and / or the protein encoded by the polynucleotide fragment are capable of providing ploid sporophyte formation function to a plant or plant cell.

2. The polynucleotide fragment according to claim 1, wherein the expression product is an RNA molecule.

3. The polynucleotide fragment according to claim 2, wherein the expression product is an mRNA molecule, siRNA, or miRNA molecule.

4. An isolated polypeptide fragment consisting of the amino acid sequence of SEQ ID NO:

12.

5. The polypeptide fragment according to claim 4, wherein the polypeptide fragment is capable of providing ploid sporophyte formation function to plants or plant cells.

6. The polypeptide fragment of claim 4, wherein the polypeptide fragment provides the plant or plant cell with a ploid sporophyte formation function when introduced into the plant or plant cell.

7. A chimeric gene comprising a polynucleotide fragment according to any one of claims 1-3.

8. A genetic construct comprising a polynucleotide fragment according to any one of claims 1-3 or a chimeric gene according to claim 7.

9. A genetic construct comprising, in the 5'-3' direction, an open reading frame polynucleotide encoding a polypeptide fragment according to any one of claims 4-6.

10. A nucleic acid vector comprising an active promoter sequence in a plant cell, said promoter sequence being operatively linked to a polynucleotide fragment according to any one of claims 1-3, a chimeric gene according to claim 7, or a genetic construct according to claim 8 or 9.

11. The nucleic acid vector according to claim 10, wherein the promoter sequence comprises the nucleic acid sequence of SEQ ID NO: 2 or SEQ ID NO: 6 or a functional fragment thereof.

12. The nucleic acid vector according to claim 11, wherein the promoter is a female ovary-specific promoter.

13. The nucleic acid vector according to claim 12, wherein the female ovary-specific promoter is at least one of a megasporocyte-specific promoter and a female gamete-specific promoter.

14. A method for producing apomixis seeds, comprising the following steps: a) Transform a plant, plant part or plant cell with at least one of the polynucleotide fragments of any one of claims 1-3, the chimeric gene of claim 7, the genetic construct of claim 8 or 9 and the nucleic acid vector of any one of claims 10-13 to produce a primary transformant; b) Growing flowering plants or flowers from the primary transformant, wherein at least one of the polynucleotides, fragments, genes, constructs, and vectors is present or expressed in at least the female ovary; and c) Pollinate the primary transformant to induce seed production.

15. The method of claim 14, wherein at least one of the polynucleotide, fragment, gene, construct, and vector is present or expressed in at least one megasporocyte and / or in female gametes.

16. The method according to claim 14 or 15, wherein the primary transformant is pollinated in step c) using pollen from a tetraploid plant or its own pollen.

17. The method of claim 14, wherein the plant, plant part or plant cell is capable of parthenogenesis.

18. The method according to claim 14, wherein the seeds are selected from the following species: *Taraxacum*, *Lactuca*, *Vaccaria*, *Capsicum*, *Solanum*, *Cucumis*, *Corn*, *Cotton*, *Soybean*, *Wheat*, *Oryza*, *Allium*, *Brassica*, *Sunflower*, *Begonia*, *Chicory*, *Chrysanthemum*, *Pennisetum*, *Rye*, *Barley*, *Alfalfa*, *Phaseolus*, *Rosa*, *Lilium*, *Coffee*, *Flaxum*, *Cannabis ... and *Cannabis*.

19. A method for producing ploid sporophytes, plant parts, or plant cells, comprising the following steps: a) Transform a plant, plant part or plant cell with at least one of the following: the polynucleotide fragment according to any one of claims 1-3, the chimeric gene according to claim 7, the genetic construct according to claim 8 or 9, and the nucleic acid vector according to any one of claims 10-13.

20. The method of claim 19, further comprising the following steps: (b) to regenerate the plant, wherein at least one of the polynucleotides, fragments, genes, constructs and vectors is present or expressed in the female ovary.

21. The method of claim 19 or 20, wherein at least one of the polynucleotide, fragment, gene, construct, and vector is present or expressed in at least one macrospore mother cell and in at least one female gamete.

22. The method according to claim 19 or 20, wherein the polynucleotide fragment according to any one of claims 1-3 is integrated into the genome of the plant, plant part or plant cell.

23. A method for producing ploid sporophytes, plant parts, or plant cells, comprising the following steps: a) Modifying an endogenous polynucleotide or polynucleotide fragment in a plant, plant part or plant cell such that the modified plant, plant part or plant cell contains the polynucleotide fragment according to any one of claims 1-3.

24. The method of claim 23, further comprising the following steps: b) Regenerate the plant.

25. The method of claim 23 or 24, wherein the endogenous polynucleotide is a polynucleotide of a vacuole protein sorting-related protein gene.

26. The method of claim 23, wherein the modified polynucleotide or fragment of a polynucleotide is expressed or encodes a polypeptide.

27. The method of claim 23, wherein the modified polynucleotide or fragment of the polynucleotide is present at least in the female ovary.

28. The method of claim 27, wherein the modified polynucleotide or fragment of the polynucleotide is present in at least one of the megaspore mother cell and the female gamete.

29. The method of claim 23, wherein the modification is performed by at least one of the following: a) Introducing or expressing at least one site-specific nuclease in the plant, plant part, or plant cell; and b) Use oligonucleotides for oligonucleotide-directed mutagenesis.

30. The method of claim 29, wherein the nuclease is selected from: Cas / RNA CRISPR nucleases, zinc finger nucleases, broad-spectrum nucleases, and TAL effector nucleases.

31. The method according to claim 29 or 30, wherein the oligonucleotide is a single-stranded oligonucleotide.

32. A method for producing hybrid plants, comprising the step of hybridizing pollen from a ploid sporophyte plant obtainable by any one of claims 19-31 with a sexually reproducing plant.

33. The method according to claim 19 or 23, wherein the plant, plant part or plant cell is derived from a species selected from the genera *Taraxacum*, *Lactuca*, *Vitis*, *Capsicum*, *Solanum*, *Cucumis*, *Corn*, *Cotton*, *Soybean*, *Wheat*, *Oryza*, *Allium*, *Brassica*, *Sunflower*, *Begonia*, *Chicory*, *Chrysanthemum*, *Pennisetum*, *Rye*, *Barley*, *Alfalfa*, *Phaseolus*, *Rosa*, *Lilium*, *Coffee*, *Flaxum*, *Cannabis ... and *Cannabis*.

34. A method for producing a clone of a hybrid plant, comprising the following steps: a) The pollen of a ploid sporophyte plant obtainable by any one of claims 19-22 is hybridized with a sexually reproducing plant to produce F1 hybrid seeds; b1) Selecting F1 plants that contain or express at least the polynucleotide fragment of any one of claims 1-3 or the polypeptide fragment of any one of claims 4-6 in the female ovary; and c) Harvest the seeds.

35. The method of claim 34, further comprising the following steps: b2) Before harvesting the seeds in step c), the selected F1 plants are pollinated to induce seed production.

36. The method of claim 34, further comprising the following steps: d) Growing hybrid clone plants from the seeds.

37. The method according to any one of claims 34-36, wherein the polynucleotide, fragment or polypeptide is contained in at least one of the megasporocyte and in the female gamete.

38. The method according to any one of claims 34-36, wherein the selected F1 plant is pollinated with pollen from a tetraploid plant in step c.

39. The method of claim 34, wherein the plant, plant part or plant cell is capable of parthenogenesis and wherein the clone is an apomixis clone.

40. Use of the polynucleotide fragment according to any one of claims 1-3 in inducing ploid sporophyte formation in plants.

41. Use of the polynucleotide according to any one of claims 1-3 for gene stacking.

42. Use of the polynucleotide fragment according to any one of claims 1-3 for developing or identifying markers of the ploidy sporophyte formation trait.

43. Use of the polynucleotide fragment according to any one of claims 1-3 for inducing polyploidization in plants.

44. Use of the polynucleotide fragment according to any one of claims 1-3 for inducing apomixis in plants.

Citation Information

Patent Citations

  • An RNA plant virus vector or portion thereof, a method of construction thereof, and a method of producing a gene derived product therefrom

    EP0067553A2

  • Process for the introduction of expressible genes into plant cell genomes and agrobacterium strains carrying hybrid Ti plasmid vectors useful for this process

    EP0116718A1

  • A process for the incorporation of foreign DNA into the genome of dicotyledonous plants; a process for the production of Agrobacterium tumefaciens bacteria

    EP0120515A2

  • Water-agglomeration method for depeptide sweetened products

    EP0120561A1

  • Regulation of gene expression by employing translational inhibition utilizing mRNA interfering complementary RNA

    EP0140308B1