A molecular marker, primer and method for identifying hybrid purity of zucchini squash variety and application thereof

CN122833197APending Publication Date: 2026-09-29HUNAN VEGETABLE RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611043508.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-14
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

主要有SNP芯片平台、样本高通量的原位扫描平台、位点和样本高通量的靶向测序技术等,但这些技术平台的费用较高,经济适用性较差

Benefits of technology

本发明以嫩早南瓜杂交种双亲的芯片测序数据为基础,找到双亲之间的差异的SNP,开发了3对KASP引物,采用3个SNP标记,利用荧光检测技术,快速的区分父本、母本和杂交种,可以直接用于早熟嫩早南瓜商品种纯度的鉴定,可依赖该分子标记进行商品种真假杂交种的评定,且能够有效保护该品种的权益。通过早期利用该分子标记可以快速鉴定该品种的纯度,有效减小因为假杂种造成的损失,减少了可能的纠纷,提高了鉴定效率和准确性并为品种的保护提供了证据。因此,本申请在早熟嫩早南瓜品种杂交种的鉴定及品种保护中具有重大应用价值。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122833197A_ABST
    Figure CN122833197A_ABST
Patent Text Reader

Abstract

This invention discloses a molecular marker, primers, method, and application for identifying the purity of hybrids of early-maturing pumpkin varieties. Based on microarray sequencing data of the parents of an early-maturing pumpkin hybrid, this method identifies the distinguishing SNPs between the parents, develops three pairs of KASP primers, and uses three SNP markers. Utilizing fluorescence detection technology, it rapidly distinguishes between the paternal, maternal, and hybrid parents. This method can be directly used to identify the purity of commercial early-maturing pumpkin varieties. It can be relied upon to evaluate true and false hybrids in commercial varieties and effectively protect the rights and interests of the variety. Early use of this molecular marker allows for rapid identification of variety purity, effectively reducing losses caused by false hybrids, minimizing potential disputes, improving identification efficiency and accuracy, and providing evidence for variety protection. Therefore, this application has significant application value in the identification and protection of hybrids of early-maturing pumpkin varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of plant breeding technology, and specifically relates to a molecular marker, primer, method and application for identifying the purity of hybrid varieties of early-maturing pumpkin. Background Technology

[0002] Early-maturing squash, also known as tender-fruit squash, is a type of Chinese squash primarily grown for its young fruit. It is generally an early or very early-maturing variety, with strong cold and heat resistance. When stir-fried, the young squash has a pumpkin aroma, a sweet and crisp taste, and is of higher quality than American squash or zucchini. Early-maturing squash has a smaller fruit size, making it particularly suitable for small and medium-sized families in urban and rural areas. Currently, provinces with large cultivation areas for early-maturing squash include Hunan, Hubei, Guizhou, Yunnan, Sichuan, and Guangdong. The cultivation area of ​​early-maturing squash in southern China accounts for about 30% of the total squash cultivation area in China. In recent years, many regions in northern China have gradually begun to cultivate and consume early-maturing squash.

[0003] Artificial emasculation and pollination are the main methods for hybrid seed production of early-maturing pumpkins. However, the purity of the seeds cannot be guaranteed. Therefore, it is necessary to identify the purity of early-maturing pumpkin hybrids to ensure that superior varieties can be applied more efficiently in production.

[0004] Seed purity reflects the degree of varietal uniformity, serving as a crucial indicator of seed quality and a primary basis for seed grading, and is positively correlated with economic benefits. Therefore, purity assessment of hybrid seeds is an indispensable step in the actual production of hybrid varieties and is key to ensuring the better utilization of superior varieties.

[0005] Conventional hybrid purity identification primarily involves field cultivation and morphological testing of plants. This method is time-consuming, labor-intensive, and susceptible to environmental and human factors, resulting in low accuracy. With the development of genome sequencing technology, SNP marker sites targeting both parents have been developed. Currently, SNP detection, as the latest molecular marker detection method, has been recommended as a method for variety identification at the DNA level. Key technologies include SNP microarray platforms, high-throughput in-situ scanning platforms, and targeted sequencing technologies for both sites and samples. However, these technologies are expensive and lack economic viability. Summary of the Invention

[0006] This application is made in view of the above-mentioned problems, and its purpose is to provide a molecular marker, primers, method and application for identifying the purity of hybrids of early-maturing pumpkin varieties. The molecular marker can be used to identify the purity of hybrids of early-maturing pumpkin varieties.

[0007] Specifically, the first aspect of this application provides the application of molecular markers in the following aspects, including: (1) Predict, screen and / or identify the purity of early-maturing pumpkin hybrids; And / or, (2) Improvements to early-maturing squash; And / or, (3) accelerate the selection of hybrid plants in offspring young early pumpkins through molecular marker-assisted selection breeding; The molecular marker includes at least one of a first molecular marker, a second molecular marker, and a third molecular marker; The first molecular marker is located at base 9213402 on chromosome 3 of the squash genome; The second molecular marker is located at base 10840907 on chromosome 5 of the squash genome; The third molecular marker is located at base 10779704 on chromosome 6 of the early-maturing squash genome.

[0008] This application utilizes parental microarray sequencing data to develop parental polymorphic KASP primers, which, combined with seedling sampling, DNA extraction, and laboratory fluorescence quantitative detection, enable rapid and accurate purity identification.

[0009] A second aspect of this application provides a primer set that satisfies at least one of the following characteristics: (1) Includes a first reverse primer set, the first reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO.1; (2) Includes a first forward primer set, the first forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.2; (3) Includes a second forward primer set, the second forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.3; (4) Includes a second reverse primer set, the second reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO.4; (5) Includes a third forward primer set, said third forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.5; (6) Includes a fourth forward primer set, said fourth forward primer set comprising a nucleotide sequence as shown in SEQ ID NO. 6; (7) Includes a fifth forward primer set, said fifth forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.7; (8) Includes a third reverse primer set, said third reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO. 8; (9) Includes a fourth reverse primer set, said fourth reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO. 9; (10) Includes nucleotide sequences that encode the same protein as any of the nucleotide sequences shown in (1) to (9), but are different from any of the nucleotide sequences shown in (1) to (9) due to the degeneracy of the genetic code; (11) A nucleotide sequence obtained by substituting, deleting or adding one or more nucleotide sequences to any of the nucleotide sequences shown in (1) to (9), and which has the same or similar function as any of the nucleotide sequences shown in (1) to (9); (12) Includes a nucleotide sequence having at least 90% sequence homology with any of the nucleotide sequences described in (1) to (9); (13) includes a nucleotide sequence having at least 95% sequence homology with any of the nucleotide sequences described in (1) to (9); (14) includes a nucleotide sequence having at least 99% sequence homology with any of the nucleotide sequences described in (1) to (9); (15) A nucleotide sequence comprising the nucleotide sequence described in any one of (1) to (9) with 1 to 3 base deletions and / or substitutions; (16) Includes a connector sequence, as shown in SEQ ID NO.10 or SEQ ID NO.11.

[0010] As a further aspect of the present invention, the primer set includes: (1) A first reverse primer set, comprising a nucleotide sequence as shown in SEQ ID NO.1; (2) A first set of forward primers, comprising a nucleotide sequence as shown in SEQ ID NO.2; (3) A second set of forward primers, comprising a nucleotide sequence as shown in SEQ ID NO.3.

[0011] As a further aspect of the present invention, the primer set includes: (1) A first reverse primer set, comprising a nucleotide sequence as shown in SEQ ID NO.1; (2) A first forward primer set, comprising a nucleotide sequence and an adapter sequence as shown in SEQ ID NO.2; (3) A second forward primer set, comprising a nucleotide sequence and an adapter sequence as shown in SEQ ID NO.3.

[0012] As a further aspect of the present invention, the primer set includes: (1) A second reverse primer set, comprising a nucleotide sequence as shown in SEQ ID NO.4; (2) A third set of forward primers, the third set of forward primers comprising a nucleotide sequence as shown in SEQ ID NO.5; (3) A fourth set of forward primers, the fourth set of forward primers comprising a nucleotide sequence as shown in SEQ ID NO.6.

[0013] As a further aspect of the present invention, the primer set includes: (1) A second reverse primer set, comprising a nucleotide sequence as shown in SEQ ID NO.4; (2) A third forward primer set, the third forward primer set comprising a nucleotide sequence and an adapter sequence as shown in SEQ ID NO.5; (3) A fourth forward primer set, the fourth forward primer set comprising a nucleotide sequence and an adapter sequence as shown in SEQ ID NO.6.

[0014] As a further aspect of the present invention, the primer set includes: (1) A fifth set of forward primers, the fifth set of forward primers comprising a nucleotide sequence as shown in SEQ ID NO.7; (2) A third reverse primer set, the third reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO.8; (3) A fourth reverse primer set, the fourth reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO.9.

[0015] As a further aspect of the present invention, the primer set includes: (1) A fifth set of forward primers, the fifth set of forward primers comprising a nucleotide sequence as shown in SEQ ID NO.7; (2) A third reverse primer set, the third reverse primer set comprising a nucleotide sequence and an adapter sequence as shown in SEQ ID NO.8; (3) A fourth reverse primer set, which includes a nucleotide sequence and an adapter sequence as shown in SEQ ID NO.9.

[0016] The third aspect of this application provides for the application of the primer set described in the second aspect in the following aspects: (1) Predict, screen and / or identify the purity of early-maturing pumpkin hybrids; And / or, (2) Improvements to early-maturing squash; And / or, (3) accelerate the selection of hybrid plants in offspring young early pumpkins through molecular marker-assisted selection breeding; And / or, (4) a kit for predicting, screening and / or identifying the purity of early-maturing pumpkin hybrids.

[0017] The fourth aspect of this application provides a reagent kit, characterized in that it includes the primer set as described in the second aspect of this application.

[0018] This application's fifth aspect provides for the application of the reagent kit described in the fourth aspect of this application in the following aspects: (1) Predict, screen and / or identify the purity of early-maturing pumpkin hybrids; And / or, (2) Improvements to early-maturing squash; And / or, (3) accelerate the selection of hybrid plants in offspring young early pumpkins through molecular marker-assisted selection breeding.

[0019] The sixth aspect of this application provides a method for identifying the purity of early-maturing pumpkin hybrids, comprising the following steps: extracting genomic DNA from early-maturing pumpkins, identifying the genotype of early-maturing pumpkins using the primer set as described in the second aspect of this application and / or the kit as described in the fourth aspect of this application, and identifying the purity of early-maturing pumpkin hybrids.

[0020] According to some embodiments of this application, the criteria for identifying the purity of early-maturing pumpkin hybrids based on the genotype are as follows: (1) If the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is CC, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is TT, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is GG, it is determined to be the paternal type. (2) If the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is TT, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is CC, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is AA, it is determined to be the maternal type. (3) If the 9213402nd base pair of chromosome 3 of the young early pumpkin genome is TC, the 10840907th base pair of chromosome 5 of the young early pumpkin genome is CT, and the 10779704th base pair of chromosome 6 of the young early pumpkin genome is AG, it is determined to be a true hybrid. (4) If any of the characteristics described in (1) to (3) are not met, it is determined to be another type of hybrid.

[0021] The seventh aspect of this application provides a method for identifying the purity of early-maturing pumpkin hybrids, comprising the following steps: extracting genomic DNA from early-maturing pumpkins, amplifying it by PCR using the primer set described in the second aspect of this application and / or the kit described in the fourth aspect of this application, and determining the purity of the corresponding variety based on the fluorescence signal of the PCR product.

[0022] According to some embodiments of this application, the standard for evaluating the purity of early-maturing pumpkin hybrids based on the fluorescence signal is as follows: (1) If, by means of fluorescence signal, the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is CC, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is TT, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is GG, it is determined to be the paternal type; (2) If, by means of fluorescence signal, the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is TT, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is CC, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is AA, it is determined to be the maternal type; (3) If, by means of fluorescence signal, the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is TC, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is CT, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is AG, it is determined to be a true hybrid. (4) If other genotype combinations appear, they are determined to be other types of hybrids.

[0023] According to some embodiments of this application, the PCR amplification reaction procedure includes: (I) Pre-denaturation at 94℃~95℃ for 14min~15min, denaturation at 94℃~95℃ for 20s~30s, annealing and extension at 65℃~56℃ for 60s~70s, 10~15 cycles, each cycle decreasing by 0.6℃~0.8℃; And / or, (II), denaturation at 94℃~95℃ for 20s~30s, annealing and extension at 55℃~57℃ for 60s~70s, 26~30 cycles.

[0024] According to some embodiments of this application, the early-maturing pumpkin variety includes early-maturing young pumpkin varieties.

[0025] According to some embodiments of this application, the early-maturing pumpkin variety includes Early-maturing No. 12.

[0026] This invention has at least the following technical effects: This invention, based on microarray sequencing data of the parents of an early-maturing pumpkin hybrid, identifies the distinguishing SNPs between the parents and develops three pairs of KASP primers. Using three SNP markers and fluorescence detection technology, it rapidly distinguishes between the paternal, maternal, and hybrid parents. This can be directly used to identify the purity of commercial early-maturing pumpkin varieties. The molecular markers can be relied upon to assess true and false hybrids in commercial varieties, effectively protecting the rights and interests of the variety. Early use of these molecular markers allows for rapid identification of variety purity, effectively reducing losses caused by false hybrids, minimizing potential disputes, improving identification efficiency and accuracy, and providing evidence for variety protection. Therefore, this application has significant application value in the identification and protection of early-maturing pumpkin hybrid varieties. Attached Figure Description

[0027] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.

[0028] Figure 1 Phenotypic diagrams of the female parent, male parent, and hybrid of "Nenzao Yichuanling"; Figure 2 Chip analysis was used to detect differences between the maternal and paternal parents of early-maturing pumpkins. Figures 3-5 Genotyping diagram of the molecular markers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 of this invention in identifying true and false hybrids of "Nenzao Yichuanling": Figure 3 This indicates that the PCR product is the fluorescence signal corresponding to primer Cmo_Chr03-9213402; Figure 4 This indicates that the PCR product is the fluorescence signal corresponding to primer Cmo_Chr05-10840907; Figure 5 This indicates that the PCR product is the fluorescence signal corresponding to primer Cmo_Chr06-10779704; In the diagram, point A represents the parent type; B indicates the parent type; C indicates the hybrid F1 ("Nenzao Yichuanling"). Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described embodiments are merely some embodiments of the invention, and not all embodiments.

[0030] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0031] The "range" disclosed in this application is defined by a lower limit and an upper limit. A given range is defined by selecting a lower limit and an upper limit, which define the boundaries of a particular range. Ranges defined in this way can include or exclude endpoints and can be arbitrarily combined; that is, any lower limit can be combined with any upper limit to form a range. For example, if ranges of 60-120 and 80-110 are listed for a specific parameter, it is expected that ranges of 60-110 and 80-120 are also included. Furthermore, if minimum range values ​​of 1 and 2 are listed, and if maximum range values ​​of 3, 4, and 5 are listed, then the following ranges are all expected: 1-3, 1-4, 1-5, 2-3, 2-4, and 2-5. In this application, unless otherwise stated, the numerical range "ab" represents a shortened representation of any combination of real numbers between a and b, where a and b are real numbers. For example, the numerical range "0-5" indicates that all real numbers between "0-5" have been listed in this article; "0-5" is simply a shortened representation of these numerical combinations. Furthermore, when a parameter is stated as an integer ≥2, it is equivalent to disclosing that the parameter is, for example, an integer such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, etc.

[0032] Unless otherwise specified, all embodiments and optional embodiments of this application can be combined to form new technical solutions.

[0033] Unless otherwise specified, all technical features and optional technical features of this application may be combined to form new technical solutions.

[0034] Unless otherwise specified, all steps in this application may be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), indicating that the method may include steps (a) and (b) performed sequentially, or it may include steps (b) and (a) performed sequentially. For example, the mention that the method may also include step (c) indicates that step (c) may be added to the method in any order. For example, the method may include steps (a), (b), and (c), or it may include steps (a), (c), and (b), or it may include steps (c), (a), and (b), etc.

[0035] Unless otherwise specified, the terms "comprising" and "including" as used in this application can be open-ended or closed-ended. For example, "comprising" and "including" can mean that other components not listed may also be included, or that only the listed components may be included.

[0036] Unless otherwise specified, the term "or" is inclusive in this application. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, the condition "A or B" is satisfied by any of the following conditions: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).

[0037] Before describing this application in detail, it should be understood that this application is not limited to the specific embodiments, which can certainly be modified. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting. As used in this specification and the appended claims, singular and singular terms, such as “an,” “a,” and “described,” include the plural referents unless explicitly stated otherwise. Thus, for example, references to “plant,” “the plant,” or “a plant” also include multiple plants; moreover, depending on the context, the use of the term “plant” can also include genetically similar or identical offspring of that plant; the use of the term “nucleic acid” optionally includes multiple copies of the nucleic acid molecule; similarly, the term “probe” optionally (and generally) covers a number of similar or identical probe molecules.

[0038] Unless otherwise specified, nucleic acids are written from left to right in a 5′ to 3′ direction. Numerical ranges mentioned in this specification include the numerical values ​​that define the range, and include every integer or any non-integer portion within the defined range. Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. Although any methods and materials similar to or equivalent to those described herein may be used in the practice of testing this application, the preferred methods and materials are those described herein. In the description and claims of this application, the following terms will be used according to the definitions listed below.

[0039] Certain definitions used in this specification and claims are provided below. To provide a clear and consistent understanding of the specification and claims (including the scope to be given by such terms), the following definitions are provided: “Agronomy,” “agronomic traits,” and “agronomic performance” refer to the traits (and underlying genetic factors) of a given plant variety that contribute to yield over the growing season. Individual agronomic traits include seedling vigor, nutrient potential, stress tolerance, disease resistance or tolerance, insect resistance or tolerance, herbicide resistance, branching, flowering, seed formation, seed size, seed density, lodging resistance, threshing rate, and fruit browning.

[0040] The term "allele" refers to any one or more alternative forms of a gene sequence. For example, in diploid cells or organisms, the two alleles of a given sequence typically occupy corresponding loci on a pair of homologous chromosomes. Regarding SNP markers, an allele is a specific nucleotide base present at a particular SNP locus in that individual plant.

[0041] An allele is "associated" with a trait when it is a DNA sequence that affects the expression of that trait, or a part of an allele, or is linked to that allele. The presence of that allele is an indicator of how the trait will be expressed.

[0042] The term "amplification" in the context of nucleic acid amplification refers to any process that results in an additional copy of a selected nucleic acid (or transcribed from it). Typical amplification methods include a variety of polymerase-based replication methods, including polymerase chain reaction (PCR), ligase-mediated methods such as ligase chain reaction (LCR), and RNA polymerase-based amplification methods (e.g., transcription). An "amplifier" is an amplified nucleic acid, for example, produced by amplifying a template nucleic acid using any available amplification method (e.g., PCR, LCR, transcription, etc.).

[0043] The term "chromosome segment" refers to a continuous linear segment of genomic DNA that exists on a single chromosome in a plant.

[0044] The term "complementary sequence" refers to a nucleotide sequence that is complementary to a given nucleotide sequence, i.e., the above sequences are associated by the Watson-Crick base pairing rule.

[0045] "Cultivated species" and "variety" are used synonymously to refer to a group of plants within the same species (e.g., early-maturing squash) that share certain genetic traits that distinguish them from other possible varieties within the same species.

[0046] "Superior strains" are agronomically superior strains that have been developed through many rounds of breeding and selection for superior agronomic performance. Many superior strains are available and are known to those skilled in the field of early squash breeding.

[0047] "Superior population" refers to a mixed population of superior individuals or strains that can be used to represent the current level of technology in terms of agronomically superior genotypes of a given crop species, such as early-maturing squash.

[0048] A "favorable allele" is an allele located at a specific locus that confers or contributes to an agronomically desired phenotype and allows for the identification of plants possessing that agronomically desired phenotype. A marked favorable allele is a marker allele that separates from the favorable phenotype.

[0049] A "gene map" is a description of gene linkages between loci on one or more chromosomes (or linkage groups) within a given species, typically represented as a graph or table. For each gene map, the distance between loci is measured by how frequently their alleles appear together within a population (their recombination frequency). Alleles can be detected using DNA or protein markers, or observable phenotypes. A gene map is the product of the mapping population, the types of markers used, and the polymorphism of each marker between different populations. The genetic distance between loci can differ between gene maps. However, by using shared markers, information can be correlated between maps. Those skilled in the art can use shared marker locations to identify the location of markers of interest and other loci on various gene maps. The order of loci should be invariant between maps; however, small variations in marker order often occur due to, for example, markers detecting alternative repetitive loci in different populations, differences in statistical methods used to orient markers, new mutations, or experimental errors.

[0050] "Genetic recombination frequency" is the frequency of exchange events (recombination) between two loci. Recombination frequency can be observed by tracking the segregation of markers and / or traits after meiosis.

[0051] A "genome" is a complete set of DNA or genes carried by chromosomes or sets of chromosomes.

[0052] "Genome type" refers to the genetic makeup of a cell or organism.

[0053] "Germanic material" refers to the genetic material that forms the physical basis of the hereditary qualities of an organism. As used herein, germplasm includes seeds and living tissues from which new plants can grow; or other plant parts that can be cultured into the whole plant, such as leaves, stems, pollen, or cells. Germanic resources provide plant breeders with a source for improving the genetic traits of commercially grown varieties.

[0054] A haplotype is an individual's genotype at multiple loci, that is, a combination of alleles. Typically, the loci described by a haplotype are physically and genetically linked, meaning they are located on the same chromosomal segment. The term "haplotype" can refer to an allele at a specific locus or to alleles at multiple loci along a chromosomal segment.

[0055] An individual is "homozygous" if it possesses only one type of allele at a given locus (e.g., a diploid individual has one copy of the same allele at a locus on each of its two homologous chromosomes). If more than one allele type exists at a given locus, the individual is "heterozygous" (e.g., a diploid individual has one copy of each of two different alleles). The term "homogeneity" indicates that members of a population share the same genotype at one or more specific loci. In contrast, the term "heterogeneity" is used to indicate that individuals within a population have different genotypes at one or more specific loci.

[0056] The term "insertion or deletion" refers to an insertion or deletion, in which one strain may be referred to as having an inserted nucleotide or DNA fragment relative to a second strain, or the second strain may be referred to as having a deleted nucleotide or DNA fragment relative to a first strain.

[0057] "Infiltration" refers to the transfer or introduction of genes, quantitative trait loci (QTLs), marker loci, haplotypes, marker profiles, traits, or trait loci from the genome of one plant into the genome of another plant.

[0058] A “strain” or “line” is a group of individuals that are similarly related, typically inbred to a certain degree and usually homozygous and homogeneous (homogeneous or nearly homozygous) at most loci. A “subline” is an inbred subgroup of offspring that is genetically different from other similar inbred subgroups originating from the same ancestor. Sublines are conventionally obtained by inbreeding seeds from a single early squash plant selected in the F3 to F5 generations until the remaining segregating loci are “fixed” or homozygous at most or all loci. Commercial early squash varieties (or strains) are typically produced by aggregating (“converging”) the self-pollinated offspring of a single F3 to F5 plant from a controlled cross between two genetically different parents. While these varieties generally appear homogeneous, the self-pollinated varieties derived from the selected plants will eventually (e.g., F8) become a mixture of homozygous plants that are heterologous in the initially selected F3 to F5 plants and can be genotyped at any locus. Marker-based sublines that differ from each other based on qualitative polymorphism at one or more specific marker loci at the DNA level are obtained by genotyping seed samples derived from single self-pollinated progeny of selected F3-F5 plants. The seed samples can be directly genotyped as seeds or as plant tissue grown from such seed samples. Optionally, seeds sharing a common genotype at specific loci (or multiple loci) are merged to provide genetically homologous sublines at the identified loci important for the trait of interest.

[0059] Linkage refers to the phenomenon where alleles on the same chromosome tend to co-segregate more frequently than would be expected by chance when their transmission is independent. Genetic recombination occurs throughout the genome at a presumed random frequency. Gene maps are constructed by measuring the frequency of recombination between pairs of traits or markers. The closer the traits or markers are to each other on the chromosome, the lower the frequency of recombination and the higher the degree of linkage. If traits or markers co-segregate substantially, they are considered linked in this paper. A recombination probability of 1 / 100 per generation is defined as a map distance of 1.0 centimoles (1.0 cM).

[0060] Genetic elements or genes located on a single chromosomal segment are physically linked. Advantageously, two loci are located close to each other, such that recombination between homologous chromosome pairs during meiosis does not occur at a high frequency between these two loci, for example, causing linked loci to co-segregate at least about 90% of the time, such as 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.75%, or more. Genetic elements located within a chromosomal segment are also genetically linked, typically within a genetic recombination distance of less than or equal to 50 centimoles (cM), such as about 49, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.75, 0.5, or 0.25 cM or less. That is, two genetic elements within a single chromosome segment recombine with each other at a frequency of less than or equal to about 50%, for example, about 49%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, or 0.25% or lower during meiosis.

[0061] In this application, the phrase "closely linked" for a locus refers to recombination between two linked loci occurring at a frequency of 10% or less (i.e., separated by no more than 10 cM on the genome map). In other words, closely linked loci co-segregate at least 90% of the time. Marker loci are particularly useful in this application when they exhibit a significant probability (linkage) of co-segregation with the desired trait. Closely linked loci, such as marker loci and second loci, are capable of exhibiting recombination frequencies of 10% or less, preferably about 9% or less, more preferably about 8% or less, more preferably about 7% or less, even more preferably about 6% or less, more preferably about 5% or less, even more preferably about 4% or less, more preferably about 3% or less, and even more preferably about 2% or less. In a highly preferred embodiment, the associated loci exhibit a recombination frequency of about 1% or less, for example, about 0.75% or less, more preferably about 0.5% or less, or more preferably about 0.25% or less. Two loci located on the same chromosome are also referred to as "neighboring" loci, meaning that recombination between them occurs at a frequency of less than 10% (e.g., approximately 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or lower). In some cases, two different markers can have the same genomic map coordinates. In such cases, the two markers are so close together that recombination between them occurs at a frequency so low as to be undetectable.

[0062] When the relationship between two genetic elements (such as a antagonistic contributing genetic element and a neighboring marker) is involved, "coupled" linkage refers to a state in which the "favorable" allele at the locus of interest is physically linked to the "favorable" allele at the corresponding linked marker locus on the same chromosomal chain. In coupled linkage, offspring that inherit both favorable alleles inherit that chromosomal chain. In repulsive linkage, the "favorable" allele at the locus of interest is physically linked to the "unfavorable" allele at the neighboring marker locus, and the two "favorable" alleles are not inherited together (i.e., the two loci are "opposite" to each other).

[0063] "Linkage disequilibrium" refers to the phenomenon where, when segregating from parents to offspring, alleles tend to remain together in linkage groups at a higher frequency than expected from their individual frequencies. A state of linkage disequilibrium implies that related loci are physically close enough along the length of the chromosome that they segregate together at a higher frequency than random (i.e., non-random). Markers exhibiting linkage disequilibrium are considered linked. Linked loci co-segregate more than 50% of the time, for example, from about 51% to about 100% of the time. In other words, two co-segregating markers have a recombination frequency of less than 50% (and, by definition, are less than 50 cM apart on the same linkage group). As used herein, linkage can occur between two markers, or between a marker and a locus influencing the phenotype. A marker locus can be "associated" (associated with) a trait. The degree of linkage between a marker locus and a locus influencing the phenotypic trait is measured, for example, as a statistical probability (e.g., an F-statistic or LOD score) of co-segregation of the molecular marker and the phenotype.

[0064] A "linkage group" (LG) is a trait or marker that is largely co-segregated. A linkage group generally corresponds to a chromosomal region containing the genetic material encoding that trait or marker. Therefore, linkage groups can be broadly attributed to a specific chromosome.

[0065] A "locus" is a defined segment of DNA. For example, it can refer to the location on a chromosome where a nucleotide, gene, sequence, or marker is located.

[0066] "Map location" is a designated position on a gene map relative to linked genetic markers where a specific marker can be found within a given species.

[0067] "Mapping" is the process of defining the linkage relationships of loci through the use of standard genetic principles such as genetic markers, marker segregation, and recombination frequencies.

[0068] "Marker," "molecular marker," or "marked locus" is a term used to denote a nucleic acid or amino acid sequence that is sufficiently unique to characterize a specific locus on the genome. Examples include restriction fragment length polymorphism (RFLP), simple repeat sequence (SSR), target region amplification polymorphism (TRAP), isoenzyme electrophoresis, random amplified polymorphic DNA (RAPD), random primer polymerase chain reaction (AP-PCR), DNA amplification fingerprint (DAF), sequence-specific amplified region (SCAR), amplified fragment length polymorphism (AFLP), and single nucleotide polymorphism (SNP). Other types of molecular markers are also known in the art, and phenotypic traits can also be used as markers in the methods described. All markers are used to define specific loci on the genome of *Cucurbita moschata*. Therefore, each marker is an indicator of a specific segment of DNA with a unique nucleotide sequence. Map locations provide a measure of the relative position of a specific marker to each other. When a trait is described as being linked to a given marker, it should be understood that the actual DNA segment whose sequence affects that trait is generally cosegregated with that marker. If markers are identified on both sides of a trait, a more precise and definitive localization of that trait can be obtained. By measuring the presence of markers in hybrid offspring, the presence of a trait can be detected by a relatively simple molecular test without actually assessing the presence of the trait itself, which can be difficult and time-consuming, as actual assessment of a trait requires the plant to grow to a stage where the trait can be expressed.

[0069] "Marker-assisted selection" refers to the process of selecting one or more plants for a desired one or more traits (wherein the nucleic acid is associated with the desired trait) by detecting one or more nucleic acids from the plant, and then selecting plants or germplasm that have the aforementioned one or more nucleic acids.

[0070] In some examples, multiple marker loci or haplotypes are used to define a “marker spectrum.” As used herein, a “marker spectrum” refers to a combination of two or more marker loci or haplotypes within the genome of a particular plant. For example, in one example, a specific combination of marker loci or haplotypes defines the marker spectrum of a particular plant.

[0071] The terms "phenotype," "phenotypic trait," or "trait" can refer to the observable expression of a gene or series of genes. A phenotype can be observed by the naked eye or by any other assessment method known in the art, such as weighing, counting, measuring (length, width, angle, etc.), microscopy, biochemical analysis, or electromechanical determination. In some cases, a phenotype is directly controlled by a single gene or locus, i.e., a "monogenous trait" or "simple inherited trait." In the absence of significant environmental variation, monogenic traits can segregate in a population, producing a "qualitative" or "discrete" distribution; that is, the phenotype is divided into discrete categories. In other cases, a phenotype is the result of multiple genes and can be considered a "polygenic trait" or "complex trait." Polygenic traits segregate in a population, producing a "quantitative" or "continuous" distribution; that is, the phenotype cannot be separated into discrete categories. Both monogenic and polygenic traits can be influenced by the environment in which they are expressed, but polygenic traits tend to have a larger environmental component.

[0072] "Favorable traits" or "favorable phenotypes," such as, for example, resistance to fruit browning, are phenotypes desired in agronomy.

[0073] The term "plant" includes the entire plant, whether immature or mature, including plants from which seeds or grains or anthers have been removed. Seeds or embryos that will produce a plant are also considered plants.

[0074] "Plant part" refers to any part or segment of a plant, including leaves, stems, buds, roots, root tips, anthers, seeds, plumules, pollen, ovules, flowers, cotyledons, hypocotyls, pods, branches, stems, tissues, tissue cultures, cells, etc.

[0075] "Polymorphism" refers to the variation or difference between two related nucleic acids. "Nucleotide polymorphism" refers to the fact that two nucleic acids, when compared for maximum similarity, are different nucleotides in a sequence compared to their related sequences.

[0076] The terms “polynucleotide,” “polynucleotide sequence,” “nucleic acid sequence,” “nucleic acid fragment,” and “oligonucleotide” are used interchangeably in this document. These terms encompass nucleotide sequences, among others. Polynucleotides can be polymers of RNA or DNA, and they can be single-stranded or double-stranded, optionally containing synthetic, non-natural, or modified nucleotide bases. Polynucleotides in the form of DNA polymers can consist of one or more strands of cDNA, genomic DNA, synthetic DNA, or mixtures thereof.

[0077] A primer is a (synthetic or naturally occurring) oligonucleotide that, when placed under conditions where it is synthesized along its complementary strand by a polymerase, serves as a starting point for nucleic acid synthesis or replication. Typically, primers are oligonucleotides of 10 to 30 nucleic acids in length, but longer or shorter sequences can be used. Primers can be provided in double-stranded form, although single-stranded form is preferred. Primers may also contain detectable markers, such as 5' end markers.

[0078] A "probe" is an oligonucleotide (synthetic or naturally occurring) that is complementary (but not necessarily perfectly complementary) to the polynucleotide of interest and forms a double-stranded structure by hybridizing with at least one strand of the polynucleotide of interest. Typically, probes are oligonucleotides of 10 to 50 nucleic acids in length, but longer or shorter sequences can be used. Probes can also contain detectable tags.

[0079] The terms "label" and "detectable label" refer to detectable molecules, including but not limited to radioactive isotopes, fluorescent agents, chemiluminescent agents, enzymes, enzyme substrates, enzyme cofactors, enzyme inhibitors, chromophores, dyes, metal ions, metal sols, semiconductor nanocrystals, ligands (e.g., biotin, avidin, streptavidin, or hapten). Detectable labels can also include combinations of reporter genes and quenchers, such as those used in FRET or TaqMan probes.

[0080] The term "reporter gene" refers to a substance or part thereof that exhibits a detectable signal that can be suppressed by a quencher. The detectable signal of the reporter gene is, for example, fluorescence within a detectable range.

[0081] The term "quencher" refers to a substance or part thereof that can inhibit, reduce, suppress, or suppress detectable signals generated by the reporter gene.

[0082] As used herein, the terms “quenching” and “fluorescent energy transfer” refer to the process in which, when a reporter gene and a quencher are in close proximity and the reporter gene is excited by an energy source, a major portion of the energy of the excited state is nonradiatively transferred to the quencher, which either dissipates nonradiatively or is emitted at a wavelength different from that of the reporter gene.

[0083] The term “quantitative trait locus” or “QTL” refers to a region of DNA associated with differential expression of a quantitative phenotypic trait in at least one genetic context (e.g., in at least one breeding population). The region of a QTL encompasses one or more genes that affect or are closely related to the trait under consideration.

[0084] A "reference sequence" or "common sequence" is a qualified sequence used as the basis for sequence comparison. A reference sequence for a PHM marker is obtained by sequencing multiple strains at that locus, aligning the nucleotide sequences in a sequence alignment program, and then obtaining the most prevalent nucleotide sequence from the alignment. Polymorphisms present between individual sequences are annotated in the common sequence. The reference sequence is typically not an exact copy of any individual DNA sequence, but rather represents a mixture of available sequences and is used to design primers and probes targeting polymorphisms within that sequence.

[0085] "Recombination frequency" is the frequency of events (recombination) that occur between two loci. Recombination frequency can be observed by tracking the segregation of markers and / or traits during meiosis.

[0086] "Self-fertilization," "self-pollination," or "self-crossing" is the process by which a breeder mates a plant with itself; for example, the second-generation hybrid F2 produces offspring named F2:3 with itself.

[0087] "SNP," or "single nucleotide polymorphism," refers to a sequence variation that occurs when a single nucleotide (A, T, C, or G) in a genome sequence is altered or mutated. When an SNP is mapped to a site on the genome of *Cucurbita moschata*, a "SNP marker" exists. Many techniques for detecting SNPs are known in the art, including allele-specific hybridization, primer extension, and real-time PCR with direct sequencing.

[0088] "Transgenic plant" refers to a plant that contains exogenous polynucleotides within its cells. Generally, the exogenous polynucleotide is stably integrated into the genome, allowing it to be passed down through successive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant expression cassette. As used herein, "transgenic" refers to any cell, cell line, callus, tissue, plant part, or plant whose genotype has been altered due to the presence of exogenous nucleic acid, including the transgenic organism or cell that initially underwent such alteration, and those transgenic organisms or cells derived from the initial transgenic organism or cell through hybridization or asexual reproduction. As used herein, the term "transgenic" does not cover genomic (chromosomal or extrachromosomal) alterations resulting from conventional plant breeding methods (e.g., hybridization) or from naturally occurring events such as random cross-fertilization, infection with a non-recombinant virus, transformation with a non-recombinant bacteria, non-recombinant transposition, or spontaneous mutation.

[0089] The marked "adverse alleles" are isolated along with the adverse plant phenotype, thus providing a marker allele for identifying plants with beneficial effects that can be removed from breeding programs or germplasm.

[0090] The term "vector" is used to refer to a polynucleotide or other molecule that transfers a nucleic acid fragment into a cell. Vectors optionally include portions that mediate the maintenance of the vector and allow for its intended use (e.g., sequences necessary for replication, genes conferring resistance to drugs or antibiotics, multiple cloning sites, operatively linked promoter / enhancer elements that enable the expression of the cloned gene, etc.). Vectors are frequently derived from plasmids, bacteriophages, or plant or animal viruses.

[0091] Here we turn to an example: Several methods well-known in the art are available for detecting molecular markers or sets of molecular markers that cosegregate with a trait of interest. The basic idea behind these methods is marker selection, for which alternative genotypes (or alleles) have significantly different average phenotypes. Therefore, one compares the magnitude or significance level of the difference between alternative genotypes (or alleles) among marker loci. The trait gene is presumed to be located closest to the marker with the largest associated genotypic difference.

[0092] Two such methods for detecting loci of the trait of interest are: 1) population-based association analysis and 2) conventional linkage analysis. In population-based association analysis, lines are derived from existing populations with multiple founders, such as superior breeding lines. Population-based association analysis relies on the decay of linkage disequilibrium (LD) and the idea that in unstructured populations, after so many generations of random mating, only the associations between genes controlling the trait of interest and markers closely linked to those genes will remain. In reality, most existing populations have substructures.

[0093] Therefore, by using data obtained from randomly distributed markers in the genome, individuals are assigned to populations, thereby minimizing the imbalances caused by population structure within individual populations (also known as subpopulations). The use of structured association methods helps control population structure. For each line in a subpopulation, phenotypic values ​​are compared to genotypes (alleles) at each marker locus. Significant marker-trait correlations indicate close proximity between the marker locus and one or more loci involved in the expression of that trait.

[0094] Conventional linkage analysis operates on the same principle; however, LOD is generated by constructing a population from a small number of founders. Founders are selected to maximize the level of polymorphism within the constructed population, and the level of co-segregation of polymorphic loci with a given phenotype is assessed. Various statistical methods have been used to identify significant marker-trait associations. One such method is the interval plotting approach, in which the probability of a gene controlling the trait of interest being located at each of many locations along a gene map (e.g., at 1 cM intervals) is tested. Genotype / phenotype data are used to calculate a LOD score (logarithm of the probability ratio) for each test location. When the LOD score exceeds a threshold, there is significant evidence that a gene controlling the trait of interest is located at that location on the gene map (which would fall between two specific marker loci).

[0095] The molecular markers used in this application for the purity of hybrids of early-maturing pumpkin varieties include a first molecular marker, a second molecular marker, and a third molecular marker.

[0096] According to some embodiments of this application, the first molecular marker is a T-to-C mutation occurring at position 9213402 on chromosome 3 of the Cucurbita genome, with the reference genome being Cucurbita_moschata / v1.

[0097] According to some embodiments of this application, the nucleotide sequence before and after the 9213402nd base on chromosome 3 of the young early pumpkin genome is shown in SEQ ID NO:12.

[0098] SEQ ID NO:12: TTCTGTATCTTCATCTCCGATGGCGAGTGCTTCGTGGCCTCGAGCACCCTCTCCGGTGGCCGGAGCGTCTTGGTCGCCCTTCCCATGCTCGCCCTCCGGGTCCACTTCTCTGGTTTCACGCCGTCTCTTTAGGTTTGTTCTCTATTACCATACATTTTCCTTCTCCAGACAGAGAAATTCGGAATCAACCATACTGCGT[T / C]GTGCAGGAGAAAGTTTGGCTACGATTCCTATAATCAAGCAATGTATTTTGAGGAGGCCTGATTTGAACATTTTGATGACAACAACCACTTATTCTGCCTTGTAAGAATCGATGTTATTTGATTTTCTCTATAAGTTTTTACATTATTAGTTCTTCTTTTCTGTAAACGAGCGAGTGCTTACTTTCTAGACATTTCTCACG.

[0099] According to some embodiments of this application, the second molecular marker is a C-to-T mutation occurring at position 10840907 on chromosome 5 of the Cucurbita genome, with the reference genome being Cucurbita_moschata / v1.

[0100] According to some embodiments of this application, the nucleotide sequence before and after the 10840907th base on chromosome 5 of the young early pumpkin genome is shown in SEQ ID NO:13.

[0101] SEQ ID NO:13: GCTGATCGGTAATACGCAGGAGAGGTACAGCCTCTTGATCTCCGATGCTGTTTCTACCGAGCAGGCTATGCTCGCAACTCAGCTCAATGATATTATTAAGACTGGACGAGTCAAGAAAGGATCAGTTATCCAGTTGATCGATTATGTTTGCAGTCCCATTAAGAGCCGCAAGTCCGTTCTCTCCCTCGTACTTCATTAAT[C / T]ATTTTTTTTCATCGTGATGGATTATTATAATTATATTTTGAAAAAATCTAGCCGTATGTGGTGTCATTATTCATTTTAATTGGCTTACTGGCCTTCTCGGACAACGTGATGCAAGTTTAGAAAACTAGGTGCGCATTTGAAGTTTTGAGATAAATCATTTAGTGATGCATAGTTGAAACGGAGAAATGGTTATTTTGCTC.

[0102] According to some embodiments of this application, the third molecular marker is an A-to-G variation occurring at the 10779704th base position on chromosome 6 of the Cucurbita genome, with the reference genome being Cucurbita_moschata / v1.

[0103] According to some embodiments of this application, the nucleotide sequence before and after the 10779704th base on chromosome 6 of the young early pumpkin genome is shown in SEQ ID NO:14.

[0104] SEQ ID NO:14: AGGAACTCTTTATTGAATTTGCTAGCTAATTCTCAGCTACAGCAATCTAAAACAATCTCAATGGATTCTGCTATGGTGGCAACCAGTGGGTGGGGGCTCCTCCATGAAACTTGATGATGGGGCCATCTAACTTTGGGGCATAGGACCGTTTGGAACTATGCGAGTTGAATCTTAGGCTCAAGTGTTCTATAGGGACAAAG[A / G]CGGATTTGTTTTGAATAGTACTTTTGGGTTTCAAATCTTCCCATTGTTCTTTTGGCACTTATAAACTAGCTGAATTCACTGTGATATTGTCTCTTGAGTCATTATGGTATCTGTAATCTATATGGTTCTGTTCTGTTTGCTCAATGAAAGGTAGGAATCAAGATATATGTCTCATGTTTAATTGAAGCATAAAATGTTCAG.

[0105] According to the "Standard for Nucleotide and / or Amino Acid Sequence Listings and Electronic Sequence Listing Documents", the "[T / C]" and "[C / T]" in the nucleotide sequences shown in SEQ ID NO.12 and SEQ ID NO.13 are replaced with "y" in the sequence listing. "y" means "t or c" and the name comes from "pyrimidine". In the nucleotide sequence shown in SEQ ID NO.14, “[A / G]” is replaced by “r” in the sequence listing. “r” means “g or a” and comes from the word “purine”.

[0106] According to some embodiments of this application, the early-maturing pumpkin variety includes early-maturing young pumpkin varieties.

[0107] According to some embodiments of this application, the early-maturing pumpkin variety includes Nenzao No. 12, also known as "Nenzao Yichuanling".

[0108] According to some embodiments of this application, the primer sequences of this application are shown in Table 1 below.

[0109] Table 1 Primer Sequences

[0110] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:3 is linked to a FAM fluorescent adapter.

[0111] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:6 is linked to a FAM fluorescent adapter.

[0112] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:9 is linked to a FAM fluorescent adapter.

[0113] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:2 is linked to a HEX fluorescent adapter.

[0114] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:5 is linked to a HEX fluorescent adapter.

[0115] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:8 is linked to a HEX fluorescent adapter.

[0116] According to some embodiments of this application, the FAM fluorescent connector sequence is shown in SEQ ID NO:10.

[0117] SEQ ID NO: 10: GAAGGTGACCAAGTTCATGCT.

[0118] According to some embodiments of this application, the HEX fluorescent connector sequence is shown in SEQ ID NO:11.

[0119] SEQ ID NO: 11: GAAGGTCGGAGTCAACGGATT.

[0120] According to some embodiments of this application, the molecular markers are used to identify true and false hybrids of the early-maturing tender pumpkin variety "Nenzao Yichuanling".

[0121] According to some embodiments of this application, the method for identifying true and false hybrids of the "Nenzao Yichuanling" early-maturing and tender pumpkin variety is as follows: (1) Using the genomic DNA of the sample to be tested as a template, amplification is performed using the three sets of primers corresponding to the molecular marker to obtain the amplification product; (2) Perform fluorescence detection on the amplification products. If the sample PCR product detects the FAM fluorescence signal corresponding to the primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704, then the corresponding detection site is C:C, T:T, G:G genotype, and it is determined to be the paternal type. If the HEX fluorescence signal corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 is detected in the PCR product of the sample, then the corresponding detection site is T:T, C:C, A:A genotype, and it is determined to be the maternal type. If both FAM and HEX fluorescence signals corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 are detected simultaneously, the detection site is a T:C, C:T, A:G genotype, and it is determined to be a true hybrid.

[0122] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:2 is linked to a FAM fluorescent adapter.

[0123] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:5 is linked to a FAM fluorescent adapter.

[0124] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:8 is linked to a FAM fluorescent adapter.

[0125] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:3 is linked to a HEX fluorescent adapter.

[0126] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:6 is linked to a HEX fluorescent adapter.

[0127] According to some embodiments of this application, the 5' end of the nucleotide sequence shown in SEQ ID NO:9 is linked to a HEX fluorescent adapter.

[0128] According to some embodiments of this application, a method for identifying the purity of a hybrid of the early-maturing, tender pumpkin variety "Nenzao Yichuanling" includes the following steps: (1) Using the genomic DNA of the sample to be tested as a template, PCR amplification was performed using the primers corresponding to the molecular marker to obtain the amplification product; (2) Perform fluorescence detection on the amplification products. If the HEX fluorescence signal corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 is detected in the PCR product of the sample, the corresponding detection site is C:C, T:T, G:G genotype, and it is determined to be the paternal type. If the sample PCR product detects the FAM fluorescence signal corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704, then the corresponding detection site T:T, C:C, A:A genotype is determined to be the maternal type. If both FAM and HEX fluorescence signals corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 are detected simultaneously, the detection site is a T:C, C:T, A:G genotype, and it is determined to be a true hybrid. If no fluorescent signal is detected, it is another genotype combination and is determined to be another type.

[0129] According to some embodiments of this application, the purity of the young early pumpkin hybrid seeds can be calculated by statistically analyzing the proportion of true hybrids to the total tested samples.

[0130] According to some embodiments of this application, PCR amplification uses Touchdown PCR; the Touchdown PCR amplification program is as follows: 94℃ for 15 min; 95℃ for 20 s; 65℃~56℃ for 60 s, 10 cycles, with the annealing extension temperature decreasing by 0.8℃ in each cycle; 94℃ for 20 s; 57℃ for 60 s, 26 cycles.

[0131] According to some embodiments of this application, the molecular markers are used in the identification of the "Nenzao Yichuanling" early-maturing tender pumpkin variety.

[0132] This application also provides the application of the above-mentioned molecular markers in identifying true and false hybrids of the early-maturing and tender pumpkin variety "Nenzao Yichuanling".

[0133] In addition, this application also discloses a method for identifying the purity of the hybrid of the early-maturing and tender early-maturing pumpkin variety "Nenzao Yichuanling", the method comprising the following steps: (1) Using the genomic DNA of the sample to be tested as a template, PCR amplification was performed using the primers corresponding to the first molecular marker, the second molecular marker and the third molecular marker to obtain the amplification product; (2) Perform fluorescence detection on the amplification products, and obtain the genotype of the corresponding detection site based on the fluorescence signals corresponding to the first molecular marker, the second molecular marker and the third molecular marker, and determine the hybrid type; (3) By statistically analyzing the proportion of true hybrids in the total tested samples, the purity of the young early pumpkin hybrid seeds can be calculated.

[0134] Preferably, Touchdown PCR is used for PCR amplification. The Touchdown PCR amplification program is as follows: 94℃ for 15 min; 95℃ for 20 s; 65℃~56℃ for 60 s, 10 cycles, with the annealing extension temperature decreasing by 0.8℃ in each cycle; 94℃ for 20 s; 57℃ for 60 s, 26 cycles. The components and amounts used for PCR amplification are shown in Table 1.

[0135] Table 1. Components and dosages used in PCR amplification

[0136] The molecular markers in this application can not only be used to determine the hybrid purity of seeds of the single variety "Nenzao Yichuanling", but also to effectively distinguish the early-maturing and tender pumpkin variety "Nenzao Yichuanling" from other varieties.

[0137] This application utilizes whole-genome microarray sequencing variation detection to screen three SNP sites that differ between two parents, and successfully developed SNP markers and three pairs of KASP primers. Using fluorescence detection technology, it can quickly distinguish between the male parent, female parent, and hybrid, namely the early-maturing tender pumpkin variety "Nenzao Yichuanling". This can be directly used for the identification of the commercial early-maturing tender pumpkin variety "Nenzao Yichuanling", and further, the evaluation of true and false hybrids of the commercial variety can be based on this molecular marker, which can effectively protect the rights and interests of this variety.

[0138] Early use of this molecular marker allows for rapid identification of the purity of the variety, effectively reducing losses caused by false hybrids, minimizing potential disputes, improving identification efficiency and accuracy, and providing evidence for variety protection. Therefore, this application has significant application value in the identification and variety protection of the early-maturing pumpkin variety "Nenzao Yichuanling" hybrid.

[0139] The present invention will be further described below with reference to specific embodiments, and the advantages and features of this application will become clearer with the description. However, unless otherwise specified, the specific experimental methods involved in the following embodiments are conventional methods or implemented according to the conditions recommended in the manufacturer's instructions.

[0140] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art. Unless otherwise specified, the experimental methods in the following embodiments are all conventional methods. Unless otherwise specified, the reagents and materials used can be purchased commercially.

[0141] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as are familiar to those skilled in the art. Furthermore, any methods and materials similar to or equivalent to those described herein may be used in this invention. The preferred embodiments and materials described herein are for illustrative purposes only.

[0142] Example 1: Obtaining molecular markers for cultivar purity identification of early-maturing tender pumpkin variety "Nenzao Yichuanling" 1. Whole-genome microarray sequencing variant detection Variation detection refers to the sequencing and differential analysis of the genome of an individual or population of a species using high-throughput sequencing technology to obtain a large number of single nucleotide polymorphism (SNP) sites, insertion / deletion sites (InDel), structural variation sites (SV), and copy number variation sites (CNV). With the significant reduction in sequencing costs and the improvement in sequencing efficiency, whole-genome microarray sequencing variation detection has become one of the fastest and most effective methods for studying human diseases and molecular breeding of plants and animals. After the sequencing data is processed, bioinformatics analysis is performed according to the following procedure.

[0143] (1) Perform quality control on the raw data to obtain clean data for analysis; the raw data is RAW data to obtain high-quality clean reads for subsequent analysis.

[0144] The sequencing data filtering steps are as follows: (a) Remove reads containing adapters; (b) Remove reads where the number of N is greater than 3; (c) Remove low-quality reads (the number of bases with a quality value of Q < 5 accounts for more than 20% of the total read).

[0145] (2) Align the Clean Data with the Reference Genome; After obtaining the clean reads, use BWA software to align the clean reads with the reference genome. The initial alignment results are in SAM format, and then use SAMtools software to convert the results to BAM format and sort them. If the results of a sample contain multiple libraries, use SAMtools to merge the BAM results of multiple libraries, use picard to mark repetitive sequences, and perform basic data information statistics; (3) SNP mutation detection was performed; GATK software (v3.8 second-generation chip sequencing mutation detection software https: / / software.broadinstitute.org / gatk / ) was used to detect SNPs; (4) SNP screening: filter out sites with QUAL value (base quality value) less than 30, MQ value less than 30, and DP value less than 2, and select SNP sites that are homozygous and differential between the parents as candidate sites.

[0146] 2. Molecular marker development Based on the differences in sequencing sequences between the parents, candidate SNP sites were selected. KASP primers (the primers corresponding to the molecular markers in this application) were designed using the online primer design software SNP Primer (www.snpway.com). These primers consisted of a pair of specific primers (Primer X and Primer Y) containing different fluorescent adapters for the SNP alleles, and a common primer (Primer C). Primer X used a FAM fluorescent adapter, and Primer Y used a HEX fluorescent adapter. Primer pairs with suitable specificity and annealing temperatures were selected. The primers were synthesized by Beijing Qingke Company.

[0147] The application of molecular markers specifically includes the following steps: (1) Using the genomic DNA of the sample to be tested as a template, Touch down PCR was performed using molecular marker amplification primers to obtain the amplification product; the Touch down PCR program was as follows: 94℃ for 15 min; 95℃ for 20 s; 65℃ for 60 s, 10 cycles, with the annealing extension temperature decreasing by 0.8℃ in each cycle; 94℃ for 20 s; 57℃ for 60 s, 26 cycles; (2) Detect and analyze the amplification products. If the HEX fluorescence signal corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 is detected in the PCR product of the sample, the corresponding detection site is C:C, T:T, G:G genotype, and it is determined to be the paternal type. If the sample PCR product detects the FAM fluorescence signal corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704, then the corresponding detection site is the T:T, C:C, A:A genotype, and it is determined to be the maternal type. If both FAM and HEX fluorescence signals corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704 are detected simultaneously, the detection site is a T:C, C:T, A:G genotype, and it is determined to be a true hybrid. If no fluorescence signal is detected, the sample is considered invalid and will not be included in the purity statistics. Resampling or repeated testing is required. In actual testing, this phenomenon may also be caused by DNA quality issues, non-target species genomes, or allele deletions due to extreme cross-pollination, but further confirmation with repeated experiments and control samples is necessary. The purity of the young early pumpkin hybrid seeds can be calculated by statistically analyzing the proportion of true hybrids in the total tested samples.

[0148] In this embodiment, the phenotypic diagrams of the maternal parent, paternal parent, and hybrid ("Nenzao Yichuanling") of "Nenzao Yichuanling" are as follows: Figure 1 As shown.

[0149] In this embodiment, the early-maturing pumpkin chip detects the differences between the maternal and paternal parents, such as... Figure 2 As shown.

[0150] A schematic diagram of the molecular marker development results in this embodiment is shown below. Figures 3-5 As shown.

[0151] In this embodiment, the parent plant 238P1-1-1-1-1-2 of "Nenzao Yichuanling" was isolated from Nenzao pumpkin collected in Huangxing Town, Changsha County in 2017.

[0152] The maternal parent 238P1-1-1-1-1-2 has a base T at position 9213402 on chromosome 3 of the genome of early pumpkin, and its nucleotide sequence before and after is shown in SEQ ID NO:12; The base at position 10840907 on chromosome 5 is C, and the nucleotide sequences before and after it are shown in SEQ ID NO:13; The base at position 10779704 on chromosome 6 is A, and the nucleotide sequences before and after it are shown in SEQ ID NO:14.

[0153] In this embodiment, the male parent 1819P1-1-1-1-1-1 was obtained by hybridizing the early-maturing pumpkin resources collected in Hengyang in 2018 with the green-skinned oval-shaped early-maturing pumpkin.

[0154] The paternal parent 1819P1-1-1-1-1-1 has a C base at position 9213402 on chromosome 3 of the genome of early pumpkin, and its nucleotide sequence before and after is shown in SEQ ID NO:12; The base at position 10840907 on chromosome 5 is T, and the nucleotide sequences before and after it are shown in SEQ ID NO:13; The base at position 10779704 on chromosome 6 is G, and the nucleotide sequences before and after it are shown in SEQ ID NO:14.

[0155] 3. Specific application examples Three SNP markers were used to test 192 seedling samples. The percentage of true hybrids in the samples was calculated based on the SNP marker results. Primers were cross-corrected. A true hybrid was identified only when all primer pairs showed complementary heterozygous banding at the three SNP marker sites; other banding patterns (including homozygous paternal, homozygous maternal, non-biparental, and single-marker heterozygous) were considered hybrids. The number of hybrids was the sum of all types of hybrids. Seed purity (%) = (Total number of valid samples – Number of hybrids) / Total number of samples × 100%. Field validation showed good primer stability and good agreement with field test results, enabling the identification of the purity of early-maturing pumpkin hybrids. The test results are shown in Tables 2 and 3.

[0156] The detection method in this embodiment can complete the seed purity identification within 2 hours, and has the advantages of being fast, low-cost, and easy to operate.

[0157] Table 2. Genotypes and determination of true and false hybrids based on marker detection of Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704

[0158] Table 3 Purity Identification Results

[0159] In summary, molecular marker identification and screening in breeding can identify maternal parent types by retaining the FAM fluorescence signals corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704; and paternal parent types by retaining the HEX fluorescence signals corresponding to primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704. If two fluorescence signals are detected simultaneously for all three primers Cmo_Chr03-9213402, Cmo_Chr05-10840907, and Cmo_Chr06-10779704, the detection sites are T:C, C:T, and A:G genotypes, indicating a true hybrid. The purity of commercial varieties can be identified and the varieties can be protected through screening of molecular markers in the early stage. This method is simple and easy to operate, which can greatly improve the efficiency for breeders.

[0160] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.

Claims

1. The application of molecular markers in the following aspects, characterized in that, include: (1) Predict, screen and / or identify the purity of early-maturing pumpkin hybrids; And / or, (2) Improvements to early-maturing squash; And / or, (3) accelerate the selection of hybrid plants in offspring young early pumpkins through molecular marker-assisted selection breeding; The molecular marker includes at least one of a first molecular marker, a second molecular marker, and a third molecular marker; The first molecular marker is located at base 9213402 on chromosome 3 of the squash genome; The second molecular marker is located at base 10840907 on chromosome 5 of the squash genome; The third molecular marker is located at base 10779704 on chromosome 6 of the early-maturing squash genome.

2. A primer set, characterized in that, The primer set satisfies at least one of the following characteristics: (1) Includes a first reverse primer set, the first reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO.1; (2) Includes a first forward primer set, the first forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.2; (3) Includes a second forward primer set, the second forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.3; (4) Includes a second reverse primer set, the second reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO.4; (5) Includes a third forward primer set, said third forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.5; (6) Includes a fourth forward primer set, said fourth forward primer set comprising a nucleotide sequence as shown in SEQ ID NO. 6; (7) Includes a fifth forward primer set, said fifth forward primer set comprising a nucleotide sequence as shown in SEQ ID NO.7; (8) Includes a third reverse primer set, said third reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO. 8; (9) Includes a fourth reverse primer set, said fourth reverse primer set comprising a nucleotide sequence as shown in SEQ ID NO. 9; (10) Includes nucleotide sequences that encode the same protein as any of the nucleotide sequences shown in (1) to (9), but are different from the nucleotide sequences shown in any of (1) to (9) due to the degeneracy of the genetic code; (11) A nucleotide sequence obtained by substituting, deleting or adding one or more nucleotide sequences to any of the nucleotide sequences shown in (1) to (9), and which has the same or similar function as any of the nucleotide sequences shown in (1) to (9); (12) Includes a nucleotide sequence having at least 90% sequence homology with any of the nucleotide sequences described in (1) to (9); (13) includes a nucleotide sequence having at least 95% sequence homology with any of the nucleotide sequences described in (1) to (9); (14) includes a nucleotide sequence having at least 99% sequence homology with any of the nucleotide sequences described in (1) to (9); (15) A nucleotide sequence comprising the nucleotide sequence described in any one of (1) to (9) with 1 to 3 base deletions and / or substitutions; (16) Includes a linker sequence having a nucleotide sequence as shown in SEQ ID NO.10 or SEQ ID NO.

11.

3. The application of the primer set as described in claim 2 in the following aspects, characterized in that, include: (1) Predict, screen and / or identify the purity of early-maturing pumpkin hybrids; And / or, (2) Improvements to early-maturing squash; And / or, (3) accelerate the selection of hybrid plants in offspring young early pumpkins through molecular marker-assisted selection breeding; And / or, (4) a kit for predicting, screening and / or identifying the purity of early-maturing pumpkin hybrids.

4. A reagent kit, characterized in that, Includes the primer set as described in claim 2.

5. The application of the kit as described in claim 4 in the following aspects: (1) Predict, screen and / or identify the purity of early-maturing pumpkin hybrids; And / or, (2) Improvements to early-maturing squash; And / or, (3) accelerate the selection of hybrid plants in offspring young early pumpkins through molecular marker-assisted selection breeding.

6. A method for identifying the purity of early-maturing pumpkin hybrids, characterized in that, The process includes the following steps: extracting genomic DNA from early-maturing pumpkin, identifying the genotype of early-maturing pumpkin using the primer set as described in claim 2 and / or the kit as described in claim 4, and identifying the purity of the early-maturing pumpkin hybrid.

7. The identification method as described in claim 6, characterized in that, According to the criteria for identifying the purity of early-maturing pumpkin hybrids based on the genotype, at least one of the following characteristics must be met: (1) If the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is CC, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is TT, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is GG, it is determined to be the paternal type. (2) If the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is TT, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is CC, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is AA, it is determined to be the maternal type. (3) If the 9213402nd base pair of chromosome 3 of the young early pumpkin genome is TC, the 10840907th base pair of chromosome 5 of the young early pumpkin genome is CT, and the 10779704th base pair of chromosome 6 of the young early pumpkin genome is AG, it is determined to be a true hybrid. (4) If any of the features described in (1) to (3) are not satisfied, it is determined to be another type.

8. A method for identifying the purity of early-maturing pumpkin hybrids, characterized in that, The process includes the following steps: extracting genomic DNA from young early pumpkins, amplifying it by PCR using the primer set as described in claim 2 and / or the kit as described in claim 4, and determining the purity of the corresponding variety based on the fluorescence signal of the PCR product.

9. The method as described in claim 8, characterized in that, The criteria for evaluating the purity of early-maturing pumpkin hybrids based on the fluorescence signal must meet at least one of the following characteristics: (1) If, by means of fluorescence signal, the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is CC, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is TT, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is GG, it is determined to be the paternal type; (2) If, by means of fluorescence signal, the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is TT, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is CC, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is AA, it is determined to be the maternal type; (3) If, by means of fluorescence signal, the 9213402nd base pair on chromosome 3 of the young early pumpkin genome is TC, the 10840907th base pair on chromosome 5 of the young early pumpkin genome is CT, and the 10779704th base pair on chromosome 6 of the young early pumpkin genome is AG, it is determined to be a true hybrid. (4) If no fluorescent signal is detected, it is another genotype combination and is determined to be another type.

10. The method as described in claim 8 or 9, characterized in that, The PCR amplification reaction procedure satisfies at least one of the following characteristics: (1) Pre-denaturation at 94℃~95℃ for 14min~15min, denaturation at 94℃~95℃ for 20s~30s, annealing and extension at 65℃~56℃ for 60s~70s, 10~15 cycles, each cycle decreasing by 0.6℃~0.8℃; (2) 94℃~95℃ denaturation for 20s~30s, 55℃~57℃ annealing and extension for 60s~70s, 26~30 cycles.