Pumpkin hybrid purity identification KASP molecular marker and application thereof
The KASP molecular marker for identifying the purity of pumpkin hybrids, along with PCR amplification and fluorescence signals to determine genotypes, solves the problems of long breeding cycles and environmental interference in traditional pumpkin breeding, achieving rapid and accurate breeding identification and improving breeding efficiency.
Patent Information
- Application Number
- CN202511423243.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional pumpkin breeding relies on phenotypic selection and artificial hybridization, which is time-consuming, inefficient, and easily affected by environmental factors. There is an urgent need for efficient, accurate, and low-cost genotype screening programs to quickly identify the paternal, maternal, and hybrid varieties.
KASP molecular markers are provided for identifying the purity of pumpkin hybrids. By locating the variant sites at specific base sites in the pumpkin genome, PCR amplification is performed using primer combinations and kits, and the genotype is determined based on the fluorescence signal, enabling rapid identification of the paternal parent, maternal parent, and hybrid.
This technology enables rapid identification of male and female parents and hybrids in pumpkin breeding, improving breeding efficiency and accuracy, reducing planting scale and workload in later identification, and solving the problems of long breeding cycles and environmental impact in traditional breeding.
Smart Images

Figure CN120924716A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of plant breeding technology, and in particular to a KASP molecular marker for identifying the purity of pumpkin hybrids and its application. Background Technology
[0002] Pumpkin is an important horticultural vegetable crop, and its breeding objectives mainly focus on key agronomic traits such as fruit characteristics, disease resistance, yield, and storage properties. Traditional breeding relies on phenotypic selection and artificial hybridization, which are time-consuming, inefficient, and easily affected by environmental factors. With the development of molecular biology, DNA-level genotyping technology has become a core means to accelerate the breeding process and improve selection accuracy. Especially in the hybrid breeding and parental origin tracing stages, there is an urgent need for an efficient, accurate, and low-cost genotyping screening scheme to quickly identify the paternal, maternal, and hybrid parent species.
[0003] Application content This application addresses the aforementioned issues and aims to provide a KASP molecular marker for identifying the purity of pumpkin hybrids and its application. This application includes predicting, screening, and / or identifying parents and their pumpkin hybrids. By locating and marking variation sites in the genotypes of pumpkin progeny populations and parents, molecular marker-assisted selection in pumpkin breeding has been achieved. This is of great significance for improving the efficiency of pumpkin hybrid purity identification and establishing an efficient pumpkin breeding system.
[0004] The first aspect of this application provides the application of molecular markers in the following aspects: (I) Predicting, screening and / or identifying pumpkins; (II) Pumpkin breeding; The molecular marker is located at base 332043 on chromosome 2 of the pumpkin genome; And / or, the molecular marker is located at base 10203990 on chromosome 3 of the pumpkin genome; And / or, the molecular marker is located at base position 10206322 on chromosome 3 of the pumpkin genome; And / or, the molecular marker is located at base 3435131 on chromosome 4 of the pumpkin genome; And / or, the molecular marker is located at base 2369572 on chromosome 6 of the pumpkin genome; And / or, the molecular marker is located at base 3554082 on chromosome 12 of the pumpkin genome.
[0005] Optionally, the molecular marker is located at base 3435131 on chromosome 4 of the pumpkin genome.
[0006] Optionally, the molecular marker is a KASP molecular marker.
[0007] Optionally, the prediction, screening, and / or identification of pumpkins may include predicting, screening, and / or identifying the parents of pumpkins and their hybrids, the purity of pumpkin hybrids, etc.
[0008] Optionally, the prediction, screening, and / or identification of pumpkins may include predicting, screening, and / or identifying agronomic traits of pumpkins, such as stress resistance, growth, development, fruit, and disease resistance.
[0009] The second aspect of this application provides a primer combination designed using the nucleotide sequence shown in SEQ ID NO.1, which is 100 bp before and after the base variation site at position 3435131 on chromosome 4 of the pumpkin genome. And / or, the primer composition comprises a combination of nucleotide sequences that are identical to or complementary to a portion of the base sequence in the nucleotide sequence shown in SEQ ID NO.1; SEQ ID NO.1: CGTCCGGTGATTTCCCCCTCCGCCGAGAATTTGAAGTATTTCAAATATGGTTTCTGGATTACGTCGTAGCTCAAAGCGAACATCTCGCCGGAAACAGG[A / G]TCGAGCTTTGGGTGAGCTATCATTGTGGATTTCAATTGACCGTCGAAATCAAACCGGCCAACGGTTTTCAAATCCCCAGACGGAGTGACACGGATTTGGT; And / or, the primer combination includes: (1) The first reverse primer set consists of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.2; SEQ ID NO.2: GATAGCTCACCCAAAGCTCGAT; And / or, (2), a second reverse primer set, consisting of a nucleotide sequence as shown in SEQ ID NO.3 and a fluorescent adapter sequence; SEQ ID NO.3: GATAGCTCACCCAAAGCTCGAC; And / or, (3) the first forward primer set, consisting of a nucleotide sequence as shown in SEQ ID NO.4; SEQ ID NO.4: CGTCGTAGCTCAAAGCGAACAT; And / or, (4), the third reverse primer set, consisting of a nucleotide sequence as shown in SEQ ID NO.5; SEQ ID NO.5: GCTCGGTTGATGCTTTTCTATGC; And / or, (5) a nucleotide sequence that encodes the same protein as any of (1) to (4), but is different from any of (1) to (4) due to the degeneracy of the genetic code; And / or, (6) a nucleotide sequence obtained by substituting, deleting or adding one or more nucleotide sequences to any of the nucleotide sequences described in (1) to (5), and a nucleotide sequence that is functionally identical or similar to any of the nucleotide sequences described in (1) to (5); And / or, (7), a nucleotide sequence having at least 90% sequence homology with any of (1) to (6).
[0010] The identical or complementary base sequences mentioned in this invention refer to identical base sequences or corresponding base sequences according to the relationship between A and T, C and G.
[0011] Optionally, the primer set includes a first reverse primer set consisting of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.2; a second reverse primer set consisting of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.3; and a first forward primer set consisting of a nucleotide sequence as shown in SEQ ID NO.4.
[0012] Optionally, the fluorescent adapter sequence has a nucleotide sequence as shown in SEQ ID NO. 6 or SEQ ID NO. 7; SEQ ID NO.6: GAAGGTGACCAAGTTCATGCT; SEQ ID NO.7: GAAGGTCGGAGTCAACGGATT.
[0013] Optionally, the first reverse primer set and the second reverse primer set are respectively connected to different fluorescent adapter sequences.
[0014] Optionally, the primer set includes a first reverse primer set, a second reverse primer set, and a first forward primer set. The nucleotide sequence and fluorescent adapter sequence of the first reverse primer set are: GAAGGTGACCAAGTTCATGCTGATAGCTCACCCAAAGCTCGAT; the nucleotide sequence and fluorescent adapter sequence of the second reverse primer set are: GAAGGTCGGAGTCAACGGATTGATAGCTCACCCAAAGCTCGAC; and the nucleotide sequence of the first forward primer set is: CGTCGTAGCTCAAAGCGAACAT.
[0015] Optionally, the primer set includes a third reverse primer set consisting of a nucleotide sequence as shown in SEQ ID NO.5; and a first forward primer set consisting of a nucleotide sequence as shown in SEQ ID NO.4.
[0016] Optionally, the primer set includes a third reverse primer set and a first forward primer set, wherein the nucleotide sequence of the third reverse primer set is GCTCGGTTGATGCTTTTCTATGC; and the nucleotide sequence of the first forward primer set is CGTCGTAGCTCAAAGCGAACAT.
[0017] A third aspect of this application provides a kit comprising the primer combination as described in the second aspect.
[0018] The fourth aspect of this application provides for the use of the primer combinations described in the second aspect or the kits described in the third aspect in the following aspects: (I) Predicting, screening and / or identifying pumpkins; (II) Pumpkin breeding.
[0019] Optionally, the application includes the following steps: extracting pumpkin genomic DNA, and obtaining the pumpkin genotype using primer combinations as described in the second aspect and / or a kit as described in the third aspect.
[0020] Optionally, the application includes a PCR amplification step, and the genotype of the pumpkin is determined based on the fluorescence signal of the PCR product.
[0021] Optionally, the criteria for determining the pumpkin genotype based on the fluorescence signal are as follows: if only the fluorescence signal corresponding to the first reverse primer set is detected, the detection site is the AA genotype; if only the fluorescence signal corresponding to the second reverse primer set is detected, the detection site is the GG genotype; if the fluorescence signals corresponding to both the first and second reverse primer sets are detected simultaneously, the detection site is the AG genotype.
[0022] Optionally, the criteria for determining the pumpkin genotype based on the fluorescence signal are as follows: if only the fluorescence signal corresponding to the first reverse primer set is detected, then the 3435131st base of chromosome 4 of the pumpkin genome is AA; if only the fluorescence signal corresponding to the second reverse primer set is detected, then the 3435131st base of chromosome 4 of the pumpkin genome is GG; if the fluorescence signals corresponding to both the first and second reverse primer sets are detected simultaneously, then the 3435131st base of chromosome 4 of the pumpkin genome is AG.
[0023] Optionally, the fluorescence signal includes at least one of FAM fluorescence signal and HEX fluorescence signal.
[0024] Optionally, the fluorescence signal corresponding to the first reverse primer set is a FAM fluorescence signal, and the fluorescence signal corresponding to the second reverse primer set is a HEX fluorescence signal.
[0025] Optionally, the PCR amplification reaction procedure includes: (D1) Treat at 94℃~95℃ for 14min~15min; treat at 94℃~95℃ for 20s~30s; treat at 65℃~57℃ for 60s~70s; 10~15 cycles, each cycle decreasing by 0.6℃~0.8℃; And / or, (D2), 94℃~95℃ treatment for 20s~30s, 55℃~57℃ treatment for 60s~70s, 26~30 cycles; And / or, (D3), 94℃~95℃ treatment for 4min~5min; 94℃~95℃ treatment for 20s~30s, 60℃~65℃ treatment for 20s~30s, 72℃~75℃ treatment for 20s~30s, 35~40 cycles; 72℃~75℃ treatment for 7min~9min.
[0026] This application has at least the following beneficial effects: The molecular markers described in this application can be directly used to identify pumpkin-related genotypes, completing the identification of paternal, maternal, and hybrid parents in a single step with high accuracy. These markers can be used for assisted breeding, effectively solving the problems of long breeding cycles and susceptibility to environmental influences in conventional breeding methods. Early use of these markers allows for rapid screening of satisfactory plants, effectively reducing the scale of planting, decreasing the workload of later identification, and improving the efficiency and accuracy of selection. This has significant implications for both pumpkin breeding research and practical applications. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of this drawing or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this drawing. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0028] Figure 1 The KASP genotyping map of the pumpkin material of this application, marked at Cmo_Chr04:3435131, shows that: A: Point A indicates that the PCR product has a fluorescent signal corresponding to the reverse primer Primer_X; G:G indicates that the PCR product has a fluorescent signal corresponding to the reverse primer Primer_Y; A:G indicates that the PCR product has two fluorescent signals, Primer_X and Primer_Y, in the reverse primer; Figure 2 This is a schematic diagram of the first-generation sequencing results of pumpkin in this application.
[0029] The purpose, features, and advantages of this accompanying drawing will be further explained in conjunction with the embodiments and with reference to the accompanying drawing. Detailed Implementation
[0030] The embodiments of this application are hereby disclosed in detail with appropriate reference to the accompanying drawings, but unnecessary details may be omitted. For example, detailed descriptions of well-known matters and repetitive descriptions of actually identical structures may be omitted. This is to avoid making the following description unnecessarily lengthy and to facilitate understanding by those skilled in the art. Furthermore, the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand this application and are not intended to limit the subject matter of the claims.
[0031] The "range" disclosed in this application is defined by a lower limit and an upper limit. A given range is defined by selecting a lower limit and an upper limit, which define the boundaries of a particular range. Ranges defined in this way can include or exclude endpoints and can be arbitrarily combined; that is, any lower limit can be combined with any upper limit to form a range. For example, if ranges of 60-120 and 80-110 are listed for a specific parameter, it is expected that ranges of 60-110 and 80-120 are also included. Furthermore, if minimum range values of 1 and 2 are listed, and if maximum range values of 3, 4, and 5 are listed, then the following ranges are all expected: 1-3, 1-4, 1-5, 2-3, 2-4, and 2-5. In this application, unless otherwise stated, the numerical range "ab" represents a shortened representation of any combination of real numbers between a and b, where a and b are real numbers. For example, the numerical range "0-5" indicates that all real numbers between "0-5" have been listed in this article; "0-5" is simply a shortened representation of these numerical combinations. Furthermore, when a parameter is stated as an integer ≥2, it is equivalent to disclosing that the parameter is, for example, an integer such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, etc.
[0032] Unless otherwise specified, all embodiments and optional embodiments of this application can be combined to form new technical solutions.
[0033] Unless otherwise specified, all technical features and optional technical features of this application may be combined to form new technical solutions.
[0034] Unless otherwise specified, all steps in this application may be performed sequentially or randomly, preferably sequentially. For example, the method includes steps (a) and (b), indicating that the method may include steps (a) and (b) performed sequentially, or it may include steps (b) and (a) performed sequentially. For example, the mention that the method may also include step (c) indicates that step (c) may be added to the method in any order. For example, the method may include steps (a), (b), and (c), or it may include steps (a), (c), and (b), or it may include steps (c), (a), and (b), etc.
[0035] Unless otherwise specified, the terms "comprising" and "including" as used in this application can be open-ended or closed-ended. For example, "comprising" and "including" can mean that other components not listed may also be included, or that only the listed components may be included.
[0036] Unless otherwise specified, the term "or" is inclusive in this application. For example, the phrase "A or B" means "A, B, or both A and B". More specifically, the condition "A or B" is satisfied by any of the following conditions: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); or both A and B are true (or exist).
[0037] Before describing this application in detail, it should be understood that this application is not limited to the specific embodiments, which can certainly be modified. It should also be understood that the terminology used herein is for the purpose of describing specific embodiments only and is not intended to be limiting. As used in this specification and the appended claims, singular and singular terms, such as “an,” “a,” and “described,” include the plural referents unless explicitly stated otherwise. Thus, for example, references to “plant,” “the plant,” or “a plant” also include multiple plants; moreover, depending on the context, the use of the term “plant” can also include genetically similar or identical offspring of that plant; the use of the term “nucleic acid” optionally includes multiple copies of the nucleic acid molecule; similarly, the term “probe” optionally (and generally) covers a number of similar or identical probe molecules.
[0038] Unless otherwise specified, nucleic acids are written from left to right in a 5′ to 3′ direction. Numerical ranges mentioned in this specification include the numerical values that define the range, and include every integer or any non-integer portion within the defined range. Unless otherwise defined, all technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. Although any methods and materials similar to or equivalent to those described herein may be used in the practice of testing this application, the preferred methods and materials are those described herein. In the description and claims of this application, the following terms will be used according to the definitions listed below.
[0039] Certain definitions used in this specification and claims are provided below. To provide a clear and consistent understanding of the specification and claims (including the scope to be given by such terms), the following definitions are provided: “Agronomy,” “agronomic traits,” and “agronomic performance” refer to the traits (and underlying genetic factors) of a given plant variety that contribute to yield over the growing season. Individual agronomic traits include seedling vigor, nutrient potential, stress tolerance, disease resistance or tolerance, insect resistance or tolerance, herbicide resistance, branching, flowering, seed formation, seed size, seed density, lodging resistance, threshing rate, and fruit browning.
[0040] The term "allele" refers to any one or more alternative forms of a gene sequence. For example, in diploid cells or organisms, the two alleles of a given sequence typically occupy corresponding loci on a pair of homologous chromosomes. Regarding SNP markers, an allele is a specific nucleotide base present at a particular SNP locus in that individual plant.
[0041] An allele is "associated" with a trait when it is a DNA sequence that affects the expression of that trait, or a part of an allele, or is linked to that allele. The presence of that allele is an indicator of how the trait will be expressed.
[0042] The term "amplification" in the context of nucleic acid amplification refers to any process that results in an additional copy of a selected nucleic acid (or transcribed from it). Typical amplification methods include a variety of polymerase-based replication methods, including polymerase chain reaction (PCR), ligase-mediated methods such as ligase chain reaction (LCR), and RNA polymerase-based amplification methods (e.g., transcription). An "amplifier" is an amplified nucleic acid, for example, produced by amplifying a template nucleic acid using any available amplification method (e.g., PCR, LCR, transcription, etc.).
[0043] The term "chromosome segment" refers to a continuous linear segment of genomic DNA that exists on a single chromosome in a plant.
[0044] The term "complementary sequence" refers to a nucleotide sequence that is complementary to a given nucleotide sequence, i.e., the above sequences are associated by the Watson-Crick base pairing rule.
[0045] "Cultivated species" and "variety" are used synonymously to refer to a group of plants within the same species (e.g., pumpkin) that share certain genetic traits that distinguish them from other possible varieties within the same species.
[0046] "Superior strains" are agronomically superior strains that have been developed through many rounds of breeding and selection for superior agronomic performance. Many superior strains are available and are known to those skilled in the field of pumpkin breeding.
[0047] "Superior population" refers to a mixed population of superior individuals or strains that can be used to represent the current level of technology in terms of agronomically superior genotypes in a given crop species such as pumpkin.
[0048] A "favorable allele" is an allele located at a specific locus that confers or contributes to an agronomically desired phenotype and allows for the identification of plants possessing that agronomically desired phenotype. A marked favorable allele is a marker allele that separates from the favorable phenotype.
[0049] A "gene map" is a description of gene linkages between loci on one or more chromosomes (or linkage groups) within a given species, typically represented as a graph or table. For each gene map, the distance between loci is measured by how frequently their alleles appear together within a population (their recombination frequency). Alleles can be detected using DNA or protein markers, or observable phenotypes. A gene map is the product of the mapping population, the types of markers used, and the polymorphism of each marker between different populations. The genetic distance between loci can differ between gene maps. However, by using shared markers, information can be correlated between maps. Those skilled in the art can use shared marker locations to identify the location of markers of interest and other loci on various gene maps. The order of loci should be invariant between maps; however, small variations in marker order often occur due to, for example, markers detecting alternative repetitive loci in different populations, differences in statistical methods used to orient markers, new mutations, or experimental errors.
[0050] "Genetic recombination frequency" is the frequency of exchange events (recombination) between two loci. Recombination frequency can be observed by tracking the segregation of markers and / or traits after meiosis.
[0051] A "genome" is a complete set of DNA or genes carried by chromosomes or sets of chromosomes.
[0052] "Genome type" refers to the genetic makeup of a cell or organism.
[0053] "Germanic material" refers to the genetic material that forms the physical basis of the hereditary qualities of an organism. As used herein, germplasm includes seeds and living tissues from which new plants can grow; or other plant parts that can be cultured into the whole plant, such as leaves, stems, pollen, or cells. Germanic resources provide plant breeders with a source for improving the genetic traits of commercially grown varieties.
[0054] A haplotype is an individual's genotype at multiple loci, that is, a combination of alleles. Typically, the loci described by a haplotype are physically and genetically linked, meaning they are located on the same chromosomal segment. The term "haplotype" can refer to an allele at a specific locus or to alleles at multiple loci along a chromosomal segment.
[0055] An individual is "homozygous" if it possesses only one type of allele at a given locus (e.g., a diploid individual has one copy of the same allele at a locus on each of its two homologous chromosomes). If more than one allele type exists at a given locus, the individual is "heterozygous" (e.g., a diploid individual has one copy of each of two different alleles). The term "homogeneity" indicates that members of a population share the same genotype at one or more specific loci. In contrast, the term "heterogeneity" is used to indicate that individuals within a population have different genotypes at one or more specific loci.
[0056] The term "insertion or deletion" refers to an insertion or deletion, in which one strain may be referred to as having an inserted nucleotide or DNA fragment relative to a second strain, or the second strain may be referred to as having a deleted nucleotide or DNA fragment relative to a first strain.
[0057] "Infiltration" refers to the transfer or introduction of genes, quantitative trait loci (QTLs), marker loci, haplotypes, marker profiles, traits, or trait loci from the genome of one plant into the genome of another plant.
[0058] A “strain” or “line” is a group of individuals that are similarly related, typically inbred to a certain degree, and are usually homozygous and homogeneous (homogeneous or nearly homogeneous) at most loci. A “subline” is an inbred subgroup that is genetically different from other similar inbred subgroups originating from the same ancestor. Sublines are conventionally obtained by inbreeding seeds from single squash plants selected in the F3 to F5 generations until the remaining segregating loci are “fixed” or homozygous at most or all loci. Commercial squash varieties (or strains) are typically produced by aggregating (“converging”) the self-pollinated offspring of single F3 to F5 plants from controlled hybridization between two genetically different parents. While these varieties generally appear homogeneous, the self-pollinated varieties derived from the selected plants will eventually (e.g., F8) become a mixture of homozygous plants that are heterologous in the initially selected F3 to F5 plants and can be genotyped at any locus. Marker-based sublines, distinguished from each other based on qualitative polymorphism at one or more specific marker loci at the DNA level, are obtained by genotyping seed samples derived from single self-pollinated progeny of selected F3-F5 plants. These seed samples can be directly genotyped as seeds or as plant tissue grown from such seed samples. Optionally, seeds sharing a common genotype at specific loci (or multiple loci) are merged to provide genetically homologous sublines at identified loci important for the traits of interest (e.g., anther extrusion, flowering time, heading time, and / or Fusarium head blight resistance, etc.).
[0059] Linkage refers to the phenomenon where alleles on the same chromosome tend to co-segregate more frequently than would be expected by chance when their transmission is independent. Genetic recombination occurs throughout the genome at a presumed random frequency. Gene maps are constructed by measuring the frequency of recombination between pairs of traits or markers. The closer the traits or markers are to each other on the chromosome, the lower the frequency of recombination and the higher the degree of linkage. If traits or markers co-segregate substantially, they are considered linked in this paper. A recombination probability of 1 / 100 per generation is defined as a map distance of 1.0 centimoles (1.0 cM).
[0060] Genetic elements or genes located on a single chromosomal segment are physically linked. Advantageously, two loci are located close to each other, such that recombination between homologous chromosome pairs during meiosis does not occur at a high frequency between these two loci, for example, causing linked loci to co-segregate at least about 90% of the time, such as 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.75%, or more. Genetic elements located within a chromosomal segment are also genetically linked, typically within a genetic recombination distance of less than or equal to 50 centimoles (cM), such as about 49, 40, 30, 20, 10, 5, 4, 3, 2, 1, 0.75, 0.5, or 0.25 cM or less. That is, two genetic elements within a single chromosomal segment recombine with each other at a frequency of less than or equal to about 50%, for example, about 49%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, or 0.25% or lower during meiosis.
[0061] In this application, the phrase "closely linked" for a locus refers to recombination between two linked loci occurring at a frequency of 10% or less (i.e., separated by no more than 10 cM on the genome map). In other words, closely linked loci co-segregate at least 90% of the time. Marker loci are particularly useful in this application when they exhibit a significant probability (linkage) of co-segregation with the desired trait. Closely linked loci, such as marker loci and second loci, are capable of exhibiting recombination frequencies of 10% or less, preferably about 9% or less, more preferably about 8% or less, more preferably about 7% or less, even more preferably about 6% or less, more preferably about 5% or less, even more preferably about 4% or less, more preferably about 3% or less, and even more preferably about 2% or less. In a highly preferred embodiment, the associated loci exhibit a recombination frequency of about 1% or less, for example, about 0.75% or less, more preferably about 0.5% or less, or more preferably about 0.25% or less. Two loci located on the same chromosome are also referred to as "neighboring" to each other, such that recombination between them occurs at a frequency of less than 10% (e.g., approximately 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or lower). In some cases, two different markers can have the same genomic map coordinates. In such cases, the two markers are so close to each other that recombination between them occurs at a frequency so low as to be undetectable.
[0062] When the relationship between two genetic elements (such as a antagonistic contributing genetic element and a neighboring marker) is involved, "coupled" linkage refers to a state in which the "favorable" allele at the locus of interest is physically linked to the "favorable" allele at the corresponding linked marker locus on the same chromosomal chain. In coupled linkage, offspring that inherit both favorable alleles inherit that chromosomal chain. In repulsive linkage, the "favorable" allele at the locus of interest is physically linked to the "unfavorable" allele at the neighboring marker locus, and the two "favorable" alleles are not inherited together (i.e., the two loci are "opposite" to each other).
[0063] "Linkage disequilibrium" refers to the phenomenon where, when segregating from parents to offspring, alleles tend to remain together in linkage groups at a higher frequency than expected from their individual frequencies. A state of linkage disequilibrium implies that related loci are physically close enough along the length of the chromosome that they segregate together at a higher frequency than random (i.e., non-random). Markers exhibiting linkage disequilibrium are considered linked. Linked loci co-segregate more than 50% of the time, for example, from about 51% to about 100% of the time. In other words, two co-segregating markers have a recombination frequency of less than 50% (and, by definition, are less than 50 cM apart on the same linkage group). As used herein, linkage can occur between two markers, or between a marker and a locus influencing the phenotype. A marker locus can be "associated" (associated with) a trait. The degree of linkage between a marker locus and a locus influencing the phenotypic trait is measured, for example, as a statistical probability (e.g., an F-statistic or LOD score) of co-segregation of the molecular marker and the phenotype.
[0064] A "linkage group" (LG) is a trait or marker that is largely co-segregated. A linkage group generally corresponds to a chromosomal region containing the genetic material encoding that trait or marker. Therefore, linkage groups can be broadly attributed to a specific chromosome.
[0065] A "locus" is a defined segment of DNA. For example, it can refer to the location on a chromosome where a nucleotide, gene, sequence, or marker is located.
[0066] "Map location" is a designated position on a gene map relative to linked genetic markers where a specific marker can be found within a given species.
[0067] "Mapping" is the process of defining the linkage relationships of loci through the use of standard genetic principles such as genetic markers, marker segregation, and recombination frequencies.
[0068] "Marker," "molecular marker," or "marked locus" is a term used to denote a nucleic acid or amino acid sequence that is sufficiently unique to characterize a specific locus on the genome. Examples include restriction fragment length polymorphism (RFLP), simple repeat sequence (SSR), target region amplification polymorphism (TRAP), isoenzyme electrophoresis, random amplified polymorphic DNA (RAPD), random primer polymerase chain reaction (AP-PCR), DNA amplification fingerprint (DAF), sequence-specific amplified region (SCAR), amplified fragment length polymorphism (AFLP), and single nucleotide polymorphism (SNP). Other types of molecular markers are also known in the art, and phenotypic traits can also be used as markers in the methods described. All markers are used to define specific loci on the pumpkin genome. Thus, each marker is an indicator of a specific segment of DNA with a unique nucleotide sequence. Map locations provide a measure of the relative position of a specific marker with respect to each other. When a trait is described as being linked to a given marker, it should be understood that the actual DNA segment whose sequence affects that trait is generally cosegregated with that marker. If markers are identified on both sides of a trait, a more precise and definitive localization of that trait can be obtained. By measuring the presence of markers in hybrid offspring, the presence of a trait can be detected by a relatively simple molecular test without actually assessing the presence of the trait itself, which can be difficult and time-consuming, as actual assessment of a trait requires the plant to grow to a stage where the trait can be expressed.
[0069] "Marker-assisted selection" refers to the process of selecting one or more plants for a desired one or more traits (wherein the nucleic acid is associated with the desired trait) by detecting one or more nucleic acids from the plant, and then selecting plants or germplasm that have the aforementioned one or more nucleic acids.
[0070] In some examples, multiple marker loci or haplotypes are used to define a “marker spectrum.” As used herein, a “marker spectrum” refers to a combination of two or more marker loci or haplotypes within the genome of a particular plant. For example, in one example, a specific combination of marker loci or haplotypes defines the marker spectrum of a particular plant.
[0071] The terms "phenotype," "phenotypic trait," or "trait" can refer to the observable expression of a gene or series of genes. A phenotype can be observed by the naked eye or by any other assessment method known in the art, such as weighing, counting, measuring (length, width, angle, etc.), microscopy, biochemical analysis, or electromechanical determination. In some cases, a phenotype is directly controlled by a single gene or locus, i.e., a "monogenous trait" or "simple inherited trait." In the absence of significant environmental variation, monogenic traits can segregate in a population, producing a "qualitative" or "discrete" distribution; that is, the phenotype is divided into discrete categories. In other cases, a phenotype is the result of multiple genes and can be considered a "polygenic trait" or "complex trait." Polygenic traits segregate in a population, producing a "quantitative" or "continuous" distribution; that is, the phenotype cannot be separated into discrete categories. Both monogenic and polygenic traits can be influenced by the environment in which they are expressed, but polygenic traits tend to have a larger environmental component.
[0072] "Favorable traits" or "favorable phenotypes" are the phenotypes desired in agronomy.
[0073] The term "plant" includes the entire plant, whether immature or mature, including plants from which seeds or grains or anthers have been removed. Seeds or embryos that will produce a plant are also considered plants.
[0074] "Plant part" refers to any part or piece of a plant, including leaves, stems, buds, roots, root tips, anthers, seeds, grains, plumules, pollen, ovules, flowers, cotyledons, hypocotyls, pods, branches, stems, tissues, tissue cultures, cells, etc.
[0075] "Polymorphism" refers to the variation or difference between two related nucleic acids. "Nucleotide polymorphism" refers to the fact that two nucleic acids, when compared for maximum similarity, are different nucleotides in a sequence compared to their related sequences.
[0076] The terms “polynucleotide,” “polynucleotide sequence,” “nucleic acid sequence,” “nucleic acid fragment,” and “oligonucleotide” are used interchangeably in this document. These terms encompass nucleotide sequences, among others. Polynucleotides can be polymers of RNA or DNA, and they can be single-stranded or double-stranded, optionally containing synthetic, non-natural, or modified nucleotide bases. Polynucleotides in the form of DNA polymers can consist of one or more strands of cDNA, genomic DNA, synthetic DNA, or mixtures thereof.
[0077] A primer is a (synthetic or naturally occurring) oligonucleotide that, when placed under conditions where it is synthesized along its complementary strand by a polymerase, serves as a starting point for nucleic acid synthesis or replication. Typically, primers are oligonucleotides of 10 to 30 nucleic acids in length, but longer or shorter sequences can be used. Primers can be provided in double-stranded form, although single-stranded form is preferred. Primers may also contain detectable markers, such as 5' end markers.
[0078] A "probe" is an oligonucleotide (synthetic or naturally occurring) that is complementary (but not necessarily perfectly complementary) to the polynucleotide of interest and forms a double-stranded structure by hybridizing with at least one strand of the polynucleotide of interest. Typically, probes are oligonucleotides of 10 to 50 nucleic acids in length, but longer or shorter sequences can be used. Probes can also contain detectable markers. The terms "marker" and "detectable marker" refer to detectable molecules, including but not limited to radioactive isotopes, fluorescent agents, chemiluminescent agents, enzymes, enzyme substrates, enzyme cofactors, enzyme inhibitors, chromophores, dyes, metal ions, metal sols, semiconductor nanocrystals, ligands (e.g., biotin, avidin, streptavidin, or haptenogenin), etc. Detectable markers can also include combinations of reporter genes and quenchers, such as those used in FRET probes or TaqMan probes.
[0079] The term "reporter gene" refers to a substance or part thereof that exhibits a detectable signal that can be suppressed by a quencher. The detectable signal of the reporter gene is, for example, fluorescence within a detectable range.
[0080] The term "quencher" refers to a substance or part thereof that can inhibit, reduce, suppress, or suppress detectable signals generated by the reporter gene.
[0081] As used herein, the terms “quenching” and “fluorescent energy transfer” refer to the process in which, when a reporter gene and a quencher are in close proximity and the reporter gene is excited by an energy source, a major portion of the energy of the excited state is nonradiatively transferred to the quencher, which either dissipates nonradiatively or is emitted at a wavelength different from that of the reporter gene.
[0082] The term “quantitative trait locus” or “QTL” refers to a region of DNA associated with differential expression of a quantitative phenotypic trait in at least one genetic context (e.g., in at least one breeding population). The region of a QTL encompasses one or more genes that affect or are closely related to the trait under consideration.
[0083] A "reference sequence" or "common sequence" is a qualified sequence used as the basis for sequence comparison. A reference sequence for a PHM marker is obtained by sequencing multiple strains at that locus, aligning the nucleotide sequences in a sequence alignment program (such as Sequencher), and then obtaining the most prevalent nucleotide sequence from the alignment. Polymorphisms present between individual sequences are annotated in the common sequence. The reference sequence is typically not an exact copy of any individual DNA sequence, but rather represents a mixture of available sequences and is used to design primers and probes targeting polymorphisms within that sequence.
[0084] "Recombination frequency" is the frequency of events (recombination) that occur between two loci. Recombination frequency can be observed by tracking the segregation of markers and / or traits during meiosis.
[0085] "Self-fertilization," "self-pollination," or "self-crossing" is the process by which a breeder mates a plant with itself; for example, the second-generation hybrid F2 produces offspring named F2:3 with itself.
[0086] "SNP," or "single nucleotide polymorphism," refers to a sequence variation that occurs when a single nucleotide (A, T, C, or G) in a genome sequence is altered or mutated. When an SNP is mapped to a site on the pumpkin genome, an "SNP marker" exists. Many techniques for detecting SNPs are known in the art, including allele-specific hybridization, primer extension, direct sequencing, TM, and real-time PCR such as TaqMan assays.
[0087] "Transgenic plant" refers to a plant that contains exogenous polynucleotides within its cells. Generally, the exogenous polynucleotide is stably integrated into the genome, allowing it to be passed down through successive generations. The exogenous polynucleotide can be integrated into the genome alone or as part of a recombinant expression cassette. As used herein, "transgenic" refers to any cell, cell line, callus, tissue, plant part, or plant whose genotype has been altered due to the presence of exogenous nucleic acid, including the transgenic organism or cell that initially underwent such alteration, and those transgenic organisms or cells derived from the initial transgenic organism or cell through hybridization or asexual reproduction. As used herein, the term "transgenic" does not cover genomic (chromosomal or extrachromosomal) alterations resulting from conventional plant breeding methods (e.g., hybridization) or from naturally occurring events such as random cross-fertilization, infection with a non-recombinant virus, transformation with a non-recombinant bacteria, non-recombinant transposition, or spontaneous mutation.
[0088] The marked "adverse alleles" are isolated along with the adverse plant phenotype, thus providing a marker allele for identifying plants with beneficial effects that can be removed from breeding programs or germplasm.
[0089] The term "vector" is used to refer to a polynucleotide or other molecule that transfers a nucleic acid fragment into a cell. Vectors optionally include portions that mediate the maintenance of the vector and allow for its intended use (e.g., sequences necessary for replication, genes conferring resistance to drugs or antibiotics, multiple cloning sites, operatively linked promoter / enhancer elements that enable the expression of the cloned gene, etc.). Vectors are frequently derived from plasmids, bacteriophages, or plant or animal viruses.
[0090] Here we turn to an example: Identification of markers. Haplotypes and / or marker profiles associated with the trait of interest.
[0091] Several methods well-known in the art are available for detecting molecular markers or sets of molecular markers that cosegregate with a trait of interest. The basic idea behind these methods is marker selection, for which alternative genotypes (or alleles) have significantly different average phenotypes. Therefore, one compares the magnitude or significance level of the difference between alternative genotypes (or alleles) among marker loci. The trait gene is presumed to be located closest to the marker with the largest associated genotypic difference.
[0092] Two such methods for detecting loci of the trait of interest are: 1) population-based association analysis and 2) conventional linkage analysis. In population-based association analysis, lines are derived from existing populations with multiple founders, such as superior breeding lines. Population-based association analysis relies on the decay of linkage disequilibrium (LD) and the idea that in unstructured populations, after so many generations of random mating, only the associations between genes controlling the trait of interest and markers closely linked to those genes will remain. In reality, most existing populations have substructures.
[0093] Therefore, by using data obtained from randomly distributed markers in the genome, individuals are assigned to populations, thereby minimizing the imbalances caused by population structure within individual populations (also known as subpopulations). The use of structured association methods helps control population structure. For each line in a subpopulation, phenotypic values are compared to genotypes (alleles) at each marker locus. Significant marker-trait correlations indicate close proximity between the marker locus and one or more loci involved in the expression of that trait.
[0094] Conventional linkage analysis operates on the same principle; however, LOD is generated by constructing a population from a small number of founders. Founders are selected to maximize the level of polymorphism within the constructed population, and the level of co-segregation of polymorphic loci with a given phenotype is assessed. Various statistical methods have been used to identify significant marker-trait associations. One such method is the interval plotting approach, in which the probability of a gene controlling the trait of interest being located at each of many locations along a gene map (e.g., at 1 cM intervals) is tested. Genotype / phenotype data are used to calculate a LOD score (logarithm of the probability ratio) for each test location. When the LOD score exceeds a threshold, there is significant evidence that a gene controlling the trait of interest is located at that location on the gene map (which would fall between two specific marker loci).
[0095] Example 1: Pumpkin whole genome resequencing 1. Genomic DNA was extracted from the maternal parent N219, paternal parent N213-4, hybrids, and 196 pumpkin leaf samples to be tested for the Hongmi 4 variety. Whole-genome resequencing was performed on three samples: maternal parent N219, N213-4, and the Hongmi 4 hybrid. Specific steps included: 1.1 DNA extraction from the sample tissue: (1) Take 20-30 mg of pumpkin leaf material and put it into a 96-well plate, and freeze dry it using a freeze dryer; (2) Use a bead separator to add two 4mm steel balls to the 96-hole plate of the vacuumed blade, and cover it with the matching silicone film. (3) Place the 96-well plate in a high-throughput tissue grinder, adjust the speed to 1400 rpm and grind for 3 minutes. The grinding time can be increased until the leaf is crushed. (4) Take out the 96-well plate and add 600 μL of CTAB extraction solution to each well with a pipette. After sealing with heat-sealing film, place it on a vortex shaker and shake to mix well. (5) Place the sealed 96-well plate in a water bath at 65°C for 1-1.5 hours, and take it out several times during this period and place it on a vortex shaker to mix it properly. (6) After the warm bath, take out the sealed 96-well plate and place it in a refrigerated centrifuge. Set the speed to 4000 rpm and the temperature to 4°C. Centrifuge for 10 minutes. (7) Take out the centrifuged 96-well plate and use a semi-automatic 96-well pipette to transfer 400 μl of supernatant to 2 ml of purified 96-well plate; (8) Add an equal volume of magnetic bead mixing solution to a 2 ml purification 96-well plate containing 400 μl, and place it into the extraction instrument ME480; (9) Place the washing buffer plate in positions 2-3 after extraction, and place the DNA dissolving plate containing 150 μL of elution buffer in position 4. (10) Run the ME-480 extraction program. The extraction process takes about 30 minutes. Cover the DNA plate with a membrane for storage and complete the extraction process.
[0096] 1.2 Quality Control: The concentration of DNA samples was detected using a Qubit fluorescence quantitative PCR instrument; the integrity of DNA samples was detected by 1% agarose gel electrophoresis, and samples that passed quality control were used for library preparation. After quality control, the DNA samples are used for library construction according to the library construction workflow. The main steps in constructing a DNA library include: 2.1. DNA fragmentation, end repair, and quality control DNA samples were fragmented using a fragmentation enzyme, the enzyme ends were repaired, and an A base was added to the 3' end. The effectiveness of DNA fragmentation was detected by 2% agarose gel electrophoresis, and samples with a clear bright band at 300-500 bp were used for subsequent reactions.
[0097] 2.2. Connector Connection and Quality Control Sequencing adapters were ligated to fragmented DNA, and the ligation products were purified using magnetic beads. The concentration of the purified products was detected using a Qubit real-time fluorescence analyzer, and samples with acceptable concentrations were used for subsequent reactions.
[0098] 2.3. Fragment amplification, selection, and quality control The ligation products were amplified using PCR, and fragments were screened using magnetic beads. The fragment concentrations of the screened products were detected using a Qubit quantitative PCR instrument, the fragment sizes were determined by 2% agarose gel electrophoresis, and the fragment sizes were verified using a Qsep400 bioanalyzer.
[0099] 2.4. Cycloning and Quality Control The linear library was denatured into single strands and then circularized. Uncircularized linear DNA molecules were digested to obtain a single-stranded circular library. The concentration of the single-stranded circular library was detected using a Qubit quantitative PCR instrument; if the concentration was within acceptable limits, subsequent reactions were performed.
[0100] 2.5. DNB Preparation and Quality Control Single-stranded circular DNA molecules replicate via rolling circle to form a DNA nanosphere (DNB) containing more than 300 copies. The DNB concentration is detected using a Qubit fluorescence quantitative quantification system, and subsequent reactions proceed only if the concentration is within acceptable limits.
[0101] 2.6. Sequencing DNB was loaded into a sequencing chip using the MGIDL-T7 loading device, and sequencing was performed using a combined probe-anchored polymerization technique. The sequencing strategy was PE150, the sequencing depth was 10×, and 8Gb was sequenced for each strain.
[0102] 3. Bioinformatics and Data Analysis Sentieon was used to align and detect variants in three resequencing datasets. The analysis workflow is as follows: (1) Use Sentieon to align reads to the corresponding pumpkin reference genome (Cmoschata_genome_v1.fa.gz (http: / / cucurbitgenomics.org / ), sort the positions, and mark duplicate reads.
[0103] (2) Use Sentieon to detect variant sites for each sample and obtain the gVCF of each sample.
[0104] (3) Joint-calling was performed using Sentiion to conduct joint analysis of gVCF for all samples, and the variation results for each individual in the population were obtained. To ensure the accuracy of SNPs, the SNP sites obtained after joint analysis were initially hard-filtered (SNP hard-filter criteria: "QD<2.0 || FS>60.0 || MQ<40.0 || SOR>3.0 || MQRankSum<-12.5 || ReadPosRankSum<-8.0"), and the variation results for each individual in the population were obtained.
[0105] Example 2: Mining of Pumpkin-Related Molecular Markers 1. Thirty functional genes of pumpkin were screened. High-quality SNP sites with variations in gene regions were screened from the resequencing data, totaling 120 SNP sites, as shown in Table 1.
[0106] Table 1. Gene types, number of genes, and number of markers
[0107] 2. Select suitable sites for KASP marker development from all variant sites. The selection criteria are as follows: a. The GC content in the 30bp sequence upstream and downstream of the site is 30%-65%. b. The copy number of the 30bp sequence upstream and downstream of the site is <10. c. There are no variable sites in the 50bp sequence upstream and downstream of the site.
[0108] A total of 24 SNP sites were identified.
[0109] 3. The indicators of the 24 SNP variant sites were screened, and the screening criteria were as follows: a. The parents are homozygous, and the offspring do not have the same genotype as the parents. b. The missing data rate is 0. c. MAF (minor allele frequency) value ≥ 0.3, d. Sequencing depth ≥ 10, e. The functional sites do not overlap with those in the Pumpkin 20K chip; A total of 6 SNP sites were screened (as shown in Table 2). These markers can be used for marker-assisted selection of important traits in pumpkin.
[0110] Table 2 SNP variant site indices
[0111] Example 3: KASP Tag Development 1. KASP primer design One of the six sites (Cmo_Chr04: 3435131) with good performance indicators and high-quality KASP primer design was selected for KASP marker development. Based on the 100bp sequence before and after this SNP, sequence information is represented by SEQ ID NO.1. KASP primers were designed using BatchPrimer3 (http: / / probes.pw.usda.gov / batchprimer3 / ) (see Table 3) for later material validation.
[0112] Table 3. Site Information Table
[0113] Note: In accordance with the "Standard for Nucleotide and / or Amino Acid Sequence Listings and Electronic Sequence Listings", the "[A / G]" in the nucleotide sequence shown in SEQ ID NO.1 is replaced by "r" in the sequence listing. "r" means "g or a" and the name comes from "purine".
[0114] Table 4 shows the KASP marker alleles (Allele_X, Allele_Y) and primer sequences for the detection of pumpkin-related molecular markers. The KASP marker consists of three primers: two allele reverse primers X (Primer_X) and Y (Primer_Y), and one forward primer C (Primer_C). The 5' ends of the reverse primers are connected to the LGC KASP reaction-specific fluorescent groups FAM and HEX, respectively. If only FAM fluorescence is detected in the sample, the genotype is homozygous allele X (Allele_X); if only HEX fluorescence is detected, the genotype is homozygous allele Y (Allele_Y); if both FAM and HEX fluorescence are detected, the genotype is heterozygous (carrying both alleles X and Y).
[0115] Table 4. Alleles (Allele_X, Allele_Y) and primer sequences for KASP markers
[0116] The primer sequence for the reverse primer Primer_X is as follows: GAAGGTGACCAAGTTCATGCTGATAGCTCACCCAAAGCTCGAT consists of a nucleotide sequence as shown in SEQ ID NO.2 and a fluorescent adapter sequence as shown in SEQ ID NO.6.
[0117] The primer sequence for the reverse primer Primer_Y is as follows: GAAGGTCGGAGTCAACGGATTGATAGCTCACCCAAAGCTCGAC is composed of a nucleotide sequence as shown in SEQ ID NO.3 and a fluorescent adapter sequence as shown in SEQ ID NO.7.
[0118] 2. KASP tag verification (1) KASP reaction procedure KASP-tagged reactive sequencing was performed using the Douglas Scientific ArrayTape system. The ArrayTape genotyping platform includes NEXAR for PCR amplification system assembly, SOELLEX for PCR amplification, ARAYA for fluorescence signal scanning, and INTELLICS for data analysis.
[0119] PCR reaction system: The PCR amplification system was automatically assembled using NEXAR, and the PCR reaction system is shown in Table 5 below.
[0120] Table 5 PCR reaction system for KASP marker genotyping
[0121] PCR amplification: PCR amplification was performed using SOELLEX under the following conditions: 94℃ pre-denaturation for 15 minutes; 94℃ denaturation for 20 seconds, followed by annealing at 65℃-57℃ for 60 seconds (with the annealing temperature decreasing by 0.8℃ per cycle), for 10 cycles; 94℃ denaturation for 20 seconds, followed by annealing at 57℃ for 60 seconds, for 30 cycles.
[0122] Signal scanning and genotyping: After the PCR reaction was completed, the fluorescence signal of the reaction system was scanned using ARAYA; then genotyping and data analysis were performed using INTELLICS.
[0123] (2) KASP test results To assess the specificity and applicability of the markers, the target SNP markers were detected using the genomes of 196 pumpkin leaf samples.
[0124] The genotyping results of the Cmo_Chr04:3435131 marker KASP are shown in Table 6, and the genotyping diagram is shown below. Figure 1 As shown. Verification revealed that the KASP markers were divided into three distinct and compact clusters. Three samples showed homozygous A:A alleles (N219-maternal parents (N219_1, N219_2) and sample 1_32); 186 samples contained heterozygous A:G alleles (Hongmi 4 hybrid); five samples showed homozygous G:G alleles [N213-4-paternal parents (N213_4_1, N213_4_2) and samples 1_16, 1_95, and 2_86]; and two samples (samples 2_58 and 2_96) did not show any alleles.
[0125] Table 6. KASP marker genotyping results
[0126] The above identification results indicate that molecular markers can be used to identify, screen, or breed pumpkin materials with specific genotypes in breeding: by retaining materials that detect the FAM fluorescence signal corresponding to primer Primer_X, pumpkins with homozygous alleles A:A can be bred; by retaining materials that detect the HEX fluorescence signal corresponding to primer Primer_Y, pumpkins with homozygous alleles G:G can be bred; and by retaining materials that detect the FAMHEX fluorescence signal (including the fluorescence signals corresponding to the previous two primers), pumpkin hybrids with heterozygous alleles A:G can be bred. Early molecular marker screening can reduce the workload of later screening and identification, accelerating the pumpkin breeding process.
[0127] Example 4: First-generation sequencing verification 1. First-generation sequencing primer design Sequencing primers were designed using BatchPrimer3 (http: / / probes.pw.usda.gov / batchprimer3 / ) (see Table 7) for later material validation. Each sequencing primer consists of two primers: a forward amplification primer (Primer_F) and a reverse amplification primer (Primer_R).
[0128] Table 7 Sequencing Primer Sequence Information
[0129] 2. First-generation sequencing amplification The target fragment in the sample was amplified using specific primers. The PCR reaction system is shown in Table 8.
[0130] Table 8 PCR Reaction System
[0131] The amplification conditions were as follows: pretreatment at 94℃ for 5 minutes; denaturation at 94℃ for 20 seconds, annealing at 60℃ for 20 seconds, extension at 72℃ for 30 seconds, for 35 cycles; final extension at 72℃ for 7 minutes. PCR products were detected by 1% agarose gel electrophoresis. After passing the detection, first-generation sequencing was performed using amplification primers. The sequencing peak results were compared with those of the maternal parent N219, paternal parent N213, and hybrids No. 1_14 and No. 1_15 of the Red Honey No. 4.
[0132] 3. First-generation sequencing results The PCR amplification products of samples with different genotypes in the KASP test results were screened and validated by first-generation sequencing. The sample numbers and genotypes are shown in Table 9. A schematic diagram of the first-generation sequencing test results is shown in the figure. Figure 2 As shown.
[0133] Table 9 Sequencing Primer Sequence Information
[0134] It should be noted that this application is not limited to the above-described embodiments. The above embodiments are merely examples, and any embodiments with the same structure and effect as the technical concept within the scope of this application are included in the technical scope of this application. Furthermore, various modifications that can be conceived by those skilled in the art to the embodiments, and other ways of constructing by combining some of the constituent elements of the embodiments, without departing from the spirit of this application, are also included in the scope of this application.
Claims
1. Applications of molecular markers in any of the following aspects: (I) Predicting, screening and / or identifying pumpkins; (II) Pumpkin breeding; The molecular marker is located at base 332043 on chromosome 2 of the pumpkin genome; And / or, the molecular marker is located at base 10203990 on chromosome 3 of the pumpkin genome; And / or, the molecular marker is located at base position 10206322 on chromosome 3 of the pumpkin genome; And / or, the molecular marker is located at base 3435131 on chromosome 4 of the pumpkin genome; And / or, the molecular marker is located at base 2369572 on chromosome 6 of the pumpkin genome; And / or, the molecular marker is located at base 3554082 on chromosome 12 of the pumpkin genome.
2. A primer combination, characterized in that, The primer combination comprises a combination of nucleotide sequences that are identical to or complementary to a portion of the base sequence in the nucleotide sequence shown in SEQ ID NO.1; And / or, the primer combination includes: (1) The first reverse primer set consists of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.2; And / or, (2), a second reverse primer set consisting of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.3; And / or, (3) the first forward primer set, consisting of a nucleotide sequence as shown in SEQ ID NO.4; And / or, (4), the third reverse primer set, consisting of a nucleotide sequence as shown in SEQ ID NO.5; And / or, (5) a nucleotide sequence that encodes the same protein as any of (1) to (4), but is different from any of (1) to (4) due to the degeneracy of the genetic code; And / or, (6) a nucleotide sequence obtained by substituting, deleting or adding one or more nucleotide sequences to any of the nucleotide sequences described in (1) to (5), and a nucleotide sequence that is functionally identical or similar to any of the nucleotide sequences described in (1) to (5); And / or, (7), a nucleotide sequence having at least 90% sequence homology with any of (1) to (6).
3. The primer combination as described in claim 2, characterized in that, The primer set includes a first reverse primer set consisting of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.2; a second reverse primer set consisting of a nucleotide sequence and a fluorescent adapter sequence as shown in SEQ ID NO.3; and a first forward primer set consisting of a nucleotide sequence as shown in SEQ ID NO.
4.
4. The primer combination as described in claim 2, characterized in that, The fluorescent adapter sequence has a nucleotide sequence as shown in SEQ ID NO.6 or SEQ ID NO.
7.
5. A reagent kit, characterized in that, Includes the primer combinations as described in any one of claims 2 to 4.
6. The use of the primer combination as described in any one of claims 2 to 4 or the kit as described in claim 5 in any of the following aspects: (I) Predicting, screening and / or identifying pumpkins; (II) Pumpkin breeding.
7. The application as described in claim 6, characterized in that, The process includes the following steps: extracting pumpkin genomic DNA, and obtaining the pumpkin genotype using the primer combination as described in any one of claims 2 to 4 and / or the kit as described in claim 5.
8. The application as described in claim 7, characterized in that, This includes the PCR amplification step, and the genotype of pumpkin is determined based on the fluorescence signal of the PCR product.
9. The application as described in claim 8, characterized in that, The criteria for determining the pumpkin genotype based on fluorescence signals are as follows: if only the fluorescence signal corresponding to the first reverse primer set is detected, the detection site is the AA genotype; if only the fluorescence signal corresponding to the second reverse primer set is detected, the detection site is the GG genotype; if the fluorescence signals corresponding to both the first and second reverse primer sets are detected simultaneously, the detection site is the AG genotype.
10. The application as described in claim 8 or 9, characterized in that, The PCR amplification reaction procedure includes: (D1) Treat at 94℃~95℃ for 14min~15min; treat at 94℃~95℃ for 20s~30s; treat at 65℃~57℃ for 60s~70s; 10~15 cycles, each cycle decreasing by 0.6℃~0.8℃; And / or, (D2), 94℃~95℃ treatment for 20s~30s, 55℃~57℃ treatment for 60s~70s, 26~30 cycles; And / or, (D3), 94℃~95℃ treatment for 4min~5min; 94℃~95℃ treatment for 20s~30s, 60℃~65℃ treatment for 20s~30s, 72℃~75℃ treatment for 20s~30s, 35~40 cycles; 72℃~75℃ treatment for 7min~9min.
Citation Information
Patent Citations
Chinese pumpkin SNP (Single Nucleotide Polymorphism) molecular marker combination, SNP chip and application thereof
CN118497397A
Application of gene, protein and molecular marker related to edible rate of pumpkin fruit
CN119020525A