Eggplant core SNP molecular marker set and application
By developing the core SNP molecular marker set of eggplant and using resequencing and KASP technology, the problems of germplasm resource distinction and variety identification in eggplant breeding are solved, and rapid and accurate variety identification and seed purity detection are achieved, and breeding preparation is optimized.
Patent Information
- Application Number
- CN202510554711.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
During the existing eggplant breeding process, it is impossible to effectively distinguish between germplasm resources, variety identification and authenticity detection, and there are phenomena such as "forexual objects of the same name" and "forexual objects of the same name".
A set of core SNP molecular markers for eggplant was developed, and 23 SNP markers were screened through resequencing technology, and genotyped based on genomic resequencing data, combined with KASP technology to form a DNA fingerprint map for variety identification and seed purity detection.
It has achieved rapid and accurate distinction between germplasm resources, simplified the variety identification process, improved authenticity detection efficiency, overcome the phenomena of "foreign objects with the same name" and "foreign objects with the same name", and optimized breeding work preparations.
Smart Images

Figure CN120485410A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular genetic breeding, and in particular to an eggplant core SNP molecular marker set and application thereof. Background Art
[0002] With the rapid development of biological sequencing technology, various DNA markers, such as RFLP, AFLP, SSR, InDel, and SNP, have been developed for germplasm screening, variety identification, and intellectual property protection. Compared with these morphological markers, the most prominent and practical are the new generation of DNA sequence-based SNP markers, which are numerous, widely distributed, highly variable, highly polymorphic, stable, and often codominant in phenotype. They are also suitable for large-scale selection and identification of breeding materials. These markers have been applied to variety identification and seed purity testing of vegetable crops such as peppers, cucumbers, and gourds. With the development of high-throughput sequencing technology, whole-genome resequencing (WGS) based on second-generation high-throughput sequencing has emerged. This technology utilizes high-throughput sequencing platforms to indiscriminately sequence the entire genomic DNA of a species, covering the entire genome (including coding and non-coding regions, regulatory elements, etc.), enabling the precise detection of various genetic variations, such as single nucleotide polymorphisms (SNPs), insertions and deletions (InDels), and structural variations (SVs), across the entire genome. Its core advantage lies in the high-resolution analysis of the genome achieved through deep sequencing, which features strong data integrity, comprehensive variant detection, and high positioning accuracy. The DNA fingerprint markers developed using whole-genome resequencing technology are derived from a large number of highly polymorphic and stable markers screened across the entire genome. These markers, when combined through algorithmic optimization, form a uniquely identifiable "molecular identity card" that can accurately distinguish between different materials within the same species, and even identify closely related varieties with highly similar genetic backgrounds. With its full genome coverage and high-precision identification characteristics, this technology has demonstrated widespread application value in crop germplasm resource research.
[0003] DNA fingerprinting technology is an effective method for studying the genetic diversity and seed purity of plant germplasm resources. In recent years, the technology has gradually matured and can quickly and accurately detect the authenticity of species, as well as identify and classify them. SNP molecular marker technology is currently the most widely researched and applied genetic marker. Its high polymorphism, large number of markers, and dominant or co-dominant expression make it very suitable for molecular marker-assisted breeding analysis. KASP technology for SNP detection is characterized by low cost, high accuracy, and suitability for detecting SNP loci in large numbers of samples. It has high application value in vegetable genetic diversity analysis and fingerprinting construction.
[0004] Eggplant (Solanum melongena L., 2n=24), a plant of the genus Solanum in the Solanaceae family, originates in tropical southern Asia and is widely cultivated worldwide. Eggplant boasts a wide adaptability, high yield, and rich nutrition profile. Rich in vitamins, trace elements, and various alkaloids, it has been shown to lower cholesterol and inhibit the proliferation of digestive tract tumor cells, making it of significant agricultural economic value and health benefits. Diverse geographical environments and eggplant consumption preferences in my country result in significant regional variations in consumer demand for eggplant shape and color. To facilitate the breeding of new eggplant varieties, diverse parent selection is crucial. However, due to local customs and regional seed industry exchanges, eggplant resources often have different names and names for the same species. Therefore, effectively distinguishing germplasm resources and rapidly identifying and verifying varieties are paramount for breeders.
[0005] Therefore, in the current eggplant breeding process, there are problems such as the inability to effectively distinguish germplasm resources, variety identification and authenticity testing. Summary of the Invention
[0006] In order to solve the above technical problems existing in the existing eggplant breeding process, the present invention provides an eggplant core SNP molecular marker set and its application, which has the characteristics of being able to effectively distinguish germplasm resources, facilitate variety identification and authenticity detection.
[0007] The first technical solution of the present invention: an eggplant core SNP molecular marker set, wherein the molecular marker set comprises 23 SNP markers; the site information of the 23 SNP markers is:
[0008]
[0009]
[0010] . The eggplant core SNP molecular marker set of the present invention is developed based on resequencing technology, which can quickly, accurately and effectively identify and evaluate seed resources, form a DNA fingerprint map for each resource, and can effectively solve the problems of germplasm mixing and duplication of preservation and management in resources; based on genome resequencing data, the present invention develops a set of eggplant candidate core SNP molecular marker sets that evenly cover the eggplant genome, have high polymorphic information content and are suitable for typing. This molecular marker set is of great significance for eggplant variety identification, eggplant seed purity detection, eggplant SNP fingerprint map construction and eggplant molecular marker-assisted breeding. The present invention is suitable for cultivating new varieties that suit the preferences of various regions, and can accurately and quickly screen out parents with large trait differences and strong compatibility. It can effectively overcome the "same name different things" and "same thing different names" phenomena of existing eggplant varieties in my country, and make sufficient preparations for the accurate identification of germplasm resources before breeding work. The core molecular marker set can be applied to eggplant variety identification, SNP fingerprint map construction, genetic diversity analysis and molecular marker-assisted breeding. It has the characteristics of effectively distinguishing germplasm resources, simplifying variety identification process, and improving authenticity detection efficiency.
[0011] The second technical solution of the present invention: a method for obtaining an eggplant core SNP molecular marker set, comprising the following steps:
[0012] (S01) Screening of representative eggplant germplasm for resequencing;
[0013] (S02) Based on the resequencing data of step (S01), the eggplant HQ-1315V1.0 genome is used as a reference genome, and 4,869,502 SNP markers are obtained as the original SNP marker dataset SNPs;
[0014] (S03) Based on the material discrimination, genomic distribution uniformity, specificity strength, heterozygosity, gene diversity index, polymorphic information content (PIC), and minimum allele frequency (MAF) of the SNP marker, the original SNP marker dataset SNPs in step (S02) were screened to obtain a first SNP molecular marker set consisting of 225 SNP markers;
[0015] (S04) randomly selecting 96 SNP molecular markers from the first SNP molecular marker set in step (S03), performing genotyping verification on the 96 SNP molecular markers using eggplant resource materials, and screening 32 candidate SNP molecular markers based on their uniform distribution and polymorphism on the chromosome;
[0016] (S05) performing genetic diversity assessment and cluster analysis on the typing results corresponding to the 32 candidate SNP molecular markers in step (S04);
[0017] (S06) After performing population structure analysis based on the genetic diversity assessment results and cluster analysis results of step (S05), 23 core SNP molecular markers are screened again, and the 23 core SNP molecular markers are the eggplant core SNP molecular marker set. The present invention first screens out representative eggplant germplasm from the eggplant resource material for resequencing, while ensuring the comprehensiveness of the original SNP marker data set SNPs obtained, reducing the workload of resequencing; using the eggplant HQ-1315V1.0 genome as the reference genome, the original SNP marker data set SNPs are accurately obtained by comparison; based on the material discrimination of the SNP marker, the uniformity of distribution on the genome, the specificity strength, heterozygosity, gene diversity index, polymorphic information content PIC and minimum allele frequency MAF principle, the multi-aspect principle data are comprehensively considered, A first SNP molecular marker set consisting of 225 SNP markers was accurately screened; the present invention used the first SNP molecular marker set to randomly screen 96 SNP molecular markers in the first SNP molecular marker set based on physical distance for genotyping verification, and based on the uniformity of distribution on the chromosome and the degree of polymorphism, retained markers with polymorphic information content PIC>0.25, retained markers with heterozygosity 0.2≤≤0.6, excluded markers with extreme deviations, and prioritized retention of gaps through sliding window statistics (e.g., at least one marker per 5Mb interval). The SNPs with uniform chromosome distribution in the region were eliminated, and the typing success rate was >95%. After development and screening, 32 candidate SNP molecular markers were obtained. The present invention clarified the genetic variation degree of eggplant resource materials and quantified the diversity level through parameters such as material discrimination, uniform distribution on the genome, specificity strength, heterozygosity, gene diversity index, polymorphic information content PIC and minimum allele frequency MAF. Markers with PIC>0.3 (high polymorphism) were retained, and low information content sites were eliminated to provide a reliable data base for subsequent analysis. Based on the above, high-value markers were screened; eggplant resource materials were divided into several subgroups (such as cultivated species, wild species, and geographical variety groups), and genetic distances were visually displayed to reveal population stratification; genetic differences between distant materials (such as wild resources S. incanum) and cultivated species were clarified to provide a basis for hybrid parent selection and guide breeding design; branch lengths in phylogenetic trees reflect the degree of genetic differentiation (such as the obvious differentiation between Asian cultivated eggplant and African wild eggplant), and the output was visualized; the genetic structure of the population was determined and the optimal number of subgroups (K value) was determined; materials with mixed genetic backgrounds (such as Q value < 0.8 individuals may be hybrid offspring), identify gene exchange; avoid false positive associations in GWAS analysis (improve accuracy by correcting population structure); screen 23 core marker sets with high discrimination and low redundancy from 32 candidate SNP molecular markers, covering key sites of the whole genome, streamlining the number of markers; reduce the cost of primer synthesis for subsequent detection (for example, 23 markers can achieve a 99% variety differentiation rate), and optimize cost efficiency; the entire acquisition method of the present invention is scientific and reasonable, developed based on resequencing technology, and can quickly, accurately and effectively identify and evaluate seed resources, forming a DNA fingerprint map for each resource, which can well solve the problem of resource varietal differentiation. The researchers used primer information corresponding to 23 candidate SNP molecular markers to generate eggplant DNA fingerprints, thereby rapidly identifying eggplant germplasm, reducing breeding workload, saving breeding time, and improving breeding efficiency. The resulting 23 markers can distinguish 99% of varieties (e.g., varieties with genetic similarity <95%) among 280 accessions, making them well suited for seed purity testing (sensitivity down to 1% impurities) and variety rights protection (DNA fingerprint library construction). Compared to SSR markers, the KASP core set reduces detection time by 70% (2 hours vs. 6 hours) and costs by 60%. The representative eggplant accessions include 45 accessions, while the resource accessions include 280 accessions. The resequencing was performed by Beijing Novogene Technology Co., Ltd.
[0018] Preferably, the step (S02) is:
[0019] (S02-1) Using fastaqc software to perform raw data quality control on the resequencing data in step (S01), and using fastp software to delete low-quality, contaminated or invalid sequences and adapters;
[0020] (S02-2) Using the eggplant HQ-1315V1.0 genome as the reference genome, the resequencing data processed in step (S02-1) were compared with the eggplant HQ-1315V1.0 genome using BWA software, and the detected base variations were separated using GATK software to obtain 4,869,502 SNP markers as the original SNP marker dataset SNPs. Fastaqc can quickly screen high-quality sequences from resequencing data, accurately and quickly control the quality of raw data, remove low-quality sequences, and prepare for alignment with the reference genome and further screening. After inputting the resequencing data into the fastaqc software, the proportion of bases with Q30 (error rate ≤ 0.1%) should be significantly improved (e.g., from 90% to above 95%) after filtering, indicating that low-quality sequences have been effectively removed. The base quality value (Q30+ percentage) should be improved and consistent with the expected distribution of the reference genome, excluding contamination or technical deviation. After filtering, the sequence length should be normalized, and short or broken sequences should be removed. An HTML report is generated to show the comparison of data quality before and after filtering. Adapter sequences are automatically identified and removed using a built-in database (such as Illumina TruSeq). The adapter detection rate should be close to 100% (the adapter trimmed read ratio is displayed in the log). Under the default parameters, bases with a quality <20 and reads with an N ratio >5% are removed. The proportion of filtered reads in the output statistics reflects the stringency of the filtering (usually retaining 80-90% of the data is ideal), generating clean data. FASTQ) provides high-quality input for subsequent alignment. The eggplant HQ-1315 V1.0 genome, serving as a reference genome, is highly representative. Comparison with BWA software allows for rapid and accurate determination of base variation in high-quality sequences. GATK software allows for rapid separation of the 4,869,502 detected base variations, serving as the raw SNP marker dataset for further screening.
[0021] Preferably, the step (S03) is:
[0022] (S03-1) performing initial filtering on the original SNP marker dataset SNPs in step (S02);
[0023] (S03-2) Using vcftools software according to standard parameters, the SNP markers obtained from the initial filtration in step (S03-1) are subjected to a secondary filtration to obtain a first SNP molecular marker set consisting of 225 SNP markers. The primary filtration can filter out SNP marker data with chaotic variation; the secondary filtration is continued using vcftools software according to standard parameters to quickly and accurately obtain the first SNP molecular marker set consisting of 225 SNP markers.
[0024] Preferably, the initial filtering method in step (S03-1) is:
[0025] (S03-1-1) Using a Python script to extract the sequences 100 bp upstream and downstream of the identified SNP markers in the original SNP marker dataset SNPs of step (S02); the extracted sequences are based on the absence of any other variant positions in the region;
[0026] (S03-1-2) Using GATK software, the SNP markers extracted in step (S03-1-1) were initially filtered according to the parameters QD < 2.0, MQ < 40.0, MQRankSum < -12.5, and ReadPosRankSum < -8.0. Using Python scripts, the 100bp upstream and downstream sequences of SNP markers identified in the original SNP marker dataset and without any other variant positions in the region were quickly and accurately extracted for subsequent processing. Using GATK software, the parameters QD < 2.0, MQ < 40.0, MQRankSum < -12.5, and ReadPosRankSum < -8.0 were used to quickly and accurately perform an initial filtering of the extracted SNP markers, filtering out SNP marker data with chaotic variations.
[0027] Preferably, the standard parameters in step (S03-2) are -mac3-maf0.05-maxangles2-minangles2-meanDP4minQ30-maxmissing0.5. After the defined standard parameters can be quickly and accurately screened, a core SNP molecular marker set of 225 SNP markers in eggplant is obtained.
[0028] Preferably, the step (S04) is to randomly select 96 SNP molecular markers in the first SNP molecular marker set of step (S03), extract 100bp sequences upstream and downstream of the 96 SNPs, design and develop KASP primers, and use KASP technology and eggplant resource materials to genotype the 96 randomly selected SNP molecular markers in the first SNP molecular marker set, and further screen and obtain 32 candidate SNP molecular markers based on the typing results. Using the first SNP molecular marker set, 96 SNP molecular markers in the first SNP molecular marker set randomly screened based on physical distance were used for genotyping verification for subsequent further development; the 100bp sequences upstream and downstream of the selected 96 SNPs were extracted and genotyped using KASP technology. A preliminary genotyping experiment was first performed using 44 eggplant germplasm resources, and 54 markers with poor typing results were first eliminated. The remaining 42 KASP markers were used to complete the typing of 236 eggplant resource materials, and the poor typing results were again eliminated to obtain the corresponding typing results of 32 candidate SNP molecular markers for 280 eggplant resource materials; the genomic DNA of the eggplant resource materials was extracted using the CTAB method.
[0029] Preferably, during genotyping using KASP technology, two allele-specific primers and one universal primer are designed for each KASP marker. The design of two allele-specific primers and one universal primer enables precise typing through competitive fluorescence signal amplification, utilizing high specificity to accurately distinguish single-base differences in SNPs. This high-throughput assay is more adaptable to breeding and seed quality inspection needs, and the universal primer also reduces consumable costs. These primers were synthesized by Zhejiang Yihe Gene Co., Ltd.
[0030] Preferably, the parameters for developing candidate KASP marker primer information are: GC% less than 60%, product length should not exceed 120 bp, and annealing temperature within the range of 55° C. to 62° C. These defined parameters can more accurately develop candidate core KASP marker primers.
[0031] Preferably, the method of performing genotyping using KASP technology comprises the following steps:
[0032] (A01) Conducting a preliminary genotyping experiment using 44 eggplant germplasm resources;
[0033] (A02) Eliminate KASP markers with poor typing results after the preliminary genotyping experiment in step (A01); 54 markers were eliminated;
[0034] (A03) using the remaining KASP markers from step (A02) to complete typing of the eggplant resource materials; the remaining KASP markers are 42; and the eggplant resource materials are 236;
[0035] (A04) After removing the KASP markers with poor typing results in step (A03) again, 32 candidate KASP markers were obtained. There are 280 eggplant resource materials. First, 44 eggplant germplasm resources were selected from the eggplant resource material library for preliminary genotyping experiments, and the feasibility and polymorphism of SNP markers were preliminarily screened. Through small-scale sample (44 copies) verification, invalid markers caused by primer design defects (such as secondary structure, non-specific binding) were eliminated, and inefficient markers were quickly eliminated to reduce the consumption of invalid consumables in subsequent large-scale testing (such as 54 markers were eliminated to avoid wasting resources in 236 materials). The small sample pre-experiment can cover the common allele frequencies to ensure that the markers are representative in the population; 54 KASP markers with poor typing results, such as low polymorphism (PIC <0.2), low amplification efficiency (such as Ct value >30), and non-specific signals (miscellaneous peaks) were eliminated to improve the quality of markers and ensure the quality of markers. The remaining 42 markers had a high amplification success rate (>95%) and clear clustering (such as high FAM / HEX signal separation), which optimized subsequent experiments, reduced data noise, and improved the accuracy of typing 236 materials. The remaining KASP markers were used to complete the typing of eggplant resource materials, and the sample size was expanded to verify the stability and population applicability of the markers, evaluate the generalization ability of the markers, and test them in a wider genetic background (such as different geographical varieties and wild relatives) to avoid markers being only applicable to specific subpopulations. Evenly distributed markers were screened, and chromosome coverage analysis (such as whether the distribution of the 42 markers was balanced) was performed to ensure that the subsequent 32 markers were free of genomic bias. After further eliminating KASP markers with poor typing results, 32 candidate KASP markers were obtained.
[0036] Preferably, the CTAB method is used to extract genomic DNA from the eggplant resource material. The CTAB method can efficiently remove polysaccharides and polyphenols from eggplant, resulting in higher-purity extracted DNA and preventing polysaccharide inhibition of enzyme activity. Because the cell wall structure and metabolite composition of different eggplant parts (leaves, stems, fruits, and seeds) vary significantly, adjusting the CTAB concentration (2%-3%) and lysis time (30-60 minutes in a 65°C water bath) can effectively lyse various tissues. The CTAB method can remove RNA interference and reduce DNA breakage, yielding high-molecular-weight DNA. DNA extracted using the CTAB method is highly pure and can be directly used in KASP reactions (without additional purification), meeting the DNA integrity and purity requirements of platforms such as Illumina and PacBio.
[0037] Preferably, in the step (S05), the genetic diversity of the eggplant resource material is assessed by typing the corresponding 32 candidate SNP molecular markers using PowerMarker software;
[0038] Cluster analysis was performed using PowerMarker V3.25 software on 32 candidate SNP molecular markers in eggplant accessions, and a phylogenetic tree was constructed using the Nei 1983 method. The genetic variation and diversity levels of the 280 accessions were determined using parameters such as material discrimination, genomic distribution uniformity, specificity strength, heterozygosity, gene diversity index, polymorphic information content (PIC), and minimal allele frequency (MAF). Markers with a PIC > 0.3 (high polymorphism) were retained, while low-information content sites were removed to provide a reliable data foundation for subsequent analysis and to screen high-value markers. PowerMarker V3.25 software was used to divide the 280 accessions into several subgroups (e.g., cultivated species, wild species, and geographic variety groups), visually displaying genetic distances and revealing population stratification. Genetic differences between distantly related accessions (e.g., wild resource S. incanum) and cultivated species were identified, providing a basis for hybrid parent selection and guiding breeding design. Branch lengths in the phylogenetic tree, reflecting the degree of genetic differentiation (e.g., the clear differentiation between Asian cultivated eggplant and African wild eggplant), were visualized.
[0039] Preferably, the genetic diversity in step (S05) includes material differentiation, uniformity of distribution on the genome, specificity strength, heterozygosity, gene diversity index, polymorphism information content PIC and minimum allele frequency MAF. The ability of markers to distinguish different eggplant varieties can be assessed through population structure analysis (such as PCA or NJ trees). Markers should be able to cluster eggplants according to geographical origin / phenotype. The uniformity of genomic distribution can be determined through sliding window statistics (such as a coefficient of variation of the number of markers per 5-Mb interval <30%). For example, 32 markers cover all 12 chromosomes with a maximum spacing of <10 Mb. Specificity refers to the uniqueness of a marker in the target species. The observed heterozygosity (Ho) ranges from 0 (homozygous) to 1 (completely heterozygous), reflecting the heterozygous state of the population. The gene diversity index (He) measures the expected value of genetic variation. The polymorphic information content (PIC) is divided into low polymorphism (PIC < 0.25), medium polymorphism (0.25 ≤ PIC ≤ 0.5), and high polymorphism (PIC > 0.5). The screening criterion for the minimum allele frequency (MAF) is MAF ≥ 0.05 (to avoid low-frequency invalid markers). Seven indicators are used to cover the multi-dimensional characteristics of genetic diversity, avoiding the bias of a single parameter and having good comprehensiveness; quantitative data supports molecular breeding decisions (such as parent selection or marker elimination) with good accuracy; through genetic diversity assessment, the value of the KASP marker set can be maximized, promoting the transition of eggplant from resource protection to molecular design breeding.
[0040] Preferably, the heterozygosity of the 32 candidate SNP molecular markers in the genetic diversity of step (S05) corresponding to the candidate KASP marker primer information varies between 0.016 and 0.719. More preferably, the average heterozygosity of the 32 candidate SNP molecular markers in the genetic diversity of step (S05) corresponding to the candidate KASP marker primer information is 0.283. The heterozygosity ranges from 0.016 to 0.719, with 0.016 likely due to inbred material or typing error, and 0.719 close to the theoretical maximum (0.75, biallelic balanced population), indicating high heterozygosity. The average value of 0.283 is close to the expected value of natural populations (plants are usually 0.2-0.4), consistent with the neutral evolution hypothesis.
[0041] Preferably, the gene diversity index of the candidate KASP marker primer information corresponding to the 32 candidate SNP molecular markers in the genetic diversity of step (S05) fluctuates between 0.293 and 0.597. More preferably, the average gene diversity index of the candidate KASP marker primer information corresponding to the 32 candidate SNP molecular markers in the genetic diversity of step (S05) is 0.460. The gene diversity index fluctuates between 0.293 and 0.597, with a lower limit of 0.293 indicating that some markers are under selection pressure (such as being located in conserved functional regions), an upper limit of 0.597 close to the theoretical maximum difference value (0.5 is a biallelic equilibrium state), and an average value of 0.460 significantly exceeding the threshold of 0.3, indicating that the marker set is highly polymorphic as a whole.
[0042] Preferably, among the candidate KASP marker primer information corresponding to the 32 candidate SNP molecular markers in the genetic diversity of step (S05), 28 candidate KASP marker primer information have a polymorphic information value (PIC) exceeding 0.3. More preferably, the overall average polymorphic information value (PIC) of the 32 candidate SNP molecular markers in the genetic diversity of step (S05) corresponding to the candidate KASP marker primer information is 0.355. 28 markers, accounting for 87.5% (28 / 32), have a PIC > 0.3, far exceeding the conventional screening standard (usually requiring > 50%); an average PIC of 0.355 is suitable for QTL mapping (a PIC > 0.3 improves efficiency by 30%).
[0043] Preferably, the minor allele frequencies of the 32 candidate SNP molecular markers in the genetic diversity of step (S05) corresponding to the candidate KASP marker primer information are between 0.441 and 0.821. More preferably, the average minor allele frequency of the 32 candidate SNP molecular markers in the genetic diversity of step (S05) corresponding to the candidate KASP marker primer information is 0.62. The minor allele frequency (MAF) ranges from 0.441 to 0.821, with a MAF lower limit > 0.4, significantly higher than the conventional MAF ≥ 0.05 standard, indicating forced selection of high-frequency minor alleles; the average value is 0.62. When MAF > 0.5, the minor allele is actually the dominant allele, which may reflect traces of artificial selection and indicate good genetic variation strength. This range and average value indicate that there is a large amount of genetic variation between these markers.
[0044] Preferably, in step (S06), based on the typing results corresponding to the candidate KASP marker primer information of the 32 candidate SNP molecular markers of the eggplant resource material, the population structure analysis is performed using Structure software, and the results are further processed through the Structure Harvester website; 23 core SNP molecular markers are obtained using SNPT software, which is the eggplant core SNP molecular marker set. The population genetic structure is judged by LnP (D) and ΔK value (calculated by Structure Harvester) (e.g., K=3 means that the material is divided into 3 categories), and the optimal number of subpopulations (K value) is determined; materials with mixed genetic backgrounds are detected (e.g., individuals with Q value <0.8 may be hybrid offspring) to identify gene exchange; false positive associations in GWAS analysis are avoided (accuracy is improved by correcting population structure); 23 high-discrimination, low-redundancy core marker sets are screened from the 32 candidate markers using SNPT software, covering key sites across the genome and streamlining the number of markers; the cost of primer synthesis for subsequent testing is reduced (e.g., 23 markers can achieve a 99% variety differentiation rate), thereby optimizing cost efficiency.
[0045] The third technical solution of the present invention: application of core KASP marker primer information corresponding to eggplant core SNP molecular markers in eggplant genomic DNA fingerprint analysis,
[0046] Based on the genotyping results of eggplant resource materials using the genetic analysis core marker set, a fingerprint map was constructed using R Studio. Typing results for 23 markers (e.g., AA / GG / AG) were converted into binary codes (e.g., 0 / 1 / 2) to generate a unique fingerprint map and digitize the variety ID. Similarities between varieties were visualized using heat maps or PCA plots, supporting infringement identification and germplasm resource management. Seed companies can quickly identify counterfeit varieties by comparing fingerprint maps (a variety is considered infringing if its similarity to the database exceeds 95%). Through the standardized operation of "analysis-screening-visualization", this application converts KASP markers into practical molecular tools, providing full-chain technical support for the eggplant seed industry from resource protection to commercial breeding; the eggplant core SNP molecular marker set developed based on resequencing technology can quickly, accurately and effectively identify and evaluate seed resources, forming a DNA fingerprint map for each resource, which can solve problems such as germplasm mixing, duplication of preservation and management, and intellectual property disputes in resources; based on this core SNP molecular marker set, the eggplant core germplasm and fingerprint map were constructed, and an accurate, simple and rapid method for eggplant resource, variety identification and purity detection was established, which can accurately analyze the genetic structure between eggplant germplasm materials.
[0047] The fourth technical solution of the present invention: Application of the eggplant core SNP molecular marker set in any of the following aspects:
[0048] (1) Eggplant variety identification;
[0049] (2) Eggplant variety purity testing;
[0050] (3) Eggplant core germplasm screening;
[0051] (4) Construction of SNP fingerprints of eggplant varieties;
[0052] (5) Analysis of genetic diversity of eggplant;
[0053] (6) Molecular marker-assisted breeding.
[0054] The present invention is developed based on genome resequencing data and has the characteristics of uniform coverage of the eggplant genome, high polymorphic information content and suitability for typing;
[0055] Used for eggplant variety identification, this technology effectively differentiates eggplant resource materials, resolving the current problem of "same name different species" and "same species different names" in eggplant resources due to local customs and regional seed industry exchanges. Core SNP markers can accurately distinguish morphologically similar varieties (such as long or round eggplant varieties from different regions). Genotype comparison can be completed within 2 hours, surpassing traditional morphological identification that requires an entire growth cycle.
[0056] Used for eggplant variety purity testing, it has high sensitivity and can detect less than 1% of impure plants (such as self-pollinated plants in F1 hybrids). It can test hundreds of samples at a time, replacing field planting testing, shortening the testing cycle (from several months to several days), and reducing testing costs.
[0057] Used for eggplant core germplasm screening, it can screen core germplasm covering more than 95% of genetic variation through whole-genome SNP data, avoid repeated preservation of resources with similar genetic backgrounds, optimize germplasm resource management, and improve breeding efficiency.
[0058] Used for constructing SNP fingerprints for eggplant varieties, generating a DNA fingerprint for each resource, this technology can address issues such as germplasm contamination, duplicated conservation management, and intellectual property disputes. Each variety corresponds to a unique SNP fingerprint, which can be generated into a QR code or database entry. Compatible with different platforms (such as Illumina and Thermo Fisher), it facilitates updates and expansions, ensuring long-term availability, and supports global variety rights protection and market regulation.
[0059] Used for analyzing eggplant genetic diversity, it reveals precise stratification of gene flow between cultivated and wild species (such as S. incanum) through PCA or population structure analysis (such as STRUCTURE); combined with phenotypic data, it locates diversity hotspots; and provides a molecular basis for hybrid parent selection to avoid inbreeding depression.
[0060] Used for molecular marker-assisted breeding, precise trait tracking, such as the Rfo-sa1 gene-associated SNP to accelerate disease resistance tracking in bacterial wilt-resistant breeding, such as SNP markers that control peel color (ANS gene) or shape (OVATE gene) fruit trait tracking; facilitates early selection, screening target genotypes at the seedling stage, shortening the breeding cycle by more than 50%; breaking through the traditional breeding's reliance on phenotypes and improving the efficiency of targeted improvement.
[0061] The present invention has the following beneficial effects:
[0062] (1) The eggplant core SNP molecular marker set was developed based on resequencing technology, which can quickly, accurately, and effectively identify and evaluate seed resources, form a DNA fingerprint of each resource, and effectively solve problems such as germplasm mixing and duplication of conservation management in resources;
[0063] (2) Based on genome resequencing data, a set of candidate core SNP molecular markers for eggplant was developed, which evenly covers the eggplant genome, has high polymorphic information content, and is suitable for typing. This molecular marker set is of great significance for eggplant variety identification, eggplant seed purity detection, eggplant SNP fingerprint construction, and eggplant molecular marker-assisted breeding.
[0064] (3) It is suitable for breeding new varieties that suit the preferences of various regions, and can accurately and quickly screen out parents with large differences in traits and strong compatibility. It can effectively overcome the "same name, different things" and "same thing, different names" phenomena of existing eggplant varieties in my country, and make adequate preparations for the accurate identification of germplasm resources before breeding. This core molecular marker set can be applied to eggplant variety identification, SNP fingerprint map construction, genetic diversity analysis and molecular marker-assisted breeding. It has the characteristics of effectively distinguishing germplasm resources, simplifying variety identification process, and improving authenticity detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a graph showing the results of cluster analysis of KASP genotyping of 280 eggplant resource materials using PowerMarker software in Example 2 of the present invention;
[0066] Figure 2 The fingerprints of 280 eggplant resource materials marked by 23 core KASPs in Example 2 of the present invention are as follows;
[0067] Figure 3 Graph showing representative results of KASP genotyping analysis of the present invention;
[0068] Figure 4 This is a partial typing result diagram of the KASP core markers of the present invention. DETAILED DESCRIPTION
[0069] The present invention will be further described below with reference to the accompanying drawings and examples, but they are not intended to limit the present invention.
[0070] Eggplant core SNP molecular marker set, the molecular marker set includes 23 SNP markers; the site information of the 23 SNP markers is shown in Table 1,
[0071] Table 1: Locus information of the 23 eggplant core SNP molecular marker sets
[0072] serial number Chromosome number Location Mutation Type qz_snp_1_3 Chr01 29,954,211 C / G qz_snp_1_47 Chr01 71,560,253 C / G qz_snp_2_5 Chr02 18,360,769 G / A qz_snp_2_10 Chr02 28,838,146 A / G qz_snp_3_12 Chr03 33,292,997 T / G qz_snp_3_40 Chr03 87,724,860 T / A qz_snp_3_42 Chr03 87,727,513 C / G qz_snp_4_6 Chr04 11,861,196 T / G qz_snp_5_2 Chr05 14,818,941 G / A qz_snp_5_18 Chr05 43,468,352 T / A qz_snp_6_2 Chr06 1,805,158 G / A qz_snp_6_4 Chr06 2,867,380 G / C qz_snp_6_17 Chr06 28,589,838 T / A qz_snp_7_17 Chr07 28,589,838 G / A qz_snp_7_29 Chr07 65,061,249 T / C qz_snp_8_2 Chr08 3,695,828 A / T qz_snp_8_13 Chr08 8,006,307 G / A qz_snp_9_2 Chr09 92,631 C / T qz_snp_9_29 Chr09 48,708,904 C / T qz_snp_10_1 Chr10 4,194,569 T / C qz_E10_1 Chr10 76481266 G / A qz_snp_11_28 Chr11 59,143,459 T / C qz_snp_12_1 Chr12 4,494,768 A / C .
[0073] The method for obtaining the eggplant core SNP molecular marker set comprises the following steps:
[0074] (S01) selecting representative eggplant germplasms for resequencing; there are 45 representative eggplant germplasms, as shown in Table 8; the resequencing was completed by Beijing Novogene Technology Co., Ltd.;
[0075] Table 8: 45 representative eggplant accessions
[0076]
[0077]
[0078] (S02) Based on the resequencing data of step (S01), the eggplant HQ-1315V1.0 genome is used as a reference genome, and 4,869,502 SNP markers are obtained as the original SNP marker dataset SNPs;
[0079] (S02-1) Using fastaqc to perform raw data quality control on the resequencing data in step (S01), deleting low-quality sequences, and using fastp software to delete low-quality, contaminated or invalid sequences and adapters;
[0080] (S02-2) Using the eggplant HQ-1315V1.0 genome as the reference genome, the resequencing data processed in step (S02-1) were compared with the eggplant HQ-1315V1.0 genome using BWA software, and the detected base variations were separated using GATK software to obtain 4,869,502 SNP markers as the original SNP marker dataset SNPs;
[0081] (S03) Based on the material discrimination, genomic distribution uniformity, specificity strength, heterozygosity, gene diversity index, polymorphic information content (PIC), and minimum allele frequency (MAF) of the SNP marker, the original SNP marker dataset SNPs in step (S02) were screened to obtain a first SNP molecular marker set consisting of 225 SNP markers;
[0082] (S03-1) performing initial filtering on the original SNP marker dataset SNPs in step (S02);
[0083] The initial filtering method is,
[0084] (S03-1-1) Using a Python script to extract the sequences 100 bp upstream and downstream of the identified SNP markers in the original SNP marker dataset SNPs of step (S02); the extracted sequences are based on the absence of any other variant positions in the region;
[0085] (S03-1-2) Using GATK software, the SNP markers extracted in step (S03-1-1) were initially filtered according to the parameters QD < 2.0, MQ < 40.0, MQRankSum < -12.5, and ReadPosRankSum < -8.0;
[0086] (S03-2) Using vcftools software to perform a secondary filtration on the SNP markers obtained by the initial filtration in step (S03-1) according to standard parameters, a first SNP molecular marker set consisting of 225 SNP markers was obtained; the standard parameters were -mac3-maf0.05-max angles2-min angles2-mean DP 4minQ 30-max missing0.5;
[0087] The site information of the first SNP molecular marker set is shown in Table 2.
[0088] Table 2: Site information of the first SNP molecular marker set
[0089]
[0090]
[0091]
[0092] The locus information in the first SNP molecular marker set obtained through multi-dimensional screening can further significantly reduce the subsequent genotyping costs and quickly screen out the optimal eggplant core SNP molecular marker set;
[0093] (S04) Randomly select 96 SNP molecular markers from the first SNP molecular marker set in step (S03), perform genotyping verification on the 96 SNP molecular markers using eggplant resource materials, and screen and obtain 32 candidate SNP molecular markers as shown in Table 3 based on the distribution uniformity and polymorphism degree on the chromosome; there are 280 eggplant resource materials; Figure 3 The following are representative results of KASP genotyping analysis. The figure indicates that a KASP marker with good polymorphism was successfully developed; B indicates that a KASP marker that is difficult to classify should be discarded; C indicates that a monomorphic KASP marker should be discarded; and D indicates that a KASP marker without polymorphism should be discarded. The genomic DNA of the eggplant resource material was extracted using the CTAB method.
[0094] Using eggplant resource materials, 96 SNP molecular markers in the first SNP molecular marker set of step (S03) are randomly selected based on physical distance, and sequences 100 bp upstream and downstream of the 96 SNPs are extracted to design and develop KASP primers. The 96 SNP molecular markers randomly selected from the first SNP molecular marker set are genotyped using KASP technology and eggplant resource materials. Based on the typing results, 32 candidate SNP molecular markers are further screened and obtained. During the genotyping process using KASP technology, two allele-specific primers and one universal primer are designed for each KASP marker. Parameters for developing the candidate KASP marker primer information are: GC% less than 60%, product length should not exceed 120 bp, and annealing temperature within the range of 55° C. to 62° C.
[0095] The method of performing genotyping using KASP technology comprises the following steps:
[0096] (A01) Conducting a preliminary genotyping experiment using eggplant germplasm resources; there are 44 eggplant germplasm resources; (A02) Eliminating KASP markers with poor typing results after the preliminary genotyping experiment in step (A01); 54 markers were eliminated; (A03) Using the remaining KASP markers from step (A02) to complete genotyping of eggplant resource materials; there are 42 remaining KASP markers; there are 236 eggplant resource materials; (A04) After further eliminating the KASP markers with poor typing results in step (A03), 32 candidate KASP markers were obtained; there are 280 eggplant resource materials as shown in Table 4;
[0097] Table 3: Primer information corresponding to 32 candidate KASP markers
[0098]
[0099]
[0100]
[0101] All 32 candidate SNP molecular markers have a high typing success rate, with ≥95% of the materials being clearly typed. Subsequent candidate core KASP marker primer information can be further verified and selected from these, which is a best-of-the-best approach.
[0102] Table 4: 280 eggplant resource materials
[0103]
[0104]
[0105]
[0106] 280 eggplant resource materials, covering a wide range of genetic diversity, including representative varieties and origins. These 280 materials typically include cultivated varieties, landraces, and wild relatives, ensuring that SNP markers can be stably typed across different genetic backgrounds, improving the reliability and representativeness of marker screening and avoiding markers that are only applicable to specific subpopulations;
[0107] (S05) performing genetic diversity assessment and cluster analysis on the typing results corresponding to the 32 candidate SNP molecular markers in step (S04), as shown in Table 5;
[0108] The genetic diversity of eggplant resource materials was evaluated by using PowerMarker software to classify the SNP molecular markers corresponding to the typing results;
[0109] The SNP molecular markers of eggplant resource materials were clustered and analyzed using PowerMarker V3.25 software, and a phylogenetic tree was constructed using the Nei 1983 method.
[0110] The genetic diversity in the step (S05) includes material differentiation, uniformity of distribution on the genome, specificity strength, heterozygosity, gene diversity index, polymorphism information content PIC and minimum allele frequency MAF;
[0111] The heterozygosity of the 32 candidate KASP marker primer information in the genetic diversity of step (S05) varies between 0.016 and 0.719; the average heterozygosity of the 32 candidate KASP marker primer information in the genetic diversity of step (S05) is 0.283;
[0112] The gene diversity index of the 32 candidate KASP marker primer information in the genetic diversity of step (S05) fluctuates between 0.293 and 0.597; the average gene diversity index of the 32 candidate KASP marker primer information in the genetic diversity of step (S05) is 0.460;
[0113] Among the 32 candidate KASP marker primer information in the genetic diversity of step (S05), 28 candidate KASP marker primer information have a polymorphic information value PIC greater than 0.3; the overall average polymorphic information value PIC value of the 32 candidate KASP marker primer information in the genetic diversity of step (S05) is 0.355;
[0114] The minor allele frequencies of the 32 candidate KASP marker primers in the genetic diversity of step (S05) range from 0.441 to 0.821; the average minor allele frequencies of the 32 candidate KASP marker primers in the genetic diversity of step (S05) is 0.62; this range and average indicates that there is a large amount of genetic variation among these markers;
[0115] Table 5: Genetic diversity assessment results corresponding to 32 candidate SNP molecular markers
[0116]
[0117] Genetic diversity assessment results indicate that the candidate KASP marker primers corresponding to the 32 candidate SNP molecular markers are highly polymorphic and can effectively capture genetic variation in the target eggplant population. When the markers are evenly distributed across the genome or concentrated near functional genes, they can comprehensively assess population structure, phylogenetic relationships, and evolutionary history. This provides rich genetic background information for eggplant breeding, helps prevent inbreeding depression, aids in the identification of rare alleles, and enables the discovery of potentially superior eggplant germplasm. These 32 KASP markers possess both theoretical value (revealing diversity) and practical significance (guiding breeding), making them an ideal molecular tool combination.
[0118] (S06) After performing population structure analysis based on the genetic diversity assessment results and cluster analysis results of step (S05), the core KASP marker primer information corresponding to the 23 core SNP molecular markers as shown in Tables 6 and 7 was obtained, and the 23 core SNP molecular marker sets were the eggplant core SNP molecular marker sets; based on the typing results corresponding to the 23 core SNP molecular markers of the eggplant resource materials, the population structure analysis was performed using Structure software, and the results were further processed through the Structure Harvester website; SNPT software was used to obtain the core KASP marker primer information corresponding to the 23 core SNP molecular markers, and the 23 core SNP molecular marker sets were the eggplant core SNP molecular marker sets; Figure 4 Shown are some typing results of KASP core markers;
[0119] Table 6: Core KASP marker primer information corresponding to the 23 core SNP molecular markers
[0120]
[0121]
[0122] The resulting 23 candidate core KASP marker primers can simultaneously detect 23 loci, possessing multiplexed detection capabilities and significantly improving SNP typing throughput. They require only trace amounts of DNA (typically 10-50 ng), making them suitable for precious samples with low sample requirements. Dual primers (targeting both alleles of a SNP) ensure precise discrimination of single nucleotide polymorphisms, with the fluorescent signal directly reflecting the genotype and reducing false positive / negative results. The KASPs are cross-platform compatible with a variety of real-time PCR instruments, with primer design based on known SNP loci, eliminating the need for complex probes or microarrays. They accelerate genetic map construction for QTL localization and gene cloning, making them well-suited for the eggplant genome. The 23 markers were screened to cover key genes or functional regions, exhibiting high polymorphism and reproducibility. As a core marker set, they can represent the genetic diversity or important phenotypic associations of the target eggplant population.
[0123] Table 7: Genetic diversity assessment results corresponding to the 23 candidate core KASP marker primer information
[0124]
[0125] The corresponding genetic diversity assessment results showed that the 23 candidate KASP marker primers are highly polymorphic core markers, which can effectively capture the genetic variation of the target eggplant population. When the markers are evenly distributed in the genome or concentrated near functional genes, they can comprehensively evaluate the population structure, kinship and evolutionary history, provide rich genetic background information for eggplant breeding, avoid inbreeding depression, assist in identifying rare alleles, and explore potential excellent eggplant germplasm; these 23 candidate core KASP markers have both theoretical value (revealing diversity) and practical significance (guiding breeding), and are an ideal combination of molecular tools.
[0126] Application of the eggplant core SNP molecular marker set in eggplant variety identification. Application of the eggplant core SNP molecular marker set in eggplant variety purity detection. Application of the eggplant core SNP molecular marker set in eggplant core germplasm screening. Application of the eggplant core SNP molecular marker set in the construction of SNP fingerprints of eggplant varieties. Application of the eggplant core SNP molecular marker set in eggplant genetic diversity analysis. Application of the eggplant core SNP molecular marker set in molecular marker-assisted breeding. Application of KASP candidate core marker primer information in eggplant genomic DNA fingerprint analysis,
[0127] According to the genotyping results of eggplant resource materials using the genetic analysis core marker set, a fingerprint map was constructed using R studio.
[0128] Example 1: A method for obtaining a core SNP molecular marker set of eggplant, comprising the following steps:
[0129] (S01) 45 representative eggplant accessions were screened for resequencing; the resequencing was completed by Beijing Novogene Technology Co., Ltd.; (S02) the resequencing data were quality controlled using fastaqc to delete low-quality sequences, and low-quality, contaminated, or invalid sequences and adapters were analyzed using fastp software; the eggplant HQ-1315V1.0 genome was used as the reference genome, and the processed resequencing data were compared with the eggplant HQ-1315V1.0 genome using BWA software. The detected base variations were separated using GATK software, and 4,869,502 SNP markers were obtained as the original SNP marker dataset SNPs;
[0130] (S03) Based on the material discrimination, genomic distribution uniformity, specificity, heterozygosity, gene diversity index, polymorphic information content (PIC), and minimum allele frequency (MAF) of the SNP markers, the original SNP marker dataset SNPs were screened to obtain a first SNP molecular marker set consisting of 225 SNP markers;
[0131] The original SNP marker dataset SNPs were initially filtered; the initial filtering method was as follows: (S03-1-1) Python script was used to extract the sequences 100bp upstream and downstream of the identified SNP markers in the original SNP marker dataset SNPs; the extracted sequences were based on the absence of any other variant positions in the region; (S03-1-2) GATK software was used to perform an initial filtering of the extracted SNP markers according to the parameters QD < 2.0, MQ < 40.0, MQRankSum < -12.5, and ReadPosRankSum < -8.0; (S03-2) vcftools software was used to perform a secondary filtering of the SNP markers obtained by the initial filtering according to standard parameters to obtain a first SNP molecular marker set consisting of 225 SNP markers; the standard parameters were -mac3 -maf0.05 -maxangles2 -min angles2 -mean DP 4minQ 30 -max missing0.5;
[0132] (S04) randomly selecting 96 SNP molecular markers from the first SNP molecular marker set of step (S03), performing genotyping verification on the 96 SNP molecular markers using eggplant resource materials, and screening and obtaining 32 candidate SNP molecular markers as shown in Table 3 based on the uniformity of distribution on the chromosome and the degree of polymorphism; the eggplant resource materials have 280 accessions; the genomic DNA of the eggplant resource materials is extracted using the CTAB method;
[0133] Using eggplant resource materials, 96 SNP molecular markers in the first SNP molecular marker set of step (S03) are randomly selected based on physical distance, and sequences 100 bp upstream and downstream of the 96 SNPs are extracted to design and develop KASP primers. The 96 SNP molecular markers randomly selected from the first SNP molecular marker set are genotyped using KASP technology and eggplant resource materials. Based on the typing results, 32 candidate SNP molecular markers are further screened and obtained. Sequences 100 bp upstream and downstream of the selected 96 SNPs are extracted and genotyped using KASP technology to obtain 32 candidate SNP molecular markers. During the genotyping process using KASP technology, two allele-specific primers and one universal primer are designed for each KASP marker. Parameters for developing the candidate KASP marker primer information are: GC% is less than 60%, product length should not exceed 120 bp, and annealing temperature is within the range of 55° C. to 62° C.
[0134] The method of performing genotyping using KASP technology comprises the following steps:
[0135] (A01) Conducting a preliminary genotyping experiment using 44 eggplant germplasm resources;
[0136] (A02) KASP markers with poor typing results after genotyping pre-experiment were eliminated; 54 markers were eliminated;
[0137] (A03) Using the remaining KASP markers to complete typing of eggplant resource materials; there are 42 remaining KASP markers; there are 236 eggplant resource materials;
[0138] (A04) After removing KASP markers with poor typing results, 32 candidate KASP markers were obtained; there were 280 eggplant resource materials;
[0139] (S05) Genetic diversity assessment and cluster analysis were performed on the typing results of the 32 candidate SNP molecular markers;
[0140] The genetic diversity of eggplant resource materials was evaluated by using PowerMarker software to match the candidate KASP marker primer information with the typing results.
[0141] The candidate KASP marker primer information of eggplant resource materials corresponded to the typing results, and cluster analysis was performed using PowerMarker V3.25 software. The phylogenetic tree was constructed using the Nei 1983 method.
[0142] The genetic diversity includes material discrimination, uniformity of distribution on the genome, specificity strength, heterozygosity, gene diversity index, polymorphic information content PIC and minimum allele frequency MAF; the heterozygosity of the 32 candidate KASP marker primer information in the genetic diversity varies between 0.016 and 0.719; the average heterozygosity of the 32 candidate KASP marker primer information in the genetic diversity is 0.283; the gene diversity index of the 32 candidate KASP marker primer information in the genetic diversity fluctuates between 0.293 and 0.597; the gene diversity index of the 32 candidate KASP marker primer information in the genetic diversity is 0. The average value of the polymorphism information value (PIC) of the 32 candidate KASP marker primer information in the genetic diversity is 0.460; among the 32 candidate KASP marker primer information in the genetic diversity, 28 candidate KASP marker primer information have a polymorphism information value (PIC) greater than 0.3; the overall average polymorphism information value (PIC) of the 32 candidate KASP marker primer information in the genetic diversity is 0.355; the minor allele frequencies of the 32 candidate KASP marker primer information in the genetic diversity range from 0.441 to 0.821; the average minor allele frequency of the 32 candidate KASP marker primer information in the genetic diversity is 0.62; this range and average value indicate that there is a large amount of genetic variation among these markers;
[0143] (S06) After performing population structure analysis based on the results of genetic diversity assessment and cluster analysis, 23 core SNP molecular markers were obtained, which are the eggplant core SNP molecular marker set; based on the typing results corresponding to the candidate KASP marker primer information of the eggplant resource materials, the population structure analysis was performed using Structure software, and the results were further processed through the Structure Harvester website; 23 core SNP molecular markers were obtained using SNPT software, which are the eggplant core SNP molecular marker set.
[0144] Example 2: Application of core SNP molecular markers of eggplant corresponding to core KASP marker primer information in eggplant genomic DNA fingerprint analysis,
[0145] Genetic diversity assessment and cluster analysis were performed on the typing results corresponding to 32 candidate SNP molecular markers;
[0146] The genetic diversity of eggplant resource materials was evaluated by PowerMarker software based on the corresponding typing results of candidate KASP marker primer information. Cluster analysis was performed on the corresponding typing results of candidate KASP marker primer information of eggplant resource materials using PowerMarker V3.25 software, and a phylogenetic tree was constructed using the Nei 1983 method. Figure 1 As shown; Figure 1The eggplant germplasm resources are broadly divided into eight clades, with Subgroup 1, the outermost, consisting of wild eggplant varieties with small, round, green fruits. The outermost layer of the phylogenetic tree represents wild types, gradually transitioning inward to cultivated varieties. Subgroup 2 primarily consists of purple-black eggplants from Europe, while Subgroup 3 includes striped eggplants and elongated purple-black varieties. Subgroup 4 comprises varieties with purple-green peels. Subgroup 5 exhibits greater diversity in fruit color and shape. Subgroup 6, primarily originating from central and northern China, features a variety of fruit colors, including purple-red, purple-black, green, white, and striped varieties. Fruit shape within this subgroup is also highly diverse, primarily oblate, spherical, elliptical, and oblong, with occasional rare shapes such as heart-shaped and spherical. In Subgroup 7, fruits are typically elongated and primarily purple-black. Subgroup 8 primarily consists of eggplant varieties with purple peels or elongated fruits (or a combination of both phenotypes).
[0147] After population structure analysis based on the genetic diversity assessment results and cluster analysis results, information on 23 candidate core KASP marker primers was obtained. The genetic analysis core marker set corresponding to the 23 candidate core KASP marker primers is the eggplant core SNP molecular marker set;
[0148] Based on the typing results corresponding to the candidate KASP marker primer information of eggplant resource materials, the population structure analysis was performed using Structure software, and the results were further processed through the Structure Harvester website. The analysis showed that dividing eggplant germplasm resources into 8 categories was the best choice. This conclusion is consistent with the results of the cluster analysis, providing a consistent framework for understanding the genetic diversity and population structure of eggplant germplasm resources. The consistency of the Structure software analysis and the cluster analysis verified the research results and further enhanced the reliability of the determined genetic analysis. The minimum core marker was determined using SNPT software, which calculated that a set of 23 KASP markers constituted the most effective core marker set for genetic analysis. In order to facilitate the rapid identification of eggplant varieties, the genotyping results of 23 KASP core markers and 280 samples were used. According to the genotyping results of the eggplant resource materials based on the genetic analysis core marker set, a fingerprint map was constructed using R studio, as shown below: Figure 2 shown.
[0149] The present invention provides an eggplant core SNP molecular marker set, comprising 23 KASP markers. This molecular marker set has applications in eggplant variety identification, eggplant variety purity testing, eggplant SNP fingerprint map construction, eggplant genetic diversity analysis, and molecular marker-assisted breeding. The eggplant core SNP molecular marker set developed by the present invention based on resequencing technology can quickly, accurately, and effectively identify and evaluate seed resources, forming a DNA fingerprint map for each resource, and can solve problems such as germplasm mixing, duplication of preservation and management, and intellectual property disputes in resources. High-quality SNPs were detected using resequencing data from 45 eggplant materials, and the developed KASP markers were used to effectively distinguish eggplant resource materials, facilitate variety identification and authenticity testing, and promote research on the genetic diversity of eggplant germplasm resources. Based on genome resequencing data, a set of candidate core SNP molecular markers for eggplant was developed that uniformly covers the eggplant genome, has high polymorphic information content, and is suitable for typing. This molecular marker set is of great significance for eggplant variety identification, eggplant seed purity testing, eggplant SNP fingerprint map construction, and eggplant molecular marker-assisted breeding. The present invention constructs eggplant core germplasm and fingerprint map based on the core SNP molecular marker set, establishes an accurate, simple and rapid eggplant resource or variety identification and purity detection method, and can accurately analyze the genetic structure between eggplant germplasm materials.
Claims
1. Eggplant core SNP molecular marker set, characterized by: The molecular marker set includes 23 SNP markers; the site information of the 23 SNP markers is:
2. The method for obtaining the eggplant core SNP molecular marker set according to claim 1, wherein: Including the following step, (S01) Screening of representative eggplant germplasm for resequencing; (S02) Based on the resequencing data of step (S01), the eggplant HQ-1315V1.0 genome is used as a reference genome, and 4,869,502 SNP markers are obtained as the original SNP marker dataset SNPs; (S03) Based on the material discrimination, genomic distribution uniformity, specificity strength, heterozygosity, gene diversity index, polymorphic information content (PIC), and minimum allele frequency (MAF) of the SNP marker, the original SNP marker dataset SNPs in step (S02) were screened to obtain a first SNP molecular marker set consisting of 225 SNP markers; (S04) randomly selecting 96 SNP molecular markers from the first SNP molecular marker set in step (S03), performing genotyping verification on the 96 SNP molecular markers using eggplant resource materials, and screening 32 candidate SNP molecular markers based on their uniform distribution and polymorphism on the chromosome; (S05) performing genetic diversity assessment and cluster analysis on the typing results corresponding to the 32 candidate SNP molecular markers in step (S04); (S06) After performing population structure analysis based on the genetic diversity assessment results and cluster analysis results of step (S05), 23 core SNP molecular markers were screened again to obtain the eggplant core SNP molecular marker set.
3. The method for obtaining the eggplant core SNP molecular marker set according to claim 2, wherein: The step (S02) is: (S02-1) Use fastaqc software to perform raw data quality control on the resequencing data in step (S01), and use fastp software to delete low-quality sequences and adapters; (S02-2) Using the eggplant HQ-1315V1.0 genome as the reference genome, the resequencing data processed in step (S02-1) were compared with the eggplant HQ-1315V1.0 genome using BWA software, and the detected base variations were separated using GATK software to obtain 4,869,502 SNP markers as the original SNP marker dataset SNPs.
4. The method for obtaining the eggplant core SNP molecular marker set according to claim 2, wherein: The step (S03) is: (S03-1) performing initial filtering on the original SNP marker dataset SNPs in step (S02); (S03-2) Using vcftools software, the SNP markers obtained by the initial filtration in step (S03-1) were subjected to a secondary filtration according to standard parameters to obtain a first SNP molecular marker set consisting of 225 SNP markers.
5. The method for obtaining the eggplant core SNP molecular marker set according to claim 2, wherein: The step (S04) is: In the random selection step (S03), 96 SNP molecular markers in the first SNP molecular marker set are extracted, 100bp sequences upstream and downstream of the 96 SNPs are extracted, KASP primers are designed and developed, and KASP technology and eggplant resource materials are used to perform genotyping on the 96 SNP molecular markers randomly selected from the first SNP molecular marker set. Based on the typing results, 32 candidate SNP molecular markers are further screened and obtained.
6. The method for obtaining the eggplant core SNP molecular marker set according to claim 5, wherein: During genotyping using KASP technology, two allele-specific primers and one universal primer were designed for each KASP marker; The KASP marker primer information corresponding to the 23 core SNP molecular markers is:
7. The method for obtaining the eggplant core SNP molecular marker set according to claim 5, characterized in that: The method of performing genotyping using KASP technology comprises the following steps: (A01) Preliminary genotyping experiment using eggplant germplasm resources; (A02) Eliminate KASP markers with poor typing results after the genotyping preliminary experiment in step (A01); (A03) Use the remaining KASP markers from step (A02) to complete the typing of eggplant resource materials; (A04) After further eliminating the KASP markers with poor typing results in step (A03), 32 candidate KASP markers were obtained.
8. The method for obtaining the eggplant core SNP molecular marker set according to claim 2, wherein: In the step (S05), The genetic diversity of eggplant resource materials was evaluated by using PowerMarker software based on the typing results of 32 candidate SNP molecular markers. The PowerMarker V3.25 software was used to perform cluster analysis on the typing results of 32 candidate SNP molecular markers of eggplant resource materials, and the phylogenetic tree was constructed using the Nei 1983 method. In the step (S06), Based on the typing results of 32 candidate SNP molecular markers of eggplant resource materials, population structure analysis was performed using Structure software, and the results were further processed through the Structure Harvester website; The 23 core SNP molecular markers obtained using SNPT software are the eggplant core SNP molecular marker set.
9. The method for obtaining the eggplant core SNP molecular marker set according to claim 2, wherein: Application of the core KASP marker primer information corresponding to the eggplant core SNP molecular marker in the analysis of the eggplant genome DNA fingerprint, According to the genotyping results of the eggplant resource materials using the genetic analysis core marker set in step (S06), a fingerprint map was constructed using Rstudio.
10. Use of the eggplant core SNP molecular marker set according to claim 1 in any of the following aspects: (1) Eggplant variety identification; (2) Eggplant variety purity testing; (3) Eggplant core germplasm screening; (4) Construction of SNP fingerprints of eggplant varieties; (5) Analysis of genetic diversity of eggplant; (6) Molecular marker-assisted breeding.
Citation Information
Cited By
Method for developing shellfish germplasm resource evaluation molecular marker
CN121862201A