A SNP molecular marker combination for constructing eggplant line grouping and its application
By constructing a combination of 1009 SNP molecular markers, the problems of low accuracy and low detection throughput in eggplant strain population division were solved, and efficient and accurate eggplant strain grouping and detection were achieved, reducing costs.
Patent Information
- Application Number
- CN202410612808.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-17
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-05-17
AI Technical Summary
The prior art has problems such as low accuracy, long identification time, and susceptible to environmental influence in the group classification of eggplant strains, and the number of molecular markers used is small, the detection flux is low, and the cost is high.
It provides a combination of SNP molecular markers for constructing eggplant strain clusters, including 1009 SNP molecular markers, which are evenly distributed on 12 chromosomes of eggplant. It uses a liquid phase probe chip to perform high label density, strong automation, and high detection throughput.
It realizes accurate grouping of eggplant strains, improves detection throughput and automation, reduces detection costs, and has important theoretical value and application significance.
Smart Images

Figure CN118374628B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of molecular breeding, and particularly relates to an SNP molecular marker combination for constructing the population classification of eggplant lines and its application. Background Art
[0002] Eggplant originated in tropical Asia and is distributed all over the world. It is most cultivated in Asia, accounting for about 74% of the world's total output; followed by Europe, accounting for about 14%. Eggplant is cultivated in various regions of China and is one of the main vegetables in summer.
[0003] The conventional population classification of eggplant mainly relies on morphological identification. However, this method is restricted by factors such as low accuracy, long identification time, and susceptibility to environmental influence. Moreover, as the number of eggplant lines increases, the phenotypic differences between lines become smaller and smaller, which further exacerbates the difficulty of line differentiation. When creating new lines using Chinese materials, foreign germplasm resources are also introduced, which also increases the requirements for population classification of lines.
[0004] With the rapid development of biological sequencing technology, various molecular markers such as RFLP, AFLP, and SSR have been used for germplasm resource screening, variety identification, and variety protection. However, the above molecular markers have the following problems: 1. According to the current standards, the number of the above markers used is small, and there are certain errors in the population classification results; 2. A large number of samples cannot be detected in a short time, and the detection throughput is low; 3. The more the above markers are used, the higher the detection cost. Summary of the Invention
[0005] In view of this, the present invention provides an SNP molecular marker combination for constructing the population classification of eggplant lines and its application. The marker combination is evenly distributed on 12 chromosomes of eggplant and has relatively high polymorphism. Based on the characteristics of high marker density, strong automation, and high detection throughput of the liquid-phase probe chip based on the SNP marker, it is more conducive to establishing a population classification system for eggplant.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] One object of the present invention is to provide an SNP molecular marker combination for constructing the population classification of eggplant lines. The SNP molecular marker combination contains 1009 SNP molecular markers. The physical positions of the 1009 SNP sites are determined based on the comparison of the eggplant GUIQIE-1 (S. melongena) genome sequence. For the specific SNP molecular marker situation, please refer to Table 1 in Example 1.
[0008] Another object of the present invention is to provide a probe for identifying the SNP molecular marker combination described in Claim 1. For the specific probe, please refer to Table 1 in Example 1.
[0009] A third object of the present invention is to provide a method for determining an eggplant variety population, comprising the following steps: determining the genotypes of 1,009 SNP loci of the genomic DNA of the eggplant sample to be tested, constructing a PCA scatter plot, and determining the eggplant sample population to be tested according to the scatter plot;
[0010] Wherein, the genotypes of the 1,009 SNP loci are the SNP molecular marker combinations described in claim 1;
[0011] Each locus in the scatter plot represents a sample. The farther the distance between two samples in the figure, the greater the difference in the genetic backgrounds of the two samples, and individuals with similar genetic backgrounds will cluster into one category in the figure.
[0012] Further, the genotypes of the 1,009 SNP loci of the genomic DNA of the eggplant sample to be tested are specifically detected as follows: extracting the DNA of the eggplant sample to be tested, performing PCR amplification on the genomic DNA of the eggplant sample to be tested using the probe described in claim 2 to obtain a PCR product, detecting the PCR product, and determining the genotypes of the 1,009 SNP loci of the genomic DNA of the eggplant sample to be tested.
[0013] A fourth object of the present invention is the application of the above SNP molecular marker combination and / or the above probe in constructing the classification of eggplant lines.
[0014] A fifth object of the present invention is the application of the above SNP molecular marker combination and / or the above probe in any one of eggplant new line identification, eggplant breeding, and eggplant population genetic diversity analysis.
[0015] A sixth object of the present invention is to provide a kit for classifying an eggplant line population, and the kit contains the above probe.
[0016] A seventh object of the present invention is to provide an eggplant whole-genome gene chip, and the gene chip contains the above probe.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0018] The present invention provides a group of SNP marker combinations that can accurately distinguish eggplant lines. Based on the whole-genome resequencing data of 577 eggplant lines collected, 1,009 polymorphic SNP marker combinations are screened out from millions of SNPs, and a SNP marker combination for classifying eggplant lines is successfully constructed. This marker combination is evenly distributed on 12 chromosomes of eggplant and has high polymorphism. Based on the characteristics of high marker density, strong automation, and high detection throughput of the liquid-phase probe chip based on this SNP marker, it is more conducive to establishing an eggplant population classification system. This marker combination can provide theoretical and data support for work such as eggplant new line identification, molecular marker-assisted selection breeding, whole-genome selection, and population genetic diversity analysis, and has relatively important theoretical value and application significance. Description of the Drawings
[0019] Figure 1 It is the SNP marker density distribution map corresponding to 1k.
[0020] Figure 2 It is the PCA result map of the whole-genome SNP markers.
[0021] Figure 3 It is the phylogenetic tree result map of the whole-genome SNP markers.
[0022] Figure 4 It is the PCA result map of the SNP markers corresponding to 1k.
[0023] Figure 5 It is the phylogenetic tree result map of the SNP markers corresponding to 1k.
[0024] Figure 6 It is the PCA clustering result map of 12 samples to be tested in Example 3. Detailed Implementation Manner
[0025] The present invention will be further described in detail below in conjunction with specific embodiments, so that those skilled in the art can understand the present invention more clearly.
[0026] Sources and physicochemical parameters of key test materials:
[0027] Specific information of 577 eggplant samples is as follows:
[0028]
[0029]
[0030]
[0031]
[0032]
[0033]
[0034]
[0035] Example 1
[0036] This example provides a method for screening SNP marker combinations and designing liquid-phase probes based on the whole-genome resequencing data of 577 eggplant resources, specifically including the following processes:
[0037] (1) The genomic DNA of leaf tissues from 577 eggplant lines was extracted using the CTAB method. After the DNA passed the quality inspection, a whole-genome resequencing library was constructed using the BGI MGIEasy Universal DNA Library Preparation Reagent Kit. Fragments in the range of 260 - 360 bp were selected, enriched and amplified. After qubit quantification was stable and 2100 fragment analysis was performed, and it was determined that the library passed the inspection, then BGI DNBSEQ-T7 was used for sequencing.
[0038] (2) The clean data obtained from the sequencing in step (1) was aligned to the eggplant GUIQIE-1 (S. melongena) genome as the reference genome using bwa, and GATK variant detection was performed to obtain 42,686,847 high-quality SNP sites. The parameters were "-filter QD<2.0 --filter-name QD2 -filter QUAL<30.0 --filter-name QUAL30 -filter SOR>3.0 --filter-name SOR3 -filter FS>60.0 --filter-name FS60 -filter MQ<40.0 --filter-name MQ40 -filter MQRankSum<-12.5 --filter-name MQRankSum-12.5 -filter ReadPosRankSum<-8.0 --filter-name ReadPosRankSum-8".
[0039] (3) The SNP sites screened in step (2) were further filtered using the vcftools software according to the conditions "-minDP 3 -min-alleles 2 -max-alleles 2 -max-missing 1 -maf 0-.05" to obtain 7,254,731 SNP sites that were all detected in all samples with high quality, and population structure analysis was performed. The results are shown in Figure 1-3 .
[0040] (4) The SNP sites screened in step (3) were used to screen SNP marker combinations using a perl program script, and finally 1009 SNP marker combinations were obtained. The construction of the population structure SNP molecular markers was successful, and the marker information is shown in Table 1 for details.
[0041] Screening principle: chromosomes are evenly distributed, and the distance between two adjacent markers is greater than 500 kb; use the bedtools software to extract SNP loci and the 100-bp sequences upstream and downstream of them, and retain SNP loci with a GC content between 40-60% and no N bases in the sequence; use the blastn software to align the extracted sequences to the eggplant GUIQIE-1 (S. melongena) genome, and only retain single-copy SNP loci.
[0042] (5) Perform population structure analysis on the 1009 SNPs screened in step (4), and the results are shown in Figure 4 、 5 。 The analysis results Figure 1-5 show that: the clustering results are basically consistent with the whole-genome SNP results, and the 1009 SNP marker combinations screened can be used for eggplant population structure analysis.
[0043] (6) Design liquid-phase probe sequences for the 1009 SNPs screened in step (5) for subsequent capture sequencing. The specific information of the probes is shown in Table 1.
[0044] Table 1 Details of the liquid-phase probe sequences designed according to the position information of the 1009 screened SNPs
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137] Example 2 This example provides a method for determining an eggplant strain population, which is as follows:
[0138] (1) Genomic DNA extraction: Collect the leaves of the eggplant sample and extract the total DNA of the sample using the CTAB method.
[0139] (2) Fragmentation and end repair of genomic DNA
[0140] 2.1 Prepare the reaction system as shown in the following table in a sterile PCR tube placed on ice:
[0141] Reagent Volume Input DNA X μL Buffer 4.5 μL Enzymes 5 μL 1×TE Buffer Up to 30 μL Total Volume 30 μL
[0142] Wherein, X represents any volume not exceeding 20 μL, which, when added to 4.5 μL of Buffer and 5 μL of Enzymes, makes up the shortfall to 30 μL with 1×TE Buffer; use a pipette to blow up and down or oscillate to ensure that the system is thoroughly mixed, and then centrifuge briefly.
[0143] 2.3 Immediately place the PCR tube on a PCR instrument preheated to 32 °C, run the reaction program shown in the following table, and set the thermal cover of the PCR instrument to 75 °C; after the temperature drops to 4 °C, immediately place it on ice for the next experiment without holding at 4 °C.
[0144] Temperature Time 32℃ 20 min 65℃ 30 min 4℃ Cool to 4°C and proceed to the next step immediately
[0145] (3) Ligation of adapters
[0146] 3.1 Prepare the reaction system as shown in the following table.
[0147] Reagent Volume End Prep Reaction Mix (product of Step 2) 30 μL Ligation Buffer 15 μL <![CDATA[ddH 2 O]]> 2.5 μL Ligation Enzymes 5 μL Truncated Adaptor 2.5 μL Total Volume 55 μL
[0148] Wherein, all operations are carried out on ice. Place the PCR tube on the PCR instrument, and the reaction program is 22 °C for 15 min, without setting the thermal cover on the PCR instrument, and then store at 4 °C (immediately proceed to the next step after dropping to 4 °C).
[0149] 3.2 DNA sample purification: Add 44 μL of DNA Clean Beads to each sample, mix well; incubate at room temperature for 5 min; place the PCR tube on the magnetic rack and let it stand for 3 min. After the solution becomes clear, remove the supernatant; add 180 μL of 80% ethanol to rinse the magnetic beads, incubate for 30 s and then remove the supernatant, and repeat this operation once; keep the PCR tube on the magnetic rack, use a 10 μL pipette to remove the residual ethanol at the bottom of the tube, and dry until there is no ethanol residue; in 21 μL of ddH 2Resuspend the magnetic beads in water, and let them stand at room temperature for 1 min to fully release the DNA on the magnetic beads. Let the PCR tube stand on the magnetic stand for 2 min, and transfer 20 μL of the supernatant to a new PCR tube for library amplification.
[0150] (4) Library amplification and purification
[0151] 4.1 Prepare the PCR reaction system as shown in the following table, mix well with a pipette and centrifuge briefly.
[0152] Reagent Volume Adapter-ligated DNA 20 μL 2×PCR Mix 12.5 μL PCR Index Primer 2.5 μL Universal PCR Primer 2.5 μL Total Volume 35 μL
[0153] 4.2 Amplify the prepared reaction system under the following PCR amplification program;
[0154]
[0155] Add 35 μL of DNA Clean Beads to the amplification product, mix well, and incubate at room temperature for 5 min. Let the sample stand on the magnetic stand for 2 min. After the solution becomes clear, remove the supernatant. Add 200 μL of 80% ethanol to wash the magnetic beads, incubate for 30 s and then remove the supernatant, and repeat this step once. Keep the PCR tube on the magnetic stand, use a 10 μL pipette to remove the residual ethanol at the bottom of the tube, open the tube cap and dry until there is no ethanol residue. Resuspend the magnetic beads in 31 μL of ddH 2 O, let them stand at room temperature for 1 min to fully release the DNA on the magnetic beads. Let the sample stand on the magnetic stand for 2 min, transfer 30 μL of the supernatant to a new PCR tube, and store the library at 20 °C for subsequent library quality detection and sequencing.
[0156] (5) Hybridization of library and probe
[0157] 5.1 Take 750 ng of the library constructed in step (4) and add it to a PCR tube, and make a mark. Add purified magnetic beads to the library, gently pipette and mix well. Incubate at room temperature for 5 min, place the PCR tube on the magnetic stand for 3 min to make the solution clear. Remove the supernatant, keep the PCR tube on the magnetic stand, add 180 μL of 80% ethanol, and let it stand for 30 s. Remove the supernatant, then add 180 μL of 80% ethanol to the PCR tube again, let it stand for 30 s and then completely remove the supernatant. Let it stand at room temperature for 5 min to completely volatilize the residual ethanol;
[0158] 5.2 Add the hybridization reaction system to the PCR tube as shown in the following table;
[0159] Reagent Volume Hyb Buffer 13 μL Block 1 5 μL Block 2 2 μL Block 3 5 μL Nuclease-Free Water 3 μL Total Volume 28 μL
[0160] Gently pipette to mix evenly, let stand at room temperature for 3 min, and perform a brief centrifugation. Place the PCR tube on the magnetic stand and let stand for 3 min; aspirate 28 μL of the supernatant into a new PCR tube, add 2 μL of Target Probe, gently pipette to mix evenly, and perform a brief centrifugation; set the PCR instrument parameters as follows: hot lid temperature: 85 °C; 80 °C for 5 min; hold at 50 °C; place the PCR tube on the PCR instrument and run the above program, incubate overnight.
[0161] (6) Capture the target region DNA library
[0162] 6.1 Magnetic bead pretreatment: Take out the capture magnetic beads from 4 °C, vortex to resuspend, and place at room temperature for 30 min to equilibrate; take 50 μL of magnetic beads and add them into a new PCR tube, place on the magnetic stand for 1 min until the solution is clear, and remove the supernatant; remove the PCR tube from the magnetic stand, add 180 μL of Binding Buffer, gently pipette several times to mix evenly, and resuspend the magnetic beads; place on the magnetic stand for 1 min, and remove the supernatant; repeat this step once; remove the PCR tube from the magnetic stand, add 180 μL of Binding Buffer, gently pipette to resuspend and mix evenly the magnetic beads for use.
[0163] 6.2 Capture the target region DNA library: Keep the hybridization product on the PCR instrument, add the 180 μL of resuspended capture magnetic beads in step 1 to the hybridization product, pipette to mix evenly, and place on a rotary mixer to bind at room temperature for 30 min; place the PCR tube on the magnetic stand for 2 min to make the solution clear, and remove the supernatant; add 150 μL of Wash Buffer preheated to 50 °C, gently pipette to mix evenly, then perform a brief centrifugation, and incubate at 50 °C on a thermostatic shaking mixer for 10 min; perform a brief centrifugation, place the PCR tube on the magnetic stand for 2 min to make the solution clear, and remove the supernatant; repeat this step 2 times, for a total of 3 times of washing the magnetic beads; keep the sample on the magnetic stand, add 150 μL of 80% ethanol to the PCR tube, let stand for 30 s, and then completely remove the ethanol solution, and air dry at room temperature; add 24 μL of Nuclease-free Water to the PCR tube, remove the PCR tube from the magnetic stand, and gently pipette to resuspend and mix evenly the magnetic beads for use.
[0164] (7) Post-capture PCR amplification
[0165] 7.1 Take out the PostPCR MasterMix and PostPCRPrimer from the -20 °C refrigerator, place on an ice box to melt, and after melting, mix evenly and place on ice or at 4 °C for use.
[0166] 7.2 After capture, the DNA library needs to be PCR amplified. Prepare the reaction system according to the following table:
[0167] Reagent Volume DNA library capturing the target region in Step (6) 24 μL Post-PCR Master Mix 25 μL Post-PCR Primer (selected according to library type) 1 μL Total Volume 50 μL
[0168] Adjust the pipette to 40 μL, gently pipette and mix 6 times, and then immediately place it on the PCR instrument; Run the PCR instrument program: Hot lid temperature: 105 °C Program: 95 °C for 1 min; 98 °C for 20 s; 60 °C for 30 s for N cycles; 72 °C for 30 s; 72 °C for 5 min; Hold at 4 °C; After PCR, add 55 μL of purified magnetic beads to the sample, vortex or pipette to mix evenly, and let it stand at room temperature for 5 min; Centrifuge briefly, and place the PCR tube on the magnetic rack for 3 min until the solution is clear; Keep the PCR tube on the magnetic rack, remove the supernatant, add 180 μL of 80% ethanol solution to the PCR tube, and let it stand for 30 s; Keep the PCR tube on the magnetic rack, remove the supernatant, add 180 μL of 80% ethanol solution to the PCR tube again, and completely remove the supernatant after standing for 30 s; Let it stand at room temperature for 5 min to completely volatilize the residual ethanol; Add 25 μL of Nuclease-free water, remove the PCR tube from the magnetic rack, vortex or pipette to mix evenly 10 times, and let it stand at room temperature for 2 min; Centrifuge briefly, and place the PCR tube on the magnetic rack for 2 min until the solution is clear; Use a pipette to aspirate 23 μL of the supernatant and transfer it to a 1.5 mL centrifuge tube, and label the sample information; Take 1 μL of the library and perform quantification using the Qubit dsDNA HS Assay Kit, record the library concentration, and the library concentration is about 1 - 20 ng / μL; Take 1 μL of the sample and perform fragment length determination using the Agilent 2100 Bioanalyzer system (Agilent DNA 1000 Kit).
[0169] (8) High-throughput sequencing: Use the MGI DNBSEQ-T7 sequencing platform to perform high-throughput sequencing on the library obtained by capture and amplification in step (7) to obtain the sequencing results of the tree species genomic DNA, and perform basic cleaning processing on the obtained data.
[0170] (9) Liquid-phase probe capture rate assessment: Use the bwa software to align the cleaned sequencing data to the eggplant GUIQIE-1 (S. melongena) reference sequence, and use the GATK software to obtain the SNP genotyping data of the eggplant line; Use the SNP genotyping data to complete the fingerprint map analysis of the eggplant line.
[0171] Example 3
[0172] To verify the SNP molecular marker combination in this application, the liquid-phase probe capture efficiency of 12 eggplant lines is now evaluated, and the specific steps are as follows:
[0173] (1) Experimental materials
[0174] Twelve eggplant leaves were selected, and the total DNA of each sample was extracted by the CTAB method. The sample numbers are shown in the following table.
[0175] Number Eggplant variety Grouping Number Eggplant variety Number Test 1 E431 EUR Test 2 BS053 EUR Test 3 BS065 N. China Test 4 180CD1 EUR Test 5 E54 N. China Test 6 WDQ SEA Test 7 BS061 EUR Test 8 E450 SEA Test 9 81 SEA Test 10 E398 S. China Test 11 E454 S. China Test 12 E368 S. China
[0176] Note: EUR, Europe; SEA, Southeast Asia; N. China, North China; S. China, South China.
[0177] (2) Population genetic analysis
[0178] The clustering results of the 12 test samples are shown in Figure 6 , and it can be seen from Figure 6 that: the actual clustering results of the 12 test samples are consistent with the theoretical clustering results, indicating that the 1009 SNP loci screened can accurately divide the eggplant population.
[0179] In the present invention, the specific raw materials not described are all existing substances and can be directly purchased from the market.
[0180] The above is only a preferred implementation of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A probe combination for identifying a combination of SNP molecular markers for constructing eggplant strain clusters, characterized in that: The SNP molecular marker combination includes 1009 SNP molecular markers, and the physical positions of the 1009 SNP sites are determined based on the comparison of the eggplant GUIQIE-1 genome sequence. The specific SNP molecular marker situation and probe combination are shown in Table 1 of the specification.
2. A method for determining eggplant strain group division, characterized in that: The following steps are involved: Determine the genotypes of 1009 SNP loci of the genomic DNA of the eggplant samples to be tested, construct a PCA scatter plot, and determine the eggplant sample population to be tested based on the scatter plot; Wherein, the 1009 SNP locus genotypes are the SNP molecular marker combination described in claim 1; Each site in the scatter plot represents a sample. The farther the distance between two samples in the plot, the greater the difference in genetic background between the two samples. Individuals with similar genetic backgrounds will be clustered into one category in the plot.
3. Use of the probe combination according to claim 1 in constructing eggplant strain clustering.
4. Use of the probe combination of claim 1 in any one of identification of new eggplant strains, eggplant breeding, and analysis of genetic diversity of eggplant populations.
5. A kit for classifying eggplant strain populations, characterized in that: The kit comprises the probe combination according to claim 1.
6. An eggplant whole genome gene chip, characterized in that: The gene chip contains the probe combination as claimed in claim 1.
Citation Information
Patent Citations
KASP molecular marker related to eggplant fruit color and application of KASP molecular marker
CN116640872A
SNP molecular marker closely linked with eggplant fruit length QTL and application
CN117187431A