Molecular markers for corn variety identification and their development methods and applications
By designing 40 pairs of marker primers and using multiplex PCR amplification technology, the problems of complex operation, high cost, and low throughput in maize variety identification were solved, achieving efficient and low-cost maize variety identification and purity detection, thus improving breeding efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUBEI KANGNONG SEED IND LIMITED
- Filing Date
- 2025-03-07
- Publication Date
- 2026-05-29
AI Technical Summary
Existing maize variety identification technologies suffer from problems such as cumbersome operation procedures, long detection cycles, low detection throughput, high costs, excessive sensitivity, and easy formation of dimers, making it difficult to meet the needs of large-scale testing.
Forty pairs of labeled primers were designed. Through primer optimization and sequence retrieval, combined with multiplex PCR amplification, sequencing library construction and bioinformatics analysis, the simultaneous detection of 40 SSR sites was achieved, avoiding dimer formation. High-throughput sequencing technology was used for maize variety identification.
It enables low-cost, high-throughput, simple-to-operate, and highly sensitive maize variety identification, meeting the practical application needs of authenticity identification, purity testing, and germplasm resource identification, thereby improving breeding efficiency and shortening the breeding cycle.
Smart Images

Figure CN120099207B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biotechnology, and in particular to a set of molecular markers for maize variety identification, their development methods, and applications. Background Technology
[0002] In my country, maize is the largest crop, making accurate detection of its variety authenticity crucial. For a long time, maize variety authenticity detection has primarily relied on the SSR molecular marker method, following the standard outlined in the "Technical Specification for Maize Variety Identification: SSR Marker Method" (NY / T 1432-2014). Looking back over the past 40 years, the detection method has continuously evolved, initially using polyacrylamide gel electrophoresis, then capillary electrophoresis, and finally fluorescently labeled electrophoresis. However, these traditional methods have many shortcomings: cumbersome and complex procedures, lengthy detection cycles, and low throughput, making it difficult to meet the ever-increasing detection demands.
[0003] With technological innovation, the KASP (Competitive Amplification) detection method has emerged. Its advent has significantly increased detection throughput and made operation simpler and easier, bringing new hope to maize variety detection. However, KASP detection is highly dependent on imported reagents, resulting in high detection costs and hindering its large-scale application.
[0004] Meanwhile, next-generation sequencing (NGS) technology has shown unique advantages in the field of SNP detection, with low cost, simple operation, and extremely high throughput. However, when applied to SSR detection, it has revealed the problem of excessive sensitivity, which often requires additional correction of the test results, increasing the complexity and uncertainty of result interpretation.
[0005] Furthermore, multiplex PCR is widely used in the capture amplification process, and its unique advantages have greatly reduced the complexity of the operation, making it a highly sought-after technology. However, multiplex PCR has stringent requirements for primers; if mixed primers produce dimers, it will severely interfere with the amplification effect and affect the entire detection process. Summary of the Invention
[0006] The technical problem to be solved by this invention is to provide a set of molecular markers for maize variety identification, as well as their development methods and applications. Through in-depth primer optimization and comprehensive sequence retrieval, this invention cleverly avoids the problem of dimer formation and successfully develops and designs 40 pairs of marker primers. Based on these primers, and relying on multiplex PCR amplification, sequencing library construction, sequencing, and bioinformatics analysis, this invention can meet the practical application needs of maize variety authenticity identification, purity detection, and germplasm resource identification. At the same time, this method has the characteristics of low cost, simple operation, high throughput, and high sensitivity.
[0007] The present invention solves the above-mentioned technical problems by adopting the following technical solutions:
[0008] A set of molecular markers for maize variety identification includes 40 pairs of marker primers designed based on 40 SSR loci of maize national standard, as shown in Table 1.
[0009] Table 1. Information on marker primers
[0010]
[0011]
[0012] As one of the preferred embodiments of the present invention, the distribution of the molecular markers across the entire genome is completely consistent with the distribution of the 40 SSR loci in the national standard for maize, and on average, each chromosome corresponds to 4 pairs of markers. Specific marker location information is shown in Table 2. The amplification product size ranges from 129 to 417 bp, with 80% of the amplification products ranging from 190 to 384 bp, making them suitable for the PE-150 sequencing platform and corresponding analysis.
[0013] Table 2 Marking Location Information
[0014]
[0015]
[0016] A method for developing molecular markers for identifying the above-mentioned maize varieties includes the following steps:
[0017] (1) Extract sequence information from 40 SSR loci of maize according to national standards to form a target locus dataset; use population variation data from the International Maize Genome Database to form a variation dataset; use the high-throughput sequencing data variation data analysis software GATK to perform variation analysis and screening, with the screening condition "the frequency of smaller alleles is not less than 5%", to obtain candidate sequencing loci and variation information.
[0018] (2) Primers for amplification of each target SNP site were designed using Primer 3 software. The parameters set included: ① primer sequence length between 17 and 32 bp; ② Tm value between 60 and 64 ℃; ③ product size not less than 150 bp and not more than 500 bp; ④ sequencing reads covering the target site.
[0019] (3) For each target site, three pairs of primers were designed, and then the amplification specificity of each pair of primers was detected using ePCR software. The possibility of dimer formation between primers was detected using a script, and finally the amplification primers for 40 SSR sites were obtained, namely the 40 pairs of labeled primers shown in Table 1.
[0020] An application of the molecular marker for identifying maize varieties is characterized by using the 40 pairs of marker primers to perform mixed PCR amplification of maize genomic DNA, and sequencing and typing of the amplification products to obtain polymorphism data of maize at the 40 SSR loci of the national standard, thus constituting maize DNA fingerprint data; based on the maize DNA fingerprint data, subsequent analysis of maize varieties can be carried out.
[0021] As one of the preferred embodiments of the present invention, the corn genomic DNA is extracted from the young tissues of corn, more preferably from dry corn seeds (endosperm), fresh leaves, young ears, seedlings or young stem segments, old leaves, etc.
[0022] As one of the preferred embodiments of the present invention, the PCR amplification reaction system is a 20 μL system or a 10 μL system; wherein, the 20 μL reaction system includes: 2 μL maize genomic DNA (50-100 ng), 5 μL 10 mM dNTPs, 1 μL 10 mM forward primer, 1 μL 10 mM reverse primer, 2 μL 10x amplification buffer, 0.5 μL Taq DNA polymerase, 3 μL 10 mM MgCl2, and finally ddH2O water is added to a final volume of 20 μL; the 10 μL reaction system is reduced accordingly.
[0023] As one of the preferred embodiments of the present invention, the PCR amplification reaction conditions are as follows: 95°C pre-denaturation for 3 min; 95°C denaturation for 30 s, 55°C annealing for 60 s, 72°C extension for 20 s, for a total of 35 cycles; 72°C extension for 5 min, and then maintained at room temperature.
[0024] Based on the above method, 100 to 20,000 maize DNA samples can be detected simultaneously.
[0025] As one of the preferred embodiments of the present invention, the subsequent analysis of the maize variety refers to one or more of the following: maize variety authenticity identification, seed purity identification, molecular design breeding, backcross breeding and background selection.
[0026] As one of the preferred embodiments of the present invention, the method for identifying the authenticity of maize varieties is as follows: extracting genomic DNA from the maize variety sample to be tested and the control sample respectively; using the above-mentioned 40 pairs of marker primers to amplify and sequence the genomic DNA of the maize sample to be tested and the control sample, thereby obtaining the genotype data of the maize sample to be tested and the control sample at 40 sites for variety authenticity identification.
[0027] As a preferred embodiment of the present invention, the method for identifying the purity of maize seeds is as follows: based on the genotypic data of the parents of the hybrid seeds of the variety to be tested at 40 loci, primers for SNP sites that exhibit polymorphism in the parents are analyzed and selected; the selected polymorphic primers are used to perform PCR amplification and sequencing analysis on the F1 hybrid seeds to analyze the relationship between the F1 seeds and the parents, thereby identifying the purity of the maize variety seeds. In this invention, the number of polymorphic markers selected is preferably no less than 2 to 3.
[0028] As one of the preferred embodiments of the present invention, the method for maize molecular design breeding, backcross breeding, and background selection is as follows: Genotype data of 40 loci of the maize variety to be bred are obtained according to the above method to obtain the background genotype; a target gene donor parent is selected according to the breeding objective and cross-crossed with the variety to be bred to establish a hybrid segregating population or a backcross segregating population; in the segregating population, prospective screening is performed based on the target trait; in the selected segregating population, genotype data of 40 loci of individuals in the population are obtained according to the above method; the individual genotype is compared with the obtained background genotype, and the individual with the fewest variable loci is selected for further breeding until a stable individual containing the target trait gene and with a completely restored background genotype is obtained.
[0029] The advantages of this invention compared to the prior art are:
[0030] (1) This invention effectively avoids dimer formation through primer optimization, sequence retrieval and other methods, and successfully developed and designed 40 pairs of marker primers covering the SSR sites in the "Technical Specification for Identification of Maize Varieties - SSR Marker Method". Using these marker primers, and through multiplex PCR amplification, sequencing library construction, sequencing and bioinformatics analysis, it can meet the practical application needs of maize variety authenticity identification, purity detection, germplasm resource identification and molecular breeding background selection.
[0031] (2) The molecular marker for maize variety identification provided by this invention is based on high-throughput sequencing technology for marker detection, and achieves simultaneous detection of 40 SSR sites through a single amplification reaction. The larger the sample size, the lower the cost. At the same time, it has the characteristics of simple operation, high throughput and high sensitivity, and has extremely high promotion and application value.
[0032] (3) This invention can realize the process and informationization of molecular design breeding of maize, improve the efficiency of maize breeding, shorten the breeding cycle, and promote agricultural development. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the distribution of the 40 marker sites of jade on the genome in Example 2;
[0034] Figure 2This is the map information of some maize resources in Example 3 (in the map, the same color for the same marker indicates the same genotype, and different colors indicate different genotypes);
[0035] Figure 3 This is a cluster analysis diagram of the corn sample in Example 3. Detailed Implementation
[0036] The embodiments of the present invention are described in detail below. These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments. Furthermore, unless otherwise specified, the reagents, methods, and equipment used in the present invention are conventional reagents, methods, and equipment in this technical field.
[0037] Example 1
[0038] This embodiment uses a set of molecular markers for maize variety identification, including 40 pairs of marker primers designed based on 40 SSR loci of maize according to national standards, as shown in Table 1.
[0039] The distribution of the molecular markers across the entire genome is completely consistent with the distribution of the 40 SSR loci in the national standard for maize, with an average of 4 pairs of markers per chromosome. Specific marker location information is shown in Table 2. Furthermore, Table 2 shows that the amplification product size ranges from 129 to 417 bp, with 80% of the amplification products falling between 190 and 384 bp, making them suitable for the PE-150 sequencing platform and corresponding analysis.
[0040] Table 1. Information on marker primers
[0041]
[0042]
[0043] Table 2 Marking Location Information
[0044] Primer name chromosome physical location amplification product size GB.01 Chr.1 43914207_43914558 352bp GB.02 Chr.1 199519487_199519723 237bp GB.03 Chr.2 74090364_74090640 277bp GB.04 Chr.2 225870759_225871108 350bp GB.05 Chr.3 751790_752081 292bp GB.06 Chr.3 232377659_232377873 215bp GB.07 Chr.4 1284154_1284563 410bp GB.O8 Chr.4 172522807_172523190 384bp GB.09 Chr.5 28892818_28893138 321bp GB10 Chr.5 213923412_213923673 262bp GB.11 Chr.6 2126021_2126205 185bp GB.12 Chr.6 151092363_151092627 265bp GB.13 Chr.7 1464003_1464226 224bp GB.14 Chr.7 173953563_173953729 167bp GB.15 Chr.8 165268302_165268538 237bp GB16 Chr.8 178412077_178412292 216bp GB17 Chr.9 68851643_68852059 417bp GB18 Chr.9 121460654_121460928 275bp GB.19 Chr.10 5439596_5439816 221bp GB.20 Chr.10 134476747_134476936 190bp GB.21 Chr.1 228756898_228757063 166bp GB.22 Chr.1 279928282_279928493 212bp GB23 Chr.2 2812983_2813242 260bp GB24 Chr.2 210231389_210231619 231bp GB25 Chr.2 230692420_230692583 164bp GB26 Chr.3 205192852_205193085 234bp GB27 Chr.4 32890631_32890902 272bp GB28 Chr.4 224943201_224943390 190bp GB29 Chr.5 12794900_12795178 279bp GB.30 Chr.5 67030592_67030720 129bp GB.31 Chr.6 44432358_44432620 263bp GB.32 Chr.6 166029502_166029727 226bp GB.33 Chr.7 10457665_10457882 218bp GB.34 Chr.7 160633569_160633738 170bp GB35 Chr.8 13780331_13780522 192bp GB.36 Chr.8 179043680_179043881 202bp GB.37 Chr.9 4092504_4092703 200bp GB38 Chr.9 135613487_135613762 276bp GB.39 Chr.10 2229791_2230100 310bp GB.40 Chr.10 115709777_115710111 335bp
[0045] Example 2
[0046] The method for developing the above-mentioned molecular markers in this embodiment is as follows:
[0047] I. Label Transformation and Reference Dataset Acquisition
[0048] Based on the location and variation information of 40 pairs of SSR primers in the national standard for maize (existing technology), the genome of the published maize variety B73 was used as the "reference genome" (version: B73_v5.0, URL https: / / staging.maizegdb.org / genome / assembly / Zm-B73-REFERENCE-NAM-5.0), and the sequence information of 40 sites was extracted to form the target site dataset.
[0049] A variation dataset was constructed using population variation data from the International Maize Genome Database (https: / / staging.maizegdb.org / ). Variation analysis and screening were performed using the high-throughput sequencing data variation data analysis software GATK (version 3.7) (screening criteria: minor allele frequency not less than 5%) to obtain candidate sequencing sites and variation information.
[0050] II. Development of Specific Primers:
[0051] Primer 3 software (version 2.5.0) was used to design amplification primers for each target SNP site. The parameters set included: ① primer sequence length between 17 and 32 bp; ② Tm value between 60 and 64℃; ③ product size not less than 150 bp and not more than 500 bp; ④ sequencing reads must be able to cover the target site.
[0052] For each target site, three primer pairs were designed, and the amplification specificity of each primer pair was then detected using e-PCR software (version 2.3.12). The possibility of primer dimer formation was detected using a script, ultimately yielding amplification primers for 40 sites. These 40 sites constitute the set of markers described in Example 1. The distribution of these 40 marker sites on the genome is as follows: Figure 1 As shown.
[0053] Example 3
[0054] This embodiment describes the identification of a maize variety / resource.
[0055] Maize resources from different sources were selected as samples for variety / resource identification to verify the feasibility of the marker primers of this invention. The specific operation is as follows:
[0056] I. Extraction of Genomic DNA from Maize Varieties / Resources
[0057] DNA was extracted from maize tissue using the CTAB method.
[0058] 1. Genomic DNA was extracted from young corn tissue.
[0059] 2. Reagent preparation:
[0060] CTAB buffer: 2% CTAB, 1.4M NaCl, 100mM Tris-HCl, 10mM EDTA, pH 8.0;
[0061] Washing buffer: 75% ethanol;
[0062] TE buffer: 20 mM Tris-HCl, 1 mM EDTA, pH 8.0;
[0063] Pre-cooling anhydrous ethanol: Store anhydrous ethanol at -20℃ for 2 hours.
[0064] 3. Extracting DNA from corn tissue:
[0065] Young corn tissue was minced and ground in liquid nitrogen. A volume equivalent to the tissue volume of preheated CTAB was added, and the mixture was immediately placed in a 65°C water bath for 30 min to 1 h, shaking every 5 min. After centrifugation at 4°C and 12000 rpm for 10 min, the supernatant was removed and an equal volume of chloroform and isoamyl alcohol (chloroform to isoamyl alcohol volume ratio 24:1) was added and mixed well. After centrifugation at 4°C and 12000 rpm for 15 min, the supernatant was removed and two volumes of ice-cold anhydrous ethanol were added. The mixture was then placed at -20°C for 1 h. After centrifugation at 4°C and 12000 rpm for 10 min, the supernatant was discarded, and the precipitate was washed with 75% ethanol and air-dried. 5 μL to 100 μL of TE buffer was added and the precipitate was thoroughly dissolved. DNA quality was assessed by agarose gel electrophoresis, and concentration and purity were determined using a UV spectrophotometer. The DNA was stored at -20°C.
[0066] II. Library Construction and Sequencing
[0067] Mixed PCR amplification of maize genomic DNA was performed using the 40 pairs of labeled primers described in Table 1 of Example 1. The PCR amplification reaction system (20 μL) included: 2 μL maize genomic DNA (50–100 ng), 5 μL 10 mM dNTPs, 1 μL 10 mM forward primer, 1 μL 10 mM reverse primer, 2 μL 10x amplification buffer, 0.5 μL Taq DNA polymerase, 3 μL 10 mM MgCl2, and finally ddH2O water to a final volume of 20 μL. The reaction conditions were: 95°C pre-denaturation for 3 min; 95°C denaturation for 30 s, 55°C annealing for 60 s, 72°C extension for 20 s, for a total of 35 cycles; 72°C extension for 5 min, followed by maintenance at room temperature.
[0068] After PCR was completed, the recovered products were inspected and found to be of good quality. Paired-end (PE) sequencing libraries were constructed according to the Illumina library construction protocol. Equal amounts of DNA from each sample were used to construct a PE library, and PE150 sequencing was performed on an Illumina Hiseq sequencer.
[0069] III. Genotyping of Target Loci
[0070] The raw sequencing data underwent quality control to obtain high-quality clean data. Then, BWA software was used to break down the clean data into individual target loci, resulting in SAM format alignment results. The SAM files were then converted to BAM format using samtools software. Finally, the reads in the BAM files were sorted using SortSam in Picard, yielding the final BAM file. GATK was then used to determine the genotypes of each target locus, thus constructing the maize germplasm DNA fingerprint.
[0071] Test results:
[0072] Figure 2 This is a partial map of maize resources. Figure 3 This is a cluster analysis diagram of the corn samples.
[0073] Combination Figure 2 and Figure 3 It can be seen that the 39 samples tested in this embodiment are all maize resources from different sources. Among them, SCL23 and SCL35 are relatively closely related, with only 6 markers showing differences.
[0074] Example 4
[0075] This embodiment describes a method for identifying the authenticity of a maize variety.
[0076] Referring to the method in Example 3, genomic DNA was extracted from the maize variety sample to be tested and the control sample, respectively; and the genomic DNA of the maize variety sample to be tested and the control sample was amplified and sequenced using 40 pairs of marker primers (Example 1), so as to obtain the genotype data of the maize variety sample to be tested and the control sample at 40 sites for comparison and variety authenticity identification.
[0077] Example 5
[0078] This embodiment describes a method for identifying the purity of corn seeds.
[0079] Referring to the method in Example 3, and based on the genotype data of the parents of the hybrid seeds of the variety to be tested at 40 loci, primers for SNP sites exhibiting polymorphism in the parents were analyzed and selected. The selected polymorphic primers were used to perform PCR amplification and sequencing analysis on the F1 hybrid seeds to analyze the relationship between the F1 seeds and the parents, thereby identifying the purity of the maize variety seeds. In this example, the number of polymorphic markers selected was no less than 2 to 3.
[0080] Example 6
[0081] This embodiment describes a method for molecular design breeding, backcross breeding, and background selection of maize.
[0082] Following the method in Example 3, genotype data at 40 loci of the maize variety to be bred were obtained to obtain the background genotype. Based on the breeding objectives, a target gene donor parent was selected and crossbred with the variety to be bred to establish a hybrid segregating population or a backcross segregating population. In the segregating population, prospective selection was performed based on the target trait. In the selected segregating population, genotype data at 40 loci of individuals were obtained according to the above method. The individual genotypes were compared with the obtained background genotypes, and the individuals with the fewest variant loci were selected for further bred development until stable individuals containing the target trait gene and with fully restored background genotypes were obtained.
[0083] In summary, this invention cleverly avoids the challenge of dimer formation through innovative methods such as in-depth primer optimization and comprehensive sequence retrieval, and successfully developed 40 pairs of marker primers. Based on these primers, and relying on multiplex PCR amplification, sequencing library construction, sequencing, and bioinformatics analysis, it can meet the practical application needs of maize variety authenticity identification, purity detection, and germplasm resource identification. At the same time, this method has the characteristics of low cost, simple operation, high throughput, and high sensitivity.
[0084] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. The application of primers for molecular markers in maize variety identification in one or more of the following: maize variety authenticity identification, seed purity identification, backcross breeding, and background selection, characterized in that... Forty pairs of labeled primers were used to perform mixed PCR amplification of maize genomic DNA, and the amplification products were sequenced and genotyped. The 40 pairs of labeled primers are shown below: 。 2. The application according to claim 1, characterized in that, The PCR amplification reaction system is either a 20 μL system or a 10 μL system; wherein, the 20 μL reaction system includes: 2 μL maize genomic DNA, 5 μL 10 mM dNTPs, 1 μL 10 mM forward primer, 1 μL 10 mM reverse primer, 2 μL 10x amplification buffer, 0.5 μL Taq DNA polymerase, 3 μL 10 mM MgCl2, and finally add ddH2O water to a final volume of 20 μL; the 10 μL reaction system is reduced accordingly.
3. The application according to claim 1, characterized in that, The PCR amplification reaction conditions were as follows: 95℃ pre-denaturation for 3 min; 95℃ denaturation for 30 s, 55℃ annealing for 60 s, 72℃ extension for 20 s, for a total of 35 cycles; 72℃ extension for 5 min, and then maintained at room temperature.