A molecular marker and its use in breeding soybean varieties with high oil content and high protein content
By developing a molecular marker based on SNP site 9020004 and its KASP primer combination, the problem of simultaneously increasing the oil and protein content of soybeans in traditional breeding has been solved, realizing efficient and precise breeding of high-oil and high-protein soybean varieties, which is suitable for high-throughput screening and large-scale breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHEAST INST OF GEOGRAPHY & AGRIECOLOGY C A S
- Filing Date
- 2026-01-29
- Publication Date
- 2026-05-29
AI Technical Summary
Traditional breeding techniques struggle to simultaneously increase the oil and protein content of soybean seeds, and existing molecular markers suffer from insufficient linkage strength and poor environmental universality, resulting in long breeding cycles, high costs, and low efficiency.
We developed a molecular marker based on SNP site 9020004 and its KASP primer combination for high-throughput detection of soybean seedling genotypes, and designed specific primer pairs to achieve simultaneous enhancement of high oil and high protein content.
It achieves simultaneous improvement in soybean oil and protein content, shortens the breeding cycle by 6-8 months, reduces costs by more than 50%, improves breeding efficiency and screening accuracy, and is suitable for high-throughput screening and large-scale breeding.
Smart Images

Figure CN122104975A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of plant breeding technology, and in particular to a molecular marker and its application in breeding soybean varieties with high oil and high protein content. Background Technology
[0002] Soybeans are an important oilseed crop and source of plant protein. Their seed oil content (SOC) and protein content (SPC) are core agronomic traits that determine their economic value. Long-term breeding practices have revealed a significant negative correlation between soybean seed oil content and protein content. This antagonistic relationship between traits makes it extremely challenging to simultaneously increase both contents using traditional breeding methods, limiting breakthroughs in the overall quality of soybean varieties.
[0003] Traditional breeding identification mainly relies on phenotypic observation of mature seeds. This method is not only greatly affected by environmental fluctuations, but also requires screening only after the plants have fully matured. This results in a long breeding cycle, and when the hybrid offspring population is large, the costs of field management and phenotypic identification are extremely high, severely limiting the efficiency of the aggregation of superior traits.
[0004] In recent years, marker-assisted breeding (MAS) technology has provided technical support for accurate prediction and screening of early generations. Among them, competitive allele-specific PCR (KASP) technology has become an important means of identifying key functional sites due to its high-throughput characteristics in genotyping transformation and the advantages of closed-tube operation in the detection process. However, since oil content and protein content are complex quantitative traits controlled by multiple genes, existing molecular markers often suffer from insufficient linkage strength or poor environmental universality. Therefore, developing specific molecular markers that can accurately and synergistically control dual traits and are suitable for high-throughput detection platforms, and constructing corresponding detection primer systems, has important application value for breeding high-quality soybean varieties. Summary of the Invention
[0005] To address the technical problems existing in the prior art, this invention provides a molecular marker and its application in breeding soybean varieties with high oil and high protein content.
[0006] In a first aspect, the present invention provides a SNP site based on the genome version number Wm82.a2.v1, wherein the SNP site is located at position 9020004 on chromosome 8 of soybean and has a polymorphism of T / A.
[0007] Secondly, the present invention provides a molecular marker comprising a nucleic acid with a nucleotide sequence as shown in SEQ ID NO.1, wherein the 26th position exhibits a polymorphism of T / A.
[0008] The nucleotide sequence shown in SEQ ID NO.1: AGCTCTGTTTATTTTCCTTCCTTTC T TTCTTTCTTAGGACAGAAAAAACCAAAACCATAGAGCAAAATGAAGCCCTTCTGCAGATTCTTGATTCTTGTATCGCTATTTTCTCTGGTGGAATCTTCATGCAACAGCAGTGAAGAGCATGACTT.
[0009] Furthermore, regarding the aforementioned SNP sites and molecular markers, A represents high oil content and protein content, while T represents low oil content and protein content.
[0010] Thirdly, the present invention provides a primer pair for amplifying the aforementioned molecular marker.
[0011] The primer pair design method described in this invention can be a conventional method of this invention. Technicians can design primer pairs (including primer pairs or KASP primer combinations) of different lengths based on existing primer design rules and primer design software (such as Primer) to amplify the aforementioned molecular markers.
[0012] Fourthly, the present invention provides a KASP primer pair comprising nucleotide sequences as shown in SEQ ID NO.2-SEQ ID NO.4.
[0013] (1) F1 (SEQ ID NO.2, T allele primer): 5'-AGCTCTGTTTATTTTCCTTCCTTTCT-3' (2) F2 (SEQ ID NO.3, A allele primer): 5'-AGCTCTGTTTATTTTCCTTCCTTTCA-3'.
[0014] (3) R (SEQ ID NO.4, universal downstream primer): 5'-AAGTCATGCTCTTCACTGCTGTTGCAT-3'.
[0015] The primers of SEQ ID NO.2 and SEQ ID NO.3 above include a fluorescent tag sequence at the 5' or 3' end, which can be one or more of FAM, TET, HEX, ROX, Cy3, Cy5, Alexa Fluor, SYBR Green, DAPI, FITC or Texas Red.
[0016] For example, GAAGGTGACCAAGTTCATGCT (FAM fluorescent tag sequence), GAAGGTCGGAGTCAACGGATT (HEX fluorescent tag sequence).
[0017] Fifthly, the present invention provides a kit comprising the aforementioned molecular markers or the aforementioned KASP primer combination.
[0018] Furthermore, the kit also includes: KASP Master Mix, A allele homozygous positive control DNA, T allele homozygous positive control DNA, a negative control, and one or more of the ingredients specified in the instruction manual. This kit can be directly used for large-scale laboratory testing, is easy to operate, and provides reliable results.
[0019] Sixthly, the present invention provides the application of the aforementioned SNP sites, or the aforementioned molecular markers, as targets in any of the following: (1) Predict or detect the oil or protein content of soybeans; (2) Identify or breed soybean varieties with high oil or protein content; (3) Molecular marker-assisted breeding of soybeans; (4) Improvement of soybean varieties related to oil or protein content; (5) Improvement of soybean germplasm resources.
[0020] The targets described in this invention include existing conventional methods and reagents for detecting nucleotides, such as gene sequencing, primer design for amplification, and probe design for targeted detection.
[0021] In a seventh aspect, the present invention provides the use of the aforementioned primer pairs, or the aforementioned KASP primer combinations, or the aforementioned kit in any of the following: (1) To predict or detect the oil or protein content of soybeans, or to prepare reagents for predicting or detecting the oil or protein content of soybeans; (2) To identify or cultivate the oil or protein content of soybeans, or to prepare reagents for identifying or cultivating the oil or protein content of soybeans; (3) Molecular marker-assisted breeding of soybeans; (4) Improvement of soybean varieties related to oil or protein content; (5) Improvement of soybean germplasm resources.
[0022] Eighthly, the present invention provides a method for detecting the oil content or protein content of soybeans, comprising: The polymorphism of molecular markers was detected in the soybean sample to be tested, as described above, and the oil content or protein content of the soybean was determined based on the genotype detection results.
[0023] Furthermore, the detection method includes one or more of the following: gene sequencing, molecular probes, liquid phase capture, or mass spectrometry.
[0024] Furthermore, the determination of the oil content or protein content of the soybean to be tested based on the genotype detection results includes: peanuts with genotype detection results of AA have higher oil and protein content than peanuts with detection results of TT.
[0025] Preferably, in the aforementioned KASP primer combination, F1 is connected to FAM and F2 is connected to HEX. In this case, the primer with only FAM signal is TT, and the primer with only HEX signal is AA.
[0026] As a preferred embodiment, the present invention provides a method for detecting the oil content or protein content of soybeans, comprising: (1) Extract genomic DNA from the soybean sample to be tested; (2) PCR amplification was performed based on the aforementioned KASP primer combination; (3) Determine the genotype detection results based on the fluorescence signal.
[0027] Preferably, the OD260 / OD280 of the genomic DNA is between 1.8 and 2.0, and the concentration is 20 to 100 ng / μL.
[0028] Preferably, the PCR amplification procedure includes: Pre-denaturation at 92~98℃ for 8~20 min; Denaturation at 92~98℃ for 15~30 seconds + annealing at 61~55℃ for 30~60 seconds (decreasing by 0.6℃ per cycle, 10 cycles); Denaturation at 92~98℃ for 15~30 seconds + annealing at 52~58℃ for 30~60 seconds (25~40 cycles).
[0029] Preferably, the PCR amplification system comprises, in a total volume of 2 μL: Soybean sample DNA template, 4~6 ng / μl, 0.7~1.5 μL; 2x Master Mix for ASPCR V1 0.7~1.5μL; KASP Assay Mix, F1:F2:R = 1:1:3, 0.02~0.05μL balance is water.
[0030] Ninthly, the present invention provides a method for breeding soybeans with high oil or protein content, comprising: during the soybean breeding process, selecting soybeans with the aforementioned molecular marker genotype AA in the offspring.
[0031] The present invention has the following beneficial effects: 1. The KASP molecular marker and its supporting technology provided by this invention have achieved, for the first time, the simultaneous improvement of soybean oil content and protein content at the genetic level. This has successfully broken through the negative genetic bottleneck of the two core traits in traditional breeding, which are difficult to balance, and significantly improved the comprehensive economic value of soybean varieties.
[0032] 2. The molecular markers involved in this invention possess excellent specificity and stability. Verified through 340 natural soybean populations and multi-environment field trials, the homozygous AA allele phenotype showed a very high degree of agreement with the high-oil, high-protein phenotype, and the detection results were unaffected by environmental factors, ensuring the accuracy of breeding screening.
[0033] 3. The detection method provided by this invention significantly improves breeding efficiency and reduces research and development costs. This technology performs genotyping during the soybean seedling stage, eliminating the need to wait for plant maturity, effectively shortening the breeding cycle by 6-8 months and reducing overall breeding costs by more than 50%.
[0034] 4. The detection process based on KASP technology in this invention is adapted to high-throughput screening requirements. This process eliminates the need for an electrophoresis step, the detection time for a single sample does not exceed 2 hours, and it can achieve parallel detection of 96-well or 384-well plates, which can fully meet the actual needs of large-scale industrial breeding.
[0035] 5. The application of this invention fills the technical gap in molecular marker breeding of soybean oil and protein dual superior traits, and provides a key technical tool for the precise breeding of high-oil and high-protein soybean varieties and the rapid evaluation of germplasm resources. It has significant practical value and broad industrial promotion value. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0037] Figure 1 This is a Manhattan plot of genome-wide association analysis (GWAS) of soybean oil-protein traits provided in this embodiment of the invention; where the horizontal axis represents soybean chromosomes 1-20, and the vertical axis represents the significance of association (-log). 10 (P) values, with different colored and shaped dots representing different GWAS analysis models (GLM, MLM, CMLM, etc.); a significant high P-value appears in the chromosome 8 region of the figure. 10 (P) Peak value, corresponding to the 9020004bp target SNP site of the present invention, indicates that this site is strongly associated with the oil-protein trait.
[0038] Figure 2 This is a QQ plot of GWAS analysis of soybean oil-protein traits provided in this embodiment of the invention; where the horizontal axis represents the expected -log 10 (P) value, the vertical axis represents the actual observed -log 10 (P) values, with different colored curves corresponding to different GWAS models; most points are close to the diagonal, indicating low false positive interference in the analysis. At the same time, the observed values of some sites are significantly higher than the expected values (such as the deviation point in the upper right corner), further verifying the true association between the 9020004bp site and the target trait.
[0039] Figure 3 The soybeans provided in the embodiments of the present invention Glyma.08G117000 Gene-associated SNP information table; where the "Gene" column is the corresponding gene name ( Glyma.08G117000 The column “Chr” indicates the chromosome on which the gene is located (Chr08), the column “Hap” indicates the haplotype number (H1, H2), the column “POS” indicates the physical location of the SNP site (9020004), and the column “ALLELE” indicates the nucleotide polymorphism type (T / A) of the site.
[0040] Figure 4 This is a schematic diagram of the target SNP site flanking sequence and primer design provided in an embodiment of the present invention; the SNP site location and primer binding region are marked. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of this invention, not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0042] Unless otherwise specified, the experimental methods involved in the following embodiments are conventional methods in the art. For example, you can refer to the experimental manual in the art or follow the conditions recommended in the manufacturer's instructions.
[0043] Unless otherwise specified, all experimental materials and reagents used in the following examples are commercially available.
[0044] Example 1: Screening and Validation of Soybean Oil-Protein Synergistic SNP Sites Three hundred and forty genetically diverse soybean varieties (covering major producing areas such as Northeast China, the Huang-Huai-Hai Plain, and the Yangtze River Basin) were selected and planted in five different provinces. A randomized complete block design with three replicates was used, and conventional agronomic management was implemented. After soybean maturity, seed oil content (SOC) was determined by Soxhlet extraction, and seed protein content (SPC) was determined by Kjeldahl method. Mean values of multiple environmental phenotypes were recorded. Simultaneously, genomic DNA was extracted from leaves of each variety at the seedling stage, and SNP data of the entire chromosome 8 region were obtained by resequencing. GWAS analysis was then used to screen loci associated with oil-protein traits. Figure 1 , Figure 2 ).
[0045] Figure 1 Significantly high -log chromosomal density was observed in the region of chromosome 8. 10 (P) Peak value, corresponding to the 9020004bp target SNP site of the present invention, indicates that this site is strongly associated with the oil-protein trait.
[0046] Figure 2 Most of the points are close to the diagonal, indicating that the false positive interference in the analysis is low. At the same time, the observed values of some sites are significantly higher than the expected values (such as the deviation point in the upper right corner), which further verifies the true association between the 9020004bp site and the target trait.
[0047] The results showed that the T / A polymorphic SNP site at 9020004 bp on chromosome 8 ( Figure 3 The association values between the locus and the oil-protein dual trait were all <0.01, and the mean SOC and SPC values of the A allele homozygous soybean were significantly higher than those of the T allele homozygous soybean, thus identifying this locus as a core functional marker locus.
[0048] Example 2: Design and Validation of Specific KASP Primer Combinations Based on the soybean reference genome Wm82.a2.v1 sequence, a KASP primer combination (SEQ ID NO.2, SEQ ID NO.3, SEQ ID NO.4) was designed targeting the T / A polymorphism at the 9020004 bp site. Following KASP technical specifications, the 3' ends of the two allele-specific upstream primers corresponded to the T / A alleles, respectively, and the 5' ends introduced different fluorescent tag sequences (FAM, HEX). The universal downstream primer was complementary to the conserved sequence downstream of the SNP site. Figure 4 ).
[0049] After purification by HPLC, the primers were verified by soybean genome BLAST comparison. No non-target regions with homology ≥85% were bound, ensuring amplification specificity. At the same time, the primer annealing temperature was verified by gradient PCR to determine the optimal reaction conditions and ensure amplification efficiency.
[0050] Example 3: Validation of the breeding application of KASP molecular markers Eighty-five genetically diverse natural soybean varieties were selected. Genomic DNA was extracted from seedling leaves and the concentration was adjusted to 50 ng / μL (OD260 / OD280 = 1.8-2.0). A 2 μL KASP amplification system (containing 1 μL DNA template, 1 μL 2×KASPMaster Mix, and 0.04 μL primer mixture) was used for detection. The amplification reaction conditions are as follows: Pre-denaturation at 95℃ for 10 min; Denaturation at 95℃ for 20 seconds, annealing at 61–55℃ for 40 seconds, decreasing by 0.6℃ per cycle, for a total of 10 cycles; Denaturation at 95℃ for 20 seconds, annealing at 55℃ for 40 seconds, for a total of 30 cycles.
[0051] After amplification, genotypes were determined using quantitative real-time PCR, resulting in the selection of 42 homozygous A-allele varieties and 43 homozygous T-allele varieties. Soybean SOC and SPC were measured after maturity. The results showed that the 42 homozygous A-allele varieties had a mean SOC of 20.26% and a mean SPC of 41.86%, both meeting the criteria for high oil and high protein. The 43 homozygous T-allele varieties had a mean SOC of 18.93% and a mean SPC of 39.58%, exhibiting a low oil and low protein phenotype. The genotypes highly matched the dual-superior traits, demonstrating that this KASP molecular marker can be efficiently applied to the screening and breeding of high-oil and high-protein soybean germplasm resources.
[0052] Table 1. Results of genotype-phenotype association analysis of 85 soybean natural varieties
[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A SNP site, characterized in that, Based on the genome version number Wm82.a2.v1, the SNP site is located at position 9020004 on chromosome 8 of soybean, and the polymorphism is T / A.
2. A molecular marker, characterized in that, The molecular marker comprises a nucleic acid with a nucleotide sequence as shown in SEQ ID NO.1, wherein position 26 is polymorphic, and the polymorphism is T / A.
3. A KASP primer combination, characterized in that, The primer pair includes: F1: 5'-AGCTCTGTTTATTTTCCTTTCCTTTCT-3'; F2: 5'-AGCTCTGTTTATTTTCCTTTCCTTTCA-3'; R: 5'-AAGTCATGCTCTTCACTGCTGTTGCAT-3'.
4. A reagent kit, characterized in that, Includes the molecular markers as described in claim 1 or 2, or the KASP primer combinations as described in claim 3 or 4.
5. The application of the SNP site of claim 1, or the molecular marker of claim 2, as a target in any of the following: (1) Predict or detect the oil or protein content of soybeans; (2) Identify or breed soybean varieties with high oil or protein content; (3) Molecular marker-assisted breeding of soybeans; (4) Improvement of soybean varieties related to oil or protein content; (5) Improvement of soybean germplasm resources.
6. The use of the KASP primer combination of claim 3, or the kit of claim 4, in any of the following: (1) To predict or detect the oil or protein content of soybeans, or to prepare reagents for predicting or detecting the oil or protein content of soybeans; (2) To identify or cultivate the oil or protein content of soybeans, or to prepare reagents for identifying or cultivating the oil or protein content of soybeans; (3) Molecular marker-assisted breeding of soybeans; (4) Improvement of soybean varieties related to oil or protein content; (5) Improvement of soybean germplasm resources.
7. A method for detecting the oil content or protein content of soybeans, characterized in that, include: The polymorphism of the molecular markers described in claim 1 or 2 is detected in the soybean sample to be tested, and the oil content or protein content of the soybean to be tested is determined based on the genotype detection results.
8. The method according to claim 7, characterized in that, The detection method includes: One or more of the following methods: gene sequencing, molecular probes, liquid phase capture, or mass spectrometry.
9. The method according to claim 7 or 8, characterized in that, The determination of the oil content or protein content of the soybean to be tested based on the genotype detection results includes: Peanuts with a genotype of AA have higher oil and protein content than peanuts with a genotype of TT.
10. A method for breeding soybeans with high oil or protein content, characterized in that, include: During soybean breeding, the offspring are selected from soybeans with the genotype AA of the molecular marker described in claim 1 or 2.