InDel molecular marker related to isoflavone content of soybean seeds and application of InDel molecular marker
By developing InDel molecular markers at chromosome 57221341-57221412 loci of soybean genome 18, specific primer pairs were designed for PCR amplification and electrophoresis detection, the problem of scarcity of soybean varieties was solved, and high-efficiency breeding and rapid identification of soybean isoflavones content was achieved.
Patent Information
- Application Number
- CN202510521941.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-24
AI Technical Summary
In the prior art, high isoflavone soybean varieties are scarce and excellent germplasm resources are limited, which leads to lag in research on soybean isoflavone, which is difficult to meet market demand, and lacks effective molecular markers for rapid identification and breeding.
A InDel molecular marker was developed, located at the chromosome 18 57221341-57221412 site of the soybean genome version Glycine_max_Wm82.a2.v1, and designed specific primer pairs for PCR amplification. The isoflavone content was detected by agarose gel electrophoresis, and soybean breeding materials with high isoflavone content were screened.
It has achieved rapid and accurate selection and breeding of soybean varieties with high isoflavones content, which has improved the efficiency and success rate of crop breeding, and met the market's demand for soybean isoflavones.
Smart Images

Figure CN120384148A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to an InDel molecular marker related to the isoflavone content in soybean seeds and its application. Background Art
[0002] Soybean isoflavones are important phytoalexins. As secondary metabolites in the phenylalanine pathway during soybean growth, previous studies have shown that soybean isoflavones have significant effects on preventing various diseases such as cardiovascular and cerebrovascular diseases, osteoporosis, hypertension, diabetes, breast cancer, and prostate cancer. Therefore, they have attracted extensive attention from researchers in the fields of medicine and food science and technology. In multiple fields, soybean isoflavones have also been widely developed and utilized, and currently, they have become highly sought-after health products.
[0003] In addition to their value in human health, soybean isoflavones also participate in the plant disease resistance process. Some studies have shown that by increasing the isoflavone content in soybean plants, the disease resistance of soybeans can be enhanced, the demand for pesticides can be reduced, and thus the economic benefits can be improved. However, in the agricultural field, the research on soybean isoflavones lags behind. This is mainly reflected in the scarcity of high-isoflavone soybean varieties, the limited availability of excellent germplasm resources, and the relatively few studies on the genetic mechanism of soybean isoflavones. To meet the increasing market demand for soybean isoflavones, it is urgent to increase the yield of high-isoflavone soybeans and cultivate soybean varieties rich in high isoflavones. Therefore, identifying the loci related to each component of soybean isoflavones, mining candidate genes related to isoflavones, developing molecular markers related to each component of soybean isoflavones, and mining high-isoflavone soybean resources are of great significance for accelerating the molecular breeding of high-isoflavone soybeans. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide an InDel molecular marker related to the isoflavone content in soybean seeds and its application.
[0005] To solve the above technical problems, the technical solutions adopted by the present invention are as follows.
[0006] An InDel molecular marker related to the isoflavone content in soybean seeds, the marker is located at the locus of 57221341 - 57221412 on chromosome 18 of the soybean genome version Glycine_max_Wm82.a2.v1, and contains an insertion / deletion mutation of 72 bp. Among them, the high-isoflavone content material carries the A-type allele containing a 72-bp insertion fragment, and the low-isoflavone content material carries the B-type allele lacking this fragment.
[0007] Further preferably, the nucleotide sequence of the A-type allele is as shown in SEQ ID NO: 1, and the nucleotide sequence of the B-type allele is as shown in SEQ ID NO: 2.
[0008] A specific primer pair for detecting the InDel molecular marker, the upstream primer sequence is as shown in SEQ ID NO: 3; the downstream primer sequence is as shown in SEQ ID NO: 4.
[0009] A kit for detecting the content of soybean isoflavones, comprising the specific primer pair described in claim 3.
[0010] A detection method for identifying the content of soybean isoflavones, comprising the following steps:
[0011] (1) Using the DNA of the soybean material to be tested as a template, performing PCR amplification with the InDel molecular marker primer to obtain a PCR amplification product;
[0012] (2) Performing agarose gel electrophoresis detection on the amplification product and observing the electrophoresis detection result.
[0013] Further preferably, when the length of the amplification product is 544 bp, the gene marker band type of the soybean to be tested is type A; when the length of the amplification product is 472 bp, the gene marker band type of the soybean to be tested is type B; the content of isoflavones is as follows: the soybean with gene marker band type A is greater than or potentially greater than the soybean with gene marker band type B.
[0014] Further preferably, the reaction program of the PCR amplification is pre-denaturation at 95 °C for 3 min; 95 °C for 30 s, 52 °C for 30 s, 72 °C for 30 s, for 33 cycles; extension at 72 °C for 5 min.
[0015] The application of the InDel molecular marker or the specific primer pair in soybean molecular marker-assisted breeding, screening soybean breeding materials with high isoflavone content by detecting the genotype of the molecular marker.
[0016] A breeding method for soybean varieties with high isoflavone content, comprising:
[0017] 1) Using the specific primer pair described in claim 3 to detect the genotype of the InDel-57221341 marker of the parent or hybrid offspring;
[0018] 2) Preferentially selecting materials carrying the A-type allele as breeding parents or selected individual plants;
[0019] 3) Combining field trait evaluation to obtain soybean varieties with high isoflavone content.
[0020] A molecular identification method for soybean varieties, distinguishing soybean varieties with different isoflavone content characteristics by detecting the genotype of the InDel molecular marker described in claim 1.
[0021] The beneficial effects of adopting the above technical solutions are as follows: For the first time, the present invention identifies a 72bp InDel molecular marker significantly related to isoflavone content at the locus of 57221341-57221412 on chromosome 18 of the soybean genome version Glycine_max_Wm82.a2.v1. According to the InDel molecular marker, InDel primer sequences are designed for PCR amplification, and the test results are observed to identify the isoflavone content phenotype. Using the primers of the present invention can be used for the identification and screening of the isoflavone content phenotype of soybeans, and can quickly, accurately and effectively breed soybean varieties with high isoflavone content, accelerate the crop breeding process, and improve the breeding efficiency and success rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is the population structure of 290 natural populations;
[0023] Figure 2 is a schematic diagram of the electrophoresis results of materials 1 to 48;
[0024] Figure 3 is a schematic diagram of the electrophoresis results of materials 49 to 96;
[0025] Figure 4 is a schematic diagram of the electrophoresis results of materials 97 to 144;
[0026] Figure 5 is a schematic diagram of the electrophoresis results of materials 145 to 193;
[0027] Figure 6 is a schematic diagram of the electrophoresis results of materials 194 to 239;
[0028] Figure 7 is a schematic diagram of the electrophoresis results of materials 240 to 286;
[0029] Figure 8 is a schematic diagram of the electrophoresis results of materials 287 to 300;
[0030] Figure 9 is the Manhattan plot and QQ plot of the genome-wide association analysis results of soybean isoflavones,
[0031] Note: Daidzin: daidzin; Genistin: genistin; Glycitin: glycitin; Malonyldaidzin: malonyldaidzin; Malonylgenistin: malonylgenistin; Malonylglycitin: malonylglycitin; TIF: total isoflavones;
[0032] Figure 10 is the gene linkage disequilibrium and haplotype block diagram within the main locus interval of chromosome 18 of soybean isoflavones;
[0033] Figure 11 These are the PCR amplification results of primers in some soybeans;
[0034] Figure 12 These are the analyses of the differences in isoflavone content between two haplotype materials of Indel markers;
[0035] Figure 13 These are the on-site planting photos of GWAS materials. Specific implementation manners
[0036] The following examples illustrate the present invention in detail. All kinds of raw materials and various equipment used in the present invention are conventional commercially available products and can be directly obtained through market purchase. The experimental methods used in the following examples are all conventional methods unless otherwise specified.
[0037] It should be understood that when used in the specification of this application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations.
[0038] It should also be understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0039] The reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0040] In addition, in the descriptions of the specification of this application and the appended claims, the terms "first", "second", "third" etc. are only used for differentiating descriptions and cannot be understood as indicating or implying relative importance.
[0041] Next, in combination with specific embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts belong to the scope of protection of the present invention.
[0042] Example 1. Materials and Methods
[0043] 1.1 Test Materials
[0044] The natural population used in this study included 290 soybean materials collected by the research team in the early stage. Among them, the domestic materials included 108 released varieties, 60 micro-core germplasms, and 40 Huang-Huai-Hai germplasm materials; the foreign materials included 82 mainly from the United States, Japan, and South Korea. They were planted at the Diti Experimental Station in Shijiazhuang City, Hebei Province (114.48°E, 38.03°N) in three growing seasons of 2021, 2022, and 2023. Each material was sown in 3 rows, with a row length of 3 m, a row spacing of 0.5 m, and a plant spacing of 0.2 m. A completely randomized block design was adopted, with 3 replicates and normal field management (as Figure 13 shown). After the materials matured naturally, the seeds of 5 plants in the middle row were harvested for the determination of soybean isoflavone content.
[0045] 1.2 Test Instruments and Equipment
[0046] Jinteng organic phase nylon filter membrane with a diameter of 13 mm and a pore size of 0.45 μm, chromatographic grade methanol solution, chromatographic grade acetonitrile, liquid phase automatic injection vials and bottle caps were purchased from Thermo Fisher Scientific; analytical grade acetic acid was purchased from Merck; the main instruments included a micro analytical balance, a multi-functional pulverizer, Thermo Scientific TM basic vortex oscillator, centrifuge, ultrasonic water bath thermostatic oscillator, and Agilent high performance liquid chromatograph.
[0047] 1.3 Test Methods
[0048] The extraction and determination methods of soybean isoflavone content are as follows:
[0049] a. Select 30 soybean seeds, crush them through an 80-mesh sieve using a high-speed universal pulverizer, and mix well. Weigh 0.02 g of soybean powder and put it into a 2-mL centrifuge tube.
[0050] b. Add 1 mL of 70% ethanol containing 0.1% acetic acid to the powder, place it on a shaker (INNOVA42, Eppendorf, Shanghai agency, Germany) at 28 °C, and shake at 200 rpm for 12 h.
[0051] c. Centrifuge at 12,000 r / min for 10 min at 4 °C, take the supernatant, filter it through a 0.45-μm organic phase syringe filter, and inject it into a liquid phase injection vial. Store it at 4 °C for standby or at -20 °C for long-term storage, and wait for detection by a high-performance liquid chromatograph.
[0052] d. Determine using an Agilent 1260 liquid chromatograph. The chromatographic column is YMC-Pack ODS-AM-303 (250 mm × 4.6 mm I.D., S-5 μm, YMC Co., Kyoto, Japan). The column temperature is 35 °C. The mobile phases A and B are 0.1% acetic acid-distilled water and acetonitrile, respectively. The injection volume is 20 μL, the flow rate of the mobile phase is 1.0 mL·min-1, the column temperature is 35 °C, and a linear gradient of 13–35% acetonitrile (V / V) is used for a running time of 70 min and maintained at 35 °C. Detect isoflavones by ultraviolet absorption at a wavelength of 260 nm.
[0053] e. Considering that soybeans mainly contain six isoflavone components, namely daidzin, glycitin, genistin, malonyldaidzin, malonylglycitin, and malonylgenistin. Therefore, this study mainly detects the above six isoflavone components, and the total isoflavone content is defined as the sum of the contents of the above 6 individual isoflavones.
[0054] 1.4 Isoflavone phenotype data analysis
[0055] In a high-performance liquid chromatography (HPLC) system, the isoflavone concentrations of the standard and sample are determined according to the corresponding retention times and peak areas. Perform a linear regression analysis on the peak areas and concentrations of the standard to establish a linear equation. In this study, SPSS 26.0 software was used to perform descriptive statistics and correlation analysis on the phenotype data, Excel was used to organize and analyze the phenotype data of soy isoflavones, and Graphpad prism 7.0 was used to create a frequency distribution histogram of the phenotype values of soy isoflavones to observe whether the data conforms to a normal distribution.
[0056] 1.5 Extraction and detection of soybean genomic DNA
[0057] In this experiment, fresh leaves of the natural population (15 days after emergence) were used, and a DNA extraction kit (Nuclean Plant Genomic DNA Kit) from Kangwei Biotech Co., Ltd. was used for extraction. The specific operations are as follows:
[0058] a. Put about 100 mg of fresh leaves into a 2 mL centrifuge tube, add liquid nitrogen, and then use a high-throughput tissue grinder to break and thoroughly grind the leaves.
[0059] b. Add 400 μL of Buffer LP1 and 6 μL of RNase A to the ground powder and mix well. After shaking on a vortex oscillator for 2 min, let it stand at room temperature for 10 - 15 minutes to ensure complete lysis.
[0060] c. Then add 130 μL of Buffer LP2 to the lysed solution, mix well and shake on a vortex oscillator for 2 min; centrifuge at 12,000 rpm for 5 min on a centrifuge, and then transfer the supernatant after centrifugation to a new centrifuge tube.
[0061] e. Add 1.5 volumes of Buffer LP3 to it and immediately mix well.
[0062] f. Pour all the solution and precipitate into an adsorption column equipped with a collection tube, centrifuge at 12,000 for 1 min, then pour out the waste liquid in the tube, and further place the adsorption column in a 2 mL centrifuge tube.
[0063] g. Add 500 μL of Buffer GW2, centrifuge at 12,000 on a centrifuge for 1 min, then pour out the waste liquid in the tube, and put it back into a 2 mL centrifuge tube.
[0064] h. Repeat step 6
[0065] i. Centrifuge at 12,000 for 2 min on a centrifuge and then pour out the waste liquid in the tube, and then place the adsorption column at room temperature for 10 min until it is completely dry.
[0066] j. Transfer the adsorption column to a new 1.5 mL centrifuge tube, add 50 - 100 μL of Buffer GE or sterilized water inside the adsorption column in a suspended manner, then let it stand at room temperature for 5 min, and centrifuge at 12,000 rpm for 1 min to collect the DNA solution.
[0067] k. After measuring the concentration of the DNA solution using a Thermo Scientific Nano Drop 2000c (Thermo, USA) instrument, store the DNA at -20 °C.
[0068] 1.6 Whole-genome resequencing data processing
[0069] The genotypes of the GWAS population in this study were sequenced using the BGI MGI-2000 / MGI-T7 sequencing platform and the whole-genome resequencing technology on the soybean genomes of 290 materials from natural populations. The sequencing mode was PE150. A total of 15,049,485 SNP locus data were included. Quality control was performed using the vcftools
[100] software, removing loci with a minor allele frequency (MAF) less than 0.05, loci with a heterozygous proportion greater than 30%, and loci with a missing rate greater than 30%. After final quality control, there were 5,136,169 high-quality SNP loci. The quality-controlled clean reads were aligned with the reference genome sequence using the BWA software. The obtained VCF file was used for population structure analysis with Structure 2.0.
[0070] 1.7 Genetic structure analysis of natural populations
[0071] Population structure analysis of 290 materials was performed using structure 2.3.4. When the number of subgroups K was 6, the cross-validation error rate (CV error) was at the minimum. Therefore, this population was divided into 6 subgroups, and the 6 subgroups contained 55, 61, 19, 67, 70, and 18 materials respectively (as Figure 1 ).
[0072] 1.8 Genome-wide association study
[0073] GWAS analysis was performed by adopting the mixed linear model (MLM) built into the TASSEL software. If the -Log10(p) value of an SNP reached or exceeded 4.0, then this SNP marker was considered to be significantly associated with the trait. The Kinship value was estimated using the TASSEL software. Manhattan plots and QQplot plots were drawn using the pheatmap package and CMplot package of the R software.
[0074] 1.9 Candidate gene identification
[0075] Based on the association analysis, within the candidate intervals where the significantly associated loci were located, candidate genes related to soybean isoflavone content were predicted based on gene function annotations in the Phytozome database (https: / / phytozome.org) and the Soybase database (https: / / www.Soybase.Org).
[0076] 1.10 Development and screening of InDel markers
[0077] Using the whole-genome re-sequencing data of natural populations, InDels present in the target interval were screened, and the base sequences in the regions near the InDels were downloaded from the Phytozome database (https: / / phytozome.org) for primer design. Primers were designed using the BioXM 2.7.1 software with the following parameter settings: primer length was 20 - 25 bp, annealing temperature was 55 - 65 °C, product size was 80 - 160 bp, GC content was 40% - 60%, the number of hairpin structures was less than 6, and the number of primer dimers was less than 8. The designed primers were subjected to BLAST alignment using the Phytozome website to ensure that the primers were in the conserved region of the DNA sequence and were specific.
[0078] Primer sequences: FP: TTAAGGCAAATTCTACCGTAAGGC (SEQ ID NO: 3); RP: TTCTAACCAAAAGTGCTGCCAACT (SEQ ID NO: 4).
[0079] 1.11 Validation of the effectiveness of InDel markers
[0080] Using the developed InDel marker primers, PCR amplification and agarose gel electrophoresis detection were performed on the genomic DNA of 300 soybean materials. The results were read in the following way: the band patterns of high-isoflavone materials that were the same were recorded as "A", the band patterns of low-isoflavone materials that were the same were recorded as "B", and the fuzzy or missing band patterns were recorded as "-", and thus the effectiveness of the InDel markers was verified.
[0081] The total volume of the 2×Taq MasterMix (ComWin) reaction system was 20 μL, as shown in Table 1 specifically:
[0082] Table 1 PCR reaction system
[0083]
[0084] The PCR reaction procedure was carried out using an Eppendorf PCR instrument, and the amplification reaction procedure is shown in Table 2 specifically:
[0085] Table 2 PCR amplification reaction procedure
[0086]
[0087] Preparation and detection of agarose gel electrophoresis:
[0088] Gel Preparation: Weigh 6 g of agar powder and put it into an Erlenmeyer flask. Add 100 mL of 1×TAE, mix well, and heat it in a microwave oven for 3 - 4 min until all the agar powder is dissolved. After the temperature of the agar solution drops, add 1 - 2 μL of nucleic acid dye, shake well, and slowly pour it into the gel preparation tank, avoiding the generation of bubbles. Then let it stand and cool for 15 min. When the gel becomes solid, it can be used.
[0089] Loading Samples: Slowly place the solidified gel into the electrophoresis buffer in the electrophoresis tank. After adding the corresponding Marker, load the samples in sequence, connect the power supply and the electrophoresis tank, and set the program to start running the gel.
[0090] Result Reading: After the electrophoresis is completed, wear disposable gloves and place the gel on the gel imaging instrument for development. Take a picture and save the gel image (such as Figures 2-8 ).
[0091] Example 2, Results and Analysis
[0092] 2.1 Statistical Analysis of Isoflavone Components under Different Environments
[0093] The determination of 6 main soybean isoflavone components and the statistical analysis of the total isoflavone content in 290 soybean germplasms were carried out. The results are shown in Table 3. The average total isoflavone content of the three-year soybean germplasms was 1832.30 μg / g, 1943.83 μg / g, and 1547.14 μg / g respectively. Among the contents of each single component, the average content of malonyl genistin was the highest, with the average contents of the three years being 726.58 μg / g, 956.05 μg / g, and 678.52 μg / g respectively. The average content of glycinin was the lowest, with the average contents of the three years being 95.09 μg / g, 76.73 μg / g, and 104.05 μg / g respectively. The coefficient of variation of the contents of 6 soybean isoflavone components and their total isoflavone content was 35% - 76%, indicating that this population has rich phenotypic variation. The broad-sense heritability of each trait under the three-year environment was between 49% - 71%.
[0094] Table 3 Descriptive Statistical Analysis of Soybean Isoflavone Phenotypes under Different Environments
[0095]
[0096]
[0097] 2.2 Genome-wide Association Analysis of Soybean Isoflavone Components
[0098] Using GWAS analysis, the associated loci of the total content of soybean isoflavones and its individual components in the natural population were mined, and the Manhattan Plot and QQ-Plot were drawn accordingly. When LOD ≥ 4, it was considered that the SNP / InDels was significantly associated with the total content of soybean isoflavones and each component (such asFigure 9 ) By performing association analysis on a total of 7 traits, including the total content of soybean isoflavones and the contents of 6 components, a total of 404 SNPs significantly associated with the total isoflavone content and the contents of 6 components were identified (Table 4). There were 23 significant SNPs for daidzin, mainly distributed on chromosomes 8, 9, 10, 13, 14, and 20. Among them, chromosomes 8, 10, 13, and 14 had 6, 2, 5, and 4 significant SNPs respectively, while chromosomes 9 and 20 each had 3 significant SNPs; the maximum value of -Log10(P) among these SNPs was 5.2. For genistein, 59 SNPs were mapped in two years and above, among which 34 SNPs were mapped in two years, distributed on chromosomes 3, 10, 12, 13, 14, and 18, and the maximum value of -Log10(P) was 6.2; while 25 significant SNPs were mapped in three years, distributed on chromosomes 1, 2, 7, 15, and 20 respectively, and the maximum value of -Log10(P) was 6.2. Glycitin had the most associated SNPs among all components, with a total of 133 significant SNPs, distributed on all chromosomes except chromosomes 6, 9, 15, 17, and 19. Among them, chromosome 18 had 55 significant SNPs, and the maximum value of -Log10(P) was 7.2. And 61 loci were simultaneously associated in three years, mainly distributed on chromosomes 1, 2, 3, 5, 10, 11, 13, 14, 16, and 20. Malonyl daidzin was associated with a total of 29 SNPs, mainly distributed on chromosomes 1, 4, 5, 10, 13, 18, and 19. Among them, 7 SNPs were simultaneously associated in three years, all distributed on chromosome 5. Malonyl genistin was associated with 6 SNPs in two years, which was also the least number of associated SNPs among all components, distributed on chromosomes 6 and 17. Malonyl glycitin had the widest distribution of SNPs among all components. Except for chromosomes 3, 11, and 20, it was distributed on all chromosomes. There were a total of 127 SNPs, among which 91 SNPs were simultaneously mapped in three years, mainly distributed on chromosomes 2, 4, 6, 8, 10, 13, 14, 16, 17, and 18, and the maximum value of -Log10(P) was 8.1. And the total content of soybean isoflavones had 27 SNPs, mainly distributed on chromosomes 3, 6, 8, 17, 18, and 19. Among them, the locus 56,734,042 - 57,223,667 on chromosome 18 overlapped with the locus of malonyl daidzin and was repeatedly detected under three-year environmental conditions ( Figure 9 ; Table 4). This indicates that it not only stably exists in different environments but is also an important locus controlling the content of soybean isoflavones.
[0099] Table 4. Results of association analysis of isoflavone content in natural populations
[0100]
[0101]
[0102] Note: *Only the loci that were located in all three environments are shown.
[0103] 2.3 Screening and analysis of candidate genes related to soybean isoflavone components
[0104] In this study, it was detected between the markers Chr18-56,734,042 and Gm18-57,223,667 on chromosome 18 in the total content of soybean isoflavones and in multiple environments. According to the soybean genome information in the Soybase database (https: / / www.soybase.org / ), it was found that this locus region contains 74 genes. Based on the LD distance, we selected the intervals of 200 kb upstream and downstream of the marker Chr18-57223667 with the highest P value, and used the Haploview 4.2 software to perform LD Block on all SNPs in this region and narrow the interval. The interval was narrowed to 107 Kb, and this region contains 29 genes (such as Figure 10 ).
[0105] 2.4 Candidate gene prediction
[0106] The prediction results of the Phytozome and Soybase databases showed that a total of 5 candidate genes related to the content of soybean isoflavones were obtained within the locus interval of the total content of soybean isoflavones (Table 5), mainly involving phospholipase 3, class I transglutaminase superfamily protein, Core-2 / I-branching beta-1,6-N-acetylglucosaminyltransferase family protein, related protein kinase, and peptidyl-tRNA hydrolase family protein, etc. Among them, Glyma.18G294500 encodes a related protein kinase. Combining with the analysis of re-sequencing data, there is an insertion / deletion mutation containing 72 bases inside this gene, resulting in great changes in its structure and function, and further identification of the effectiveness of InDel markers was carried out on it.
[0107] Table 5 Prediction and functional annotation of candidate genes for soybean isoflavone content
[0108]
[0109] 2.5 Detection of the effectiveness of InDel markers
[0110] Sequencing data analysis revealed that there is an insertion / deletion mutation of 72 bases within the Glyma.18G294500 gene in this region (Chr18_57221341: ATCAATTATCCCTTCCCTTCCCAAACCCTTATCTTCCTTGACCGTATCCGCTACAGCAAAA ACAACACCAAC (SEQ ID NO: 5) / -), which causes significant changes in structure and function. The homologous gene of this gene in Arabidopsis thaliana is AT1G66940.1. Gene function annotation information indicates that this gene is related to protein kinase. The upstream and downstream 200bp base sequences of InDel-57221341 were used for primer design. The primer sequences are shown in Table 6. PCR amplification and agarose gel electrophoresis detection were performed using 300 soybean DNAs as templates. According to the band types read and counted, a total of 2 allelic variations were detected ( Figure 11 ). It is predicted that the amplified band size in Williams82 is 544bp, and the subsequent PCR electrophoresis results of this type are marked as "A", and the specific sequence is shown in Sequence Listing SEQ ID NO: 1; the amplified band size of the mutant material is 472bp, and the subsequent PCR electrophoresis results of this type are marked as "B", and the specific sequence is shown in Sequence Listing SEQ ID NO: 2. Among them, 145 soybean resources have the 544bp allelic variation, and 126 soybean resources have the 472bp allelic variation. The incidence of the 544bp allelic variation is 48.33%, and the frequency of the 472bp allelic variation is 42%. There are 29 fuzzy or missing band types.
[0111] Table 6 Primer sequences of InDel markers
[0112]
[0113] Using this marker for genotyping in the natural population, the population was divided into two subpopulations. As shown in Table 7, among them, the isoflavone content of 145 materials corresponding to the "A" type band is 550.38 μg / g - 4668.57 μg / g, with an average value of 1923.46 μg / g; 126 are of type B, and the isoflavone content of the corresponding materials is 456.06 μg / g - 4484.02 μg / g, with an average value of 1674.24 μg / g. Statistical analysis shows that in the continuous three years and the three-year average value, the isoflavone content of the materials with the A band type is significantly higher than that of the B band type materials (such as Figure 12 ). It shows that using this marker can assist in screening high-isoflavone soybean offspring materials.
[0114] Table 7 Isoflavone content and mean phenotypic data of the natural population from 2021 to 2023
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122] As Figure 12 shown in Table 7 and Table 8, combined with the phenotypic data results, it shows that in 2021, the average isoflavone content of the materials with band type A is 1935.11; the average isoflavone content of the materials with band type B is 1830.48; the average isoflavone content of group A materials is greater than that of group B materials, but it does not reach a highly significant level. In 2022, the average isoflavone content of the materials with band type A is 1919.53; the average isoflavone content of the materials with band type B is 1666.27; there is a highly significant difference between group A and group B, and the P value is 0.002. In 2023, the average isoflavone content of the materials with band type A is 1637.24; the average isoflavone content of the materials with band type B is 1450.82; there is a significant difference between group A and group B, and the P value is 0.0208.
[0123] Table 8 Association analysis results of isoflavone content in the natural population from 2021 to 2023
[0124]
[0125] This marker can effectively distinguish two types of gene marker band types and has the characteristic of codominance; using this marker can assist in screening soybean offspring materials with high isoflavone content.
[0126] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these examples without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0127] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0128] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An InDel molecular marker related to the isoflavone content of soybean seeds, characterized in that, The said marker is located at the site of 57221341 - 57221412 on chromosome 18 of the soybean genome version Glycine_max_Wm82.a2.v1, containing a 72 bp insertion / deletion mutation. Among them, the high isoflavone content material carries the A-type allele with a 72 bp insertion fragment, and the low isoflavone content material carries the B-type allele lacking this fragment.
2. The InDel molecular marker according to claim 1, characterized in that, The nucleotide sequence of the said A-type allele is shown as SEQ ID NO: 1, and the nucleotide sequence of the B-type allele is shown as SEQ ID NO:
2.
3. A specific primer pair for detecting the InDel molecular marker described in claim 1, characterized in that, The upstream primer sequence is shown as SEQ ID NO: 3; the downstream primer sequence is shown as SEQ ID NO:
4.
4. A kit for detecting the content of soybean isoflavones, characterized in that, It contains the specific primer pair described in claim 3.
5. A detection method for identifying the content of soybean isoflavones, characterized in that, It includes the following steps: (1) Using the DNA of the soybean material to be tested as a template, perform PCR amplification with the InDel molecular marker primers to obtain a PCR amplification product; (2) Perform agarose gel electrophoresis detection on the amplification product and observe the electrophoresis detection results.
6. The detection method according to claim 5, wherein When the length of the said amplification product is 544 bp, the gene marker band type of the soybean to be tested is A-type; when the length of the said amplification product is 472 bp, the gene marker band type of the said soybean to be tested is B-type; the high or low isoflavone content is: the soybean with the gene marker band type A is greater than or candidate greater than the soybean with the gene marker band type B.
7. The method according to claim 5, characterized in that The reaction program of the said PCR amplification is pre-denaturation at 95°C for 3 min; 95°C for 30 s, 52°C for 30 s, 72°C for 30 s, for 33 cycles; extension at 72°C for 5 min.
8. Use of the InDel molecular marker according to claim 1 or the specific primer pair according to claim 3 in soybean molecular marker-assisted breeding, characterized in that, Screen soybean breeding materials with high isoflavone content by detecting the genotype of the said molecular marker.
9. A breeding method for a soybean variety with high isoflavone content, characterized in that, It includes: 1) Detect the InDel-57221341 marker genotype of the parents or hybrid offspring using the specific primer pair described in claim 3; 2) Prioritize selecting the materials carrying the A-type allele as breeding parents or selected individual plants; 3) Combine field trait evaluation to obtain soybean varieties with high isoflavone content.
10. A molecular identification method for a soybean variety, characterized in that, Distinguish soybean varieties with different isoflavone content characteristics by detecting the genotype of the InDel molecular marker described in claim 1.
Citation Information
Patent Citations
Molecular marker-based high-isoflavone soybean variety breeding method
CN112852990A
Quantitative character gene locus related to content of soybean isoflavone and application of quantitative character gene locus
CN115044702A
Haploid molecular marker related to isoflavone content of soybean seeds and application of haplotype molecular marker
CN119061178A