Application of in del molecular marker related to soybean hundred kernel weight

By constructing a secondary population of chromosome segment substitution lines and a high-density genetic map, the GmSW01.1 gene was precisely located and the InDel molecular marker was developed. This solved the problems of imprecise localization of soybean 100-seed weight and unclear candidate genes, and achieved a significant increase in soybean 100-seed weight and efficient breeding.

CN122484338APending Publication Date: 2026-07-31ANHUI SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI SCI & TECH UNIV
Filing Date
2026-06-22
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, the QTL loci associated with 100 soybean grain weight are not precisely located, and the candidate genes are not clearly identified, which limits the application of molecular breeding and makes it difficult to improve soybean yield.

Method used

By constructing a secondary population of chromosome segment substitution lines, using high-density genetic maps and screening recombinant individuals, the qSW01.1 site was precisely located, GmSW01.1 was identified as a key candidate gene, and InDel molecular markers were developed. Their functions were verified by combining CRISPR-Cas9 technology, and specific InDel molecular markers were developed for PCR amplification and electrophoresis detection.

Benefits of technology

This study enabled precise localization and functional verification of the 100-seed weight trait in soybeans, providing efficient and stable molecular markers that significantly improved the 100-seed weight of soybean progeny. It is suitable for large-scale population screening and supports the breeding of high-yield and high-quality soybean varieties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122484338A_ABST
    Figure CN122484338A_ABST
Patent Text Reader

Abstract

This invention relates to the field of plant breeding technology, specifically to the application of InDel molecular markers related to the 100-seed weight of soybeans. This marker is located in... GmSW01.1 The gene, as shown in SEQ ID NO.7, contains a deletion of the AAC tribase at positions 118bp-120bp. This invention, through expression analysis and sequence alignment, determined... Glyma.01g051700 As a key candidate gene, it was named GmSW01.1 Sequence analysis revealed the presence of an InDel molecular marker with a -AAC triplet deletion in the CDS sequence of GmSW01.1 within NN1138-2. Based on this InDel molecular marker, specific primers were developed, and the 100-seed weight of soybean materials was identified by PCR amplification of fragment size. This invention provides important genetic resources and technical support for marker-assisted breeding of soybean 100-seed weight and for the selection of high-yielding and high-quality varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of plant breeding technology, specifically to the application of InDel molecular markers related to the 100-seed weight of soybeans. Background Technology

[0002] Soybeans are one of the world's most important food and economic crops. Not only are soybeans a major source of plant protein for humans, but their oil is also one of the world's largest-produced vegetable oils. However, the soybean industry faces severe challenges: soybean production has stagnated for a long time, import dependence exceeds 80%, soybean arable land is nearing its limit, and supply and demand are severely imbalanced. Therefore, increasing yield per unit area has become one of the key ways to solve the soybean supply problem.

[0003] 100-grain weight is an important agronomical trait affecting soybean yield and quality, and is a quantitative trait regulated by multiple genes. From wild soybeans... Glycine soja To cultivate soybeans Glycine max The average 100-grain weight significantly increased from 2.2g to 15.9g, and the grain size also increased significantly, indicating that 100-grain weight has great potential for genetic improvement. Currently, although several QTL loci associated with 100-grain weight have been identified, most loci have not yet been finely mapped, candidate genes are unclear, and functional verification is insufficient, which limits their application in molecular breeding.

[0004] Therefore, it is urgent to explore superior allelic variations in soybeans, precisely locate and clone key genes for 100-seed weight, and develop efficient and stable molecular markers to provide theoretical guidance and practical tools for high-yield soybean breeding. Summary of the Invention

[0005] To address the above problems, this invention provides the application of InDel molecular markers related to the 100-seed weight of soybeans.

[0006] This invention is achieved through the following technical solution: Application of InDel molecular markers related to 100-seed weight of soybean, wherein the InDel molecular markers are located in GmSW01.1 The nucleotide sequence of the gene is shown in SEQ ID NO.7, with a deletion of the AAC tribase at position 118bp~120bp.

[0007] The application refers to any one of the following (1) and (2): (1) Identify the 100-seed weight of soybeans; plants with AAC triple base deletion at position 118bp~120bp of SEQ ID NO.7 have a higher 100-seed weight than those without deletion.

[0008] (2) Improve the 100-seed weight of soybean offspring; select soybean samples with AAC triple base deletion at position 118bp~120bp of SEQ ID NO.7 as parents for breeding to improve the 100-seed weight of soybean offspring.

[0009] Preferably, the GmSW01.1 The nucleotide sequence of the gene is shown in SEQ ID NO.1.

[0010] Preferably, the method for determining the weight of 100 soybeans is as follows: Genomic DNA was extracted from the soybeans to be tested; PCR amplification was performed using a primer set to obtain the amplified fragment; the primer set consisted of a forward primer and a reverse primer.

[0011] Electrophoretic analysis of the amplified fragments showed that the weight of 100 soybeans with an amplified fragment size of 181 bp was greater than that of soybeans with an amplified fragment size of 184 bp.

[0012] Preferably, the method for increasing the 100-grain weight of soybean offspring is as follows: Genomic DNA was extracted from the soybeans to be tested; PCR amplification was performed using a primer set to obtain the amplified fragment; the primer set consisted of a forward primer and a reverse primer.

[0013] Electrophoresis was performed on the amplified fragments, and soybean samples with an amplified fragment size of 181 bp were selected as parents for breeding to improve the 100-seed weight of soybean offspring.

[0014] Preferably, the amplification program is set as follows: pre-denaturation: 95℃ for 3 min; denaturation: 95℃ for 10 sec, annealing: 55℃ for 10 sec, extension: 72℃ for 30 sec, for a total of 30-35 cycles; complete extension: 72℃ for 5 min; storage: 4℃.

[0015] Preferably, the amplification system consists of 3 μL DNA template, 1.5 μL forward primer, 1.5 μL reverse primer, and 4 μL 2×taq MasterMix.

[0016] Preferably, the concentrations of both the forward and reverse primers are 10 μM.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention utilizes the chromosome segment substitution line CSSL130 and cultivated soybean NN1138-2 to construct a secondary population. Through high-density genetic mapping and screening of recombinant individuals, [the following is likely a separate, unrelated sentence:] qSW01.1 The locus was precisely located to a region of approximately 310 kb on chromosome 1, and uniquely identified through expression analysis and CDS sequence alignment. Glyma.01g051700Key candidate genes were identified and named. GmSW01.1 This invention overcomes the shortcomings of existing technologies, such as large QTL site localization intervals and unclear candidate genes. Furthermore, this invention employs CRISPR-Cas9 technology to... GmSW01.1 Knockout was performed, resulting in three homozygous knockout lines. Phenotypic analysis showed that the 100-seed weight decreased significantly by 7.98%–8.06% after knockout, and seed length, width, thickness, and number of branches were all significantly reduced, directly proving that... GmSW01.1 This gene is a positive regulator of soybean 100-seed weight, providing an important supplement to the current lack of sufficient functional verification in existing technologies. Based on the -AAC triple base deletion between cultivated soybean NN1138-2 and the wild substitution line CSSL130, this invention developed a specific InDel molecular marker and its primer set. Through simple PCR amplification and electrophoresis detection, 181bp vs 184bp, superior haplotype Hap can be efficiently distinguished. NN Compared with ordinary haplotype Hap CSSL It is simple to operate, low in cost, and has good stability, making it suitable for large-scale population screening. Haplotype and 100-grain weight correlation analyses in secondary and natural populations of Sichuan and Chongqing materials both showed that those carrying Hap... NN The weight per 100 grains of haplotype material is significantly higher than that of Hap. CSSL Materials confirm Hap NN This invention identifies and verifies an excellent haplotype for controlling soybean grain weight. In summary, this invention has discovered and validated a superior haplotype from wild soybean. GmSW01.1 The function of the gene was identified, its superior allelic variation types were clarified, and molecular markers that can be applied in industry were developed. This provides theoretical guidance and gene resources for molecular marker-assisted breeding of soybean 100-seed weight and the selection of high-yield and high-quality varieties, and has potential economic and social value for alleviating the contradiction between soybean supply and demand in my country. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the genome structure of the parent CSSL130 and a comparison diagram of the seed phenotypes of the parents in this invention; Figure 1In the diagram, A is a schematic diagram of the genome structure of the secondary population parent CSSL130; black boxes represent wild segments from wild soybean N24852, white boxes represent segments from cultivated soybean NN1138-2, green boxes represent heterozygous segments, and red circles indicate the locations of wild segment loci associated with the 100-seed weight trait; B is a comparison of mature seed size between parents NN1138-2 and CSSL130, with a scale bar of 1 cm; C shows a significant difference in the 100-seed weight of mature seeds between the secondary population parents NN1138-2 and CSSL130; D shows the seed length of parents NN1138-2 and CSSL130; E shows the seed width of parents NN1138-2 and CSSL130; F shows the seed thickness of parents NN1138-2 and CSSL130; error bars represent the mean ± standard deviation of three independent replicates, asterisks indicate significant differences between the two parents, and * indicates... p <0.05, ** indicates p <0.01.

[0020] Figure 2 This is a diagram showing the construction of the secondary population and the phenotypic analysis of 100-grain weight in this invention; Figure 2 In the diagram, A represents the roadmap for constructing the soybean 100-seed weight secondary population CL130BDP by crossing CSSL130 with the recurrent parent NN1138-2; dashed wide arrows represent multiple generations, solid wide arrows represent one generation, single arrows represent selected families within the population, open crosses represent hybridization, and circled crosses represent self-pollination; B is the distribution of 100-seed weight among the secondary populations under three environments; the position of the arrows indicates the parental seed weight. 2022DT, 2022WH, and 2024WH represent three independent environments: Dangtu, Anhui in 2022, Wuhe, Anhui in 2023, and Wuhe, Anhui in 2024, respectively.

[0021] Figure 3 For the present invention qSW01.1 Detailed location map of the site; Figure 3 In the middle, A is F 2:3 Preliminary localization of 100-grain heavy loci in a population (485 families) qSW01.1 A diagram showing the relationship between the InD01-28-72 and InD01-29 markers located on chromosome 1; B shows the phenotypes of the linkage group in three environments using the QTL IciMapping software, and the loci are... qSW01.1 The results of locating between InD01-28-58 and InD01-28-72 are shown in Figure C; C represents the fine localization of chromosome 1 using a local high-density genetic linkage map constructed and combined with genotypic and phenotypic data from 8 recombination-exchange individual plants. qSW01.1Locus plot; the numbers under the bars indicate the physical location (Mb) of the marker; white boxes represent the genotype of NN1138-2, black boxes represent the genotype of N24852, and green bars represent heterozygous genotypes; the right side shows the 100-grain weight of the parents and recombinant families; error bars represent the mean ± standard deviation of three independent biological replicates; letters a, b, and c indicate significant differences. p <0.05.

[0022] Figure 4 This is a diagram of candidate gene analysis for the present invention; Figure 4 In the diagram, A shows the expression pattern FPKM of 13 genes in the localization region, with data from https: / / phytozome.jgi.doe.gov / ; B shows the expression pattern of GmSW01.1 in grains, with qRT-PCR analysis of GmSW01.1 expression levels at 28DAF during the mid-stage of grain development in parents NN1138-2 and CSSL130; C represents the gene expression pattern. Glyma.01g051400 Expression patterns in different tissues of the parent; D is... Glyma.01g051700 Expression patterns in different tissues of the parent organism; qRT-PCR analysis GmSW01.1 The expression levels of 14DAF, 28DAF, and 42DAF in different tissues and at different stages of grain development in parental NN1138-2 and CSSL130 soybean were analyzed. UBI3 The gene is an internal reference gene; the error bar represents the mean ± standard deviation of three independent biological replicates; E is a graph of seed size at different developmental stages of parental NN1138-2 and CSSL130; the scale bar is 1 cm; F is the fresh weight of soybean seeds at different developmental stages of parental NN1138-2 and CSSL130 at 14 DAF, 28 DAF, and 42 DAF; G is the fresh weight of soybean seeds at different developmental stages of parental NN1138-2 and CSSL130 at 14 DAF, 28 DAF, and 42 DAF. The length of fresh seeds at 14 DAF and 42 DAF is represented by H; the width of fresh seeds at 14 DAF, 28 DAF, and 42 DAF is represented by H at different developmental stages of soybeans from parents NN1138-2 and CSSL130; the thickness of fresh seeds at 14 DAF, 28 DAF, and 42 DAF is represented by I at different developmental stages of soybeans from parents NN1138-2 and CSSL130; error bars represent the mean ± standard deviation of three independent biological replicates, ns indicates no significant difference, asterisk indicates significant difference, and * indicates... p <0.05, ** indicates p <0.01, T detection.

[0023] Figure 5 This is a diagram illustrating the functional analysis of candidate genes in this invention. Figure 5 In the diagram, A represents the relationship between different materials. GmSW01.1Gene sequence comparison diagram; the blue area in the structure diagram represents the domain region, the numbers above the structural boxes represent the amino acid sequence, - indicates a missing base, and the numbers at the tail indicate the position of the base; B is the GmSW01.1 protein structure model prediction and physicochemical property analysis diagram; the GmSW01.1 protein structure was predicted using the SWISS-Model website, and the changes in the instability index and hydrophobic index of the parental NN1138-2 and CSSL130 protein sequences were predicted using the Expasy-ProtParam website; C shows the changes in the instability index and hydrophobicity index of different species. GmSW01.1 Phylogenetic analysis diagram of transcription factor homologous genes; D is the subcellular localization map of GmSW01.1; GmSW01.1-YFP fusion protein is located on the nucleus of tobacco cells; its yellow fluorescence was observed using a Zeiss LSM780 fluorescence microscope 48 hours after Agrobacterium infection, with empty vector 35S:YFP as a control.

[0024] Figure 6 This is a phenotypic analysis diagram of homozygous knockout plant lines in this invention; Figure 6 In the diagram, A represents the sequence analysis of three soybean knockout lines; Will82 refers to the Williams82 sequence, and four target sequences Target1~Target4 were set based on its sequence. CR#1, CR#45, and CR#73 are T2 generation homozygous lines. The green crossed-out bases indicate the sequences missing compared to the control, and the numbers on the right indicate the number of missing bases. B is a comparison of seed size between the transgenic homozygous soybean lines and Williams82. C is a graph showing the 100-seed weight of cultivated soybeans Williams82, CR#1, CR#45, and CR#73. D is a graph showing seed length. E is a graph showing seed width. F is a graph showing seed thickness. G is a graph showing the number of branches. H is a graph showing plant height. I is a graph showing the number of main stem nodes. The error bars represent the mean ± standard deviation of 6 independent biological replicates. ns indicates no significant difference, an asterisk indicates a significant difference, and * indicates... p <0.05, ** indicates p <0.01, T detection.

[0025] Figure 7 For the present invention GmSW01.1 Figure showing haplotype analysis results in natural populations; Figure 7 In the diagram, A represents the molecular marker band of this subpopulation, the red numbers represent the bands of the two parents, "1" represents the NN1138-2 genotype band, "2" represents the CSSL130 genotype band, and the white numbers represent the genotypes in the population. B represents the single-marker analysis of the 100-grain weight of this secondary population, Hap NN Hap CSSLThe two haplotypes of the parents are represented, and n is the number of individuals of each haplotype. The error bar represents the mean ± standard deviation of three independent biological replicates. ns indicates no significant difference, an asterisk indicates a significant difference, and * indicates... p <0.05, ** indicates p <0.01, T detection; C represents the molecular marker bands of the Sichuan-Chongqing materials, the red numbers represent the bands of the two parents, "1" represents the NN1138-2 genotype band, "2" represents the CSSL130 genotype band, and the white numbers represent the genotypes in the population; D represents the single-marker analysis of the 100-grain weight of the Sichuan-Chongqing materials, Hap NN Hap CSSL The two haplotypes of the parents are represented, and n is the number of individuals of each haplotype. The error bar represents the mean ± standard deviation of three independent biological replicates. ns indicates no significant difference, an asterisk indicates a significant difference, and * indicates... p <0.05, ** indicates p <0.01, T detection. Detailed Implementation

[0026] To facilitate understanding of the present invention, a more comprehensive description is provided below, along with preferred embodiments. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the present invention.

[0027] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used in this invention and in its specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.

[0028] The beneficial effects of the present invention will be illustrated below through specific embodiments.

[0029] Example 1: Constructing a graphing secondary population In the early stages of this invention, wild soybean N24852 and cultivated soybean Nannong 1138-2 (NN1138-2 for short) were used as parents to construct a chromosome segment substitution line population, SojaCSSP1, covering the entire wild soybean genome. Detection revealed a locus on chromosome 1 of the CSSL130 family within this population that is significantly associated with the 100-seed weight trait. qSW01.1, Furthermore, the CSSL130 family showed significant differences in seed size and grain weight compared to its parent NN1138-2, such as... Figure 1 As shown.

[0030] To precisely locate the loci of the 100-grain weight trait qSW01.1A secondary population containing 485 families was constructed by crossing the family line CSSL130 with the parent NN1138-2. Figure 2 As shown in A, statistical analysis of the 100-grain weight trait in the secondary population under different environments showed that it exhibited an approximately normal distribution and a unimodal distribution, as indicated by Figure 1. Figure 2 As shown in B, this indicates that the 100-seed weight trait of soybean is a quantitative trait regulated by multiple genes. Further analysis of the 100-seed weight phenotypic variation under three environments revealed significant differences in the 100-seed weight trait between parents and the population under different environments, with small coefficients of variation and high broad-sense heritability, as shown in Table 1. This indicates that this secondary population is suitable for genetic analysis of the 100-seed weight trait.

[0031] Table 1. Statistical analysis of 100-grain weight of parent and secondary populations under three environments. Note: "-" indicates that this item is not available.

[0032] Example 2: Fine-tuning of 100 key grain sites qSW01.1 To determine the 100-grain weight site qSW01.1 This invention designs different Indel markers based on the genomic sequences of chromosome segment substitution line populations N24852 and NN1138-2 and family CSSL130 to perform single-marker analysis on different substitution segments, and identifies 100-grain heavy loci. qSW01.1 Preliminary localization is between InD01-28-72 and InD01-29 on chromosome 1, such as Figure 3 As shown in A in the diagram. To narrow down the localization interval, genotypic analysis of the population pedigree was first performed using polymorphic Indel markers designed for the localization interval. Then, linkage analysis was conducted using IciMapping combined with 100-grain weight phenotypic data from the secondary population under three different environments. The mapping results showed that the localization results were consistent across the three different environments, with the peak located between IND01-28-58 and IND01-28-72, with a physical distance of approximately 1350 kb. Figure 3 As shown in B in the diagram. To further shorten the localization interval, this invention first selects heterozygous individuals in the IND01-28-58 and IND01-28-72 intervals to construct a segregating population of approximately 4500 F3 generations. Families exhibiting recombination within the intervals are then screened. Next, a locally high-density genetic map is constructed using designed Indel markers for the intervals, and the genotypes of the recombinant families are identified. Finally, combining the genotypes and phenotypes of the recombinant families and the parental materials NN1138-2 and CSSL130, the intervals are ultimately anchored between IND01-28-605 and IND01-28-63, with a physical distance of approximately 310 kb, containing 13 ORFs, as shown in the diagram. Figure 3 As shown in C.

[0033] Example 3: Candidate gene analysis of localization intervals To identify key candidate genes, this invention first analyzed the expression patterns of 13 candidate genes within a given interval based on data from the Phytozome webpage, and found that only... Glyma.01g051400 , Glyma.01g051500 , Glyma.01g051700 , Glyma.01g051800 , Glyma.01g052200 , Glyma.01g052300 and Glyma.01g052400 It is expressed in the grain, while the other genes are not expressed, such as Figure 4 As shown in A, these 7 genes were initially selected as candidate genes for soybean 100-seed weight. To further confirm the candidate genes, RNA was extracted from seeds of parental materials NN1138-2 and CSSL130, and the relative expression levels of the 7 genes in the parents were analyzed. The results showed... Glyma.01g051400 and Glyma.01g051700 There were significant differences in the relative expression levels among the parent seeds, such as Figure 4 As shown in Figure B; subsequently, the expression patterns of these two genes in the leaves, flowers, apical meristems, and grains at 14, 28, and 42 days after flowering of the parental materials were analyzed. Significant differences were found in the relative expression levels of both genes in different tissues and at different stages of grain development, as shown in the figure. Figure 4 As shown in C and D, the expression patterns of the two genes at different stages of grain development are consistent with the phenotypic changes of fresh grains, such as... Figure 4 As shown in E~I in the diagram. Therefore, Glyma.01g051400 and Glyma.01g051700 It has been identified as a key candidate gene.

[0034] To further identify key candidate genes, specific primers were designed to amplify the genes from the parental genes NN1138-2 and CSSL130. Glyma.01g051400 and Glyma.01g051700 CDS sequences were found to contain only Glyma.01g051700 The CDS sequence has a 3-base (-AAC) deletion between the two parents, encoding one asparagine (N), such as Figure 5 As shown in A. However, this deletion is not located in the domain, so it does not affect the protein's structure, but only changes the protein's instability coefficient; it also changes the protein's hydrophobicity, such as... Figure 5 As shown in B in the diagram. Therefore Glyma.01g051700 Identified as a 100-grain heavy site qSW01.1 The key candidate gene was named GmSW01.1 . Glyma.01g051700 This gene encodes 317 amino acids and belongs to the MYB transcription factor family. Evolutionary analysis across different species revealed that its evolutionary differentiation is synchronous with the taxonomic evolution of monocotyledons and dicotyledons in plants, and that the gene exhibits strong conservation in legume dicotyledons, such as... Figure 5As shown in C. Subcellular localization results indicate that it expresses this only in the cell nucleus, such as... Figure 5 As shown in D, the results are consistent with the predicted characteristics of transcription factors.

[0035] GmSW01.1 The CDS sequence of the gene is shown in SEQ ID NO.1, and is as follows: ATGGGAAGGATGCCATGCTGTGAGAAGGGAGGGTTGAAGAAAGGACTATGGACACCCGAAGAAGATAAGAAGCTCGTTGCTTATGTTGAAAAGCATGGCCATGGAAACTGGCGTTCAGTGCCTGACAAAGCAGGTCTTGAAAGATGTGGAAAGAGTTGCAGATTGAGGTGGATTAACTACCTCAAGCCAGATATAAAACGAGGAAACTTCAGCATGGAGGAAGACCACACCATTATTCAACTTCATGCTCTTCTGGGAAACAAATGGTCAATCATAGCAGCTCACTTGCCCAGGAGAACAGATAATGAGATCAAGAACTACTGGAACACCAACGTCAAGAAAAGACTTATCAGAATGGGCTTAGATCCCGTTACTCACAAACCAATAAAACCCAATACGTTTGAACGCTATGGTGGTGGCCATGGCCAGTTCAAGAACACCATCAATACCAACCACGTGGCTCAATGGGAGAGTGCTCGAATGGAAGCCGAAGCAAGAGGATCCGTGTTGCAAGTTGGATCTCACTCCTCACATCAACCTCAGCTAGTCTTAAGCAAAATCCCAACTCAACCTTGTCCTTCATCATCAGATTCAGTATCAACCAAACACAACACAGTGTATAACATGTATGCCCTCGTGCTTGCCACAAATCATGACCCTTTATGGCCAGTATCCCCGTTGAGCATCCCCGGTTGGAAGGTTCCTGCAGTCTCTACCAATGTTGGACAATTCACCAACACGGGAAGCTCACTCTCTTATGAGAGTGATGTTAACGTCACCGAAACTAATAGTCAAATTCAAAAGATACCAGAAAATTACTTGTCAAACTTACAAGATGAAGATATTATGGTGGCTGTGGAAGCCTTTAGAACAATAAGATGTGAGAGTATTCTAGAATTGTTTAGGGGGTCGAATGACATGGAAGGCCTCAATAAACAATCATTCTTTTGA。

[0036] GmSW01.1The amino acid sequence of the protein expressed by the gene is shown in SEQ ID NO.2. MGRMPCCEKGGLKKGLWTPEEDKKLVAYVEKHGHGNWRSVPDKAGLERCGKSCRLRWINYLKPDIKRGNFSMEEDHTIIQLHALLGNKWSIIAAHLPRRTDNEIKNYWNTNVKKRLIRMGLDPVTHKPIKPNTFERYGGGHGQFKNTINTNHVAQWES ARMEAEARGSVLQVGSHSSHQPQLVLSKIPTQPCPSSSDSVSTKHNTVYNMYALVLATNHDPLWPVSPLSIPGWKVPAVSTNVGQFTNTGSSLSYESDVNVTETNSQIQKIPENYLSNLQDEDIMVAVEAFRTIRCESILELFRGSNDMEGLNKQSFF.

[0037] Example 4 GmSW01.1 Gene function verification To verify GmSW01.1 Regarding gene function, this invention first utilizes CRISP-Cas9 technology for functional verification. According to... GmSW01.1 Four knockout target sequences were designed, as shown in SEQ ID NO.3~SEQ ID NO.6. Using Williams 82 as the recipient material, gene knockout of these four knockout target sequences yielded three different knockout type lines: CR#1, CR#45, and CR#73, as shown below. Figure 6 As shown in A; analysis of yield-related traits in T2 homozygous lines revealed significant changes in grain size compared to wild-type Williams 82, such as... Figure 6 As shown in B, the 100-grain weight decreased by 7.98%–8.06% compared to the control, and the grain length, width, and thickness were all significantly reduced. Figure 6 As shown in C~F, the number of branches also decreases significantly, as... Figure 6 As shown in G; however, plant height and the number of nodes on the main stem did not change significantly compared to the control, such as Figure 6 The values ​​of H and I are shown in the figure. Based on the above results, it is determined that... GmSW01.1 Genes are key genes that regulate the 100-seed weight trait in soybeans.

[0038] The nucleotide sequence of target sequence 1 is shown in SEQ ID NO.3, which is: ATACGTTTGAACGCTATGGT.

[0039] The nucleotide sequence of target sequence 2 is shown in SEQ ID NO.4, which is: TCAATACCAACCACGTGGCT.

[0040] The nucleotide sequence of target sequence 3 is shown in SEQ ID NO.5, which is: TCATGGGAGAGTGCTCGAA.

[0041] The nucleotide sequence of target sequence 4 is shown in SEQ ID NO.6, which is: AGAGGATCCGTGTTGCAAGT.

[0042] Example 5: Development and Utilization of Molecular Markers This invention uses the Williams 82 genome sequence as a reference, CSSL130 as the wild-type substitution family, and NN1138-2 as the cultivated soybean sequence. Sequence alignment showed that the CDS sequence of NN1138-2, compared to the CDS sequence of the wild-type substitution line CSSL130, contains an InDel molecular marker with a -AAC triplet deletion, such as... Figure 5 As shown in A; The reference genome for soybean Williams 82 is Glycine max Wm82.a2.v1 (Phytozomegenome ID:275• NCBI taxonomy ID:3847).

[0043] The InDel molecular marker is located on the forward strand of soybean Williams 82 chromosome 1 between Chr01:6252994..6254510. The nucleotide sequence of the InDel molecular marker is shown in SEQ ID NO.7 and SEQ ID NO.8, with an AAC triplet deletion at position 118bp~120bp.

[0044] SEQ ID NO. 7: AGAGGATCCGTGTTGCAAGTTGGATCTCACTCCTCACATCAACCTCAGCTAGTCTTAAGCAAAATCCCAACTCAACCTTGTCCTTCATCATCAGATTCAGTATCAACCAAACACAACACAGTGTATAACATGTATGCCCTCGTGCTTGCCACAAATCATGACCCTTTATGGCCAGTATCCC.

[0045] SEQ ID NO. 8: AGAGGATCCGTGTTGCAAGTTGGATCTCACTCCTCACATCAACCTCAGCTAGTCTTAAGCAAAATCCCAACTCAACCTTGTCCTTCATCATCAGATTCAGTATCAACCAAACACAACACAGTGTATAACATGTATGCCCTCGTGCTTGCCACAAATCATGACCCTTTATGGCCAGTATCCC.

[0046] To analyze the application of the InDel molecular marker in soybean 100-seed weight breeding, this invention designs specific InDel primers to amplify the InDel molecular marker sequence based on the sequence variation between two parents and their upstream and downstream sequences. The specific InDel primer sequences are shown in SEQ ID NO.9 and SEQ ID NO.10.

[0047] Forward primer SEQ ID NO.9: 5'-AGAGGATCCGTGTTGCAAGTT-3'.

[0048] The reverse primer SEQ ID NO.10: 5'-GGGATACTGGCCATAAAGGGT-3' amplifies the target fragment containing the missing region.

[0049] The PCR reaction system and amplification procedure are as follows: The PCR reaction system uses a 10μL system specification, and then liquid paraffin is added instead of sealing film to prevent the system from evaporating; the reaction system consists of 3μL DNA template, 10μM 1.5μL forward primer, 10μM 1.5μL reverse primer and 4μL 2×taqMasterMix.

[0050] Place the sample into the PCR instrument and set the corresponding amplification program as shown in Table 2 below: Table 2 PCR amplification program Note: "-" indicates that this item is not available.

[0051] Add 2.5 μL of bromophenol blue solution to the PCR amplification product, centrifuge for 1 min, and store at 4°C for later use.

[0052] Genotyping was performed using 8% polyacrylamide gel electrophoresis.

[0053] ① First, clean the two glass plates for making the adhesive to ensure they are clean and free of impurities. Then, place them on a glass plate rack to air dry naturally for later use in making the adhesive.

[0054] ② Snap the glass plates with rough edges together with the smooth glass plates, and then use 1mm adhesive strips to fix them in place.

[0055] ③ Take a 1% agarose solution after heating and boiling, and carefully apply it to the joint between the adhesive strip and the glass plate to seal it. Wait for it to solidify at room temperature for about 10 minutes. At the same time, prepare the matching sample comb for subsequent operations.

[0056] ④ Measure 26.6 mL of 30% polyacrylamide gel and 20 mL of 5×TBE, then add 900 μL of 10% ammonium persulfate (AP) and 45 μL of tetramethylethylenediamine (TEMED) in sequence. After stirring thoroughly, slowly and continuously pour the mixture into a glass trough until it is full. Quickly insert a comb with the teeth facing outwards into the gel and let it stand at room temperature for 30 minutes until the gel is completely solidified.

[0057] ⑤ Add 2 μL of the sample to be tested to the gel well, and add 2 μL of marker to the first well. Use 1×TBE electrolyte for electrophoresis. According to the specific size of the sample fragment, set the voltage and current parameters for electrophoresis appropriately, and control the electrophoresis time. Run the electrophoresis at 250V and 250A for 45 minutes.

[0058] ⑥ Before the electrophoresis is completed, prepare the silver nitrate staining solution in advance. This solution can be reused. Weigh 2.5g of silver nitrate, pour it into 2L of distilled water, and stir thoroughly until completely dissolved.

[0059] ⑦ After electrophoresis, carefully separate the gel from the glass plate with a blade and scrape it into the silver nitrate solution, being careful not to break the gel. Place the container containing the gel on a shaker and oscillate at 50 rpm for about 12 minutes.

[0060] ⑧ Preparation of colorimetric solution: In a fume hood, weigh 40g of sodium hydroxide solid, add 2L of distilled water and stir to dissolve, then add 6mL of formaldehyde solution and mix thoroughly.

[0061] ⑨ Rinse the stained gel three to four times with distilled water, then place it on a shaker at 50 rpm and add an appropriate amount of developing solution. Shake for 10 minutes until clear bands appear on the gel. Finally, rinse the gel thoroughly with water to complete the gel preparation process and store it properly for later use. After completing the above steps, read the experimental data and photograph the gel pattern for future reference.

[0062] Haplotype determination based on electrophoretic band pattern: The alleles with the -AAC triple deletion in the amplified fragment sequence were named Hap. NNIt is consistent with cultivated soybean NN1138-2; the allele lacking this deletion is named Hap. CSSL Consistent with the wild replacement line CSSL130; the heterozygous type is named Hap. NN / CSSL ,like Figure 7 As shown in A in the diagram.

[0063] Because the deletion of this three bases leads to differences in the length of PCR amplified fragments, genotyping can be performed using electrophoretic banding. InDel molecular markers are used to label the banding patterns with 1, 2, 3, and "-", respectively, with the specific criteria as follows:

[0064] Deletion-type Hap with banding consistent with NN1138-2 NN Marked with 1, the amplified fragment size is 181bp; Insert-type Hap with CSSL130 tape design CSSL The fragment size was 184 bp, labeled with 2. Hybrid band type and Hap NN / CSSL Two bands appear, marked with 3; Missing bands are denoted as "-".

[0065] The phenotype of soybean 100-seed weight in the secondary population was analyzed. InDel molecular marker analysis was performed on the 100-seed weight data of the two haplotypes and the population. Figure 7 As shown in B, the results indicate that the secondary population contains [a certain element] under all three environments. Hap NN The 100-grain weight of families with sequence sequences was significantly higher than that of families containing heterozygous fragments. Hap NN / CSSL and Hap CSSL The family lineage weighs 100 grains, and contains heterozygous fragments. Hap NN / CSSL and Hap CSSL There was no significant difference in 100-grain weight among family lines, indicating that it originated from wild bean variety N24852. Glyma.01g051700 The allele is a dominant gene, while the allele from cultivated bean NN1138-2 is a recessive gene.

[0066] Seventy-seven samples from natural populations in Chongqing and Sichuan, with significant differences in 100-grain weight, were selected as samples. Allelic variations of this gene in the natural populations of Chongqing and Sichuan were analyzed, such as... Figure 7 As shown in C and D, allelic variations were found. Hap NN The 100-seed weight of soybean materials was significantly higher than that of materials containing allelic variation. Hap CSSL The material, this indicates Hap NNIt may be an excellent haplotype for regulating soybean grain weight, providing theoretical guidance and genetic resources for molecular improvement of soybean yield traits and breeding of high-yield and high-quality varieties.

[0067] It should be noted that when numerical ranges are mentioned in the claims of this invention, it should be understood that the two endpoints of each numerical range and any value between the two endpoints can be selected. To avoid redundancy, the present invention describes preferred embodiments.

[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0069] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of this invention should be determined by the appended claims.

Claims

1. Use of an InDel molecular marker associated with soybean hundred kernel weight, characterized in that, The InDel molecular marker is located in GmSW01.1 The nucleotide sequence of the gene is shown as SEQ ID NO. 7, and there is an AAC three-base deletion at the site of 118bp~120bp. The application refers to any one of the following (1) and (2): (1) Identify the 100-seed weight of soybeans; plants with AAC triple base deletion at position 118bp~120bp of SEQ ID NO.7 have a higher 100-seed weight than those without deletion; (2) Improve the 100-seed weight of soybean offspring; select soybean samples with AAC tribase deletion at position 118bp~120bp of SEQ ID NO.7 as parents for breeding to improve the 100-seed weight of soybean offspring.

2. Use according to claim 1, characterized in that, The GmSW01.1 The nucleotide sequence of the gene is shown in SEQ ID NO.

1.

3. Use according to claim 1, characterized in that, The method for determining the weight of 100 soybeans is as follows: Genomic DNA was extracted from the soybeans to be tested; PCR amplification was performed using a primer set to obtain the amplified fragment; the primer set consisted of a forward primer and a reverse primer. Electrophoretic analysis of the amplified fragments showed that the weight of 100 soybeans with an amplified fragment size of 181 bp was greater than that of soybeans with an amplified fragment size of 184 bp.

4. Use according to claim 1, characterized in that, The following methods can be used to increase the 100-grain weight of soybean offspring: Genomic DNA was extracted from the soybeans to be tested; PCR amplification was performed using a primer set to obtain the amplified fragment; the primer set consisted of a forward primer and a reverse primer. Electrophoresis was performed on the amplified fragments, and soybean samples with an amplified fragment size of 181 bp were selected as parents for breeding to improve the 100-seed weight of soybean offspring.

5. Use according to claim 3 or claim 4, characterised in that, The nucleotide sequence of the forward primer is shown in SEQ ID NO.9, and the nucleotide sequence of the reverse primer is shown in SEQ ID NO.

10.

6. Use according to claim 3 or claim 4, characterised in that, The amplification program was set as follows: pre-denaturation: 95℃ for 3 min; denaturation: 95℃ for 10 sec, annealing: 55℃ for 10 sec, extension: 72℃ for 30 sec, for a total of 30-35 cycles; complete extension: 72℃ for 5 min; storage: 4℃.

7. Use according to claim 3 or claim 4, characterised in that, The amplification system consisted of 3 μL DNA template, 1.5 μL forward primer, 1.5 μL reverse primer, and 4 μL 2×taq MasterMix.

8. Use according to claim 3 or claim 4, characterised in that, The concentrations of both the forward and reverse primers were 10 μM.