Application of wild soybean MADS-box family gene GsAGL62 in soybean quality traits

By overexpressing the wild soybean MADS-box family gene GsAGL62, the problem of insufficient regulation of seed protein content in existing technologies was solved, and the effect of increasing the protein content of soybean seeds was achieved without reducing the fat content.

CN120989092APending Publication Date: 2025-11-21NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510986493.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, the superior wild soybean allele GsAGL62 has been rarely used to regulate soybean seed protein content, and overexpression may affect seed fat content.

Method used

By overexpressing the wild soybean MADS-box family gene GsAGL62, plant expression vectors were constructed using enhanced or inducible promoters to transform plant cells and increase seed protein content. At the same time, transformed plants were screened by selective marker genes or phenotypic traits to ensure that fat content was not reduced.

Benefits of technology

It significantly increases the protein content of soybean seeds without affecting or increasing the fat content, providing a high-protein genetic engineering application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120989092A_ABST
    Figure CN120989092A_ABST
Patent Text Reader

Abstract

The invention discloses an application of a gene GsAGL62 of a wild soybean MADS-box family. The nucleotide sequence of the wild soybean GsAGL62 protein coding gene GsAGL62 is as shown in SEQ ID NO. 1. A constructed plant overexpression vector pBA002-GsAGL62 is expressed in a soybean receptor JACK, and it is found that a transgenic plant is changed in the aspects of quality characters, and the protein content is increased. The gene can be introduced into a plant as a target gene, and the transgenic plant affects the protein content of seeds and improves the quality character through the expression of the GsAGL62 gene. Therefore, the wild soybean GsAGL62 protein coding gene GsAGL62 provided by the invention can be applied to improvement of quality characters and increase of protein content through genetic engineering.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of plant genetic engineering, and particularly relates to application of a wild soybean MADS-box protein family gene GsAGL62 or a coding protein thereof in regulation of soybean quality traits. BACKGROUND

[0002] Soybean (Glycine Max (L.) Merr.) is rich in protein, fat and isoflavones and other beneficial substances, and is a high-quality plant protein and oil source. With the improvement of people's living standards and consumption level, breeding of soybean varieties with high nutritional value has become one of the important breeding goals of breeders. Wild soybean (Glycine soja Siebold & Zucc.) is a close ancestor of cultivated soybean, and has excellent high-protein characteristics compared with cultivated soybean. Mining excellent genes related to protein content in wild soybean will help the breeding of high-protein varieties and enrich the genetic resources of soybean.

[0003] Diers et al. (1992) first applied the excellent allelic variation of protein content in wild soybean to QTL genetic analysis. They used wild soybean PI 468916 and cultivated soybean breeding line A81-356022 to construct an F2 mapping population, and constructed a linkage map consisting of 252 RFLP markers. Eight wild fragments associated with seed protein content were detected on linkage groups B2, E, G, L, and I, which could explain 12-42% of the phenotypic variation. Among them, four QTLs were detected on linkage group I with the largest effect. Nichols et al. (2006) further fine-mapped the locus on linkage group I using the soybean linkage map published by Song et al. (2004), and finally located the locus within a 3 cM range between SSR marker Satt239 and AFLP marker ACG9b. In recent years, with the rapid development of high-throughput sequencing, many soybean seed protein content-related QTLs and functional genes have been located using whole-genome association analysis and gene chip technology. Bolon et al. (2010) used wild soybean and cultivated soybean to construct a set of NILs population for the large-effect QTL locus on soybean Chr20, and finally located the QTL within a 8.4 Mb range between markers Sat_174 and ssrpqtl_38 based on BAC sequencing map. Thirteen candidate genes within the interval were detected to be significantly different between NILs by Soy GeneChip. Wang et al. (2020) studied the effect of GmSWEET10a, a homolog of SWEET, on soybean seed quality and quantity during domestication. The variation of GmSWEET10a during soybean domestication led to changes in soybean seed size, oil content, and protein content. Duan et al. (2022) found that GmST05, a homolog of TFL1 and FT (MFT) proteins, may affect the protein and oil content of soybean by regulating the transcription of GmSWEET10a. Although hundreds of QTLs associated with soybean protein content have been reported on the soybean database Soybase, only about 26 QTLs come from wild soybean, and many excellent alleles of wild soybean still need to be further explored.

[0004] The MADS-box protein family genes are involved in encoding various transcription factors related to biological functions in eukaryotes, including flowering time control, meristem, flower organ characteristics, fruit ripening, embryonic development, and development of vegetative organs such as roots and leaves (Gramzow et al., 2010; Yu et al., 2014). MADS-box proteins are usually composed of a conserved MADS-box domain and K-, C- and I- domains, and can be divided into type I and type II according to the sequence homology of the domain. GsAGL62 mentioned in the present application belongs to type II subgroup, and the MIKC type subgroup plays a key role in regulating genes necessary for the vegetative and reproductive stages of plants (Zhang et al., 2024). For example, Arabidopsis MADS-box gene AGL21 promotes root elongation, while SOC1 and AGL24 act as inhibitors of root growth and show redundant effects (Castañón-Suárez et al., 2024). In rice, OsMADS50 activates flowering-related genes such as OsMADS14, while OsMADS51 acts as a flowering activator and transmits the signal of OsGI to Ehd1 under short-day conditions to regulate flowering time (Kim et al., 2007). GmAGL20 (SOC1a) can promote flowering of soybean under both long-day and short-day conditions and affect stem node number and yield (Kou et al., 2022). These studies all show that MADS-box protein family genes play an important role in plant growth and development, but there are few reports on their role in regulating seed quality in crops. SUMMARY

[0005] The purpose of the present application is to disclose a wild soybean gene GsAGL62.

[0006] Another purpose of the present application is to provide the genetic engineering application of the gene in regulating soybean protein content.

[0007] The purpose of the present application can be achieved by the following technical solutions:

[0008] The wild soybean MADS-box family gene GsAGL62 has a nucleotide sequence of SEQ ID NO. 1.

[0009] The protein encoded by the wild soybean MADS-box family gene GsAGL62 has an amino acid sequence of SEQ ID NO. 2.

[0010] Application of wild soybean GsAGL62 protein coding gene in improving seed protein content.

[0011] As a preferred embodiment of the present application, overexpression of the wild soybean GsAGL62 protein-encoding gene can increase the protein content of soybean seeds.

[0012] As a preferred embodiment of the present application, overexpression of the wild soybean GsAGL62 protein-encoding gene can increase the protein content of soybean seeds.

[0013] An expression vector containing the wild soybean MADS-box family gene GsAGL62 according to the present application.

[0014] When a plant expression vector is constructed using GsAGL62, any one of an enhanced promoter or an inducible promoter can be added before the transcription initiation nucleotide. In order to facilitate identification and screening of transgenic plant cells or plants, the plant expression vector used can be processed, such as adding a selectable marker gene (GUS gene, GFP gene, etc.) or an antibiotic marker (gentamicin marker, kanamycin marker, hygromycin marker, etc.) that can be expressed in plants. In consideration of the safety of transgenic plants, no selectable marker gene can be added, and the transformed plants can be directly screened by phenotypic traits.

[0015] The wild soybean MADS-box family gene GsAGL62 according to the present application is used in genetic engineering for increasing the protein content of soybeans.

[0016] The transgenic soybeans overexpressing GsAGL62 according to the present application have a significantly increased soybean protein content.

[0017] The plant expression vector carrying GsAGL62 according to the present application can be used to transform plant cells or tissues by using a Ti plasmid, a Ri plasmid, a plant virus vector, direct DNA transformation, microinjection, electroporation, Agrobacterium-mediated transformation, etc. The transformed plant host can be a monocotyledonous plant such as rice, wheat, and corn, or a dicotyledonous plant such as tobacco, Arabidopsis, soybean, rape, cucumber, tomato, poplar, turf grass, and alfalfa.

[0018] The recombinant expression vector according to the present application is used in increasing the protein content of soybean seeds.

[0019] As a preferred embodiment of the present application, the recombinant expression vector is used in increasing the protein content of soybean seeds without reducing the fat content of soybean seeds.

[0020] Beneficial effects:

[0021] The GsAGL62 in the application belongs to the MADS-box family and contains a MADS_MEF2 type domain. The GsAGL62 gene is mainly expressed in soybean leaves and seeds, and the subcellular localization result shows that the GsAGL62 protein is located in the nucleus. The GsAGL62 of the application is introduced into a receptor material Jack by using a plant overexpression vector pBA002-GsAGL62, and transgenic plants are obtained. Compared with the receptor material Jack, overexpression of the GsAGL62 gene can increase the protein content of soybean. The application discloses that the GsAGL62 as a target gene is overexpressed by a genetic engineering method to change the protein content of soybean and improve the seed quality traits of soybean. BRIEF DESCRIPTION OF DRAWINGS

[0022] The application is further described below in combination with the drawings and examples.

[0023] Figure 1 . Agarose gel electrophoresis diagram of GsAGL62 after PCR cloning. The target fragment size is 522 bp. Marker: DL2000

[0024] Figure 2 . Tissue expression pattern of GsAGL62. N=3.

[0025] Figure 3 . Subcellular localization of GsAGL62.

[0026] Figure 4. Positive identification of T3 generation overexpression transgenic lines and relative expression level of GsAGL62. OE-2, OE-3, OE-4: 3 overexpression GsAGL62 transgenic lines. Jack: recipient material, control. (A) PCR detection of the target gene in T3 generation transgenic soybean lines. Partial vector sequence and GsAGL62 CDS sequence were amplified using specific detection primers. M: Marker, DL2000. Lane 1: plasmid, positive control. Lanes 2, 4, 5: OE-2, N=3. Lane 3: Jack, negative control. Lanes 6-8: OE-3, N=3. Lanes 9-11: OE-4, N=3. Lane 12: water, blank control. (B) PCR detection of the Bar gene in T3 generation transgenic soybean lines. Partial vector sequence was amplified using specific detection primers. M: Marker, DL2000. Lane 1: Jack, negative control. Lanes 2-4: OE-2, N=3. Lanes 5-7: OE-3, N=3. Lanes 8-10: OE-4, N=3. Lane 11: plasmid, positive control. Lane 12: water, blank control. (C) Relative expression level is the relative expression of the control (relative expression level in Jack = 1) after standardization by tubulin gene. *: P <0.05 level is significant. **: P <0.01 level is significant. Error bars represent ±SD.

[0027] Figure 5 . Statistical analysis of seed protein content and oil content of transgenic soybean and recipient material Jack. (A) Statistical analysis of seed protein content after harvest. **: P <0.01 level is relatively significant. Error bars represent ±SD. (B) Statistical analysis of seed oil content after harvest. ns: P>0.05 level is not significant. *: P <0.05 level is significant. Error bars represent ±SD. DETAILED DESCRIPTION

[0028] The present application will be further described in details below with reference to the accompanying drawings and examples, and further with reference to the data. These examples are only for illustrating the present application, and do not limit the scope of the present application in any way. In the following examples, various processes and methods that are not described in details are conventional methods known in the art. The primers used are all indicated at the first occurrence, and the same primers used thereafter are all the same as indicated at the first occurrence.

[0029] Example 1 Cloning and identification of wild soybean GsAGL62 and its encoding gene

[0030] GsAGL62-F: ATGCCAGACTTGAACGGTGTCG; (SEQ ID NO. 3), GsAGL62-R: TCAGTTGAGGTTCCCACCTTTT. (SEQ ID NO. 4). The GsAGL62 gene was amplified from total RNA of soybean seed organs by RT-PCR. The soybean pod tissue was ground with a mortar, added to a 1.5 mL EP tube containing lysis solution, shaken thoroughly, and then transferred to a glass homogenizer. After homogenization, it was transferred to a 1.5 mL EP tube, and total RNA was extracted using a plant total RNA extraction kit (TIANGEN DP404). The quality of total RNA was identified by formaldehyde denaturation gel electrophoresis, and then the RNA content was determined on a spectrophotometer. The obtained total RNA was used as a template, and the first strand of cDNA was synthesized according to the instructions of the reverse transcription kit provided by Takara. PCR amplification reaction was performed. The PCR reaction system was: 2 μl of cDNA (0.05 μg), 2 μl of upstream and downstream primers (10 μM) each, 25 μl of 2x Phanta Max Buffer, 1 μl of dNTP (10 mM), and 1 U of Phanta Max Super-Fidelity DNA polymerase (Vazyme), and the volume was made up to 50 μl with ultrapure water. The PCR program was as follows: performed on a Bio-RAD PTC200 PCR instrument, the program was as follows: 94°C pre-denaturation for 3 min; 94°C denaturation for 15 s, 58°C annealing for 15 s, 72°C extension for 45 s, a total of 30 cycles; then 72°C extension for 5 min to terminate the reaction, and 4°C storage. The PCR product was recovered and cloned into pClone007 vector, and the vector was named T-GsAGL62. After sequencing, the cDNA sequence of soybean gene GsAGL62 with complete coding region was obtained, SEQ ID NO. 1, the full length was 462 bp, and encoded 173 amino acids represented by SEQ ID NO. 2.

[0031] Example 2 Expression characteristics of GsAGL62 in different organs of wild soybean

[0032] The RNA of the root, stem, leaf, flower, pod, and seed of the HAAS_187 material was extracted, and reverse-transcribed into cDNA for RT-PCR analysis. The total RNA was extracted as in Example 1. The soybean constitutive expression gene Tubulin was used as an internal reference gene, and the amplification primers were Tubulin forward primer sequence: GGAGTTCACAGAGGCAGAG (SEQ ID NO. 5), and Tubulin reverse primer sequence: CACTTACGCATCACATAGCA (SEQ ID NO. 6). The cDNA from different tissues or organs of soybean was used as a template for real-time fluorescent quantitative PCR analysis. The amplification primers of GsAGL62 were: GsAGL62-qPCR-F: TCGGGTGACATTCTCGAAGC (SEQ ID NO. 7), and GsAGL62-qPCR-R: ACATCCACGTCACAAAGGGT (SEQ ID NO. 8). The results (Figure 2) showed that the expression of GsAGL62 in leaves and seeds was relatively high, indicating that GsAGL62 might be related to soybean seed development.

[0033] Example 3 Subcellular localization of GsAGL62

[0034] The subcellular localization was performed by the method of Nicotiana benthamiana transient expression, and the vector used was P2, and the primers were GsAGL62-P2-F: ACAAATCTATCTCTCTCGAGATGCCAGACTTGAACGGTGTCG (SEQ ID NO. 9), and GsAGL62-P2-R: GCTCACCATGGATCCGTTGAGGTTCCCACCTTTT (SEQ ID NO. 10). After the target band was correctly amplified by PCR, it was recovered by gel cutting, and the gel recovery product was connected to the vector by homologous recombination to construct the subcellular localization vector P2-GsAGL62 (the gene at the N-terminal of GFP). After the transient expression of tobacco was expressed and dark culture for 48 h, green fluorescent signals were generated after laser irradiation by a laser confocal microscope (Zeiss, LSM780), the protein was located, and observation and photography were performed. The results are shown in Figure 3, and the GsAGL62:GFP fusion protein was distributed in the nucleus, indicating that GsAGL62 might function in the nucleus.

[0035] Example 4 Genetic engineering of GsAGL62

[0036] The vector is constructed by double enzyme digestion method, which is divided into two steps of enzyme digestion reaction and recombination reaction. The primer sequence is: GsAGL62-pBA002-F: CGCGCCGGGCCCAGGCCTACGCGTATGCCAGACTTGAACG (SEQ ID NO. 11); GsAGL62-pBA002-R: ATCGGGGAAATTCGAGCTCTCAGTTGAGGTTCC (SEQ ID NO. 12). The pBA002 vector is double enzyme-digested by Mlul and SacI, the target fragment is amplified from the T-GsAGL62 vector constructed in Example 1 by adding a linker primer, and then recombined and connected. The product is the expression vector pBA002-GsAGL62. The E. coli TOP10 is transformed, and the transformation liquid is coated on the LB solid medium containing 50 mg / L Kana to screen positive clones. After sequencing verification, the plasmid is extracted to obtain the pBA002-GsAGL62 plant overexpression vector, which is transformed into Agrobacterium tumefaciens strain EHA105 by freeze-thaw method. The pBA002-GsAGL62 is transformed into soybean by Agrobacterium strain EHA105 mediation. The Jack seeds are sterilized with chlorine gas in a fume hood for 6 h, and then the seeds are germinated in SG4 solid medium for 1 night. The germinated seeds are peeled, the seeds are divided into two parts, the lower hypocotyls are removed, and 7-8 wounds are made at the leaf axils. At the same time, the bacterial liquid for infection is prepared, and the Agrobacterium tumefaciens liquid containing the pBA002-GsAGL62 recombinant vector is shaken at 28°C, 200 rpm to OD600 value of 0.8-0.9, centrifuged at 4000 rpm for 10 min, the supernatant is discarded, and then resuspended in CCM solution to OD600 value of about 0.5-0.6. The prepared explants are placed in the bacterial liquid, 28°C, 140-150 rpm, and infected for 30-40 min. Then the explants are taken out and plated on CCM solid medium with the wounds downward, and incubated at 25°C for 4-5 d. Then the explants are washed with sterilized ultrapure water and Wish-Liquid, and then the explants are inserted into SIM medium without glufosinate at 45°C. The explants are cultured at 26°C under light for 15 d to induce buds. After 15 days, the large buds are cut off and replaced into SIM medium with 6 mg / L glufosinate for induction of bud emergence. After two weeks, the cotyledons and yellow leaves are removed and replaced into SEM medium with 4 mg / L glufosinate. Then the subculture is carried out every 15 days, and the dose of glufosinate is gradually reduced. When the buds of the explants grow to about 6 cm, the multiple buds are transferred into rooting medium and cultured for about 10 d to induce roots. When the root system grows well, the explants are transplanted into sterilized mixed nutrient soil and cultured in an artificial incubator (16 h light / 8 h dark, 25°C).

[0037] DNA-level positivity was assessed in the progeny of overexpression-positive seedlings using gene sequence-specific primers F: GCTCTACAAATGCCATCATTGC (SEQ ID NO. 13) and R: TCAGTTTGAGGTTCCCACCT (SEQ ID NO. 14), respectively. Figure 4 A) and the vector sequence-specific primers are Bar-F: AAACCCACGTCATGCCAGTTC (SEQ ID NO.15) and Bar-R: CGAGACAAGCACGGTCAACTT (SEQ ID NO.16). Figure 4 B). Real-time quantitative PCR results showed that, compared with the control, the expression level of the GsAGL62 gene in the transgenic material was significantly increased. Figure 4 C). The primer sequences for quantitative fluorescence and the primer sequences for the internal reference gene Tubulin are the same as in Example 2.

[0038] Phenotypic observation was performed on the T3 generation 35S::GsAGL62 transgenic soybean. Under net-house growth conditions, compared with the control Jack, the seed protein content of the three overexpression lines OE-2, OE-3, and OE-4 of the 35S::GsAGL62 transgenic soybean was significantly increased. Figure 5 A), and the oil content of OE-2 and OE-4 did not decrease significantly ( Figure 5 B). This provides a new and excellent genetic resource for increasing soybean protein content without sacrificing soybean oil content.

Claims

1. Application of wild soybean GsAGL62 protein coding gene in improving seed protein content, wherein the nucleotide sequence of the wild soybean GsAGL62 protein coding gene is SEQ ID NO.

1.

2. Use according to claim 1, characterized in that, Overexpression of the wild soybean GsAGL62 protein coding gene can improve the seed protein content of soybean.

3. Use according to claim 1, characterized in that, Application of overexpression of the wild soybean GsAGL62 protein coding gene in improving the seed protein content of soybean without reducing the seed fat content of soybean.

4. The recombinant expression vector containing the wild soybean GsAGL62 protein coding gene according to claim 1, characterized in that, The wild soybean GsAGL62 protein coding gene in claim 1 is inserted into the MIUI and SacI enzyme cutting sites of pBA002 vector.

5. Application of the recombinant expression vector in claim 4 in improving the seed protein content of soybean.

6. Use according to claim 1, characterized in that, Application of the recombinant expression vector in claim 4 in improving the seed protein content of soybean without reducing the seed fat content of soybean.