Application of soybean GmCDF3 in regulating and controlling contents of seed protein, amino acid and raffinose
By overexpressing the GmCDF3 gene in soybean and Arabidopsis thaliana, the problems of low protein content and high raffinose content in soybean seeds were solved, resulting in increased protein and amino acid content and decreased raffinose content, thus promoting the breeding of high-yield and high-quality soybeans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENTER FOR AGRICULTURAL TECHNOLOGY NORTHEAST INSTITUTE OF GEOGRAPHY & AGROECOLOGY
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, soybean seeds have low protein content, which is negatively correlated with yield, and high raffinose content, which affects the improvement of soybean quality. There is a lack of genes that synergistically improve grain weight, protein, and raffinose, which limits the breeding of high-yield and high-quality soybeans.
By overexpressing the GmCDF3 gene in soybean and Arabidopsis thaliana, and introducing the gene using a recombinant expression vector, the protein and amino acid content of seeds was increased, while the sugar content of raffinose was decreased. Transformation was carried out using methods such as Ti plasmid, Ri plasmid, plant virus vector, direct DNA transformation, microinjection, electroporation, and Agrobacterium-mediated transformation.
It significantly increased the protein content of soybean and Arabidopsis seeds, decreased the sugar content of raffinose, while having no significant effect on oil content and grain weight, providing potential for breeding new varieties of high-protein, high-quality soybeans.
Smart Images

Figure CN121950833A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the application of soybean GmCDF3. Background Technology
[0002] Soybean seeds are rich in protein and have high nutritional value. Their amino acid composition is similar to that of essential amino acids for the human body, especially rich in lysine, which is lacking in grains. Sulfur-containing amino acids participate in metabolism, including hormones, and play important roles in cellular functions such as anti-inflammation, anti-oxidation, and enzyme activity. However, the sulfur-containing amino acid content in soybean seeds is 0.6% lower than the reference content of the Food and Agriculture Organization of the United Nations (FAO), making them limiting amino acids in soybeans. Furthermore, soybean seeds are rich in raffinose, which is indigestible and can cause bloating. Therefore, from a breeding perspective, selecting varieties with low raffinose levels is of great significance (Aguilera et al., 2009). In recent years, with the continuous optimization of people's dietary structure, the demand for soybeans has been increasing. However, the negative correlation between soybean protein and yield means that currently cultivated high-yield soybeans have low protein content, which is detrimental to the cultivation of high-yield and high-quality soybeans. At the same time, protein is a complex trait regulated by multiple genes, and currently only a few genes have been identified. Currently, no genes have been reported that have the potential to synergistically improve important yield and quality traits such as grain weight, protein, and raffinose, which limits the application of molecular breeding in the cultivation of high-yield and high-quality soybeans. Therefore, identifying genes that synergistically increase soybean protein content without affecting grain weight and reduce raffinose content is of great significance for elucidating the molecular regulatory mechanism and providing a theoretical basis for marker-assisted precision breeding of high-protein and high-quality soybeans.
[0003] CDF3 is a DNA-binding one zinc finger (DOF) factor that oscillates during transcription under constant light conditions. DOF proteins play important roles in plant metabolic regulation, seed development, and tissue differentiation (Noguero et al., 2013). Arabidopsis thaliana AtCDF1 is a component of a regulatory network involved in nitrogen response (Alvarez et al., 2020). Its closest homolog, AtCDF3, shows higher yield and improved fruit sugar content when overexpressed in tomato (Renau-Morata et al., 2017). To date, there are no reports of GmCDF3 genes co-regulating yield, protein, amino acid, and raffinose in plants. Summary of the Invention
[0004] This invention provides an application of soybean GmCDF3 in regulating the content of seed protein, amino acids and raffinose.
[0005] Application of a soybean GmCDF3 gene in regulating seed protein, amino acid, and raffinose content. The soybean GmCDF3 gene increases the protein content of soybean and Arabidopsis seeds, as well as the content of cysteine and four essential amino acids in soybean seeds; it decreases the raffinose content in soybean seeds.
[0006] Furthermore, the application of the soybean GmCDF3 gene in regulating seed protein, amino acid, and raffinose content can be achieved by using the soybean GmCDF3 gene to cultivate plant seeds, thereby improving plant quality.
[0007] Further, methods to improve plant quality include: introducing the GmCDF3 gene, which can be expressed, into recipient plants to obtain transgenic plants with increased GmCDF3 gene expression. This is achieved by introducing a recombinant expression vector carrying the GmCDF3 gene into the recipient plant. A method for cultivating plant varieties with improved seed quality involves introducing the GmCDF3 gene, which can be expressed, into recipient plants to obtain transgenic plants with increased GmCDF3 gene expression. Compared to the recipient plant, the transgenic plants exhibit increased seed protein and amino acid content and decreased raffinose content. Specifically, overexpression of the GmCDF3 gene is used to increase the seed protein content of soybeans and Arabidopsis thaliana, increase the content of 12 amino acids in soybean seeds, and decrease raffinose content. Haplotype analysis of the soybean GmCDF3 gene is performed, along with comparative analysis of protein content, oil content, 100-seed weight, and yield for each haplotype, and the identification of the optimal haplotype. Introducing the GmCDF3 gene into recipient plants involves introducing a recombinant expression vector carrying the GmCDF3 gene into the recipient plant. Specifically, this is achieved by transforming plant cells or tissues using conventional biological methods such as Ti plasmids, Ri plasmids, plant virus vectors, direct DNA transformation, microinjection, electroporation, and Agrobacterium-mediated transformation, followed by culturing the transformed plant tissues into plants.
[0008] Furthermore, the recipient plant is a dicotyledonous or monocotyledonous plant; the dicotyledonous plant is a legume or a cruciferous plant; the legume is soybean, and the cruciferous plant is Arabidopsis thaliana.
[0009] Furthermore, the protein encoded by the GmCDF3 gene is any one of the following:
[0010] (A1) A protein with the amino acid sequence shown in SEQ ID NO.1;
[0011] (A2) A protein having the same function as the amino acid sequence shown in SEQ ID NO.1, by substitution and / or deletion and / or addition of one or more amino acid residues;
[0012] Proteins that have 99%, 95%, 90%, 85%, or 80% homology and the same function with any of the amino acid sequences defined in (A3) and (A1)-(A2);
[0013] (A4) A fusion protein obtained by attaching a protein tag to the N-terminus and / or C-terminus of any of the proteins defined in (A1)-(A3).
[0014] Furthermore, the protein tag refers to a polypeptide or protein fused with a target protein using in vitro DNA recombination technology for expression, to facilitate the expression, detection, tracing, and / or purification of the target protein. The protein tag may be a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag.
[0015] Furthermore, the soybean GmCDF3 gene includes any of the following DNA molecules:
[0016] (B1) The DNA molecule shown in SEQ ID NO.2;
[0017] (B2) A DNA molecule that hybridizes under stringent conditions to the DNA molecule defined by the sequence shown in SEQ ID NO.2 and encodes the GmCDF3 protein;
[0018] A DNA molecule encoding the GmCDF3 protein that has 99%, 95%, 90%, 85%, or 80% or more homology with the DNA sequence defined by (B3) and (B1) or (B2).
[0019] This invention demonstrates through experiments that overexpressing the soybean GmCDF3 gene in Arabidopsis thaliana and obtaining three transgenic lines significantly increased the protein content of Arabidopsis seeds by 10.5%-21.6% compared to wild-type Arabidopsis seeds. Simultaneously, overexpressing the GmCDF3 gene in soybean and obtaining three transgenic lines significantly increased the protein content of soybean seeds by 6.49%-8.63% compared to untransformed soybean DN50. Furthermore, cysteine content significantly increased by 6.49%-13.67%, and the four essential amino acids significantly increased by 3.15%-8.67% (leucine), 2.59%-8.19% (lysine), 3.56%-7.77% (phenylalanine), and 4.28%-8.97% (threonine), respectively. Raffinose content significantly decreased by 9.24%-11.50%, while grain weight and oil content remained largely unaffected. This indicates that the GmCDF3 gene can regulate seed protein quality. The GmCDF3 gene can be used to improve soybean protein, amino acids, and raffinose while maintaining soybean yield and oil content. Therefore, GmCDF3 has the potential for application in the breeding of new soybean varieties and the creation of new germplasm.
[0020] This invention utilizes molecular techniques to construct a GmCDF3 overexpression vector. Through Agrobacterium-mediated genetic transformation, overexpression of the soybean GmCDF3 gene in soybean significantly increases the content of 12 amino acids in soybean seeds, including cysteine, leucine, lysine, phenylalanine, and threonine. It also significantly increases seed protein content without affecting oil content or grain weight, and significantly reduces raffinose content. Simultaneously, overexpression of the soybean GmCDF3 gene in Arabidopsis thaliana also significantly increases the protein content of Arabidopsis seeds. This discovery of gene function provides a new approach for studying the regulatory mechanisms of soybean protein and can also accelerate the molecular breeding process for high-yielding and high-quality soybean varieties. Attached Figure Description
[0021] Figure 1 The results are for the PCR amplification products of GmCDF3 in a 1% agarose gel.
[0022] Figure 2 Analysis of the relative expression levels of GmCDF3 in different tissues.
[0023] Figure 3 Subcellular localization results for GmCDF3.
[0024] Figure 4 This is a schematic diagram of the overexpression vector for GmCDF3.
[0025] Figure 5 The results of bar test strip detection and PCR detection for GmCDF3 overexpressing soybeans are shown.
[0026] Figure 6 The results show the transcriptional level of GmCDF3 in soybean overexpression.
[0027] Figure 7 The results are for detecting GmCDF3 overexpressing Arabidopsis thaliana plants.
[0028] Figure 8 The results show the comparison of protein content in GmCDF3-overexpressing soybean seeds and control seeds.
[0029] Figure 9 The results show the comparison of oil content and 100-seed weight in GmCDF3-overexpressing soybeans and control seeds.
[0030] Figure 10 The results show the comparison of protein content in GmCDF3-overexpressing Arabidopsis thaliana seeds with control seeds.
[0031] Figure 11 The results show the comparison of raffinose in GmCDF3-overexpressing soybean seeds and control seeds.
[0032] Figure 12The results show the amino acid composition of GmCDF3-overexpressing soybeans compared to control seeds.
[0033] Figure 13 The results show the comparison of protein content, oil content, 100-grain weight, and yield of the GmCDF3 haplotype. Detailed Implementation
[0034] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0036] Specific implementation method one: This implementation method is applied to the regulation of seed protein, amino acid and raffinose content by the soybean GmCDF3 gene.
[0037] In this embodiment, the soybean GmCDF3 gene increases the protein content of soybean and Arabidopsis seeds, as well as the content of cysteine and four essential amino acids in soybean seeds; and decreases the raffinose content in soybean seeds.
[0038] Specific Implementation Method Two: This implementation method differs from Specific Implementation Method One in that the soybean GmCDF3 gene application method utilizes the GmCDF3 gene to cultivate plant seeds, thereby improving plant varieties. Everything else is the same as in Specific Implementation Method One.
[0039] Specific Implementation Method 3: This implementation method differs from Specific Implementation Method 2 in that the method for improving the plant variety involves introducing a GmCDF3 gene capable of expression into the recipient plant, thereby obtaining a transgenic plant with increased GmCDF3 gene expression. Everything else is the same as in Specific Implementation Method 2.
[0040] Compared to the recipient plant, the transgenic plants introduced with the GmCDF3 gene in this embodiment exhibited increased seed protein and amino acid content, and decreased raffinose content. Specifically, overexpression of the GmCDF3 gene increased the seed protein content of soybean and Arabidopsis thaliana, increased the content of 12 amino acids in soybean seeds, and decreased raffinose content. Haplotype analysis of the soybean GmCDF3 gene was performed, along with comparative analysis of protein content, oil content, 100-seed weight, and yield for each haplotype, and the optimal haplotype was identified.
[0041] Specific Implementation Method Four: This implementation method differs from Specific Implementation Method Three in that the GmCDF3 gene capable of expression is introduced into the recipient plant via a recombinant expression vector carrying the GmCDF3 gene. Everything else is the same as in Specific Implementation Method Three.
[0042] Specifically, this invention involves transforming plant cells or tissues using biological methods such as Ti plasmids, Ri plasmids, plant virus vectors, direct DNA transformation, microinjection, electrocoagulation, and Agrobacterium-mediated transformation, and then cultivating the transformed plant tissues into plants.
[0043] Specific Implementation Method Five: This implementation method differs from Specific Implementation Methods Three or Four in that the recipient plant is a dicotyledonous or monocotyledonous plant. Everything else is the same as in Specific Implementation Methods Three or Four.
[0044] Specific Implementation Method Six: This implementation method differs from Specific Implementation Method Five in that the dicotyledonous plants are leguminous plants and cruciferous plants; the leguminous plant is soybean, and the cruciferous plant is Arabidopsis thaliana. Everything else is the same as in Specific Implementation Method Five.
[0045] Specific Implementation Method Seven: This implementation method differs from Specific Implementation Method One in that the protein encoded by the GmCDF3 gene is any of the following proteins:
[0046] (A1) A protein with the amino acid sequence shown in SEQ ID NO.1;
[0047] (A2) A protein having the same function as the amino acid sequence shown in SEQ ID NO.1, by substitution and / or deletion and / or addition of one or more amino acid residues;
[0048] Proteins that have 99%, 95%, 90%, 85%, or 80% homology and the same function with any of the amino acid sequences defined in (A3) and (A1)-(A2);
[0049] (A4) A fusion protein obtained by attaching a protein tag to the N-terminus and / or C-terminus of any of the proteins defined in (A1)-(A3). The rest is the same as in Specific Embodiment 1.
[0050] In this embodiment, a protein tag refers to a polypeptide or protein that is fused with a target protein using in vitro DNA recombination technology, so as to facilitate the expression, detection, tracing, and / or purification of the target protein.
[0051] Specific Implementation Method Eight: This implementation method differs from Specific Implementation Method Seven in that the protein tag is a Flag tag, His tag, MBP tag, HA tag, myc tag, GST tag, and / or SUMO tag. Everything else is the same as in Specific Implementation Method Seven.
[0052] Specific Implementation Method Nine: This implementation method differs from Specific Implementation Method One in that the soybean GmCDF3 gene includes any of the following DNA molecules:
[0053] (B1) The DNA molecule shown in SEQ ID NO.2;
[0054] (B2) A DNA molecule that hybridizes under stringent conditions to the DNA molecule defined by the sequence shown in SEQ ID NO.2 and encodes the GmCDF3 protein;
[0055] (B3) is a DNA molecule encoding the GmCDF3 protein that has 99%, 95%, 90%, 85%, or 80% homology with the DNA sequence defined by (B1) or (B2). Everything else is the same as in Specific Embodiment 1.
[0056] Example 1: Cloning, gene tissue expression analysis, and subcellular localization of the soybean GmCDF3 gene described in this invention.
[0057] I. Cloning of the soybean GmCDF3 gene
[0058] Primers for amplifying GmCDF3 were designed based on the Phytozome database. The upstream primer sequence is shown in SEQ ID NO.3 (5'-GCCTGAGCCAAAATAGAGA-3'), and the downstream primer sequence is shown in SEQ ID NO.4 (5'-ACTATGTAGGATGAAGCTAGGTC-3'). Flower tissue from the soybean variety "Williams 82" (W82) was collected, ground with liquid nitrogen, and the powder was placed in a 1.5 mL EP tube. 1 mL of lysis buffer was added, and the mixture was vortexed until homogeneous. RNA extraction was then performed according to the kit instructions (Total RNA Kit, Tiangen, China). RNA integrity was assessed by 1% agarose gel electrophoresis. Using the obtained total RNA as a template, cDNA synthesis was performed according to the instructions of the Vazyme HiScript 1st Strand cDNA Synthesis Kit (Nanjing, China). PCR amplification was then performed using the cDNA as a template. The PCR reaction mixture was as follows: 2 μL template, 2 μL each of forward and reverse primers, 25 μL 2×Phanta Max Master Mix, and finally, ddH2O was added to bring the volume to 50 μL. The PCR program was as follows: 95 ℃ pre-denaturation for 3 min; 95 ℃ denaturation for 15 sec, 58 ℃ annealing for 15 sec, 72 ℃ extension for 45 sec, for a total of 35 cycles; final extension at 72 ℃ for 5 min, followed by incubation at 4 ℃ for 30 minutes. Figure 1 The image shows the PCR amplification product of GmCDF3 in a 1% agarose gel, where M represents the DL2000 DNA Marker. After gel recovery, product purification, and T-vector ligation, the product was transformed into an E. coli enrichment plasmid. Sequencing yielded the complete soybean GmCDF3 coding sequence (CDS), which is 1440 bp in length and is shown in SEQ ID NO.2.
[0059] II. Tissue Expression Analysis of Soybean GmCDF3 Gene
[0060] The seeds of soybean variety "Dongnong 50" (DN50) were sown in flowerpots (peat moss: vermiculite, 2:1) and placed in a greenhouse for cultivation at 25℃ for 16 h / 8 h (light / dark). Roots (V1 stage), stems (V1 stage), leaves (V1 stage), flowers, and seeds of DN50 were collected, and the samples were quick-frozen with liquid nitrogen and stored in an ultra-low temperature freezer at -80 ℃ for later use.
[0061] Total RNA was extracted using a plant total RNA extraction kit (Tiangen), and RNA integrity was assessed by 1% agarose gel electrophoresis. cDNA synthesis was performed according to the instructions of the reverse transcription kit (Vazyme). The expression level of GmCDF3 was detected using quantitative real-time PCR (qRT-PCR). The upstream primer sequence is shown in SEQ ID NO. 5 (5'-TCTCCTGATGAACCTGCTCC-3'), and the downstream primer sequence is shown in SEQ ID NO. 6 (5'-CAGCATCATATTCCGTTGTGG-3'). Figure 2 To analyze the relative expression levels of GmCDF3 in different tissues, error bars represent mean ± standard deviation. In the figure, Root: root tissue; Stem: stem tissue; Leaf: leaf tissue; Flower: flower tissue; Seed: seed. Figure 2 It can be seen that GmCDF3 is expressed at low levels in roots, stems, leaves and seeds, but at high levels in flowers.
[0062] III. Subcellular localization of the soybean GmCDF3 gene
[0063] Specific primers were designed based on the known sequence of the GmCDF3 gene and the cloning site of the plant expression vector pBSK. The primers are as follows: the upstream primer is SEQ ID NO.7 (5'-CCGGAATTCATGTCTGAGGTTGTTA-3'), and the downstream primer is SEQ ID NO.8 (5'-CGCGGATCCTGTCCTCTCATGGAAGA-3'). The PCR program was as follows: 95 ℃ pre-denaturation for 3 min; 95 ℃ denaturation for 15 sec, 56 ℃ annealing for 15 sec, 72 ℃ extension for 45 sec, for a total of 35 cycles; and a final extension at 72 ℃ for 5 min, followed by isothermal incubation at 4 ℃ for 30 minutes. The target gene fragment amplified by PCR with enzyme digestion adapters was identified by gel electrophoresis, and the product was recovered from the gel. The gel-recovered product was double-digested with BamH1 and EcoR1, reacted at 37 ℃ for 1 h, and then recovered by PCR. The target fragment and the linearized vector pBSK, which had been double-digested with BamH1 and EcoR1, were ligated using the T4 DNA ligase (Vazyme). The reaction mixture consisted of 5 μL of vector, 3 μL of target fragment, 1 μL of 10 × T4 Buffer, and 1 μL of T4 DNA ligase. The reaction program was 16 °C for 2 h. The recombinant product was transformed into DH5α competent cells using a freeze-thaw method. The cells were plated, single colonies were picked, and colony PCR sequencing was performed for verification. Plasmids from correctly sequenced colonies were extracted and named pBSK-GmCDF3. Arabidopsis protoplasts were transiently transformed using the PEG method. The transformed protoplasts were cultured at 28 °C in the dark for 16 h, and the fluorescence excitation signal intensity of the protoplasts was observed using a Leica THUNDER system (Leica Microsystems, Wetzlar, Germany). Figure 3 The images show the subcellular localization of GmCDF3. The first row shows the microscopic images of the GmCDF3-GFP protein, from left to right: the green fluorescence channel (GFP), the chloroplast fluorescence channel (ChlorophyII), the bright field, and a merged image of the three channels. The second row shows the microscopic images of the empty vector, with the same left-to-right arrangement as the first row. The results indicate that the empty plasmid vector exhibits green fluorescence signals in all tissues, while the GmCDF3 protein fluorescence signal is primarily localized in the cell nucleus.
[0064] Example 2: Application of the soybean GmCDF3 gene described in this invention
[0065] I. Construction of plant overexpression vectors
[0066] Specific primers were designed based on the cloned CDS sequence of the GmCDF3 gene and the cloning site of the plant expression vector pTF101. The upstream primer sequence is shown in SEQ ID NO.9 (5'-CGGGGGACTCTAGAAACAGAGGATCCATGTCTGAGGTTGTTATTAA-3'), and the downstream primer sequence is shown in SEQ ID NO.10 (5'-TCTGTTTCTAGAGTCCCCCGACTAGTTGTCCTCTCATGGAAGACAA-3'). The PCR program was as follows: 95 ℃ pre-denaturation for 3 min; 95 ℃ denaturation for 15 sec, 70 ℃ annealing for 15 sec, 72 ℃ extension for 45 sec, for a total of 35 cycles; final extension at 72 ℃ for 5 min, followed by isothermal incubation at 4 ℃ for 30 minutes. The amplified target gene fragment was identified by gel electrophoresis, and the product was recovered from the gel. An overexpression vector was constructed using homologous recombination. The specific steps are as follows: The target fragment and the linearized vector pTF101, which had been double-digested with BamH1 and Spe1, were ligated using the homologous recombinase ClonExpress (Vazyme). The reaction system was as follows: 1.5 μL vector, 1 μL target fragment, and 2.5 μL 2× ClonExpress Mix. The reaction program was 50℃ for 30 min. The recombinant product was transformed into DH5α competent cells using a freeze-thaw method. The cells were plated, single colonies were picked, and colony PCR was performed for verification. The upstream primer sequence is shown in SEQ ID NO.11 (5'-CATTTCATTTGGAGAGAACACG-3'), and the downstream primer sequence is shown in SEQ ID NO.12 (5'-AGCGGATAACAATTTCACACAG-3'). The recombinant plasmid pTF101-GmCDF3, which overexpresses the soybean GmCDF3 gene, was obtained and stored at -20℃ for later use. Figure 4 This is a schematic diagram of the GmCDF3 overexpression vector. The GmCDF3 gene is represented by a green rectangular arrow, and its 5' end is connected to the CaMV 35 promoter.
[0067] II. Transformation of Agrobacterium tumefaciens with overexpression vector
[0068] The recombinant plasmid obtained in step one of this embodiment was transformed into Agrobacterium tumefaciens EHA105 and GV3101 using the freeze-thaw method. The specific experimental procedure is as follows:
[0069] ① Use a pipette to aspirate 3 µL of the recombinant plasmid and transfer it into 50 µL of EHA105 and GV310 competent cells, then gently mix.
[0070] ② Place the mixture on ice for 5 min, in liquid nitrogen for 5 min, incubate in a 37°C water bath for 5 min, and then immediately transfer it to an ice bath for 3 min.
[0071] ③ In a clean bench, use a pipette to take 800 µL of YEP culture medium and revive it for 2 h in a constant temperature shaker at 28 ℃ and 200 rpm.
[0072] ④ Centrifuge at 5000×g for 1 min at room temperature, resuspend the bacterial cells with an appropriate amount of supernatant, take 100 µL of the revived bacterial solution and spread it evenly on the screening plate medium, and incubate upside down at 28 ℃ for 2 days.
[0073] III. Agrobacterium-mediated transformation of soybean cotyledonary nodes and Arabidopsis inflorescences
[0074] ① Soybean cotyledon node transformation
[0075] High-quality DN50 soybean seeds were selected and sterilized in a fume hood using chlorine sterilization for 16 hours. The sterile seeds were then soaked in sterile water and cultured in a tissue culture room at 25 °C for 24 hours. After imbibition, the germinating soybeans were longitudinally split along the midline, and the apical and axillary buds were removed. Three incisions were made at the cotyledonary node. The transformed Agrobacterium EHA105 bacterial suspension was shaken to OD... 600 Centrifuge at 0.5 μL, 12000 rpm, resuspend in the infusion liquid, and place the prepared cotyledonary explants in it, allowing them to stand at 28 °C for 30 min. Dry the infused cotyledonary explants with the bacterial solution, place them adaxially on CCM solid medium, arrange them neatly, and incubate in the dark for 3 days. Wash the cotyledonary nodes three times in sterile water, blot off the liquid with sterilized filter paper, and insert them into recovery solid medium, incubating in the tissue culture room for 15 days. Cut off resistant shoots that have grown to 1-2 cm and insert them into elongation medium, incubating in the tissue culture room for 15 days. Once the elongated seedlings are robust, insert them into rooting medium and culture upright for 10 days. When the elongated seedlings have developed sufficient roots, remove them and place them in humus mixed with vermiculite, then cultivate soybean transformation plants in a greenhouse.
[0076] ② Arabidopsis inflorescence transformation
[0077] Before transformation, remove all pods and exposed flowers from the Arabidopsis thaliana plants. The transformed Agrobacterium GV3101 bacterial suspension is shaken to OD0.05. 600 =0.8, centrifuged at 6000 rpm for 5 min, and resuspended to OD using a 5% sucrose resuspension containing 0.02% Silwet L-77. 600 =0.6-0.8. Dip unopened Arabidopsis flower buds into the bacterial solution for 30 seconds, then blot off excess solution. Incubate in the dark for 24 hours. After one week, cut off uninfected flowers and branches.
[0078] IV. Identification of Positive Transformation and Gene Expression Analysis
[0079] ① Identification and gene expression analysis of soybean transformant plants
[0080] The overexpression tissue culture seedlings transplanted into soil were first subjected to bar gene detection (PAT / bar test strips), and the results are shown below. Figure 5 As shown in Figure a, the control (DN50) only showed one control band, with no detection band, while the three transgenic overexpression lines (OE1, OE2, and OE3) all showed detection bands (red arrows). DNA was extracted from leaves of the overexpression tissue culture seedlings after the initial detection. PCR identification was performed using Novizan polymerase (2×Rapid Taq Master Mix). The PCR system was as follows: 10 μL 2×Rapid Taq Master Mix, 0.6 μL Primer 1 (10 μM), 0.6 μL Primer 2 (10 μM), 1 μL Template DNA, and finally, ddH2O to a final volume of 15 μL. The reaction program was as follows: 95 ℃ pre-denaturation for 3 min; 95 ℃ denaturation for 15 sec, 60 ℃ annealing for 15 sec, 72 ℃ extension for 15 sec, for a total of 30 cycles; 72 ℃ final extension for 5 min. Seedlings with PCR bands matching the specified size were considered positive for overexpression. Figure 5 The figures show the results of bar assays and PCR detection for GmCDF3-overexpressing soybeans. In figure b, P represents the positive control expression vector pTF101-GmCDF3; M represents the DL2000 DNA Marker. Further analysis of the GmCDF3-overexpressing transgenic soybean plants will be conducted using qRT-PCR. Figure 6 The results show the transcriptional level of GmCDF3 in soybean overexpression. Error bars represent mean ± standard deviation. Statistical analysis was performed using a two-tailed t-test. *, P < 0.05; **, P < 0.01. The results indicate that the expression level of GmCDF3 in GmCDF3-overexpressing transgenic soybean plants was significantly higher than that in the control (…). Figure 6 PAT / bar test strips, PCR molecular detection, and qRT-PCR experiments all confirmed that the recombinant plasmid was successfully transferred into soybeans and successfully expressed.
[0081] ② Identification and gene expression analysis of Arabidopsis thaliana transformants
[0082] T1 generation transgenic Arabidopsis seeds were planted in 1 / 2 MS medium containing 6 mg / L phosmet (PPT): seeds were sterilized by vortexing with 10% NaClO for 8 min; washed twice with ddH2O, then spotted onto 1 / 2 MS medium and vernalized at 4℃ for 2 days; then transferred to an artificial climate chamber for growth. Plants that turned yellow were considered bar-resistant, while those that grew normally were considered bar-resistant. These were then transferred to culture soil for further cultivation, and leaf DNA was extracted for PCR identification. The PCR system was as follows: 10 μL 2×Rapid Taq Master Mix, 0.6 μL Primer 1 (10 μM), 0.6 μL Primer 2 (10 μM), 1 μL Template DNA, and finally adjusted to 15 μL with ddH2O. The reaction procedure was as follows: 95 ℃ pre-denaturation for 3 min; 95 ℃ denaturation for 15 sec, 60 ℃ annealing for 15 sec, 72 ℃ extension for 15 sec, for a total of 30 cycles; 72 ℃ final extension for 5 min. Plants that tested positive by PCR were further cultured to harvest T2 generation seeds. T2 generation seeds were planted in 1 / 2 MS medium containing 6 mg / L PPT, and the segregation ratio was calculated. Positive plants were further cultured until T3 generation seeds were harvested. Non-segregating lines from the T3 generation were selected for PCR identification and gene expression analysis. Figure 7 The figures show the detection results of GmCDF3-overexpressing Arabidopsis thaliana plants. Figure a represents the PPT screening results of the T3 generation of GmCDF3-overexpressing Arabidopsis thaliana lines, with OE7 / 8 / 11 being homozygous lines; b represents the PCR detection results of GmCDF3-overexpressing Arabidopsis thaliana plants; c represents the GmCDF3 transcriptional level detection results in overexpressing Arabidopsis thaliana. WT in b and c represents wild-type Arabidopsis thaliana; P in b represents the positive control expression vector pTF101-GmCDF3; and M represents the DL2000 DNA Marker. The results indicate that OE7, OE8, and OE11 are successfully transformed and expressed transgenic Arabidopsis thaliana (…). Figure 7 Next, T4 generation seeds were harvested for protein content determination.
[0083] V. Determination of Protein Content in Transgenic Soybean and Arabidopsis Seeds
[0084] 1. Determination of protein content in transgenic soybean seeds. Harvested T3 generation transgenic soybean seeds were dried in a 30 ℃ oven for one week. The protein content of mature seeds from control materials DN50 and GmCDF3 overexpressing soybean plants was then determined. Three independent transgenic lines (OE1, OE2, and OE3) were used in the experiment, with 5-8 individual plants from each line measured. Protein content was determined using an Antaris II Fourier transform near-infrared spectrometer (Thermo Fisher Scientific, America). Figure 8The comparison results of protein content in GmCDF3 overexpressing soybean seeds and control seeds are shown in the figure; a represents protein content, and b represents water-soluble protein content; error bars represent mean ± standard deviation. Statistical analysis was performed using one-way ANOVA. *, P < 0.05; ***, P < 0.001; ****, P < 0.0001. The results indicate that ( Figure 8 The protein content and water-soluble protein content of the three transgenic lines overexpressing GmCDF3 were significantly higher than those of the untransformed recipient soybean DN50. The protein content of OE1, OE2, and OE3 increased by 8.63%, 7.34%, and 6.49%, respectively; the water-soluble protein content increased by 6.15%, 6.59%, and 16.13%, respectively. Meanwhile, Figure 9 The figure shows the comparison of oil content and 100-seed weight in GmCDF3-overexpressing soybean seeds and control seeds. Figure a represents oil content, and figure b represents 100-seed weight. Error bars represent mean ± standard deviation. Statistical analysis was performed using one-way ANOVA; ns indicates no significant difference. Figure 9 It can be seen that the oil content and 100-grain weight of the converted soybeans were not significantly different from those of the unconverted soybeans (DN50).
[0085] 2. Protein content determination of transgenic Arabidopsis seeds. 1 mg of transgenic Arabidopsis seeds were ground in 350 μL of protein extraction solution. The mixture was centrifuged at 14000 rpm and 4 ℃ for 10 min, and the supernatant was used for protein quantification. A standard curve was constructed using bovine serum albumin (BSA) according to the instructions of NanoOrange (Thermo Fisher Scientific, America). Protein quantification was performed using a microplate reader, with fluorescence excitation and emission set to 485 nm and 590 nm, respectively. The protein content of the samples was calculated using the standard curve. Figure 10 The results show the comparison of protein content in GmCDF3-overexpressing Arabidopsis thaliana seeds and control seeds; WT represents wild-type Arabidopsis thaliana, and error bars represent mean ± standard deviation. Statistical analysis was performed using a two-tailed t-test. *, P < 0.05; ***, P < 0.001. The results indicate that ( Figure 10 The seed protein content of the three Arabidopsis transgenic lines (OE7, OE8, and OE11) that heterologously overexpressed GmCDF3 was significantly higher than that of the untransformed wild-type Arabidopsis, with OE7, OE8, and OE11 showing increases of 21.6%, 10.5%, and 10.5%, respectively.
[0086] VI. Determination of Raffinose and Amino Acid Content in Genetically Modified Soybean Seeds
[0087] The content of raffinose and amino acids in dried seeds of transgenic soybean T3 generation was determined using a Perten DA7250 near-infrared analyzer (Sweden). Three independent transgenic lines (OE1, OE2, and OE3) were used in the experiment, and 5-8 individual plants from each transgenic line were measured. Figure 11 The results of comparing raffinose content in GmCDF3-overexpressing soybean seeds with those in the control seeds are shown. Error bars represent mean ± standard deviation. Statistical analysis was performed using a two-tailed t-test. *, P < 0.05; ****, P < 0.0001. The results showed that the raffinose content in OE1, OE2, and OE3 seeds was significantly lower than that in unconverted recipient soybean DN50 by 11.50%, 10.17%, and 9.24%, respectively. Figure 11 ).at the same time, Figure 12 The results show the comparison of amino acids in GmCDF3 overexpressing soybeans and control seeds. Error bars represent the mean ± standard deviation. In the figure, a1 represents alanine, aspartic acid, cysteine, glutamic acid, glycine, proline, serine, tyrosine, leucine, lysine, phenylalanine, and threonine, respectively. Statistical analysis was performed using one-way ANOVA. *, P < 0.05; **, P < 0.01; ***, P < 0.001; ****, P < 0.0001.
[0088] The contents of 12 amino acids (alanine, aspartic acid, cysteine, glutamic acid, glycine, proline, serine, tyrosine, leucine, lysine, phenylalanine, and threonine) in OE1, OE2, and OE3 were significantly increased compared to those in unconverted recipient soybean DN50. Figure 12 Among them, the content of sulfur-containing amino acids (cysteine) increased by 6.49%-13.67%. Figure 12 c), the essential amino acids leucine, lysine, phenylalanine, and threonine increased by 3.15%-8.67%, respectively. Figure 12 i), 2.59%-8.19% Figure 12 j), 3.56%-7.77% Figure 12 k), 4.28%-8.97% Figure 12 l).
[0089] Example 3: Effects of soybean GmCDF3 gene haplotype on soybean seed protein, oil content, 100-seed weight and yield
[0090] Genotyping and molecular markers are helpful in screening superior soybean germplasm resources and are of great significance for soybean molecular breeding. To analyze the effects of different haplotypes of the GmCDF3 gene on soybean quality traits, we used the R package geneHapR to analyze 1247 soybean accessions containing genotype data (Zhang et al., 2022, 2023). Figure 13The figures show comparisons of protein content, oil content, 100-grain weight, and yield for haplotype GmCDF3; figure a represents the genotyping results of GmCDF3; figures be and be represent comparisons of oil content, protein content, 100-grain weight, and yield for the three haplotypes of GmCDF3, respectively; statistical analysis employed one-way ANOVA with multiple comparisons labeled by letters; the numbers below the box plots indicate the number of materials for each haplotype used in phenotypic analysis. Figure 13 The results showed that, based on nucleotide variations in the CDS sequence of the GmCDF3 gene, we identified three haplotypes from 1247 materials. Figure 13 a). By comparing the oil content, protein content, 100-grain weight, and yield of the three haplotypes, we found that the oil content, 100-grain weight, and yield of haplotype HAP1 of GmCDF3 were all higher than those of HAP2 and HAP3. Figure 13 However, HAP3 has a higher protein content than HAP1 and HAP2. Figure 13 c), which indicates that nucleotide variations at positions 5766756 and 5766804 of the GmCDF3 gene can be used for marker-assisted selection breeding.
Claims
1. Application of a soybean GmCDF3 gene in regulating seed protein, amino acid and raffinose content.
2. The application of the soybean GmCDF3 gene in regulating seed protein, amino acid, and raffinose content according to claim 1, characterized in that, This application method utilizes the soybean GmCDF3 gene to cultivate plant seeds, thereby improving plant quality.
3. The application of the soybean GmCDF3 gene in regulating seed protein, amino acid, and raffinose content according to claim 2, characterized in that... Methods to improve plant quality: Introduce the GmCDF3 gene that can be expressed into the recipient plant to obtain transgenic plants with increased GmCDF3 gene expression.
4. The application of the soybean GmCDF3 gene according to claim 3 in regulating seed protein, amino acid, and raffinose content, characterized in that... Introducing the GmCDF3 gene into the recipient plant: by introducing a recombinant expression vector carrying the GmCDF3 gene into the recipient plant.
5. The application of the soybean GmCDF3 gene according to claim 3 or 4 in regulating seed protein, amino acid, and raffinose content, characterized in that... The recipient plant is a dicotyledonous or monocotyledonous plant.
6. The application of the soybean GmCDF3 gene according to claim 5 in regulating seed protein, amino acid, and raffinose content, characterized in that... Dicotyledonous plants include legumes and cruciferous plants; the legume is soybean, and the cruciferous plant is Arabidopsis thaliana.
7. The application of the soybean GmCDF3 gene according to claim 1 in regulating seed protein, amino acid, and raffinose content, characterized in that... The protein encoded by the GmCDF3 gene is any one of the following: (A1) A protein with the amino acid sequence shown in SEQ ID NO.1; (A2) A protein having the same function as the amino acid sequence shown in SEQ ID NO.1, by substitution and / or deletion and / or addition of one or more amino acid residues; Proteins that have 99%, 95%, 90%, 85%, or 80% homology and the same function with any of the amino acid sequences defined in (A3) and (A1)-(A2); (A4) A fusion protein obtained by attaching a protein tag to the N-terminus and / or C-terminus of any of the proteins defined in (A1)-(A3).
8. The application of the soybean GmCDF3 gene according to claim 7 in regulating seed protein, amino acid, and raffinose content, characterized in that... Protein tags include Flag tags, His tags, MBP tags, HA tags, myc tags, GST tags, and / or SUMO tags.
9. The application of the soybean GmCDF3 gene according to claim 1 in regulating seed protein, amino acid, and raffinose content, characterized in that, The soybean GmCDF3 gene comprises any of the following DNA molecules: (B1) The DNA molecule shown in SEQ ID NO.2; (B2) A DNA molecule that hybridizes under stringent conditions to the DNA molecule defined by the sequence shown in SEQ ID NO.2 and encodes the GmCDF3 protein; A DNA molecule encoding the GmCDF3 protein that has 99%, 95%, 90%, 85%, or 80% or more homology with the DNA sequence defined by (B3) and (B1) or (B2).