Method for screening soybean rhizoctonia web blight response genes based on transcriptome analysis

Transcriptome analysis was used to screen soybean stem rot response genes, and four key genes were identified. This solved the problem of the inability to effectively screen response genes in existing technologies and enabled dynamic regulation of soybean resistance to stem rot.

CN120853675BActive Publication Date: 2026-03-03聊城市农业科学院 +6
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies are insufficient to fundamentally address the damage caused by soybean stem spot disease to crops. Detection methods mainly focus on pathogen primers or kits, which cannot effectively screen out response genes.

Method used

By selecting soybean pods from both infected and healthy plants grown in the field, RNA was extracted and transcriptome sequencing was performed. Differentially expressed genes were screened using HTSeq and DESeq2, and key genes were identified by combining GO term/KEGG pathway functional enrichment analysis, protein interaction analysis, and StringTie transcript assembly.

Benefits of technology

Four key genes (GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_12G226400) were screened out. These genes dynamically optimize the soybean disease resistance response through a dual transcriptional and post-transcriptional regulatory mechanism and synergistically regulate resistance to stem rot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853675B_ABST
    Figure CN120853675B_ABST
Patent Text Reader

Abstract

The application provides a method for screening soybean Cercospora sojina blight response genes based on transcriptome analysis, which comprises the following steps: selecting pods of soybean Cercospora sojina blight infected plants and healthy plants planted in the field, performing transcriptome sequencing, screening differential expression genes of soybean responding to Cercospora sojina blight, analyzing biological pathways of the differential expression genes in the Cercospora sojina blight infection process, performing protein interaction analysis, screening interaction core genes, comparing sequences, performing transcript from scratch assembly, analyzing differential variable splicing events and exon usage differences through integration analysis, analyzing post-transcriptional regulation network dynamic characteristics of key candidate genes of soybean responding to Cercospora sojina blight, and screening Cercospora sojina blight response genes of soybean. Four Cercospora sojina blight response genes of soybean are obtained, and expression characteristics in different soybean tissues are analyzed, which provides a theoretical basis for soybean Cercospora sojina blight resistance breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of disease resistance gene screening technology, specifically involving a method for screening soybean stem rot response genes based on transcriptome analysis. Background Technology

[0002] Soybean stem rot (Phomopsis longicolla) is one of the three major diseases affecting summer soybeans in the Huang-Huai-Hai Plain. It causes browning at the base of the stem and stem wilting, with a high incidence rate; in severe cases, the disease rate can reach over 90%, even leading to complete crop failure. This disease is particularly severe during seed maturity, often accompanied by a "greening" phenomenon (green plants at maturity with shriveled and rotten seeds). The pathogen of soybean stem rot can be transmitted through seeds, leading to disease spread the following year and creating a vicious cycle.

[0003] Most existing research focuses on primers or kits for detecting soybean seed rot. While these methods can detect the pathogen, they cannot fundamentally solve the problem of the damage caused by this pathogen to soybean crops. Summary of the Invention

[0004] The technical problem to be solved by this invention is to address the shortcomings of the prior art by providing a method for screening soybean stem rot response genes based on transcriptome analysis. By selecting pods from soybean plants infected with stem rot and pods from healthy plants grown in the field, RNA extraction and transcriptome sequencing are performed on them respectively to identify key genes and important pathways in soybean response to stem rot stress. By analyzing and comparing differentially expressed genes and differential splicing events, response genes of soybean stem rot can be obtained.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a method for screening soybean stem rot response genes based on transcriptome analysis, comprising the following steps:

[0006] S1. Select pod samples from soybean plants infected with pseudostem spot seed rot and healthy soybean plants grown in the field, extract RNA samples, perform transcriptome sequencing, and obtain sequencing data.

[0007] S2. Using HTSeq and DESeq2 software, differentially expressed genes in soybean responding to *Pseudomonas stolonifera* were screened based on the original expression levels (Read Count) of the sequencing data in S1.

[0008] S3. Through GO term / KEGG pathway functional enrichment analysis and gene set enrichment analysis, analyze the biological pathways in which the differentially expressed genes described in S2 participate in the process of susceptibility to stem rot. Use statistical methods to test whether the pre-set gene set is enriched at the top or bottom of the sorting table to determine whether the pathway is activated or inhibited.

[0009] S4. Based on the STRING database, protein interaction analysis was performed to screen out the core interaction genes, which can be used as design objects for multi-target editing or modular breeding.

[0010] S5. The aligned sequences obtained from RNA-seq sequencing were assembled from transcripts de novo using StringTie software. By integrating and analyzing differential alternative splicing events and exon usage differences, the dynamic characteristics of the post-transcriptional regulatory network of key candidate genes in soybean response to pseudostem rot were analyzed.

[0011] S6. Analyze the expression characteristics of key candidate genes in soybean response to pseudostem rot in different soybean tissues.

[0012] S7. Screening out response genes for soybean pseudostem rot.

[0013] Preferably, the conditions for using the DESeq2 software in S2 to screen differentially expressed genes in soybean in response to pseudostem rot are: fold change in expression |log2FoldChange|>1, and significance P-value<0.05.

[0014] Preferably, in S3, GO term / KEGG pathway functional enrichment analysis and gene set enrichment analysis are performed. Specifically, clusterProfiler software is used to perform GO term / KEGG pathway functional enrichment analysis and gene set enrichment analysis, and the criterion for significant enrichment is P-value < 0.05.

[0015] Preferably, the method for protein interaction analysis in S4 is as follows: using the STRING database, protein interaction relationships are screened with a high confidence threshold of Score > 0.95.

[0016] Preferably, the specific steps in S5 for analyzing the dynamic characteristics of the post-transcriptional regulatory network of key candidate genes in soybean response to pseudostem rot are as follows: using StringTie software to perform de novo transcript assembly of the obtained aligned sequences, using rMATS software to analyze differentially spliced ​​events, and using the R language package DEXSeq to analyze exon usage differences.

[0017] Preferably, in step S6, the expression characteristics of key candidate genes in soybean response to pseudostem rot in different soybean tissues are analyzed on the SoyOmics website to obtain response genes for soybean pseudostem rot.

[0018] Compared with the prior art, the present invention has the following advantages:

[0019] 1. In the technical solution of this invention, RNA is extracted from the pod tissues of soybean plants infected with *S. spp.* and healthy plants, respectively, to construct transcriptome libraries and perform high-throughput sequencing. The obtained raw sequencing data undergoes strict quality control and filtering to obtain high-quality qualified sequences. Subsequently, these sequences are aligned to the soybean reference genome. Based on the alignment results, the expression level of each gene is calculated, and 200 key differentially expressed genes (top200DEGs) are analyzed and screened.

[0020] 2. This invention integrates StringTie transcript de novo assembly, rMATS differential alternative splicing analysis, and DEXSeq exon usage difference detection to systematically analyze the dynamic characteristics of the post-transcriptional regulatory network of key candidate genes in soybean response to *Pseudomonas aeruginosa*.

[0021] 3. Based on transcriptome analysis, this invention screened four key genes in soybean response to stem spot rot: GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_12G226400. These four genes all exhibited significant changes in expression levels and differential splicing events, indicating that they dynamically optimize soybean disease resistance response through a dual transcriptional and post-transcriptional regulatory mechanism.

[0022] 4. This invention analyzed the expression levels of four soybean stem rot response genes in different soybean tissues and found that each gene exhibited a specific high expression pattern in different soybean tissues. This reflects the functional division of different genes and indicates that the four genes synergistically regulate the resistance of soybean to stem rot through tissue-specific functional division and metabolic pathways.

[0023] 5. This invention analyzes the response mechanism between four genes (GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_12G226400) and soybean pseudostem rot, as well as their biological relationship in the disease resistance process. This helps researchers solve the damage of pseudostem rot to soybean crops at its root through biobreeding methods.

[0024] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0025] Figure 1 The images show soybean plants and pods infected with pseudostem rot and healthy soybean plants and pods in Example 1 of this invention.

[0026] Figure 2 This is a volcano diagram of differentially expressed genes in Example 1 of the present invention.

[0027] Figure 3 This is a statistical analysis of the distribution of differentially expressed genes on chromosomes in Example 1 of the present invention.

[0028] Figure 4 This is the GO enrichment analysis of the top 200 DEGs in Example 1 of the present invention.

[0029] Figure 5 This is a KEGG enrichment analysis of the top 200 DEGs in Example 1 of the present invention.

[0030] Figure 6 This is a cluster analysis of the top 200 DEGs differential genes in Example 1 of the present invention.

[0031] Figure 7 This is an analysis of the interaction between the differentially expressed gene PPI protein in Example 1 of the present invention.

[0032] Figure 8 This is an example of transcript splicing and new transcript analysis in Embodiment 1 of the present invention. In the figure, A represents the annotation analysis of spliced ​​transcripts in different databases, and B represents the statistical distribution of classification codes of newly assembled transcripts.

[0033] Figure 9 This refers to the differential variable shear analysis in Embodiment 1 of the present invention.

[0034] Figure 10 This is a heatmap analysis of the expression of four response genes in different soybean tissues in Example 1 of the present invention. Detailed Implementation

[0035] Example 1

[0036] The method for screening soybean stem rot response genes based on transcriptome analysis in this embodiment is as follows:

[0037] S1. Selection of soybean RNA sequencing samples

[0038] The soybean variety selected in this embodiment is "He Dou 36";

[0039] Planting location: Liaocheng City, Shandong Province (N36.45°, E115.98°);

[0040] Planting date: June 19, 2024;

[0041] Selection time of sequencing samples: Soybean pod samples were collected on September 1, 2024. Healthy, uninfected soybean plants and pods infected with *S. seed rot* were selected and randomly divided into two groups, designated Group A and Group B respectively. Figure 1 As shown, 6-7 pods were randomly collected from each group. After the samples were selected, they were immediately flash-frozen in liquid nitrogen and stored at -80℃ for later use.

[0042] S2. RNA extraction and determination from bean pods

[0043] (1) RNA extraction from normal, uninfected bean pods in group A

[0044] a. Homogenization: Take the bean pod sample from group A and place it in a 2 mL grinding tube. Grind it at 60 Hz for 60 s using a grinder to grind it into a powdered sample. Keep it at 4 ℃ and add 1 mL of Invitrogen TRIzol extraction reagent. Mix it immediately and continue grinding at 55 Hz for 30 s. Then incubate it at 25 ℃ for 5 min to facilitate the complete separation of ribosomes in the homogenized sample. Then centrifuge at 12000 rpm for 5 min and take the supernatant into another clean centrifuge tube.

[0045] b. Phase separation: Add 0.2 mL of chloroform to the supernatant obtained in step a, shake the tube vigorously for 30 seconds, let it stand, incubate at 25°C for 2-3 minutes, and then centrifuge at 12000 rpm for 15 minutes at 4°C to obtain a mixture of three phases: upper, middle and lower. The middle and upper layers are colorless aqueous phases, and the lower layer is a red phenol-chloroform phase. Take the colorless aqueous phases of the middle and upper layers into a clean test tube No. 1, and dissolve the RNA to be tested in the colorless aqueous phase.

[0046] c. RNA precipitation: Take 400 μL of the colorless aqueous phase from the clean tube 1 obtained in b) into the clean tube 2, add 500 μL of isopropanol, mix by inverting, incubate at room temperature for 10 min, then centrifuge at 12000 rpm for 10 min at 4℃, remove the supernatant, and obtain the RNA precipitate.

[0047] d. RNA precipitate dissolution: Add 1 ml of 75% ethanol solution to the RNA precipitate obtained in c, mix well, centrifuge at 12000 rpm for 5 min at 4℃, remove the supernatant, add 40~50 μL of RNase-free ribonuclease aqueous solution to obtain the RNA sample of group A to be tested.

[0048] (2) RNA extraction from group B pods infected with pseudostem rot

[0049] The RNA extraction method for the pods infected with pseudostem rot was the same as that for the healthy pods in group A, resulting in RNA samples for group B to be tested.

[0050] (3) RNA assay in bean pods

[0051] The concentration and purity of RNA samples from groups A and B were determined using a Thermo Scientific NanoDrop 2000 ultra-micro spectrophotometer.

[0052] The integrity of the RNA samples from groups A and B was determined using 2.0% (p / v) RNA-specific agarose gel electrophoresis or the Agilent 2100 Bioanalyzer system and RNA 6000 Nano kit 5067-1511.

[0053] The manufacturer and model number of the Thermo Scientific micro-volume spectrophotometer is ThermoScientific, Waltham, Massachusetts, USA; the manufacturer and model number of the Agilent 2100 Bioanalyzer System and RNA 6000 Nano kit 5067-1511 is Agilent Technologies Inc, California, USA.

[0054] S3. Transcriptome RNA-seq Library Construction and Quality Control

[0055] (1) Construction of transcriptome RNA-seq library

[0056] ≥1 μg of RNA samples were taken from groups A and B respectively. Using a strand-specific library preparation kit (NEBNextUltra II RNA Library Prep Kit for Illumina), mRNA with polyA tails was enriched using Oligo(dT) magnetic beads. Then, the mRNA was randomly fragmented using divalent cations via ion fragmentation. Double-stranded cDNA was synthesized using the fragmented mRNA as a template and random oligonucleotides as primers. The double-stranded cDNA was then purified, followed by end repair and the introduction of an "A" base at the 3' end, ligating a sequencing adapter. Double-stranded cDNA of 400–500 bp was screened using AMPure XP beads reagent, and PCR amplification was performed. The PCR amplified product was then purified again using AMPure XP beads reagent, ultimately obtaining transcriptome RNA-seq libraries for groups A and B respectively.

[0057] The NEBNext Ultra II RNA Library Prep Kit for Illumina is manufactured by New England Biolabs Inc; Ipswich, Massachusetts, USA.

[0058] (2) Quality control of transcriptome RNA-seq library

[0059] The quality of transcriptome RNA-seq libraries in groups A and B was assessed using the Agilent 2100 Bioanalyzer and Agilent High Sensitivity DNA Kit, respectively. The total concentration of transcriptome RNA-seq libraries in groups A and B was determined using Pico Green, Quantifluor-STfluorometer, and Promega, respectively. The effective library concentration of transcriptome RNA-seq libraries in groups A and B was quantitatively determined using the Quant-iTPicoGreen dsDNA Assay Kit and StepOnePlus Real-Time PCR Systems, respectively.

[0060] Transcriptome RNA-seq libraries from groups A and B were quantitatively diluted with 10 mM Tris-HCl (pH 8.5) containing 0.1% Tween 20, and then sequenced in PE150 mode on an Illumina sequencer.

[0061] Among them, the Agilent High Sensitivity DNA Kit is manufactured by Agilent Technologies Inc., California, USA, 5067-4626; the Pico green detection library is manufactured by Madison, Wisconsin, USA, E6090; the Quant-iTPicoGreen dsDNA Assay Kit is manufactured by Invitrogen, California, USA, P7589; and the StepOnePlus Real-Time PCR Systems are manufactured by Thermo Scientific, Waltham, Massachusetts, USA.

[0062] S4. Quality control and comparative analysis of machine data from transcriptome RNA-seq sequencing.

[0063] (1) Quality control of transcriptome RNA-seq data

[0064] The transcriptome RNA-seq libraries from groups A and B were sequenced to obtain image files, which were then converted by the sequencing platform's built-in software to generate FASTQ raw data (i.e., unsequential data). This raw data contains reads with adapters and low quality sequences, which can significantly interfere with subsequent information analysis. Therefore, further filtering of the sequencing data is necessary. The filtering criteria mainly include: using FASTP (0.22.0) to remove sequences with 3' adapters; and removing reads with an average quality score lower than Q20. All subsequent analyses are high-quality analyses based on clean data.

[0065] After quality testing of the transcriptome RNA-seq libraries in groups A and B, the RNA sample concentration and integrity met the requirements for constructing RNA-Seq sequencing cDNA libraries. The clean data of all samples reached 6.89 Gb, the GC content was above 42.02%, the Q20 base percentage was above 98.72%, and the Q30 base percentage was above 96.47%. After filtering, the number of high-quality sequence bases reached 6.78 Gb, the percentage of high-quality sequence reads in the total sequencing reads was above 98.57%, and the percentage of high-quality sequence bases in the total sequencing bases was above 98.33%. These results indicate that the sequencing data quality is high and meets the experimental requirements, and the data volume is sufficient for further analysis.

[0066] In the above-mentioned processes of RNA extraction, amplification, and quality control of normal, uninfected, and infected pods with pseudostem rot, if it is necessary to prevent the selected samples from failing to meet the requirements for RNA sequencing, 2 to 4 subgroups can be added under both group A and group B before RNA extraction and determination.

[0067] (2) Comparison and analysis of transcriptome RNA-seq data

[0068] The reference genome and gene model annotation files were downloaded directly from the genome website. The reference genome index was constructed using HISAT2 (v2.1.0), and the paired-end clean reads were aligned with the reference genome Glycine_max.Glycine_max_v2.1 using HISAT2. HISAT2 was chosen as the alignment tool because it can generate a spliced ​​database based on the gene model annotation files, thus providing better alignment results than other non-spliced ​​alignment tools.

[0069] Sequence alignment of transcriptome RNA-seq libraries from groups A and B revealed that 96.96-97.12% of clean reads in group A aligned to the reference genome, while 18.04-33.87% of clean reads in group B aligned to the reference genome. For both groups, the percentage of clean reads aligned to a single location was 97.30-98.17%, those aligned to multiple locations were 1.83-2.70%, those aligned to gene regions were 97.47-98.21%, those aligned to intergenic regions were 1.79-2.60%, and those aligned to exon regions were 93.71-95.85%.

[0070] The distribution of reads aligned to different regions of the genome, including the CDS (coding region), introns, UTR (5' and 3' untranslated regions), TSS_up (upstream of transcription start site), and TES_down (downstream of transcription termination site), was statistically analyzed in the transcriptome RNA-seq libraries of groups A and B. The soybean genome annotation is relatively complete. In both groups A and B, the number of reads aligned to the CDS (coding region) was the highest, followed by the UTR, introns, TES_down, and TSS_up. Specifically, the number of aligned reads gradually increased within 1kb, 5kb, and 10kb of distance from TES_down and TSS_up.

[0071] S5. Differential gene screening and annotation analysis

[0072] (1) Gene expression analysis

[0073] Using HTSeq (v0.9.1), the Read Count values ​​of each gene in the transcriptome RNA-seq libraries of groups A and B were used as the original gene expression levels. To ensure the comparability of gene expression levels between different genes and samples, FPKM (Fragments Per Kilo bases per Million fragments) / TPM (Transcripts per Million) were used to standardize the expression levels (Nosrmalization).

[0074] (2) Differentially expressed gene screening analysis

[0075] Gene expression was compared between transcriptome RNA-seq libraries of groups A and B using DESeq2 (v1.38.3) for differential expression analysis. The criteria for screening differentially expressed genes were: fold change |log2FoldChange| > 1, and a significance P-value < 0.05. Figure 2 As shown, a volcano plot was drawn based on the differential expression results. The differential gene expression was displayed based on the fold change and significance results. The total number of differentially expressed genes between the transcriptome RNA-seq libraries of group A and group B was 8555. The left side of the figure shows 4967 downregulated genes in bean pod samples infected with *Pseudomonas aeruginosa* compared to normal uninfected bean pod samples. The right side of the figure shows 3588 upregulated genes in bean pod samples infected with *Pseudomonas aeruginosa* compared to normal uninfected bean pod samples. Furthermore, based on the analysis results, the top 200 genes with the highest significance in the differential expression analysis were screened, as shown in Table 1.

[0076] like Figure 3 As shown, using the R language's Circlize package, differentially expressed genes are labeled on the genome based on genomic information and gene differential expression analysis results, and a genome circle map is drawn, reflecting the uneven distribution of differentially expressed genes on chromosomes.

[0077] Table 1. Top 200 genes with the highest statistical significance in differential expression analysis.

[0078]

[0079]

[0080]

[0081] (3) Differential gene enrichment analysis of top 200 DEGs

[0082] Enrichment analysis was performed using clusterProfiler (v4.6.0). The P-value was calculated using the hypergeometric distribution method (generally, the standard for significant enrichment is P-value < 0.05) to identify the GOterm / KEGG pathways that were significantly enriched in differentially expressed genes (all / up / down), thereby determining the main biological functions of the differentially expressed genes.

[0083] like Figure 4As shown, GO annotation analysis of differentially expressed genes revealed significant differences in pathways related to plant cell wall biosynthesis, carbohydrate metabolism, and oxidative stress response after the occurrence of soybean stem rot, which may be a key biological effect induced by the pathogen; the P-value (-log) of plant-type cell wall organization or biogenesis was also analyzed. 10 The most significant difference was observed at (P-value = 40), indicating that *Sterculia vesicularis* may strongly affect the structural reconstruction or biosynthesis of plant cell walls; cellular glucan metabolic process (P = 6.54) e-44 It also showed extremely high significance, indicating its core role and suggesting the active adjustment of cell wall polysaccharide metabolism caused by *Pseudomonas stolonifera*.

[0084] like Figure 5 As shown, KEGG annotation analysis of differentially expressed genes revealed three highly significant pathways. Glyoxylate and dicarboxylate metabolism (-log10(P)=7) showed the highest significance and the largest number of differentially expressed genes (DEG=70). This pathway is associated with energy metabolism such as the glyoxylate cycle and photorespiration, indicating that it is strongly regulated in soybean response to *S. seed rot* and may be the core of metabolic reregulation. Amino sugar and nucleotide sugar metabolism (-log10(P)=6) showed outstanding significance, possibly related to cell wall synthesis or the generation of signaling molecules such as UDP-glucose, suggesting significant changes in the synthesis or breakdown of amino sugars and nucleotide sugars. Carbon fixation in photosynthetic... The significant enrichment of organisms (-log10(P)=5) indicates that photosynthetic carbon assimilation processes such as the Calvin cycle may be affected by diseases, and special attention should be paid to changes in light energy utilization efficiency. The distribution of up-regulated and down-regulated genes, as shown in the circular diagram, shows that the enrichment factor of some pathways is close to 0.8, suggesting that soybean response to *S. stem rot* may activate pathways such as phenylpropanoid synthesis while inhibiting other pathways such as sterol synthesis. In addition, the KEGG network diagram also shows that Glyoxylate metabolism is closely related to Glycine / serine metabolism and TCA cycle through the glyoxylate branch, indicating that carbon skeleton reuse plays a pivotal role in metabolic regulation. The direct connection between Starch and sucrose metabolism and Glycolysis suggests the synergistic regulation of carbohydrate storage and decomposition.

[0085] (4) Cluster analysis of differential genes in the top 200 DEGs

[0086] Bidirectional clustering analysis was performed using the Pheatmap and ComplexHeatmap packages in R to calculate the union of differentially expressed genes in normal, disease-free samples (group A) and diseased samples (group B). The Euclidean method was used by default to calculate distances, and hierarchical clustering was performed using the Complete Linkage method. Figure 6 As shown, a clustering dendrogram was generated based on all differentially expressed genes from the transcriptome RNA-seq libraries of groups A and B, describing the expression level of the same gene in different samples. The results showed that the top 200 differentially expressed genes were classified into 9 clusters. This not only indicates that genes within the same cluster are actually related in certain biological processes or metabolic and signaling pathways, but also visually demonstrates the changes in the expression abundance of different types of genes among samples, narrowing the scope of analysis and focusing on key genes.

[0087] (5) Analysis of protein network interactions of top 200 DEGs

[0088] Protein-protein interaction analysis was performed using the STRING database (https: / / string-db.org / ) to reveal the interactions between target genes. When the STRING database contained PPI information for the species, PPI pairs containing differentially expressed genes and with a score > 0.95 were directly selected from the database based on the results of gene differential expression analysis. The score value could be adjusted if the network was too large or too small. When the STRING database did not contain PPI information for the species, protein sequences from similar species were compared with those of the species to obtain the interactions between proteins in the species.

[0089] like Figure 7 As shown, 117,055 protein-protein pairwise interaction pairs were obtained by screening with a Score>0.95 threshold. Among them, four genes with interacting proteins, namely GLYMA_01G197500, GLYMA_12G226400, GLYMA_17G128000, and GLYMA_19G114500, were selected as the top 200 DEGs genes.

[0090] S6, top 200 DEGs differential gene set enrichment analysis (GSEA)

[0091] Differential gene set enrichment analysis (GSEA) was performed using clusterProfiler (v4.6.0) software.

[0092] (1) GO GSEA analysis

[0093] GO-GSEA analysis, as shown in Table 2, identified 23 core enriched genes (CORE ENRICHMENT=Yes), including two top 200 differentially expressed genes (DEGs): GLYMA_06G050100 and GLYMA_18G061100. These genes are significantly enriched in specific biological pathways or functions and may be key genes driving functional changes in differentially expressed genes (DEGs). Functional annotation of the core genes indicates that the significantly enriched biological processes may include stress response (oxidative stress, heavy metal detoxification), metabolic regulation (nitrogen metabolism, sulfur metabolism), transcriptional regulation (transcription factors such as NAC1 and RAV), and signal transduction and development (NAC family genes are involved in organ development). Analysis of gene expression trends shows that most core genes, such as GLYMA_06G050100 and GLYMA_18G061100, are positively enriched genes (RANK METRIC SCORE). The RANKMETRIC SCORE < 0 indicates that these genes may be significantly upregulated in these pathways; GLYMA_02G297700 and GLYMA_11G003500 are negatively enriched genes (RANKMETRIC SCORE < 0), which may be related to repressive regulation.

[0094] Table 2. GO and GSEA analysis of the top 200 DEGs genes.

[0095]

[0096] (2) KEGG GSEA analysis

[0097] KEGG-GSEA analysis, as shown in Table 3, identified two core enriched genes (CORE ENRICHMENT=Yes), including one differentially expressed gene from the top 200 DEGs, GLYMA_06G050100. This indicates that these genes are significantly enriched in specific KEGG pathways and may be key genes driving differential expression in metabolic or signaling pathways. Gene expression trend and function association analysis showed that the RANK METRIC SCORE of both core genes were positive, suggesting that they may be significantly upregulated in plant stress resistance or metabolic regulation, driving the activation of related pathways.

[0098] Table 3. Analysis of the top 200 DEGs genes (KEGG, GSEA)

[0099]

[0100] S7, new transcript splicing and differentially expressed alternative splicing and exon analysis of differentially expressed genes in the top 200 DEGs.

[0101] (1) New transcript splicing and gene structure annotation

[0102] Based on the reference genome annotation system, StringTie (v2.2.1) software was used to perform de novo transcriptome assembly from the aligned sequences (mapped reads) obtained by RNA-seq sequencing. The results are as follows: Figure 8 As shown;

[0103] To analyze the functional characteristics of novel transcripts, a multi-dimensional functional annotation system was constructed by integrating five authoritative databases: eggNOG, GO, KEGG, NR, and SwissProt. The following annotation results were obtained ( Figure 8 (A)

[0104] Functional classification: 36,117 genes were annotated in the eggNOG database, revealing the evolutionary classification and metabolic function of their orthologous genes;

[0105] Biological processes: The GO database successfully annotated 25,182 genes, clarifying the biological processes they participate in (such as stress response and metabolic regulation).

[0106] Pathway analysis: 17,066 genes were annotated in the KEGG database and located to pathways such as plant hormone signal transduction and secondary metabolism synthesis;

[0107] Sequence homology: The NR database has the most extensive annotation coverage (36,996 genes), suggesting sequence similarity to proteins from known species;

[0108] High-quality annotation: 30,940 genes are annotated using the SwissProt database, providing accurate protein function descriptions;

[0109] By systematically comparing with known annotated transcripts, the following four types of novel transcripts were screened out ( Figure 8 B): Unannotated transcripts containing Class Code "j" (novel splicing variant), "i" (intronic region transcript), and "u" (unknown intergenic region transcript); and potential antisense transcripts with Class Code "x" (overlapping with known antisense strands).

[0110] (2) Differential alternative splicing analysis of the top 200 DEGs genes based on new transcripts

[0111] Differential alternative splicing events were analyzed using rMATS (v4.0.1) software. The main types of alternative splicing events analyzed included A3SS (alternative 3'splice site) exon 5' alternative splicing (i.e., the 3' splicing site of the intron preceding it is variable), A5SS (alternative 5'splice site) exon 3' alternative splicing site (i.e., the 5' splicing site of the intron following it is variable), skipped exon (SE), retained intron (RI), and mutually exclusive exons (MXE).

[0112] Differential alternative splicing analysis was performed based on gene structure annotation files, and the results are as follows: Figure 9 As shown in Table 4, a total of 11,545 variable shearing events were found, including A3SS, A5SS, SE, RI, and MXE.

[0113] Using FDR < 0.01 as the significance threshold, 195 known alternative splicing events showed significant differences, while 4228 showed no significant differences; among the new alternative splicing events, 95 showed significant differences, while 7027 showed no significant differences. Therefore, the results indicate a total of 290 significantly differentially splicing events, both known and new, suggesting that *Sterculia vesicularis*-specifically triggers splicing reprogramming, significantly affecting the post-transcriptional regulatory layer and increasing transcript diversity through selective splicing. New events accounted for 32.75%, reflecting that plants may adapt to environmental changes by evolving new splicing sites or activating silencing sites. Furthermore, two genes, GLYMA_07G073500 and GLYMA_12G024100, which showed significant differential splicing, were among the top 200 differentially expressed genes, indicating that these genes are regulated at both the transcriptional (expression level) and post-transcriptional (splicing pattern) levels, and may be core regulatory hubs for phenotypic changes.

[0114] Table 4. Variable Shear Difference Analysis

[0115]

[0116] (3) Differential analysis of exons of the top 200 DEGs based on new transcripts (DEXSeq)

[0117] Differences in exon usage can lead to different protein isoforms. The DEXSeq (v1.44.0) package was used to analyze differences in exon usage in RNA-seq experimental data. Here, differences in exon usage refer to relatively different exon usages due to experimental conditions. Screening with a padj < 0.05 threshold revealed significant differences in exon regions of 2067 genes, indicating that *Stipa rot* not only affects the overall expression level of soybean genes but also finely regulates transcript diversity through alternative splicing, potentially altering protein function or regulatory activity. Among these, GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_03G0... Ten genes, including 73500, GLYMA_20G187400, GLYMA_07G014500, GLYMA_03G083200, GLYMA_12G226400, GLYMA_14G013400, and GLYMA_12G141000, were identified as the top 200 differentially expressed genes among the top 200 DEGs. This suggests that these genes are subject to dual regulation (changes in expression level + alterations in splicing pattern), and may be key hubs for adapting to the environment or phenotypic changes, providing a new perspective for elucidating the molecular mechanisms of complex traits.

[0118] Identification and Characterization of Key Candidate Response Genes for S8 Soybean Stem Rot Disease

[0119] (1) Identification of response genes for soybean stem rot

[0120] The analysis results in S1-S7 show that the GLYMA_18G061100 gene is upregulated in soybean response to *S. spp.* rot. The candidate gene with the most significant difference in expression level (sort: 1) has a significant interaction with the key nitrogen metabolism enzyme gene GLYMA_05G046700 and membrane transport-related genes GLYMA_08G043400 and GLYMA_17G128600 (Score>0.95). Further GO-GSEA gene set enrichment analysis confirmed that the GLYMA_18G061100 gene is a core enriched gene (CORE ENRICHMENT=Yes), significantly enriched in metabolic networks centered on amino acid metabolism and nitrogen cycling, involving multiple enzyme activities (transferases, carboxylases) and membrane system functions. It may respond to *S. spp.* rot through metabolic pathway obstruction, increased oxidative stress, and dysregulation of defense signals.

[0121] The GLYMA_17G128000 gene was upregulated in soybean response to *Stipa rot*, with significant differences in expression levels (sort: 11). GO enrichment analysis showed that it responds to *Stipa rot* through core mechanisms of metabolic competition, toxin synthesis, and disease resistance in the pathogen-host interaction. It was comparable to GLYMA_18G116900, GLYMA_19G107300, GLYMA_01G019200, GLYMA_08G273100, GLYMA_02G015700, GLYMA_02G015800, GLYMA_06G091500, GLYMA_05G074000, and GLYM... Genes A_06G303200, GLYMA_08G302600, GLYMA_09G100400, GLYMA_09G203500, GLYMA_10G016200, GLYMA_12G100500, and GLYMA_16G044600 form a high-confidence interaction network (Score>0.95) with 15 key genes. Among them, GLYMA_18G116900 (a homolog of the disease-associated protein PR1), GLYMA_05G074000 (an ABC transporter), and GLYMA_09G100400 (a flavonoid synthase) are directly involved in disease resistance and the synthesis of secondary metabolites.

[0122] The GLYMA_06G050100 gene was upregulated in soybean response to *S. spp.*, with a significant difference in expression levels (sort: 47). It formed a high-confidence interaction with the GLYMA_13G207900 gene (annotated as thioredoxin reductase) (Score > 0.95). This gene is involved in redox homeostasis regulation, suggesting that the two may scavenge pathogen-induced reactive oxygen species (ROS) through a synergistic "antioxidant enzyme-substrate" model, preventing cell damage. Further enrichment analysis using GO-GSEA and KEGG-GSEA gene sets confirmed that the GLYMA_06G050100 gene was a core enriched gene (CORE ENRICHMENT = Yes), involved in the synthesis of defense compounds, regulation of oxidative stress, competition for metabolic resources, and transmembrane signaling and transport functions. It was significantly enriched in nicotinate and nicotinamide metabolism, glutathione metabolism, valine, leucine, and isoleucine degradation, valine, leucine, and isoleucine biosynthesis, and cysteine ​​and methionine metabolic pathways, driving the activation of these pathways.

[0123] The GLYMA_12G226400 gene was downregulated in soybean in response to the seed rot disease, with significant differences in expression levels (sort: 138). It responded to the seed rot disease through metabolic pathway blockage, increased oxidative stress, and dysregulation of defense signals, and had a significant interaction with the GLYMA_20G179400 gene (Score>0.95).

[0124] In summary, based on multi-dimensional data analysis, four genes—GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_12G226400—were identified as key regulatory hubs in soybean's response to pseudostem rot disease. All four genes exhibited significant changes in expression levels and differential splicing events, indicating that they dynamically optimize disease resistance responses through a dual transcriptional and post-transcriptional regulatory mechanism, precisely balancing the intensity and energy consumption of the defense response.

[0125] (2) Tissue expression characteristics of four soybean stem rot response genes

[0126] The expression characteristics of four genes, GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_12G226400, in cotyledons, hypocotyls, shoot tips, leaves, flowers, pods (seeds), seeds, lateral buds, and roots were analyzed on the SoyOmics website (https: / / ngdc.cncb.ac.cn / soyomics / index). The pods (seeds) were immature green pods containing seeds. The results are as follows: Figure 10 As shown, the results indicate that each gene is highly expressed in specific tissues, reflecting their functional division of labor. This may enhance the resistance of soybean to stem rot by synergistic interaction between tissue-specific functional division of labor and metabolic pathways.

[0127] Among them, the tissues in which the GLYMA_18G061100 gene was highly expressed were cotyledons, hypocotyls, flowers, and roots; the tissues in which the GLYMA_17G128000 gene was highly expressed were cotyledons and leaves; the tissues in which the GLYMA_06G050100 gene was highly expressed were cotyledons; and the tissues in which the GLYMA_12G226400 gene was highly expressed were pods and roots.

[0128] (3) Variation analysis of four soybean stem rot response genes

[0129] SNP and InDel variant analyses were performed on four soybean stem rot response genes: GLYMA_18G061100, GLYMA_17G128000, GLYMA_06G050100, and GLYMA_12G226400. The results are shown in Table 5.

[0130] SNP and InDel sites were obtained using the Varscan (v2.3.9) program. The filtering criteria were: SNP site base Q>20; number of reads covering the site>8; number of reads supporting the mutation site>2; 4) p-value of SNP site<0.01;

[0131] Among them, the GLYMA_06G050100 gene was found to have 3 SNPs and 2 InDel variants: 3798193 (T / T→G / G) and 3798053 (C / C→CA / CA) in the UTR3 region, and 3799019 (A / A→T / T), 3799065 (G / G→A / A), and 3798988 (AC / AC→A / A) in the intron region; notably, the GLYMA_18G061100 gene was found to have a key SNP variant (G / ) at splice site 5543283. The variants G→G / A were detected, and the intron region InDel (C / C→C / CT) was detected at position 5543649. Among these variants, the UTR region variants may affect gene transcription regulation, the splice site SNP may lead to abnormal mRNA processing, and the intron region variants may play a role by affecting splice regulatory elements or non-coding RNA. The genotypic differences between healthy plants and those infected with *S. spp.* provide important clues for elucidating the molecular mechanism of resistance. No variant sites were found in the GLYMA_12G226400 and GLYMA_17G128000 genes.

[0132] Table 5. SNP and InDel variation analysis of four soybean stem rot response genes.

[0133]

[0134] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any way. Any simple modifications, alterations, and equivalent changes made to the above embodiments based on the inventive essence shall still fall within the protection scope of the present invention.

Claims

1. A method for screening soybean stem rot response genes based on transcriptome analysis, characterized in that, Includes the following steps: S1. Select pod samples from soybean plants infected with pseudostem spot seed rot and healthy soybean plants grown in the field, extract RNA samples, perform transcriptome sequencing, and obtain sequencing data. S2. Using HTSeq and DESeq2 software, differentially expressed genes in soybean responding to *Pseudomonas stolonifera* were screened based on the original expression levels (Read Count) of the sequencing data in S1. S3. Through GO term / KEGG pathway functional enrichment analysis and gene set enrichment analysis, analyze the biological pathways in which the differentially expressed genes described in S2 participate in the process of susceptibility to stem rot. Use statistical methods to test whether the pre-set gene set is enriched at the top or bottom of the sorting table to determine whether the pathway is activated or inhibited. S4. Based on the STRING database, protein interaction analysis was performed to screen out the core interaction genes, which can be used as design objects for multi-target editing or modular breeding. S5. The aligned sequences obtained from RNA-seq sequencing were assembled from transcripts de novo using StringTie software. By integrating and analyzing differential alternative splicing events and exon usage differences, the dynamic characteristics of the post-transcriptional regulatory network of key candidate genes in soybean response to pseudostem rot were analyzed. S6. Analyze the expression characteristics of key candidate genes in soybean response to pseudostem rot in different soybean tissues. S7. Screening out response genes for soybean pseudostem rot.

2. The method for screening soybean stem rot response genes based on transcriptome analysis according to claim 1, characterized in that, The conditions for using the DESeq2 software described in S2 to screen differentially expressed genes in soybean in response to pseudostem rot are: fold change in expression |log2FoldChange|>1, and significance P-value<0.

05.

3. The method for screening soybean stem rot response genes based on transcriptome analysis according to claim 1, characterized in that, In S3, GO term / KEGG pathway functional enrichment analysis and gene set enrichment analysis were performed. Specifically, clusterProfiler software was used to perform GO term / KEGG pathway functional enrichment analysis and gene set enrichment analysis. The criterion for significant enrichment was P-value < 0.

05.

4. The method for screening soybean stem rot response genes based on transcriptome analysis according to claim 1, characterized in that, The method for protein interaction analysis in S4 is as follows: using the STRING database, protein interaction relationships are screened with a high confidence threshold of Score > 0.

95.

5. The method for screening soybean stem rot response genes based on transcriptome analysis according to claim 1, characterized in that, The specific steps for analyzing the post-transcriptional regulatory network dynamics of key candidate genes in soybean response to pseudostem rot in S5 are as follows: StringTie software was used to perform de novo transcript assembly of the obtained aligned sequences; rMATS software was used to analyze differentially spliced ​​events; and the R language package DEXSeq was used to analyze exon usage differences.

6. The method for screening soybean stem rot response genes based on transcriptome analysis according to claim 1, characterized in that, In S6, the expression characteristics of key candidate genes in soybean response to pseudostem rot in different soybean tissues were analyzed on the SoyOmics website to obtain response genes for soybean pseudostem rot.

Citation Information

Patent Citations

  • SE11545C1

  • Method for constructing gene expression profiles of two anti-infection konjac materials in different Pcc infection stages

    CN118186062A

  • Highly efficient self-assembling programmable genome and epigenome editing tools

    WO2024191816A1