Key gene identification method related to tobacco nitrogen response

By combining field experiments and high-throughput transcriptome sequencing with WGCNA analysis, key genes were identified, solving the problem of unknown response mechanisms of plant nitrogen metabolism systems under high nitrogen conditions, and achieving systematic analysis of gene expression patterns and improvement of nitrogen use efficiency.

CN120989217APending Publication Date: 2025-11-21YUNNAN TOBACCO COMPANY YUXI PREFECTURE COMPANY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511134265.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies have failed to effectively reveal the response mechanism of plant nitrogen metabolism systems under high nitrogen conditions, and have been unable to screen key regulatory genes, resulting in the inability to optimize flue-cured tobacco fertilization strategies and low nitrogen use efficiency.

Method used

A completely randomized block field experiment was conducted with four nitrogen fertilizer gradient treatments. High-throughput transcriptome sequencing and weighted gene co-expression network analysis (WGCNA) were combined to screen for significantly differentially expressed genes and construct co-expression networks. Low-expression genes were removed, and candidate genes with the highest co-expression weights were extracted and verified by RT-qPCR.

Benefits of technology

The system accurately identified candidate core genes highly related to nitrogen assimilation, transport, and signal transduction, enabling the capture and module mining of gene expression patterns under different growth conditions and varietal backgrounds, thereby improving nitrogen use efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120989217A_ABST
    Figure CN120989217A_ABST
Patent Text Reader

Abstract

The invention discloses a key gene identification method related to tobacco nitrogen response, which comprises the following steps: S1, setting four nitrogen fertilizer gradient treatments on the basis of same phosphorus and potassium fertilization by adopting a field experiment of a completely random block; s2, randomly taking the 6th to 8th leaves of the five plants in each area, and dividing a sample into two parts: quickly freezing one part with liquid nitrogen, and storing at-80 DEG C for RNA (Ribonucleic Acid) extraction and transcriptome sequencing; one part is used for measuring the nitrogen content, and after baking and drying treatment, a KjeltecTM8100 automatic nitrogen analyzer is used for measuring; s3, nitrogen content determination: determining the nitrogen content in the treated sample by using an automatic nitrogen analyzer KjeltecTM8100; according to the method, high-throughput transcriptome sequencing and weighted gene co-expression network analysis (WGCNA) are combined, so that not only can gene expression maps of flue-cured tobacco leaves treated by different nitrogen fertilizers be comprehensively captured, but also gene modules with similar expression modes can be mined from a global perspective.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of plant nitrogen metabolism regulation, and particularly to a key gene identification method related to tobacco nitrogen response. BACKGROUND

[0002] In the field of plant nitrogen metabolism regulation, previous studies mainly focused on screening of key genes under nitrogen starvation conditions, and relatively less attention was paid to the transformation process of plant nitrogen metabolism under high nitrogen conditions. In addition, only a few studies have performed RNA-seq and WGCNA analysis on flue-cured tobacco under different nitrogen application rates and leaf retention numbers, and identified some genes that may be involved in nitrogen signal regulation. However, these studies did not analyze the nitrogen application rate alone, and the nitrogen application gradient was within the low nitrogen or normal nitrogen range, which failed to reveal the response mechanism of plant nitrogen metabolism system under high nitrogen conditions. It is difficult to screen key regulatory genes, which is not conducive to optimizing the fertilization strategy of flue-cured tobacco, and cannot effectively improve the nitrogen utilization efficiency. SUMMARY

[0003] The present application aims to provide a key gene identification method related to tobacco nitrogen response to solve the problems in the background.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a key gene identification method related to tobacco nitrogen response, comprising the following steps:

[0005] S1, a completely randomized block field experiment is adopted, and four nitrogen fertilizer gradient treatments are set on the basis of the same phosphorus and potassium fertilization;

[0006] S2, randomly take the 6th-8th leaves of 5 plants in each plot, and divide the samples into two parts:

[0007] One part is quickly frozen in liquid nitrogen and stored at -80℃ for RNA extraction and transcriptome sequencing;

[0008] The other part is used for nitrogen content determination, and after drying treatment, the Kjeltec TM 8100 automatic nitrogen analyzer is used for determination;

[0009] S3, nitrogen content determination: the Kjeltec TM 8100 automatic nitrogen analyzer is used to determine the nitrogen content in the treated samples;

[0010] S4, RNA extraction and high-throughput sequencing;

[0011] S5, data quality control and alignment, expression calculation and differential expression analysis, and screening of significant differential genes;

[0012] S6. Construct a co-expression network based on the WGCNA method, remove low-expression genes, and extract candidate genes with the highest co-expression weight;

[0013] S7. RT-qPCR verification of core differentially expressed genes.

[0014] Preferably, the content of the same phosphorus and potassium is 105 kg / hm² for phosphorus. 2 And potassium 210 kg / hm 2 The four nitrogen fertilizer gradients include:

[0015] T1: No nitrogen application (0 kg / hm) 2 )

[0016] T2: Low nitrogen (60 kg / hm) 2 )

[0017] T3: Standard nitrogen (105 kg / hm) 2 )

[0018] T4: High nitrogen (150 kg / hm) 2 ).

[0019] Preferably, S4 includes transcriptome sequencing, which includes:

[0020] Total RNA was extracted from tobacco leaf samples that were flash-frozen and pulverized in liquid nitrogen using TRIzol reagent;

[0021] RNA purity (A260 / A280) and concentration were determined using a UV spectrophotometer, and the quality of the constructed cDNA library was assessed using an Agilent 2100 Bioanalyzer to ensure RIN ≥ 7.0.

[0022] mRNA was enriched using oligo dT magnetic beads and fragmented (200–300 bp) at 94 °C.

[0023] First and second strands were synthesized using random primers, followed by end repair, A-tailing, adapter ligation, and PCR enrichment. The distribution of library fragments was then assessed again using an Agilent 2100 Bioanalyzer.

[0024] 150bp paired-end sequencing was performed using the Illumina NovaSeq6000 platform, generating approximately 40,750 raw reads per sample.

[0025] Preferably, step S5 includes differential gene expression analysis, and the method for differential gene expression analysis is as follows:

[0026] Use the FASTP software to filter out low-quality and connector-contaminated sequences;

[0027] Alignment to the tobacco reference genome using HISAT2 allows for a maximum of 2 base mismatches;

[0028] The FPKM values ​​of each gene were calculated and differential expression analysis was performed. Differentially expressed genes (DEGs) were identified with FDR < 0.05 and Fold change ≥ 1.5 as thresholds.

[0029] Preferably, step S6 includes weighted gene co-expression network analysis, and the method for weighted gene co-expression network analysis is as follows:

[0030] Use the WGCNA package in R software to read the gene expression matrix of all samples;

[0031] Genes with expression standard deviation SD ≤ 0.5 were removed, while hypervariable genes with SD > 0.5 were retained;

[0032] Construct an adjacency matrix and convert it into a topology reconstruction matrix; dynamically merge similar modules into larger units.

[0033] Based on gene function annotation, modules related to nitrogen uptake and utilization were identified, and the top 6 co-expression weighted genes were selected as candidate genes.

[0034] Preferably, the RT-qPCR verification step in S7 includes: designing primers using Primer Premier, with the amplified fragment lengths of the target gene and the internal reference gene L25 being 100–200 bp;

[0035] The qPCR reaction system was 20 μL, and the reaction program consisted of 40 cycles. Each sample was triple-replicated. L25 was used as the internal reference gene. Finally, the expression levels of the nine candidate genes were calculated using the 2^-ΔΔCT method.

[0036] Compared with existing technologies, the advantages of this invention are as follows: This invention combines high-throughput transcriptome sequencing with weighted co-expression network analysis (WGCNA), which can not only comprehensively capture the gene expression profiles of flue-cured tobacco leaves under different nitrogen fertilizer treatments, but also discover gene modules with similar expression patterns from a global perspective; This invention can accurately identify candidate core genes highly related to nitrogen assimilation, translocation, and signal transduction through dynamic tree pruning and module merging, combined with a core network screening strategy based on weight thresholds; This method is compatible with different growth conditions and varietal backgrounds, and the resulting analytical framework can be extended to other crops or environmental stress studies, realizing the systematic analysis of plant molecular response networks to different factors, and has broad application prospects. Attached Figure Description

[0037] Figure 1 Differential gene expression levels under different nitrogen treatments;

[0038] Figure 2 A diagram of a co-expression network constructed using nitrogen uptake-related genes;

[0039] Figure 3 This is the result of RT-qPCR validation. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Please see Figures 1-3 This invention provides a technical solution: a method for identifying key genes related to tobacco nitrogen response, comprising:

[0042] S101, Field phenotypic observation and leaf nitrogen content determination;

[0043] S102, RNA extraction and library construction;

[0044] S103, High-throughput sequencing and data quality control;

[0045] S104, Differentially expressed gene (DEG) screening;

[0046] S105 and KEGG pathway enrichment analysis;

[0047] Construction of S106 and WGCNA co-expression network;

[0048] S107, Core gene screening and functional annotation;

[0049] S108, RT-qPCR verification.

[0050] In one embodiment, in step S101, the experimental material preparation uses the tobacco variety "Yunyan 87", the soil is yellow soil, and the experimental site is Chunhe Community, Hongta District, Yuxi City, Yunnan Province. Tobacco seedlings are cultivated using the standard floating bed method and transplanted on April 28, 2022, when they are at the four-leaf stage, healthy and free from pests and diseases. A nitrogen fertilizer gradient and field management are set: firstly, the phosphorus and potassium fertilizer application conditions are fixed at 105 kg / hm². 2 (phosphate fertilizer) and 210 kg / hm 2 Under (potassium fertilizer) conditions, four nitrogen fertilizer treatment gradients were set up, namely: T1: 0 kg / hm 2 (No nitrogen application), T2: 60 kg / hm 2 (Nitrogen deficiency), T3: 105 kg / hm2 (Standard nitrogen application), T4: 150 kg / hm 2 (Excess nitrogen); secondly, the field experiment adopted a completely randomized block design, with each plot area of ​​66.7m². 2 The experiment was repeated three times, with a planting density of 16,492 plants / hm². 2 (Row spacing 1.10m, plant spacing 0.55m), with protective rows around the perimeter. For fertilizer application, nitrogen and potassium were applied using a "70% basal application + 30% topdressing" method, while all phosphorus fertilizer was applied basally. The following samples were used for field leaf nitrogen content determination: Leaves from leaves numbered 6–8 of 5 plants in each block were randomly selected and divided into two portions. One portion was flash-frozen in liquid nitrogen and stored at -80℃ for nitrogen content determination; the other portion was dried, pulverized, and analyzed using Kjeltec... TM The 8100 fully automatic Kjeldahl nitrogen analyzer is used to determine the total nitrogen content.

[0051] In one embodiment, step S102, which involves RNA extraction and library construction from the tissue sample to obtain high-throughput sequencing data, specifically includes the following steps:

[0052] S301. Take leaf samples stored at -80℃ and extract total RNA using TRIzol reagent (Invitrogen);

[0053] S302. RNA purity and integrity were assessed using a UV spectrophotometer and an Agilent 2100 Bioanalyzer, with A260 / A280 = 1.8–2.1 and RIN ≥ 7.0 as the acceptance criteria.

[0054] S303. mRNA was enriched using oligodT magnetic beads and fragmented to 200–300 bp at 94 °C.

[0055] S304. Synthesize double-stranded cDNA, perform end repair, add A tail, ligate adapters, and enrich PCR (8–12 rounds);

[0056] S305. Verify the distribution of library fragments again using Agilent 2100 Bioanalyzer;

[0057] S306. Perform library quality assessment on the prepared library;

[0058] S307. Detect cDNA integrity;

[0059] S308, high-throughput sequencing.

[0060] In one embodiment, the preprocessing of high-throughput sequencing data of leaf samples in step S103 specifically includes the following steps:

[0061] S401. Filter with FASTP, which involves aligning the FASTQ file to the reference genome, removing PCR duplicates, removing low-quality and adapter sequences, and generating clean reads.

[0062] 1. Data Reading

[0063] Following the standard FASTP workflow for processing RNA sequencing data, the sequencing R1 / R2 files for each sample were read and converted into valid BAM objects. Each sample contained 2.7–7.31 Gb of valid data, with a Q30 ≥ 94.72% and a GC content of approximately 43.1%.

[0064] 2. Data comparison

[0065] First, clean reads are aligned to the tobacco reference genome using HISAT2 (allowing ≤2 mismatches). Then, featureCounts or StringTie are used to calculate the FPKM of each gene to facilitate subsequent genome-wide operations.

[0066] S403, differentially expressed gene (DEG) screening;

[0067] 1. Difference screening

[0068] DESeq2 was used for differential analysis, and DEG was selected based on FDR < 0.05 and Fold change ≥ 1.5. Then, comparisons between groups were performed: T2 vs T1, T3 vs T1, and T4 vs T1.

[0069] 2. KEGG pathway enrichment analysis

[0070] The top 10 pathways of each control group were extracted from the screened DEGs in the KEGG database.

[0071] Construction of S405 and WGCNA co-expression network;

[0072] 1. Screening of candidate key genes

[0073] First, genes with SD≤0.5 were removed from the whole genome set, leaving 7454 hypervariable genes;

[0074] 2. Construction of co-expression networks

[0075] In the R environment, the WGCNA package (v1.6.6) was called, and the soft threshold power=6 was determined using pickSoftThreshold. Then, a topology reconstruction matrix (TOM) was constructed, and the initial modules were identified using dynamic pruning (minModuleSize=30). Similar modules were merged with mergeCutHeight=0.25, resulting in 31 modules. The correlation between each module and the leaf nitrogen content phenotype was calculated, and black and lightcyan1 were identified as key response modules. In the above modules, the top 50 genes were extracted according to their co-expression weight (Weight>0.1) to construct the module subnetwork.

[0076] 3. Locating transcription factors

[0077] Multiple modules extract the corresponding transcription factors.

[0078] S406, Key Gene Screening;

[0079] Using nitrogen metabolism-related DEGs as bait, key genes were further screened from the co-expression network module subnetwork.

[0080] S407, Key Gene Function Annotation;

[0081] PlantCARE was used to predict the response elements on its promoter and to analyze the distribution of low temperature, hormones (ABA, MeJA, gibberellin) and light-responsive elements.

[0082] S408, RT-qPCR verification;

[0083] 1. Design primers

[0084] Primers for the target gene and internal control L25 (amplifying fragment 100–200 bp) were designed using Primer Premier 5.0.

[0085] 2. qPCR reaction system and procedure

[0086] The total reaction volume was 20 μL: 10 μL of 2×Talent qPCR PreMix, 1 μL each of primers (10 μM), 1.6 μL of cDNA template, and ddH2O supplementation. The reaction program was 95℃ for 30 s; 95℃ for 5 s / 60℃ for 30 s, for a total of 40 cycles.

[0087] Example 2

[0088] Based on the method in Example 1, this embodiment conducts phenotypic observation, nitrogen content determination, RNA-seq, differential gene analysis, and WGCNA module mining on flue-cured tobacco samples under different nitrogen fertilizer treatments. The main results are as follows.

[0089] Step 1: Differential genes in sample specimens under different nitrogen treatments

[0090] a. As the amount of nitrogen applied increases, the leaf color changes from light to dark. The T4 treatment shows the strongest "green preservation" characteristics and the largest leaf area. The total nitrogen content of leaves is 48% higher in T4 than in T3, 91% of that in T2, and the lowest in T1, which is only 70% of that in T3.

[0091] b. A total of 70.80 Gb of raw RNA-seq data was generated from 12 samples (4 treatments, 3 replicates per treatment). The effective data volume of each sample was 2.7–7.31 Gb, the Q30 base ratio was 94.72%–95.8%, and the average GC content was 43.08%, which met the requirements for high-quality sequencing.

[0092] c. Figure 1 The distribution of differentially expressed genes was shown: 1458 differentially expressed genes (DEGs) were identified in T2vsT1, of which 987 were upregulated and 471 were downregulated; 1877 DEGs were identified in T3vsT1, of which 1044 were upregulated and 833 were downregulated; and 2147 DEGs were identified in T4vsT1, of which 1344 were upregulated and 803 were downregulated. This indicates that the number of DEGs increases with nitrogen application, and the proportions of upregulated genes in the T2vsT1, T3vsT1, and T4vsT1 groups were 67.7%, 55.62%, and 62.6%, respectively.

[0093] Step 2: The WGCNA module identifies and screens key genes related to nitrogen metabolism. The results are shown in the attached table. Figure 2 As shown:

[0094] a. The differentially expressed genes identified in the first step were clustered into 31 modules. Among them, four genes involved in nitrogen metabolism pathways appeared in the black module and two appeared in the lightcyan1 module. This indicates that the black and lightcyan1 modules are key modules for tobacco leaves to respond to different nitrogen levels.

[0095] The b.black module contains 169 genes, while the lightcyan1 module contains 281 genes. The gene co-expression network of the black module includes the MYB family transcription factor NtSRM1 (Nitab4.5_0000153g0080), and the network of the lightcyan1 module includes two classes of transcription factors: the AP2 / ERF-ERF family NtRAP2-10 (Nitab4.5_0001603g0050) and the WRKY family NtWRKY3 (Nitab4.5_0001591g0020).

[0096] c. Figure 2Using differentially expressed genes related to nitrogen absorption as "bait," a co-expressed gene subnetwork was constructed within the black and lightcyan1 modules, and five key regulatory genes (glutamine synthase-like, GDSL lipase, SGNH hydrolase, etc.) were identified.

[0097] d. These key regulatory genes play a central role in nitrogen assimilation and translocation. The promoters of all five genes are rich in light-response, hormone (ABA, GA, MeJA), stress (drought, low temperature) and development (meristematic tissue, mesophyll differentiation, diurnal rhythm) related elements, indicating that these key genes have multiple synergistic regulatory roles in environmental adaptation and growth and development, and also provide potential targets for molecular breeding.

[0098] Step 3: RT-qPCR verification, results are attached. Figure 3 As shown.

[0099] a. Nine representative genes (including five core genes and four accessory genes related to nitrogen metabolism or stress response) were selected from the differentially expressed genes. L25 was used as an internal reference gene, and the comparison threshold cycling method (2^–ΔΔCT method) was employed for validation. Figure 3 As shown, the expression trends of these key genes are highly consistent with RNA-seq data, demonstrating the accuracy and reliability of this method.

[0100] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for identifying key genes related to tobacco nitrogen response, characterized in that: Includes the following steps: S1. A field experiment with completely randomized block design was conducted, with four nitrogen fertilizer gradient treatments set up on the basis of the same phosphorus and potassium fertilization. S2. Randomly select the 6th–8th leaves from 5 plants in each area, and divide the samples into two parts: One sample was flash-frozen in liquid nitrogen and stored at -80°C for RNA extraction and transcriptome sequencing; One sample was used for nitrogen content determination, and was baked and dried before being used with Kjeltec. TM Measured by 8100 automated nitrogen analyzer; S3. Nitrogen content determination: using an automated nitrogen analyzer, Kjeltec. TM 8100 was used to determine the nitrogen content in the treated samples; S4. RNA extraction and high-throughput sequencing; S5. Data quality control and comparison: Gene alignment, expression level calculation and differential expression analysis are performed to screen out significantly differentially expressed genes. S6. Construct a co-expression network based on the WGCNA method, remove low-expression genes, and extract candidate genes with the highest co-expression weight; S7. RT-qPCR verification of core differentially expressed genes.

2. The method for identifying key genes related to tobacco nitrogen response according to claim 1, characterized in that: The content of phosphorus and potassium is the same, with phosphorus at 105 kg / hm². 2 And potassium 210 kg / hm 2 The four nitrogen fertilizer gradients include: T1: No nitrogen application (0 kg / hm) 2 ) T2: Low nitrogen (60 kg / hm) 2 ) T3: Standard nitrogen (105 kg / hm) 2 ) T4: High nitrogen (150 kg / hm) 2 ).

3. The method for identifying key genes related to tobacco nitrogen response according to claim 1, characterized in that: S4 includes transcriptome sequencing, which includes: Total RNA was extracted from tobacco leaf samples that were flash-frozen and pulverized in liquid nitrogen using TRIzol reagent; RNA purity (A260 / A280) and concentration were determined using a UV spectrophotometer, and the quality of the constructed cDNA library was assessed using an Agilent 2100 Bioanalyzer to ensure RIN ≥ 7.

0. mRNA was enriched using oligo dT magnetic beads and fragmented (200–300 bp) at 94 °C. First and second strands were synthesized using random primers, followed by end repair, A-tailing, adapter ligation, and PCR enrichment. The distribution of library fragments was then assessed again using an Agilent 2100 Bioanalyzer. 150bp paired-end sequencing was performed using the Illumina NovaSeq6000 platform, generating approximately 40,750 raw reads per sample.

4. The method for identifying key genes related to tobacco nitrogen response according to claim 1, characterized in that: S5 includes differential gene expression analysis, and the method for differential gene expression analysis is as follows: Use the FASTP software to filter out low-quality and connector-contaminated sequences; Alignment to the tobacco reference genome using HISAT2 allows for a maximum of 2 base mismatches; The FPKM values ​​of each gene were calculated and differential expression analysis was performed. Differentially expressed genes (DEGs) were identified with FDR < 0.05 and Fold change ≥ 1.5 as thresholds.

5. The method for identifying key genes related to tobacco nitrogen response according to claim 1, characterized in that: S6 includes weighted gene co-expression network analysis, and the method for weighted gene co-expression network analysis is as follows: Use the WGCNA package in R software to read the gene expression matrix of all samples; Genes with expression standard deviation SD ≤ 0.5 were removed, while hypervariable genes with SD > 0.5 were retained; Construct an adjacency matrix and convert it into a topology reconstruction matrix; dynamically merge similar modules into larger units. Based on gene function annotation, modules related to nitrogen uptake and utilization were identified, and the top 6 co-expression weighted genes were selected as candidate genes.

6. The method for identifying key genes related to tobacco nitrogen response according to claim 1, characterized in that: The RT-qPCR verification step in S7 includes: using Primer Premier to design primers, and amplifying fragments of the target gene and internal reference gene L25 with a length of 100–200 bp; The qPCR reaction system was 20 μL, and the reaction program consisted of 40 cycles. Each sample was triple-replicated. L25 was used as the internal reference gene. Finally, the expression levels of the nine candidate genes were calculated using the 2^-ΔΔCT method.

Citation Information

Cited By

  • Method for identifying key gene module of sugar acid metabolism of prunus mume fruits and application thereof

    CN121629027A

  • A method for identifying key gene modules of sugar and acid metabolism in plum fruit and its application

    CN121629027B