A method for constructing a pan-transcriptome within a plant genus and its application
By integrating the second and third generation transcriptome data and building the pan-transcriptionome within the plant genus based on ORF length classification, the problem of pan-transcriptionome construction of different species within the genus was solved, and the systematic analysis of small peptides and non-coding RNA was realized, providing multi-dimensional research data support.
Patent Information
- Application Number
- CN202510822132.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art lacks a method for building a pan-transcriptionome of different species within the plant genus, and fails to systematically analyze the transcriptional characteristics and cross-species conservatism of small peptides and non-coding RNA.
By integrating the second- and third-generation transcriptome sequencing data of different species within the plant genus, combining open reading frame length classification transcripts, bioinformatics methods were used to identify gene conservatism among species, and a genus-level pan-transcriptionome containing core transcriptomes, optional transcriptomes and specific transcriptomes were constructed.
Cross-species transcriptome research was achieved, systematically analyzing the conservatism and differences between protein-coding genes, small peptides and non-coding RNAs, providing a multi-dimensional data basis for plant functional genome research.
Smart Images

Figure CN120356516B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to plant transcriptomics, bioinformatics data processing, and a pan-transcriptome construction method and applications thereof. Background Art
[0002] Some plant pan-transcriptome construction methods primarily focus on diverse populations within a single species. For example, in rice, software such as minimap2, TAMA, and SQANTI3 have been used to integrate single-species transcriptome data and annotate novel transcripts (Zhong et al., 2024). However, there is a lack of pan-transcriptome construction strategies for diverse species within a plant genus. Furthermore, current plant transcriptome analyses focus on protein-coding genes. However, small peptides (with varying length definitions, such as 5–100 amino acids) and non-coding RNAs (functional RNAs that do not encode proteins) that are widely present in plants have not been systematically analyzed due to unknown functions and a lack of analytical methods (Choi et al., 2019; Khavinson et al., 2021; Ning et al., 2018). With the advancement of high-throughput sequencing technologies, integrated analysis methods are urgently needed to analyze massive amounts of transcriptome data to reveal the conservation and diversity of protein-coding genes, small peptides, and non-coding RNAs among species within a genus.
[0003] Therefore, a plant transcriptomics, bioinformatics data processing and pan-transcriptome construction method is urgently needed to solve the problem that the existing technology lacks a pan-transcriptome construction method for different species within a plant genus and has not systematically analyzed the transcriptional characteristics and cross-species conservation of small peptides and non-coding RNAs. Summary of the Invention
[0004] The purpose of the present invention is to address the above-mentioned problems existing in the prior art and provide a method for constructing a pan-transcriptome within a plant genus. By integrating the second and third generation transcriptome sequencing data of different species within the plant genus, combining the open reading frame (ORF) length to classify transcripts (protein encoding, small peptide transcripts, non-coding RNA), and using bioinformatics methods to identify gene conservation between species, a genus-level pan-transcriptome comprising a core transcriptome, an optional transcriptome, and a specific transcriptome is constructed.
[0005] In order to achieve the above application objectives, the present invention adopts the following technical solution: A method for constructing a pan-transcriptome in a plant genus comprises the following steps:
[0006] (1) Data preparation and quality control: Obtain second-generation and third-generation transcriptome sequencing data of different species within the plant genus, including samples from different tissues and developmental stages; Use quality control software to perform quality control on the sequencing data and remove low-quality bases and adapter sequences;
[0007] (2) Transcript splicing and screening: The second and third generation transcriptome sequencing data after quality control are aligned to the reference genome of the corresponding species, and the second and third generation splicing results are merged using transcript splicing software, and transcripts above the preset coverage threshold are retained; valid transcripts are screened by comparing with reference gene annotations;
[0008] (3) Transcript classification based on open reading frames: The open reading frames of the effective transcripts of each species are identified, and the transcripts are divided into protein-coding transcripts, small peptide-coding transcripts and non-coding RNA according to the length of the open reading frames; and based on the transcript classification results, the gene loci are determined to be protein-coding genes, small peptide-coding genes or non-coding genes;
[0009] (4) Identification of gene conservation among species: For genes encoding proteins or small peptides, core genes, optional genes, and specific genes are identified based on amino acid sequence homology analysis; for non-coding genes, core non-coding genes, optional non-coding genes, and specific non-coding genes are identified based on nucleotide sequence homology analysis;
[0010] (5) Pan-transcriptome integration construction: core genes, optional genes, specific genes and their corresponding transcripts are combined with core non-coding genes, optional non-coding genes, specific non-coding genes and their corresponding transcripts to construct a genus-level pan-transcriptome.
[0011] Furthermore, the second-generation transcriptome sequencing data in step (1) is the paired-end sequencing data generated by the Illumina or BGI platform, and the third-generation transcriptome sequencing data is selected from the long-read sequence data generated by the PacBio Sequel II or Nanopore platform.
[0012] Furthermore, in step (2), the second-generation transcriptome sequencing data were aligned using HISAT2 software, the third-generation transcriptome sequencing data were aligned using Minimap2 software, and the transcripts were spliced using StringTie software. The preset coverage threshold was ≥5 sequencing reads.
[0013] Furthermore, the reference gene annotation comparison in step (2) was implemented using gffcompare software, and transcripts with comparison results of “=, c, x, j, o, i, u, p” were retained;
[0014] Where "=" indicates that the predicted transcript is completely identical to the reference transcript; "c" indicates that the predicted transcript is contained within the reference transcript; "x" indicates that the predicted exon overlaps with the reference transcript on the reverse strand; "j" indicates that at least one splice junction of the predicted transcript is shared with the reference transcript; "o" indicates that part of the exon of the predicted transcript overlaps with the reference transcript; "i" indicates that the predicted transcript is completely aligned to the intron of the reference transcript; "u" indicates that the predicted transcript is within the gene interval; and "p" indicates that the predicted transcript is within 2 kb of the reference transcript.
[0015] Furthermore, the criteria for transcript classification in step (3) are:
[0016] Transcripts with an open reading frame length of ≥100 amino acids were identified as protein-encoding transcripts, transcripts with an open reading frame length of 10 amino acids ≤ and <100 amino acids were identified as small peptide-encoding transcripts, and transcripts with an open reading frame length of <10 amino acids or no open reading frame were identified as non-coding RNAs.
[0017] Furthermore, the gene locus classification rule in step (3) is:
[0018] Gene loci with at least one protein-coding transcript were identified as protein-coding genes;
[0019] Gene loci that do not encode proteins but have at least one transcript encoding a small peptide are identified as genes encoding small peptides;
[0020] Gene loci where all transcripts are non-coding RNAs are considered non-coding genes.
[0021] Furthermore, in step (4), genes present in all analyzed species were considered core genes, genes present only in a single species were considered specific genes, and genes present in ≥2 species were considered optional genes.
[0022] Furthermore, in step (4), an E value of <1e-5 was used as the screening threshold and the best alignment result was taken. Genes present in all analyzed species were considered core non-coding genes, genes present only in a single species were considered specific non-coding genes, and genes present in ≥2 species were considered optional non-coding genes.
[0023] Furthermore, in step (5), the pan-transcriptome includes a core transcriptome consisting of core genes + core non-coding genes and their transcripts, an optional transcriptome consisting of optional genes + optional non-coding genes and their transcripts, and a specific transcriptome consisting of specific genes + specific non-coding genes and their transcripts.
[0024] For example, the above-mentioned construction method is applied to improve plant stress resistance traits, and regulatory elements related to adverse stress response are screened by analyzing gene expression characteristics in specific transcriptomes.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] 1. Break through the limitations of a single species and achieve genus-level pan-transcriptome construction
[0027] Existing technologies only construct pan-transcriptomes for different populations of a single species. The present invention, for the first time, provides a method for constructing pan-transcriptomes for different species within a plant genus. By integrating second- and third-generation transcriptome data from multiple species, it systematically analyzes the conservation (such as core genes) and specificity (such as species-specific genes) of protein-coding genes, small peptides, and non-coding RNAs within the genus, filling the gap in genus-level transcriptome research.
[0028] 2. Integration of second- and third-generation data improves comprehensive transcript identification
[0029] Combining the advantages of second-generation sequencing (high precision, low cost) and third-generation sequencing (long reads, identification of new transcripts), the completeness of transcript identification is significantly improved.
[0030] 3. ORF-based classification method clarifies the coding characteristics of transcripts
[0031] To address the problem of ambiguous definitions of protein-coding genes and non-coding RNA in the existing technology, the present invention uses the ORF length standard (≥100 amino acids for mRNA, 10-99 amino acids for small peptide transcripts, and <10 amino acids for ncRNA) to achieve for the first time the accurate classification of transcript coding capacity, and accordingly divide gene types (protein-coding genes, small peptide genes, non-coding genes), providing clear targets for functional research.
[0032] 4. Systematic analysis of cross-species characteristics of non-coding RNA and small peptides
[0033] Existing technologies primarily focus on coding genes, neglecting the regulatory functions of small peptides and non-coding RNAs. This present invention, through conservation analysis (based on amino acid sequences for coding genes and small peptides, and nucleotide sequences for non-coding genes), identifies core, optional, and specific non-coding genes and small peptide genes at the genus level for the first time. For example, the core non-coding genes of the Oryza genus are enriched in pathways essential for life activities, while the specific transcriptome contains genes related to stress resistance, providing new avenues for functional research.
[0034] 5. Construct a multi-dimensional pan-transcriptome to support downstream in-depth analysis
[0035] The integrated genus-level pan-transcriptome includes the core transcriptome (conserved genes shared by species), the optional transcriptome (genes present in some species) and the specific transcriptome (species-specific genes), which can be further used for gene expression quantification and functional enrichment analysis (such as the enrichment of rice core genes in the osmotic pressure response pathway), providing a systematic data basis for research on plant evolution, stress resistance mechanisms, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a graph showing the difference in transcript quantity between the second and third generation data of different species in an embodiment of the present invention;
[0037] Figure 2 is a comparison chart of the number of transcripts identified according to the present invention and reference transcripts in different species in the embodiments of the present invention;
[0038] Figure 3 is a diagram of transcript classification of different species according to an embodiment of the present invention;
[0039] Figure 4 is a diagram of the gene classification of different species according to an embodiment of the present invention;
[0040] Figure 5 This is a diagram showing the conservation classification results of genes encoding proteins and small peptides of Oryza species according to an embodiment of the present invention;
[0041] Figure 6 2. It is a diagram showing the conservative classification results of non-coding genes of Oryza species according to an embodiment of the present invention;
[0042] Figure 7 is a flowchart of steps 1 and 2 of an embodiment of the present invention;
[0043] Figure 8 Flowchart of step 3 and step 4 of an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention are within the scope of protection of the present invention.
[0045] Those skilled in the art should understand that, in the disclosure of the present invention, the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms cannot be understood as limiting the present invention.
[0046] like Figure 7 and Figure 8 As shown, the method for constructing the pan-transcriptome of this plant genus includes the following steps:
[0047] Step 1: Construct a transcript assembly process based on plant second and third generation transcriptome sequencing data. The specific steps are as follows:
[0048] (1a) Sequencing data preparation:
[0049] Download transcriptome sequencing (RNA-seq) data from public databases such as NCBI (https: / / www.ncbi.nlm.nih.gov / ) or NGDC (https: / / ngdc.cncb.ac.cn / ) for different tissues and different stages of plant species (e.g., different rice species), including both second- and third-generation sequencing data. Alternatively, prepare your own samples and generate data using high-throughput sequencing platforms such as Illumina, PacBio, Nanopore, or BGI. The purpose of using second- and third-generation transcriptome sequencing data from different stages and tissues is to fully identify candidate transcript sites. For example:
[0050] Based on four species of Oryza genus (cultivated rice Oryza sativa 、Wild rice 1 Oryza barthii 、Wild Rice 2 Oryza nivara and wild rice 3 Oryza rufipogon ) to construct the pan-transcriptome of the genus Oryza, and collected public transcriptome sequencing data and self-test data from the NCBI and NGDC databases (Appendix 1). The sequencing platforms included the second-generation Illumina PE150, the third-generation PacBio Sequel II, and the third-generation Nanopore. Oryza sativa ) data volume is about 2T; the total data volume of the three wild rice species is about 1.5T. According to the "step 1" method to identify transcripts, it was found that the third-generation transcriptome data can identify a large number of new transcripts, especially long transcripts ( Figure 1 Compared with the reference annotation information of species, the comprehensive use of large amounts of transcriptome sequencing data (including second and third generation) can identify more new transcripts ( Figure 2 , the number of transcript identifications increased by more than half).
[0051] Table 1
[0052]
[0053] Figure 1 In the figure, the horizontal axis is the length range of the transcript, in bp; "S2" represents the identification result of the second-generation sequencing data (blue), "S3" represents the identification result of the third-generation sequencing data (orange), and "S2&S3" represents the identification result of the integrated second and third-generation sequencing data (yellow). Figure 2In the figure, blue indicates the number of transcripts identified using second- and third-generation sequencing data, orange indicates the number of transcripts in the reference annotation file, and yellow indicates the number of transcripts after the two are combined.
[0054] (1b) Sequencing data quality control:
[0055] Sequencing data were quality-checked using FastQC (https: / / github.com / s-andrews / FastQC) and low-quality bases and adapter sequences were removed using software such as Trim Galore (https: / / github.com / FelixKrueger / TrimGalore), fastp (https: / / github.com / OpenGene / fastp), or fastplong (https: / / github.com / OpenGene / fastplong). Fastp is suitable for second-generation short-read sequencing data, while fastplong is suitable for third-generation long-read sequencing data.
[0056] (1c) Sequencing data alignment and transcript splicing:
[0057] Second- and third-generation transcriptome sequencing (RNA-seq) data (referred to as second-generation data and third-generation data) were aligned to the species reference genomes, with species reference genomes and annotation information obtained from databases such as NCBI and Ensembl Plants. HIAST2 (https: / / daehwankimlab.github.io / hisat2 / ) was used for second-generation data alignment, and Minimap2 (https: / / github.com / lh3 / minimap2) was used for third-generation data alignment. All alignments were assembled using StringTie (https: / / ccb.jhu.edu / software / stringtie / ) software (retaining transcripts supported by at least five sequencing reads; other similar alignment and assembly software can also be used). The "-L" parameter was used for long read processing for third-generation data during assembly. Transcripts from second- and third-generation data of the same species were merged using the merge function in StringTie.
[0058] (1d) Transcript screening:
[0059] The software gffcompare (https: / / github.com / gpertea / gffcompare) was used to compare the reference gene annotation information of the corresponding species with the transcript information obtained in (1c). Transcripts with the comparison results of "=, c, x, j, o, i, u, p" were retained. Finally, the predicted transcripts after screening were merged with the reference transcripts using the merge function of the StringTie software to obtain the union of all transcript information for each species.
[0060] Where "=" indicates that the predicted transcript is completely identical to the reference transcript; "c" indicates that the predicted transcript is contained within the reference transcript; "x" indicates that the predicted exon overlaps with the reference transcript on the reverse strand; "j" indicates that at least one splice junction of the predicted transcript is shared with the reference transcript; "o" indicates that part of the exon of the predicted transcript overlaps with the reference transcript; "i" indicates that the predicted transcript is completely aligned to the intron of the reference transcript; "u" indicates that the predicted transcript is within the gene interval; and "p" indicates that the predicted transcript is within 2 kb of the reference transcript.
[0061] Step 2: The specific steps for classifying transcripts based on open reading frames are as follows:
[0062] (2a) Open reading frame (ORF) identification:
[0063] Use open reading frame identification software such as ORFfinder or TransDecoder to identify the open reading frames of all transcripts from each species obtained in step 1. Count the lengths of all open reading frames; for any transcript, select the longest open reading frame as the likely amino acid sequence encoded by that transcript.
[0064] (2b) Classification of transcripts:
[0065] Based on the results of (2a), each transcript corresponds to 1 or 0 sequences with open reading frames. All transcripts in each species are classified according to the length of their open reading frames. Specifically, transcripts with an ORF length ≥ 100 amino acids are determined to be protein-encoding transcripts, i.e., messenger RNA (mRNA); transcripts with an ORF length ≥ 10 amino acids and an ORF length < 100 amino acids are determined to be small peptide-encoding transcripts; transcripts with an ORF length < 10 amino acids or no predicted ORF are determined to be non-protein-encoding transcripts, i.e., non-coding RNA (ncRNA).
[0066] (2b) Classification of genes:
[0067] A gene locus may produce multiple transcripts. As long as one transcript is determined to encode a protein, that is, mRNA, the gene is considered a protein-coding gene. In the absence of protein-coding transcripts (mRNA), as long as one transcript is determined to encode a small peptide, the gene is considered a small peptide-coding gene. If all transcripts do not encode proteins or small peptides, that is, ncRNA, the gene is considered a non-coding gene.
[0068] For example, when the identified transcripts were classified according to the “Step 2 (2a)” method, it was found that the most common transcripts in Oryza species were protein-encoding transcripts, followed by small peptide-encoding transcripts, and finally non-coding RNA ( Figure 3 The transcripts generated at the same site were defined as a gene locus, and the genes were classified according to the "Step 2 (2b)" method. It was found that the most common genes in rice species were those encoding proteins, followed by those encoding small peptides, and finally non-coding genes ( Figure 4 ).
[0069] in, Figure 3 Medium blue represents non-coding transcripts (ncRNA transcripts), orange represents protein-coding transcripts (mRNA transcripts), and yellow represents small peptide-coding transcripts (peptide transcripts). Figure 4 In the figure, blue represents non-coding genes (ncRNA genes), orange represents protein-coding genes (mRNA genes), and yellow represents small peptide-coding genes (peptide genes).
[0070] Step 3: The specific steps for identifying conservation of protein-coding genes and non-coding genes between species are as follows:
[0071] (3a) Conservative identification of genes encoding proteins or small peptides:
[0072] Based on the gene classification determined in step 2, for genes encoding proteins or small peptides, the amino acid sequence corresponding to the longest ORF was selected and orthologous genes between different species were identified using OrthoFinder (https: / / github.com / davidemms / OrthoFinder). Genes present in all analyzed species were considered core genes; genes present in only one species were considered specific genes; genes present in two or more species were considered accessory genes.
[0073] (3b) Conservation identification of non-coding genes:
[0074] Based on the gene classification determined in step 2, for non-coding genes, the longest transcript (i.e., nucleotide sequence) is selected. The "-evalue" parameter (i.e., E-value) in BLASTN is used to define a desired threshold to determine base sequence conservation. The E-value indicates the probability that other sequences would be more similar to the target sequence than the displayed sequence, given a random chance. Setting an E-value of less than 1e-5 is a standard and stringent standard for homologous sequence alignment and can significantly reduce nonspecific matches. The lower the E-value, the more stringent the screening criteria.
[0075] Specifically, non-coding transcript sequences from all analyzed species were paired-wise aligned, using an E-value of 1e-5 as a screening criterion to identify homologous transcripts. For each transcript, the single best alignment was retained. Non-coding genes present in all analyzed species were considered core non-coding genes; those present in only one species were considered specific non-coding genes; and those present in two or more species were considered accessory non-coding genes. Core non-coding genes may have specific biological functions and serve as candidates for subsequent gene function verification.
[0076] For example, according to the "Step 3 (3a)" method, conservation analysis of genes encoding proteins and small peptides was performed, and it was found that the number of optional genes (accessory genes) in the genus Oryza was the largest, followed by core genes (core genes), and finally specific genes ( Figure 5 Similarly, according to the "step 3 (3b)" method, the conservation analysis of non-coding genes was conducted, and it was found that the number of optional non-coding genes (accessory non-coding genes) in the genus Oryza was the largest, followed by specific non-coding genes (specific non-coding genes), and finally core non-coding genes ( Figure 6 ).
[0077] in, Figure 5 The blue color represents the core gene, the orange color represents the specific gene, and the yellow color represents the optional gene. Figure 6The blue color represents core non-coding genes, the orange color represents specific non-coding genes, and the yellow color represents accessory non-coding genes.
[0078] Step 4: Integrate information to construct the pan-transcriptome within the genus
[0079] Based on the sequence conservation of coding and non-coding genes determined in step 3, the transcriptome consisting of core coding and non-coding genes (including all transcripts corresponding to these genes) is defined as the core transcriptome; the transcriptome consisting of specific coding and non-coding genes (including all transcripts corresponding to these genes) is defined as the specific transcriptome; and the transcriptome consisting of optional coding and non-coding genes (including all transcripts corresponding to these genes) is defined as the accessory transcriptome. These three types of transcriptomes are combined to form the pan-transcriptome at the genus level. This pan-transcriptome can be used for downstream gene expression quantification and other analyses.
[0080] For example, by analyzing different rice species, we finally obtained the core transcriptome, specific transcriptome, and accessory transcriptome of rice species. The core transcriptome includes 8608 homologous groups, 111923 genes, and 200276 transcripts; the species-specific transcriptome includes 33650 genes and 60143 transcripts; the accessory transcriptome includes 73803 homologous groups, 151253 genes, and 270638 transcripts (the number of genes and transcripts is similar to that of the Figure 5 、 Figure 6 Corresponding). Functional enrichment (GO) analysis of the core Oryza genes (Appendix Table 2) revealed pathways primarily focused on calmodulin binding (GO:0005516), intracellular signal transduction (GO:0035556), and transmembrane transport activity (GO:0022857), highlighting the importance of core Oryza genes in maintaining essential plant life activities and indirectly demonstrating the reliability of this analysis process. The genes and corresponding transcripts in the Oryza-specific transcriptome are of great significance for exploring Oryza genetic resources, for example, providing genes related to stress resistance in Oryza and providing a crucial basis for exploring regulatory networks in response to adverse stresses.
[0081] Table 2
[0082]
[0083] The parts not described in detail in the present invention are prior art, so the present invention does not describe them in detail.
[0084] It is to be understood that the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the elements may be multiple, and the term "one" should not be understood as a limitation on the quantity.
[0085] Although this document uses a lot of professional terms, it does not exclude the possibility of using other terms. These terms are used only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitations is contrary to the spirit of the present invention.
[0086] The present invention is not limited to the above-mentioned optimal implementation mode. Anyone can derive other forms of products under the inspiration of the present invention. However, no matter what changes are made in the shape or structure, any technical solution that is the same or similar to the present invention falls within the scope of protection of the present invention.
Claims
1. A method for constructing a pan-transcriptome in a plant genus, characterized in that: The following steps are involved: (1) Data preparation and quality control: Obtain second-generation and third-generation transcriptome sequencing data of different species within the plant genus, including samples from different tissues and developmental stages; Use quality control software to perform quality control on the sequencing data and remove low-quality bases and adapter sequences; (2) Transcript splicing and screening: The second and third generation transcriptome sequencing data after quality control are aligned to the reference genome of the corresponding species, and the second and third generation splicing results are merged using transcript splicing software, and transcripts above the preset coverage threshold are retained; valid transcripts are screened by comparing with reference gene annotations; (3) Transcript classification based on open reading frames: The open reading frames of the effective transcripts of each species are identified, and the transcripts are divided into protein-coding transcripts, small peptide-coding transcripts and non-coding RNA according to the length of the open reading frames; and based on the transcript classification results, the gene loci are determined to be protein-coding genes, small peptide-coding genes or non-coding genes; (4) Identification of gene conservation among species: For genes encoding proteins or small peptides, core genes, optional genes, and specific genes are identified based on amino acid sequence homology analysis; For non-coding genes, core non-coding genes, optional non-coding genes, and specific non-coding genes were identified based on nucleotide sequence homology analysis; (5) Pan-transcriptome integration construction: core genes, optional genes, specific genes and their corresponding transcripts are combined with core non-coding genes, optional non-coding genes, specific non-coding genes and their corresponding transcripts to construct a genus-level pan-transcriptome.
2. The method for constructing a pan-transcriptome in a plant genus according to claim 1, characterized in that: The second-generation transcriptome sequencing data described in step (1) are paired-end sequencing data generated by Illumina or BGI platforms, and the third-generation transcriptome sequencing data are selected from long-read sequence data generated by PacBio Sequel II or Nanopore platforms.
3. The method for constructing a pan-transcriptome in a plant genus according to claim 1, characterized in that: In step (2), the second-generation transcriptome sequencing data were aligned using HISAT2 software, the third-generation transcriptome sequencing data were aligned using Minimap2 software, and the transcripts were spliced using StringTie software. The preset coverage threshold was ≥5 sequencing reads.
4. The method for constructing a pan-transcriptome in a plant genus according to claim 1, characterized in that: The reference gene annotation comparison in step (2) was performed using gffcompare software, and transcripts with comparison results of "=, c, x, j, o, i, u, p" were retained; Where "=" indicates that the predicted transcript is completely identical to the reference transcript; "c" indicates that the predicted transcript is contained within the reference transcript; "x" indicates that the predicted exon overlaps with the reference transcript on the reverse strand; "j" indicates that at least one splice junction of the predicted transcript is shared with the reference transcript; "o" indicates that some exons of the predicted transcript overlap with the reference transcript; "i" indicates that the predicted transcript is completely aligned to the introns of the reference transcript; "u" indicates that the predicted transcript is within the gene interval; and "p" indicates that the predicted transcript is within 2 kb of the reference transcript.
5. The method for constructing a pan-transcriptome in a plant genus according to claim 1, characterized in that: The criteria for transcript classification in step (3) are: Transcripts with an open reading frame length of ≥100 amino acids were identified as protein-encoding transcripts, transcripts with an open reading frame length of 10 amino acids ≤ and <100 amino acids were identified as small peptide-encoding transcripts, and transcripts with an open reading frame length of <10 amino acids or no open reading frame were identified as non-coding RNAs.
6. The method for constructing a pan-transcriptome in a plant genus according to claim 1, characterized in that: The gene locus classification rules in step (3) are: Gene loci with at least one protein-coding transcript were identified as protein-coding genes; Gene loci that do not encode proteins but have at least one transcript encoding a small peptide are identified as genes encoding small peptides; Gene loci where all transcripts are non-coding RNAs are considered non-coding genes.
7. The method for constructing a pan-transcriptome in a plant genus according to claim 1, characterized in that: In step (4), genes present in all analyzed species were considered core genes, genes present in only a single species were considered specific genes, and genes present in ≥2 species were considered optional genes.
8. The method for constructing a pan-transcriptome in a plant genus according to any one of claims 1 to 7, characterized in that: In step (4), the E value < 1e-5 was used as the screening threshold and the best alignment result was taken. Genes present in all analyzed species were considered core non-coding genes, genes present only in a single species were considered specific non-coding genes, and genes present in ≥2 species were considered optional non-coding genes.
9. The method for constructing a pan-transcriptome in a plant genus according to any one of claims 1 to 7, characterized in that: In step (5), the pan-transcriptome includes a core transcriptome consisting of core genes + core non-coding genes and their transcripts, an optional transcriptome consisting of optional genes + optional non-coding genes and their transcripts, and a specific transcriptome consisting of specific genes + specific non-coding genes and their transcripts.
10. Application of the construction method according to any one of claims 1 to 9 in improving plant stress resistance traits, characterized in that By analyzing the gene expression characteristics in specific transcriptomes, regulatory elements related to adverse stress response are screened.
Citation Information
Patent Citations
Gene and method for changing corn flowering period
CN112646014A
Compositions and methods for characterizing a complex biological sample
WO2023002325A1