Construction method and application of plant intra-genus generic transcriptome
By integrating the second and third generation transcriptome data and bioinformatics methods, the intra-general transcriptome of the plant genera is solved, and the lack of intra-general transcriptome construction in the existing technology is achieved, and the systematic analysis of small peptides and non-coding RNA is achieved, providing a cross-species conservative analysis tool.
Patent Information
- Application Number
- CN202510822132.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art lacks a method for building a pan-transcriptionome of different species within the plant genus, fails to systematically analyze the transcriptional characteristics and cross-species conservatism of small peptides and non-coding RNA, and the analysis method is missing.
By integrating the second- and third-generation transcriptome sequencing data of different species within the plant genus, combining open reading frame length classification transcripts, bioinformatics methods were used to identify gene conservatism among species, and a genus-level pan-transcriptionome containing core transcriptomes, optional transcriptomes and specific transcriptomes were constructed.
The construction of pan-transcriptionomes across species within the plant genus has been achieved, and the conservatism and differences of protein-encoded genes, small peptides and non-coding RNAs have been systematically analyzed, providing new tools for plant functional genome research.
Smart Images

Figure CN120356516A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to plant transcriptomics, bioinformatics data processing, and methods for constructing a pan-transcriptome and their applications. Background Art
[0002] Methods for constructing a plant pan-transcriptome mainly focus on different populations of a single species. For example, in rice, software such as minimap2, TAMA, and SQANTI3 are used to integrate single-species transcriptome data and annotate new transcripts (Zhong et al., 2024). However, there is a lack of a pan-transcriptome construction strategy for different species within a plant genus. In addition, current plant transcriptome analysis focuses on protein-coding genes, while small peptides (with inconsistent length definitions, such as 5 - 100 amino acids) and non-coding RNAs (functional RNAs that do not encode proteins) widely present in plants have not been systematically analyzed due to unknown functions and lack of analysis methods (Choi et al., 2019; Khavinson et al., 2021; Ning et al., 2018). With the development of high-throughput sequencing technology, a large amount of transcriptome data urgently requires integrated analysis methods to reveal the conservation and differences of protein-coding genes, small peptides, and non-coding RNAs among species within a genus.
[0003] Therefore, there is an urgent need for a method for plant transcriptomics, bioinformatics data processing, and pan-transcriptome construction to solve the problems in the prior art, which lack a pan-transcriptome construction method for different species within a plant genus and have not systematically analyzed the transcriptional characteristics and cross-species conservation of small peptides and non-coding RNAs. Summary of the Invention
[0004] The object of the present invention is to address the above-mentioned problems in the prior art and provide a method for constructing a pan-transcriptome within a plant genus. By integrating second-generation and third-generation transcriptome sequencing data of different species within a plant genus, transcripts (protein-coding, small peptide transcripts, non-coding RNAs) are classified according to the open reading frame (ORF) length, and bioinformatics methods are used to identify gene conservation among species, thereby constructing a genus-level pan-transcriptome including a core transcriptome, an optional transcriptome, and a specific transcriptome.
[0005] To achieve the above application object, the present invention adopts the following technical solutions: A method for constructing a pan-transcriptome within a plant genus includes the following steps: (1) Data preparation and quality control: Obtain second-generation and third-generation transcriptome sequencing data of different species within a plant genus, including samples from different tissues and developmental stages; use quality detection software to perform quality control on the sequencing data to remove low-quality bases and adapter sequences; (2)Transcript splicing and screening: Align the quality-controlled second- and third-generation transcriptome sequencing data to the reference genome of the corresponding species respectively, use transcript splicing software to merge the second- and third-generation splicing results, and retain transcripts above the preset coverage threshold; screen out valid transcripts by comparing with the reference gene annotation; (3)Transcript classification based on open reading frames: Identify open reading frames for the valid transcripts of each species, and classify the transcripts into protein-coding transcripts, small peptide-coding transcripts, and non-coding RNAs according to the open reading frame length; and based on the transcript classification results, determine the gene loci as protein-coding genes, small peptide-coding genes, or non-coding genes; (4)Identification of gene conservation among species: For genes encoding proteins or small peptides, identify core genes, alternative genes, and specific genes based on amino acid sequence homology analysis; for non-coding genes, identify core non-coding genes, alternative non-coding genes, and specific non-coding genes based on nucleotide sequence homology analysis; (5)Construction of pan-transcriptome integration: Merge the core genes, alternative genes, specific genes and their corresponding transcripts, with the core non-coding genes, alternative non-coding genes, specific non-coding genes and their corresponding transcripts to construct the pan-transcriptome at the genus level.
[0006] Furthermore, in step (1), the second-generation transcriptome sequencing data is paired-end sequencing data generated by the Illumina or BGI platform, and the third-generation transcriptome sequencing data is selected from the long-read sequence data generated by the PacBio Sequel II or Nanopore platform.
[0007] Furthermore, in step (2), HISAT2 software is used for aligning the second-generation transcriptome sequencing data, Minimap2 software is used for aligning the third-generation transcriptome sequencing data, StringTie software is used for transcript splicing, and the preset coverage threshold is ≥5 sequencing reads support.
[0008] Furthermore, in step (2), the reference gene annotation comparison is implemented by gffcompare software, and transcripts with comparison results of "=, c, x, j, o, i, u, p" are retained; where "=" means that the predicted transcript is exactly the same as the reference transcript; "c" means that the predicted transcript is contained in the reference transcript; "x" means that the predicted exon overlaps with the reference transcript on the reverse strand; "j" means that at least one splicing junction of the predicted transcript is shared with the reference transcript; "o" means that some exons of the predicted transcript overlap with the reference transcript; "i" means that the predicted transcript is completely aligned to the intron of the reference transcript; "u" means that the predicted transcript is in the gene interval; "p" means that the predicted transcript is within 2 kb of the reference transcript.
[0009] Further, the criteria for classifying transcripts in step (3) are as follows: Transcripts with an open reading frame length ≥ 100 amino acids are determined as protein-coding transcripts, transcripts with 10 amino acids ≤ open reading frame length < 100 amino acids are determined as small peptide-coding transcripts, and transcripts with an open reading frame length < 10 amino acids or no open reading frame are determined as non-coding RNAs.
[0010] Further, the gene locus classification rules in step (3) are as follows: Gene loci with at least one protein-coding transcript are determined as protein-coding genes; Gene loci with no protein-coding transcripts but at least one small peptide-coding transcript are determined as small peptide-coding genes; Gene loci where all transcripts are non-coding RNAs are determined as non-coding genes.
[0011] Further, in step (4), genes that exist in all analyzed species are defined as core genes, genes that exist only in a single species are defined as specific genes, and genes that exist in ≥ 2 species are defined as optional genes.
[0012] Further, in step (4), with an E-value < 1e-5 as the screening threshold and taking the best alignment result, genes that exist in all analyzed species are core non-coding genes, genes that exist only in a single species are specific non-coding genes, and genes that exist in ≥ 2 species are optional non-coding genes.
[0013] Further, in step (5), the pan-transcriptome includes a core transcriptome composed of core genes + core non-coding genes and their transcripts, an optional transcriptome composed of optional genes + optional non-coding genes and their transcripts, and a specific transcriptome composed of specific genes + specific non-coding genes and their transcripts.
[0014] For the application of the above construction method in improving plant stress resistance traits, by analyzing the gene expression characteristics in the specific transcriptome, regulatory elements related to stress response are screened.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Break through the limitation of a single species and achieve the construction of a pan-transcriptome at the genus level The prior art only constructs pan-transcriptomes for different populations of a single species, while the present invention for the first time provides a method for constructing a pan-transcriptome of different species within a plant genus. By integrating multi-species second- and third-generation transcriptome data, the conservation (such as core genes) and specificity (such as species-specific genes) of protein-coding genes, small peptides, and non-coding RNAs within the genus are systematically analyzed, filling the gap in transcriptome research at the genus level.
[0016] 2. The integration of second- and third-generation data improves the comprehensiveness of transcript identification Combining the advantages of second-generation sequencing (high precision, low cost) and third-generation sequencing (long read lengths, ability to identify new transcripts) significantly improves the completeness of transcript identification.
[0017] 3. Classification method based on ORF clarifies the coding characteristics of transcripts Aiming at the problem of ambiguous definitions of protein-coding genes and non-coding RNAs in the prior art, the present invention for the first time realizes the precise classification of the coding ability of transcripts through the ORF length standard (≥100 amino acids for mRNA, 10 - 99 amino acids for small peptide transcripts, <10 amino acids for ncRNA), and divides gene types accordingly (protein-coding genes, small peptide genes, non-coding genes), providing clear targets for functional research.
[0018] 4. Systematic analysis of the cross-species characteristics of non-coding RNAs and small peptides The prior art mainly focuses on coding genes and ignores the regulatory functions of small peptides and non-coding RNAs. Through conservation analysis (coding genes and small peptides are based on amino acid sequences, non-coding genes are based on nucleotide sequences), the present invention for the first time identifies core / optional / specific non-coding genes and small peptide genes at the genus level. For example, the core non-coding genes of the rice genus are enriched in basic life activity pathways, and the specific transcriptome contains gene resources related to stress resistance, providing new directions for functional research.
[0019] 5. Construction of a multi-dimensional pan-transcriptome to support downstream in-depth analysis The integrated pan-transcriptome at the genus level includes a core transcriptome (conserved genes shared by species), an optional transcriptome (genes present in some species), and a specific transcriptome (species-specific genes), which can be further used for gene expression quantification and functional enrichment analysis (such as the core genes of the rice genus being enriched in the osmotic stress response pathway), providing a systematic data basis for research on plant evolution, stress resistance mechanisms, etc. Description of the Drawings
[0020] Figure 1 It is a diagram showing the difference in the number of transcripts identified from second- and third-generation data of different species in the embodiments of the present invention; Figure 2 It is a diagram comparing the number of transcripts identified according to the present invention and the number of reference transcripts in different species in the embodiments of the present invention; Figure 3 It is a diagram showing the transcript classification of different species in the embodiments of the present invention; Figure 4 It is a diagram showing the gene classification of different species in the embodiments of the present invention; Figure 5 It is a diagram showing the conservation classification results of genes encoding proteins and small peptides in rice genus species in the embodiments of the present invention; Figure 6It is a diagram showing the classification results of the conservation of non-coding genes in Oryza species of the embodiments of the present invention; Figure 7 It is a flowchart of Steps 1 and 2 of the embodiments of the present invention; Figure 8 It is a flowchart of Steps 3 and 4 of the embodiments of the present invention. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0022] Those skilled in the art should understand that in the disclosure of the present invention, the orientation or positional relationships indicated by the terms "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the accompanying drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention.
[0023] As Figure 7 and Figure 8 shown, the method for constructing a pan-transcriptome within the genus of this plant includes the following steps: Step 1: Construct a transcript splicing process based on the transcriptome sequencing data of plant second-generation and third-generation sequencing data. The specific process steps are as follows: (1a) Sequencing data preparation: Download the transcriptome sequencing (RNA-seq) data of different tissues at different times of plant species (such as different Oryza species) available in public databases such as NCBI (https: / / www.ncbi.nlm.nih.gov / ) or NGDC (https: / / ngdc.cncb.ac.cn / ), including both second-generation and third-generation sequencing data; or obtain data by yourself through high-throughput sequencing platforms such as Illumina, PacBio, Nanopore, or BGI. The use of second-generation and third-generation transcriptome sequencing data of different tissues at different times is to fully identify candidate transcript loci. For example: Based on 4 Oryza species (cultivated rice Oryza sativa , wild rice 1 Oryza barthii , wild rice 2 Oryza nivara and wild rice 3 Oryza rufipogonto construct the pan-transcriptome of Oryza genus, collecting public transcriptome sequencing data and self-tested data from NCBI and NGDC databases (Appendix 1), and the sequencing platforms include second-generation Illumina PE150, third-generation PacBio Sequel II, third-generation Nanopore, etc. Among them, the data volume of cultivated rice ( Oryza sativa is about 2T; the total data volume of 3 wild rice species is about 1.5T. Identifying transcripts according to the method of "Step 1", it is found that a large number of new transcripts, especially long transcripts ( Figure 1 ), can be identified from the third-generation transcriptome data; compared with the reference annotation information of the species, more new transcripts ( Figure 2 , and the number of identified transcripts increases by more than half) can be identified by comprehensively using a large amount of transcriptome sequencing data (including second and third generations).
[0024] Table 1
[0025] Figure 1 , the abscissa is the length range of transcripts, with the unit of bp; "S2" represents the identification result of second-generation sequencing data (blue), "S3" represents the identification result of third-generation sequencing data (orange), and "S2&S3" represents the identification result of integrating second and third-generation sequencing data (yellow). Figure 2 , blue represents the number of transcripts identified using second and third-generation sequencing data, orange represents the number of transcripts in the reference annotation file, and yellow represents the number of transcripts after merging the two.
[0026] (1b) Quality control of sequencing data: Using the software FastQC (https: / / github.com / s-andrews / FastQC) to perform quality checks on the sequencing data, and using software such as Trim Galore (https: / / github.com / FelixKrueger / TrimGalore) or fastp (https: / / github.com / OpenGene / fastp) or fastplong (https: / / github.com / OpenGene / fastplong) to remove low-quality bases and adapter sequences. Among them, fastp is suitable for second-generation short-read sequencing data, and fastplong is suitable for third-generation long-read sequencing data.
[0027] (1c) Alignment of sequencing data and transcript splicing: Align the second- and third-generation transcriptome sequencing (RNA-seq) data (referred to as second-generation data and third-generation data respectively) to the species reference genome, and obtain the species reference genome and annotation information from databases such as NCBI or Ensembl Plants. For the alignment of second-generation data, use HISAT2 (https: / / daehwankimlab.github.io / hisat2 / ), and for the alignment of third-generation data, use Minimap2 (https: / / github.com / lh3 / minimap2). Use the StringTie (https: / / ccb.jhu.edu / software / stringtie / ) software to splice transcripts for all alignment results (retain transcripts supported by at least 5 sequencing reads; other similar alignment and splicing software can also be used). When splicing, use the parameter "-L" for long-read processing for third-generation data. Use the merge function of the StringTie software to merge the transcript splicing results of the second- and third-generation data of the same species.
[0028] (1d) Transcript screening: Use the software gffcompare (https: / / github.com / gpertea / gffcompare) to compare the reference gene annotation information of the known corresponding species and the transcript information obtained in (1c), and retain transcripts with comparison results of "=, c, x, j, o, i, u, p". Finally, use the merge function of the StringTie software to merge and take the union of the screened predicted transcripts and the reference transcripts to obtain all transcript information of each species.
[0029] Among them, "=" means that the predicted transcript is exactly the same as the reference transcript; "c" means that the predicted transcript is included in the reference transcript; "x" means that the predicted exon overlaps with the reference transcript on the reverse strand; "j" means that at least one splice junction of the predicted transcript is shared with the reference transcript; "o" means that some exons of the predicted transcript overlap with the reference transcript; "i" means that the predicted transcript is completely aligned to the intron of the reference transcript; "u" means that the predicted transcript is in the gene interval; "p" means that the predicted transcript is within 2 kb of the reference transcript.
[0030] Step 2: The specific process steps for classifying transcripts based on open reading frames are as follows: (2a) Open reading frame (ORF) identification: Use open reading frame identification software such as ORFfinder or TransDecoder to identify the open reading frames of all transcripts of each species obtained in step 1. Count the lengths of all open reading frames; for any one transcript, select its longest open reading frame as the possible encoded amino acid sequence of this transcript.
[0031] (2b) Classification of transcripts: Based on the results of (2a), each transcript corresponds to 1 or 0 sequences with open reading frames. All transcripts in each species are classified according to the length of their open reading frames. Specifically, transcripts with an ORF length ≥ 100 amino acids are determined to be protein-coding transcripts, i.e., messenger RNA (mRNA); transcripts with an ORF length ≥ 10 amino acids and an ORF length < 100 amino acids are determined to be small peptide-coding transcripts; transcripts with an ORF length < 10 amino acids or without a predicted ORF are determined to be non-protein-coding transcripts, i.e., non-coding RNA (ncRNA).
[0032] (2b) Classification of genes: A gene locus may produce multiple transcripts. As long as one transcript is determined to be protein-coding, i.e., mRNA, then this gene is determined to be a protein-coding gene; on the premise that there is no protein-coding transcript (mRNA), as long as one transcript is determined to be small peptide-coding, then this gene is determined to be a small peptide-coding gene; if all transcripts do not code for proteins or small peptides, i.e., ncRNA, then this gene is determined to be a non-coding gene.
[0033] For example: Classify the identified transcripts according to the method of "step 2 (2a)", and it is found that the most transcripts in rice species are protein-coding transcripts, followed by small peptide-coding transcripts, and finally non-coding RNA ( Figure 3 ). Define the transcripts produced by the same locus as a gene locus, and classify the genes according to the method of "step 2 (2b)", and it is found that the most genes in rice species are protein-coding genes, followed by small peptide-coding genes, and finally non-coding genes ( Figure 4 ).
[0034] Among them, Figure 3 blue represents non-coding transcripts (ncRNA transcripts), orange represents protein-coding transcripts (mRNA transcripts), and yellow represents small peptide-coding transcripts (peptide transcripts). Figure 4 blue represents non-coding genes (ncRNA genes), orange represents protein-coding genes (mRNA genes), and yellow represents small peptide-coding genes (peptide genes).
[0035] Step 3: The specific process steps for the conservation identification of protein-coding genes and non-coding genes between species are as follows: (3a) Conservation identification of genes encoding proteins or small peptides: Based on the gene classification determined in Step 2, for genes encoding proteins or small peptides, select the amino acid sequence corresponding to the longest ORF, and use OrthoFinder (https: / / github.com / davidemms / OrthoFinder) to identify homologous genes between different species. If a gene exists in all analyzed species, it is determined as a core gene; a gene that exists only in one species is determined as a specific gene; a gene that exists in two or more species is determined as an accessory gene.
[0036] (3b) Conservation identification of non-coding genes: Based on the gene classification determined in Step 2, for non-coding genes, select their longest transcript (i.e., nucleotide sequence), and use the "-evalue" parameter (i.e., E-value) in BLASTN to define the expected threshold to judge the conservation of their base sequences. The E-value indicates the probability that, under random circumstances, the similarity of other sequences to the target sequence is greater than the displayed sequence. Setting the E-value to less than 1e-5 is a conventional strict standard for homologous sequence alignment, which can greatly avoid non-specific matching; the smaller the E-value, the stricter the screening criteria.
[0037] Specifically, pairwise alignments are performed on the non-coding transcript sequences of all analyzed species, and the E-value of 1e-5 is used as the screening condition to determine their homologous transcripts; and for each transcript, only the unique best alignment result is retained. If a non-coding gene exists in all analyzed species, it is determined as a core non-coding gene; a non-coding gene that exists only in one species is determined as a specific non-coding gene; a non-coding gene that exists in two or more species is determined as an accessory non-coding gene. Core non-coding genes may have specific biological functions and can be used as candidates for subsequent gene function verification.
[0038] For example: According to the method of "Step 3 (3a)", the conservation analysis of genes encoding proteins and small peptides was carried out, and it was found that the number of accessory genes in the genus Oryza was the largest, followed by core genes, and finally specific genes (Figure 5 ). Similarly, according to the method of "Step 3 (3b)", the conservation analysis of non-coding genes was carried out, and it was found that the number of accessory non-coding genes in Oryza was the largest, followed by specific non-coding genes, and finally core non-coding genes ( Figure 6 ).
[0039] Among them, Figure 5 blue represents core genes, orange represents specific genes, and yellow represents accessory genes. Figure 6 blue represents core non-coding genes, orange represents specific non-coding genes, and yellow represents accessory non-coding genes.
[0040] Step 4: Integrate information to construct a pan-transcriptome within the genus Based on the judgment of the sequence conservation of coding genes and non-coding genes in Step 3, the transcriptome composed of core coding genes and core non-coding genes (including all transcripts corresponding to these genes) was determined as the core transcriptome; the transcriptome composed of specific coding genes and specific non-coding genes (including all transcripts corresponding to these genes) was determined as the specific transcriptome; the transcriptome composed of accessory coding genes and accessory non-coding genes (including all transcripts corresponding to these genes) was determined as the accessory transcriptome. The transcriptomes of these three types were combined to form the pan-transcriptome at the genus level. Based on this pan-transcriptome, downstream gene expression quantification and other analyses can be carried out.
[0041] For example: Through the analysis of different Oryza species, the core transcriptome, specific transcriptome, and accessory transcriptome of Oryza species were finally obtained. Among them, the core transcriptome includes 8608 homologous groups, 111923 genes, and 200276 transcripts; the species-specific transcriptome includes 33650 genes and 60143 transcripts; the accessory transcriptome includes 73803 homologous groups, 151253 genes, and 270638 transcripts (the number of genes and transcripts isFigure 5 , Figure 6 (corresponding). Functional enrichment (GO) analysis of the core genes of Oryza genus found (see Table 2 in the appendix) that its pathways mainly concentrated in calmodulin binding (GO:0005516), intracellular signal transduction (GO:0035556), and transmembrane transport activity (GO:0022857), reflecting the importance of the core genes of Oryza genus in maintaining the basic activities of plant life and indirectly reflecting the reliability of this analysis process. The genes and corresponding transcripts in the specific transcriptome of Oryza genus are of great significance for the exploration of Oryza gene resources. For example, they provide gene resources related to stress resistance in Oryza genus and provide an important basis for exploring the regulatory network when dealing with adversity stress.
[0042] Table 2
[0043] The parts not detailed in the present invention are prior arts, so the present invention does not detail them.
[0044] It can be understood that the term "a" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of this element can be multiple. The term "a" cannot be understood as a limitation on the number.
[0045] Although many technical terms are used in this article, the possibility of using other terms is not excluded. These terms are used only to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitation is contrary to the spirit of the present invention.
[0046] The present invention is not limited to the above-mentioned optimal implementation manner. Anyone can obtain other various forms of products under the inspiration of the present invention. However, no matter what changes are made in its shape or structure, as long as it has a technical solution identical or similar to that of the present invention, it falls within the protection scope of the present invention.
Claims
1. A method for constructing a pan-transcriptome within a plant genus, characterized in that, It includes the following steps: (1) Data preparation and quality control: Obtain second-generation and third-generation transcriptome sequencing data of different species within a plant genus, including samples from different tissues and developmental stages; Use quality detection software to perform quality control on the sequencing data, removing low-quality bases and adapter sequences; (2) Transcript splicing and screening: Align the quality-controlled second-generation and third-generation transcriptome sequencing data to the reference genome of the corresponding species respectively, use transcript splicing software to merge the second-generation and third-generation splicing results, and retain transcripts above the preset coverage threshold; Through comparison with the reference gene annotation, screen out valid transcripts; (3) Transcript classification based on open reading frames: Identify open reading frames for the valid transcripts of each species, and classify the transcripts into protein-coding transcripts, small peptide-coding transcripts, and non-coding RNAs according to the length of the open reading frames; And based on the transcript classification results, determine the gene loci as protein-coding genes, small peptide-coding genes, or non-coding genes; (4) Identification of gene conservation among species: For genes encoding proteins or small peptides, identify core genes, alternative genes, and specific genes based on amino acid sequence homology analysis; For non-coding genes, identify core non-coding genes, alternative non-coding genes, and specific non-coding genes based on nucleotide sequence homology analysis; (5) Construction of pan-transcriptome integration: Merge core genes, alternative genes, specific genes and their corresponding transcripts, with core non-coding genes, alternative non-coding genes, specific non-coding genes and their corresponding transcripts to construct a pan-transcriptome at the genus level.
2. The construction method of a pan-transcriptome within a plant genus according to claim 1, wherein, In step (1), the second-generation transcriptome sequencing data is paired-end sequencing data generated by the Illumina or BGI platform, and the third-generation transcriptome sequencing data is selected from long-read sequencing data generated by the PacBio Sequel II or Nanopore platform.
3. The construction method of a pan-transcriptome within a plant genus according to claim 1, characterized in that In step (2), the HISAT2 software is used for aligning the second-generation transcriptome sequencing data, the Minimap2 software is used for aligning the third-generation transcriptome sequencing data, and the StringTie software is used for transcript splicing. The preset coverage threshold is ≥5 sequencing reads supported.
4. The construction method of a pan-transcriptome within a plant genus according to claim 1, wherein In step (2), the comparison with the reference gene annotation is achieved through the gffcompare software, and transcripts with comparison results of "=, c, x, j, o, i, u, p" are retained; where "=" means that the predicted transcript is exactly the same as the reference transcript; "c" means that the predicted transcript is included in the reference transcript; "x" means that the predicted exon overlaps with the reference transcript on the reverse strand; "j" means that at least one splicing junction of the predicted transcript is shared with the reference transcript; "o" means that some exons of the predicted transcript overlap with the reference transcript; "i" means that the predicted transcript is completely aligned to the intron of the reference transcript; "u" means that the predicted transcript is in the gene interval; "p" means that the predicted transcript is within 2 kb of the reference transcript.
5. The construction method of a pan-transcriptome within a plant genus according to claim 1, characterized in that The criteria for transcript classification in step (3) are: Transcripts with an open reading frame length of ≥100 amino acids are determined as transcripts encoding proteins, transcripts with an open reading frame length of 10 amino acids ≤ open reading frame length < 100 amino acids are determined as transcripts encoding small peptides, and transcripts with an open reading frame length < 10 amino acids or no open reading frame are determined as non-coding RNAs.
6. The construction method of a pan-transcriptome within a plant genus according to claim 1, characterized in that, The gene locus classification rules described in step (3) are as follows: Gene loci with at least one transcript encoding a protein are determined as protein-encoding genes; Gene loci with no protein-encoding transcripts but at least one small peptide-encoding transcript are determined as small peptide-encoding genes; Gene loci where all transcripts are non-coding RNAs are determined as non-coding genes.
7. A method for constructing a pan-transcriptome within a plant genus according to claim 1, characterized in that In step (4), genes that are present in all analyzed species are core genes, genes that are present only in a single species are specific genes, and genes that are present in ≥2 species are optional genes.
8. A method for constructing a pan-transcriptome within a plant genus according to any one of claims 1-7, characterized in that, In step (4), with an E-value < 1e-5 as the screening threshold and taking the best alignment result, genes that are present in all analyzed species are core non-coding genes, genes that are present only in a single species are specific non-coding genes, and genes that are present in ≥2 species are optional non-coding genes.
9. A method for constructing a pan-transcriptome within a plant genus according to any one of claims 1-7, characterized in that, In step (5), the pan-transcriptome includes a core transcriptome composed of core genes + core non-coding genes and their transcripts, an optional transcriptome composed of optional genes + optional non-coding genes and their transcripts, and a specific transcriptome composed of specific genes + specific non-coding genes and their transcripts.
10. Use of the construction method according to any one of claims 1-9 in improving plant stress resistance traits, characterized in that By analyzing the gene expression characteristics in the specific transcriptome, regulatory elements related to stress response are screened.
Citation Information
Patent Citations
Gene and method for changing corn flowering period
CN112646014A
Eukaryote generic transcriptome annotation method
CN114373506A
Distributed transcriptional regulation network large model construction method based on transfer learning
CN119339795A
Compositions and methods for characterizing a complex biological sample
WO2023002325A1
Cited By
Non-coding region identification method and device for convergence selection
CN121354660A