A method for eukaryotic pan-transcriptome annotation
By integrating transcriptome annotation methods that depend on and do not depend on reference genomes, and utilizing multiple tools for data alignment and assembly, the problems of incompleteness and accuracy in pantranscriptome annotation have been solved, achieving more comprehensive gene and transcript annotation, which is applicable to pantranscriptome research in eukaryotes.
Patent Information
- Application Number
- CN202111671919.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-12-31
AI Technical Summary
Existing technologies are insufficient for pantranscriptome annotation, especially for alternative splicing types, mainly due to the diversity of different species and the structural variations in the genome of the same species, resulting in incomplete annotation and low accuracy.
By integrating transcriptome annotation methods that depend on and do not depend on reference genomes, and using tools such as cufflinks, SOAPdenovo, Trinity, and PASA, transcriptome data from eukaryotic species samples from multiple sources were compared and assembled. The cuffmerge and cuffcompare methods were then used for information integration and filtering to obtain more complete and accurate pantranscriptome annotations.
It significantly improves the completeness and accuracy of pantranscriptome annotation, especially at the gene and transcript levels, outperforming existing technologies. It is applicable to pantranscriptome annotation in eukaryotes and can be used in fields such as disease control and rapid detection.
Smart Images

Figure CN114373506B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of biotechnology, in particular to a eukaryotic pan-transcriptome annotation method. BACKGROUND
[0002] Pan-transcriptome is a new research field after pan-genome. It is the total collection of RNA sequences transcribed by all species, which is used to specify the pan-analysis of gene expression and transcription. Due to the huge and complex number of sequences in pan-transcriptome, the research method of annotation is still rare. At present, pan-transcriptome is relatively weak in gene structure annotation, especially in variable splicing type, the main reasons are one is the diversity of different species, and two is the variation of genome structure of the same species.
[0003] Chinese patent CN101894211B discloses a gene annotation method and system, which adopts a gene prediction method of sequence characteristics and statistical model to obtain the position of potential genes on the target genome, adopts a gene annotation method of sequence similarity to align known gene sequences and species-conserved sequences to the target genome, and uses a weighted voting method to integrate and screen the results to obtain gene annotation results containing variable splicing forms; Chinese patent CN106446601B discloses a method for large-scale annotation of lncRNA function, and provides a genome annotation method using second-generation and third-generation transcriptome sequencing data. The above patents do not involve pan-transcriptome annotation. SUMMARY
[0004] Therefore, the present application aims to provide a eukaryotic pan-transcriptome annotation method, which can obtain more complete pan-transcriptome annotation structure by integrating different methods, and can greatly improve the annotation accuracy by mutual verification of different strategies.
[0005] The technical scheme of the present application is as follows:
[0006] A eukaryotic pan-transcriptome annotation method, comprising the following steps:
[0007] (1) Transcriptome annotation dependent on reference genome: aligning transcriptome data of 3 or more eukaryotic species samples of different sources to the reference genome to obtain corresponding gene structure annotation information, integrating to obtain transcriptome annotation results dependent on the reference genome, i.e. structure annotation information gg1-gg3; wherein the reference genome of step (1) is the reference genome of the eukaryotic species sample;
[0008] (2) Reference genome-independent transcriptome annotation: the sample transcriptome data is spliced and assembled to obtain corresponding mature transcript sequence information, the mature transcript sequence information is aligned to the reference genome to obtain corresponding gene structure information, and reference genome-independent transcriptome annotation results, i.e. transcript structure annotation information gi-gi3, are obtained;
[0009] (3) Integration and filtering of annotation information: the structure annotation information 1 and the transcript structure annotation information 2 are integrated, and transcript structure alignment is performed, the transcripts of the reference genome that cannot be aligned are filtered, and new transcript annotation information G1-G3 is obtained;
[0010] (4) Integration of the new transcript annotation information G1-G3 to obtain the pan-transcriptome annotation information of the target eukaryotic species.
[0011] Further, in step (1), the sample transcriptome data is aligned to the reference genome, and the alignment is performed by using a method including cufflinks.
[0012] Further, in step (1), the integrated gene structure annotation information is obtained by using a method including cuffmerge.
[0013] Further, in step (2), the splicing and assembly are performed by using a method including SOAPdenovo or Trinity.
[0014] Further, in step (2), the mature transcript sequence information is aligned to the reference genome by using a method including PASA.
[0015] Further, in step (3), the integrated structure annotation information gg1-gg3 and the transcript structure annotation information gi-gi3 are obtained by using a method including cuffcompare.
[0016] Compared with the prior art, the present application has the following beneficial effects:
[0017] The transcriptome annotation method provided by the present application is superior to the reference genome annotation and the transcriptome annotation of the prior art in terms of the number of genes, multi-exon genes and transcripts, is suitable for pan-transcription annotation of eukaryotes, and can obtain more complete pan-transcriptome annotation structure by using only a small amount of eukaryotic species samples of different sources, and can effectively improve the accuracy of annotation. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 The figure is a schematic diagram of the steps of the pan-transcription annotation method of the eukaryote of the present application;
[0019] Figure 2Test results of the present application using the pathogenic microorganism Magnaporthe oryzae for pan-transcriptome. DETAILED DESCRIPTION
[0020] In order to better understand the technical content of the present application, the following specific examples are provided to further illustrate the present application.
[0021] The experimental methods used in the embodiments of the present application are all conventional methods unless otherwise specified.
[0022] The materials, reagents, etc. used in the embodiments of the present application can be obtained from commercial channels unless otherwise specified.
[0023] EMBODIMENT
[0024] Experiments were performed using the pathogenic microorganism Magnaporthe oryzae as a sample, and three different annotation methods were used to test Magnaporthe oryzae from three different sources (SRA database access numbers: SRX11477966, SRX11477965, and SRX11477964). The comparison was made in four aspects, namely the number of genes, the number of multi-exon genes, transcript data, and multi-exon transcript data. The comparison results included the original reference genome annotation, the annotation results obtained using the conventional transcript annotation method, and the annotation results obtained using the present method.
[0025] The pan-transcriptome annotation of the present application includes the following steps:
[0026] (1) Transcriptome annotation dependent on reference genome: The cufflinks method can be used to align the transcriptome data of Magnaporthe oryzae from three different sources to the reference genome to obtain the corresponding gene structure annotation information (see the "gg_1,2…n" part in Figure 1 The cuffmerge method is used to integrate the gene structure annotation information from different Magnaporthe oryzae to obtain the transcriptome annotation results dependent on the reference genome, i.e., the structure annotation information gg1-gg3 (see the "gg assembly" part in Figure 1
[0027] (2) Transcriptome annotation independent of reference genome: Methods such as SOAPdenovo or Trinity can be used to assemble the transcriptome data of Magnaporthe oryzae from different sources. After assembly, the mature transcript sequence information corresponding to each Magnaporthe oryzae is obtained (see the "sample gi_1,2…n" part in Figure 1 The PASA method can be used to align the mature transcript sequence information to the reference genome to obtain the corresponding gene structure information, and the transcriptome annotation results independent of the reference genome, i.e., the transcript structure annotation information gi-gi3 (see the "gi assembly" part in Figure 1
[0028] (3) Integration and filtering of annotation information: using methods such as cuffcompare, the structural annotation information gg1-gg3 and the transcript structural annotation information gi-gi3 are integrated, and transcript structural alignment is performed, and the transcripts of the reference genome that cannot be aligned are filtered to obtain new transcript annotation information G1-G3;
[0029] (4) Integration of new transcript annotation information G1-G3, and the pan-transcriptome annotation information of the pathogenic microorganism Magnaporthe oryzae can be obtained.
[0030] The reference genome is Magnaporthe oryzae 70-15 (assembly MG8), and the reference genome annotation is based on the literature (Ralph AD.; Nicholas J T.; et.al. The genome sequence of the rice blast fungus Magnaporthe grisea. [J]. Nature. 2005, 434 (7036));
[0031] The transcript annotation method is based on the literature (Minfeng Xue; Jun Yang; et.al. Comparative analysis of the genomes of two field isolates of the rice blast fungus Magnaporthe oryzae. [J]. PLoS Genet. 2012, 8 (8)).
[0032] From Figure 2 It can be known that the pan-transcriptome annotation of the present application can greatly improve the completeness of the transcript annotation of the pathogenic microorganism Magnaporthe oryzae, especially at the gene level, and can more completely annotate the gene set of the eukaryotic species, and can improve the accuracy of the transcript, and effectively reduce the false annotation results. The present application is superior to the reference genome annotation and the transcriptome annotation in the number of genes, multi-exon genes and transcripts, which shows that the pan-transcriptome annotation of the present application can obtain a more complete pan-transcriptome annotation structure, and effectively improve the accuracy of the annotation, and can be applied to the fields of pest control and rapid detection.
[0033] The above only describes the preferred embodiments of the present application, and is not used to limit the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for pan-transcriptome annotation in eukaryotes, characterized in that: The method comprises the following steps: (1) Reference genome-dependent transcriptome annotation: aligning transcriptome data of samples of eukaryotic species from more than three different sources to a reference genome to obtain corresponding gene structure annotation information, integrating the information to obtain reference genome-dependent transcriptome annotation results, i.e. structure annotation information gg1-gg3; The sample transcriptome data is aligned to the reference genome using the cufflinks method; The integrated gene structure annotation information is obtained using the cuffmerge method; (2) Reference genome-independent transcriptome annotation: splicing and assembling the sample transcriptome data to obtain corresponding mature transcript sequence information, aligning the mature transcript sequence information to the reference genome to obtain corresponding gene structure information, and obtaining reference genome-independent transcriptome annotation results, i.e. transcript structure annotation information gi-gi3; The splicing and assembling is performed using the SOAPdenovo or Trinity method; (3) Integration and filtering of annotation information: integrating the structure annotation information gg1-gg3 and the transcript structure annotation information gi-gi3, and performing transcript structure alignment and filtering of transcripts that cannot be aligned to the reference genome to obtain new transcript annotation information G1-G3; (4) Integrating the new transcript annotation information G1-G3 to obtain pan-transcriptome annotation information of the target eukaryotic species; In step (2), the mature transcript sequence information is aligned to the reference genome using the PASA method; In step (3), the integrated structure annotation information gg1-gg3 and the transcript structure annotation information gi-gi3 are obtained using the cuffcompare method.
Citation Information
Patent Citations
Gene annotation method and system
CN101894211B
A method for large-scale labeling of lncRNA function
CN106446601B
Gene annotation method and system
CN101894211A