Method for gene structure annotation
By using multiple sequence alignment and software integration, the transcriptome data of the target species was supplemented with the whole genome sequences of closely related species, which solved the problem of inaccurate gene structure annotation caused by low transcriptome data quality and achieved higher quality gene structure annotation.
Patent Information
- Application Number
- CN202311156052.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-08
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2043-09-08
AI Technical Summary
The low quality of transcriptome data affects the accuracy and completeness of gene structure annotation, and existing methods have failed to effectively address this issue.
By using multiple sequence alignment, the transcriptome data of the target species are supplemented with whole genome sequences of closely related species. Gene family clustering and annotation are performed using OrthoFinder, PASA, and EVM software, integrating gene structure annotation results from multiple sources.
It significantly improves the accuracy and completeness of gene structure annotation, enabling comprehensive identification of all genes in the genome and providing more reliable biological conclusions.
Smart Images

Figure CN119601079B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bio-information analysis, in particular to a gene structure annotation method. BACKGROUND
[0002] With the rapid development and gradual popularization of high-throughput sequencing technology, from prokaryotic microorganisms to complex multicellular organisms, more and more organism genomes have been deciphered, and now we have entered a new era of "functional genomics", and genome annotation is the basis for studying the functions of various genes of organisms. Nowadays, de novo prediction, homologous prediction and transcriptome data assisted prediction are widely used in gene structure annotation. Among them, transcriptome data is considered to be one of the important sources of gene structure annotation. However, due to the fact that some samples are precious, resulting in a small number of samples that can be collected, sequencing errors, insufficient sequencing depth, exogenous contamination and other reasons, the information of the transcriptome data is often missing or its quality is not high enough, which directly affects the quality of gene structure annotation, leading to the possibility of missing some low-expression genes when revealing biological problems, inaccurate gene positioning, and ultimately making the biological conclusion unreliable.
[0003] At present, the method for gene structure annotation mainly integrates de novo prediction, homologous prediction and transcriptome data through different weights. However, in the process of gene structure annotation, there is no method to solve the problem of insufficient quality of transcriptome data, therefore, the method for gene structure annotation has important significance and reference value.
[0004] In view of this, the present application is proposed. SUMMARY
[0005] In order to solve the above problems, the present application provides a new method for gene structure annotation, which uses the principle of multiple sequence alignment, compares the whole genome sequences of closely related species to identify core gene sequences, and screens other series of conditions to make up for the insufficient quality of the target species transcriptome data, fill in the evidence source of gene annotation, and finally obtain high-quality gene structure annotation results.
[0006] In order to achieve the above purpose, the present application provides a method for gene structure annotation, which assembles the transcriptome data of the target species, then downloads the nucleic acid and protein amino acid sequences of the closely related species from the public database for alignment and clustering, identifies the core sequences of the closely related species, and integrates the core gene series of the same closely related species into the assembled transcriptome data, thereby obtaining a high-quality reference data set as an evidence source for gene structure annotation, and finally obtaining a comprehensive and accurate gene structure annotation result. The method of the present application is carried out by the following steps:
[0007] Step 1. Transcriptome assembly: using Trinity software to assemble the transcriptome data of the target species;
[0008] Step 2. Data preparation: downloading the nucleic acid sequence of the longest transcript and the amino acid sequence file of the coding protein of the target species from the public database such as NCBI, CNGB, etc. of the close relative species of the target species;
[0009] Step 3. Gene family clustering: using OrthoFinder software to perform gene family clustering analysis on the gene set of homologous species, and identifying the orthologous gene pairs of the close relative species through the gene family clustering results;
[0010] Step 4. Determine whether there is a close relative species of the same genus and species of the target species in step 2;
[0011] Step 5. Defining the orthologous gene pairs shared by the genomes of the close relative species screened in step 3 as the core gene set, and selecting the longest transcript of the closest relative species as the candidate reference sequence;
[0012] Step 6. If there is a close relative species of the same genus and species, integrate the longest transcript sequence in the close relative species of the same genus and species into the transcriptome assembly results of step 1; if there is no close relative species of the same genus and species, use PAML software to predict the ancestral sequence of all species, and extract the ancestral sequence as the reference sequence, and then merge it with the transcriptome results of step 1 to obtain the reference transcript sequence of the target species auxiliary annotation;
[0013] Step 7. Using PASA (Program to Assemble Spliced Alignments) software to annotate the transcript set obtained in step 6 to make up for the lack of incomplete or low-quality transcriptome data;
[0014] Step 8. Using the gene structure obtained in step 7 as the skeleton, using EVM (EVidence Modeler) software, setting a reasonable weight value, integrating the results of de novo prediction, homologous annotation and transcriptome annotation, and obtaining high-quality gene structure annotation results.
[0015] In the example of the present application, the environment required for the method for gene structure annotation using software is configured as follows: a PC client with a hardware configuration of a Pentium II CPU or above, 16G or more of memory, and 80G or more of hard disk; software configured as a Linux operating system and configured with gcc 4.9.3 or above, Python 2.7.9 or above, Perl 5.22.0 or above, R 2.15.3 or above, Orthofinder (or OrthMCL, TreeFam) any version, PAML software package, ab initio prediction software August, GeneScan, Glimmer, homologous annotation software Genewise, transcriptome-assisted annotation software PASA, and integrated annotation software EVM.
[0016] Compared with the prior art, the present application has the following beneficial effects:
[0017] The method for gene structure annotation provided by the present application, based on a new gene structure annotation strategy, can effectively compensate for the negative impact of incomplete transcriptome data information on genome annotation due to various reasons, can more comprehensively identify all genes on the genome, and improve the accuracy and integrity of gene structure annotation, thereby providing comprehensive and important reference value and technical support for further comprehensively revealing various genes and their functions of a biological organism. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the main content and specific embodiments of the present application, Figure 1 A brief introduction will be given to the specific embodiments of the present application, wherein Trinity, Cufflinks, OrthoFinder, PAML, EVM, and Pasa are bioinformatics analysis software; and pep is the abbreviation of amino acid sequence; and CDS is the abbreviation of nucleic acid sequence.
[0019] Figure 1 The technical flowchart provided by the present application is shown in the following figure. DETAILED DESCRIPTION
[0020] The embodiments of the present application will be described in detail below in combination with the embodiments and specific examples, but those skilled in the art will understand that the following embodiments and examples are only used to illustrate the present application, and should not be regarded as limiting the scope of the present application.
[0021] The present application provides a method for gene structure annotation, which is performed by the following steps:
[0022] Step 1. Transcriptome assembly: using Trinity software to assemble the transcriptome data of the target species;
[0023] Step 2. Data preparation: download the nucleic acid sequence of the longest transcript of the target species and the amino acid sequence file of the coding protein from the public database such as NCBI, CNGB, etc. of the close relative species of the target species;
[0024] Step 3. Gene family clustering: use OrthoFinder software to perform gene family clustering analysis on the gene set of homologous species, and identify the orthologous gene pairs of the close relative species through the gene family clustering results;
[0025] Step 4. Determine whether there is a close relative species of the same genus and species of the target species in step 2;
[0026] Step 5. Define the orthologous gene pairs commonly existing in the genomes of the close relative species screened in step 3 as the core gene set, and select the longest transcript of the closest relative species as the candidate reference sequence;
[0027] Step 6. If there is a close relative species of the same genus and species, integrate the longest transcript sequence in the close relative species of the same genus and species into the transcriptome assembly result in step 1; if there is no close relative species of the same genus and species, use PAML software to predict the ancestral sequence of all species, and extract the ancestral sequence as the reference sequence, and then combine it with the transcriptome result in step 1 to obtain the reference transcript sequence of the target species auxiliary annotation;
[0028] Step 7. Use PASA (Program to Assemble Spliced Alignments) to annotate the transcript set obtained in step 6 to make up for the deficiency of incomplete or low-quality transcriptome data;
[0029] Step 8. Use EVM (EVidenceModeler) software to set a reasonable weight value to integrate the de novo prediction, homologous annotation and transcriptome annotation results based on the gene structure obtained in step 7, and obtain high-quality gene structure annotation results.
[0030] In some optional implementation steps, the software used for gene family clustering in step 3 can also use OrthoMCL or Treefam.
[0031] In some optional implementation steps, the prediction of the ancestral sequence in step 6 uses the anester model.
[0032] In some optional implementation steps, the software used for integrating the de novo prediction, homologous annotation and transcriptome annotation data in step 8 can also use BRAKER or MAKER3.
[0033] In some optional implementation steps, the homologous annotation software for the plant in step 8 is Exonerate.
[0034] In some optional implementation steps, after step 8, further comprising: upgrading the annotation result obtained in step 8 by using PASA software to obtain a higher-quality gene structure annotation result.
[0035] The present application is based on the principle of whole genome alignment of closely related species, and selects the nucleic acid sequences of the longest transcripts and the amino acid sequences of the encoded proteins of multiple closely related species of the target species to perform alignment to identify the core gene sequences of the closely related species, so as to obtain a reliable evidence source for gene structure assisted annotation, make up for the deficiency of the transcriptome data of the target species, and greatly improve the accuracy, integrity and comprehensiveness of gene structure annotation compared with the existing gene structure annotation method of integrating annotation information from three different sources of de novo prediction, homologous prediction and transcriptome data.
[0036] In the embodiments of the present application, the quality of gene structure annotation can be further improved by making up for the defects of the transcriptome data of the target species. The present application has been practically verified in the gene structure annotation of a ruminant and a rodent, and the present application will be further described below through specific embodiments.
[0037] Embodiment 1
[0038] Taking the gene structure annotation of a certain ruminant as an example, the annotation process is as follows:
[0039] The ruminant is identified as mammal, and the closely related species of the mammal are identified as mammal-A, mammal-B, mammal-C and mammal-D.
[0040] 1) Use the Trinity software to assemble the transcriptome data of the mammal;
[0041] 2) Download the amino acid sequence mammal-A-pep and the nucleic acid sequence mammal-A-CDS of the genome of the female individual of the ruminant, as well as the amino acid sequences and the nucleic acid sequences of the three closely related species from NCBI, which are mammal-B-pep, mammal-C-pep, mammal-D-pep, mammal-B-CDS, mammal-C-CDS and mammal-D-CDS, respectively;
[0042] 3) Use the software OrothoFinder to cluster the homologous gene sets of the species in 2) and identify the orthologous gene pairs common to the closely related species;
[0043] 4) Select the longest transcript of the core gene of the female individual of the ruminant animal from 3) as a candidate reference sequence;
[0044] 5) Integrating the transcriptome assembly results obtained in 1);
[0045] 6) Using PASA to perform gene structure annotation based on the transcript evidence obtained in 5);
[0046] 7) Using the gene structure predicted in 6) as a basic framework, using EVM software to integrate the results of de novo prediction, homologous prediction and the above-mentioned improved transcriptome annotation results of the species, to obtain the annotation results of the gene structure.
[0047] According to the above analysis method, the quality of gene structure annotation is evaluated. Using the traditional method, the results from de novo prediction, homologous prediction and transcriptome annotation are integrated by using EVM software, and a total of 25,932 protein-coding genes are annotated, and the BUSCO evaluation is 79.7% (C: 79.7% [S: 77.5%, D: 2.2%], F: 6.7%, M: 13.6%, n: 3354). Through the further improvement of the transcriptome data of the ruminant animal by the present application, the result of the gene structure annotation is improved, and a total of 26,294 protein-coding genes are identified, and the BUSCO evaluation is 92.1% (C: 92.1% [S: 89.7%, D: 2.4%], F: 3.0%, M: 4.9%, n: 3354). The above examples show that the results of the present application are reliable, and can effectively improve the accuracy and integrity of the gene structure annotation of the ruminant animal.
[0048] Example 2
[0049] Taking the gene structure annotation of a certain rodent as an example, the annotation process is as follows:
[0050] The rodent is identified as Rodent, and the identification of its close species is rhizomys sinensis, bamboo rat and Rhizomys pruinosus.
[0051] 1) Using Trinity software to assemble the transcriptome data of Rodent;
[0052] 2) Download amino acid sequences rhizomys_sinensis_pep, bamboo_rat_pep, Rhizomys_pruinosus_pep and nucleic acid sequences rhizomys_sinensis_cds, bamboo_rat_cds, Rhizomys_pruinosus_cds of rhizomys_sinensis, bamboo rat and Rhizomys_pruinosus from NCBI;
[0053] 3) Cluster homologous gene sets of the species in 2) using software OrthoFinder, and identify orthologous gene pairs of the close species;
[0054] 4) Select the longest transcript of the close species rhizomys_sinensis of the target rodent from 3) as the candidate reference sequence;
[0055] 5) Integrate the transcriptome assembly result obtained in 1);
[0056] 6) Annotate the combination of the transcripts obtained in 5) using PASA;
[0057] 7) Integrate the results of de novo prediction, homologous prediction and the improved transcriptome annotation obtained above using EVM software based on the gene structure in 6), to obtain the annotation result of the rodent gene structure.
[0058] The quality of the gene structure annotation is evaluated according to the above analysis method. Using the traditional method, the results from de novo prediction, homologous prediction and transcriptome annotation are integrated using EVM software, and a total of 31,465 protein-coding genes are annotated, with a BUSCO evaluation of 64.3% (C: 64.3% [S: 58.9%, D: 5.4%], F: 14.3%, M: 21.4%, n: 3354). Through further improvement of the transcriptome data of the ruminant by the present application, the result of the gene structure annotation is improved, and a total of 39,742 protein-coding genes are identified, with a BUSCO evaluation of 89.0% (C: 89.0% [S: 80.8%, D: 8.2%], F: 5.3%, M: 5.7%, n: 3354). The above research shows that the present application greatly improves the integrity of the gene structure annotation, and has universality, and is suitable for gene structure annotation research of various species.
[0059] Note: The number of genes in the above two examples is the evaluation result after integration, which has not been filtered for repeated sequence regions, that is, the result after PASA update.
[0060] The above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can still be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method of gene structure annotation, characterized by, The following steps are performed: Step 1. Transcriptome assembly: the transcriptome data of the target species is assembled using Trinity software; Step 2. Data preparation: the nucleic acid sequence of the longest transcript of the target species and the amino acid sequence file of the encoded protein of the target species are downloaded from public databases such as NCBI, CNGB, etc.; Step 3. Gene family clustering: the gene set of homologous species is subjected to gene family clustering analysis using OrthoFinder software, and the orthologous gene pairs of the homologous species are identified by the gene family clustering results; Step 4. Determine whether there is a same-species homologous species of the target species in the homologous species in step 2; Step 5. The orthologous gene pairs common to the genomes of the homologous species screened in step 3 are defined as the core gene set, and the longest transcript of the closest homologous species is selected as the candidate reference sequence; Step 6. If there is a same-species homologous species, the longest transcript sequence in the same-species homologous species is integrated into the transcriptome assembly results in step 1; If there is no same-species homologous species, use PAML software to predict the ancestral sequence of all species, and extract the ancestral sequence as the reference sequence, and then merge it with the transcriptome results in step 1 to obtain the reference transcript sequence for auxiliary annotation of the target species; Step 7. Use the PASA (Program to Assemble Spliced Alignments) software to annotate the transcript set obtained in step 6 to make up for the lack of transcriptome data or insufficient quality; Step 8. Use the gene structure obtained in step 7 as the skeleton, use the EVM (EVidence Modeler) software, set a reasonable weight value, integrate the de novo prediction, homologous annotation and transcriptome annotation results, and obtain high-quality gene structure annotation results.
2. The method of claim 1, wherein, In step 3, the software for gene family clustering can also use OrthMCL or Treefam.
3. The method of claim 1, wherein, In step 6, the ancestral sequence prediction uses the anester model.
4. The method of claim 1, wherein, In step 8, the software for integrating de novo prediction, homologous annotation and transcriptome annotation data is BRAKER or MAKER3.
5. The method of claim 1, wherein, The environment configuration required for the method of using software to perform gene structure annotation is: the hardware configuration is a PC client with a Pentium II CPU, 16G of memory, and an 80G hard disk; the software configuration is a Linux operating system and is configured with gcc 4.9.3 and later versions, Python 2.7.9 and later versions, Perl 5.22.0 and later versions, R 2.15.3 and later versions, Orthofinder or OrthoMCL, TreeFam any version, PAML software package, de novo prediction software August, GeneScan, Glimmer, etc., homologous annotation software Genewise, transcriptome auxiliary annotation software PASA, and integration annotation software EVM.
Citation Information
Patent Citations
Gene annotation method and system
CN101894211A
Accurate and efficient eukaryotic biological gene identification method based on transcriptome
CN108753994A