A method for identifying lineage-specific expanded gene families
The method of identifying lineage-specific amplified genes by orthologous homology analysis and genomic data screening solves the identification problem in the existing technology, provides a simple and effective method for identifying lineage-specific genes, and the screened genes have good representativeness.
Patent Information
- Application Number
- CN202111634211.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-29
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-12-29
AI Technical Summary
The lack of effective methods in the current technology to identify lineage-specific amplified gene families affects our understanding of plant evolution and adaptation.
Orthologous analysis was used to sort and screen the genomic data of the target species, and the genomic data of the experimental subfamily and the control subfamily were compared to screen for lineage-specific amplified genes. Their functions were then determined by GO enrichment analysis.
This method enables the simple and efficient identification of lineage-specific gene families, provides an analysis method that requires minimal data and is easy to operate, and the screened genes are highly representative, laying the foundation for further research.
Smart Images

Figure CN114300050B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bioinformatics analysis technology, and in particular to a method for identifying lineage-specific amplified gene families. Background Technology
[0002] Phylogenetic genes refer to a gene family exhibiting significant amplification within a particular plant species. These amplified genes are called phylogenetic genes. They generally play a crucial role in the functional traits of that species. Similarly, a gene family exhibiting significant amplification within a subfamily can be called a phylogenetic amplified gene family. The significant amplification of these genes plays a vital role in speciation within that family. Therefore, the identification of these genes is of great importance for understanding plant evolution and adaptation. Currently, no methods have been reported for identifying phylogenetic amplified gene families. Summary of the Invention
[0003] To overcome the aforementioned deficiencies in the existing technology, the present invention provides a method for identifying lineage-specific amplified gene families appearing in different subfamilies using orthologous analysis.
[0004] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0005] This invention provides a method for identifying lineage-specific amplified gene families, comprising the following steps:
[0006] (1) Obtain the genomic information of the target species and record the genomic data of the target species. Then, based on the known evolutionary information, sort the target species according to their evolutionary relationship.
[0007] (2) Orthologous genes were obtained by performing orthologous analysis on the whole proteins of the target species;
[0008] (3) Check whether the order of the evolutionary relationship of the orthologous genes is consistent with the evolutionary relationship of the target species that has been sorted in step (1). If they are inconsistent, sort them according to the evolutionary relationship of the target species that has been sorted in step (1).
[0009] (4) Using the subfamily of the target species as the experimental subfamily and the non-experimental subfamily as the control subfamily, the genomic data of the experimental and control subfamilies are compared pairwise, and preliminary screening and detailed screening are carried out in sequence.
[0010] The rough screening was based on the difference in the average number of genes compared in pairs between the experimental subfamily and the control subfamily as the first criterion for sorting. Gene data with an average difference ≤ 6 were removed, and the remaining data were carefully screened.
[0011] The careful screening refers to screening for genes that meet any of the following conditions:
[0012] Genes that are absent in the control subfamily but present in over 70% of species in the experimental subfamily;
[0013] Or genes present in less than 10% of species in the control subfamily, but present in more than 80% of species in the experimental subfamily;
[0014] Or it may be present in more than 80% of species in the experimental and control subfamilies, and the average number of genes in the experimental subfamilies is 2 to 3 times the average number of genes in the control subfamilies;
[0015] The genes that were carefully selected and retained above are those that have undergone lineage-specific amplification in the experimental subfamily.
[0016] Preferably, the number of species in the experimental subfamily and the non-experimental subfamily in step (4) is the same.
[0017] Preferably, the method further includes a step of performing GO enrichment analysis on the genes that have undergone lineage-specific amplification in the experimental subfamily obtained in step (4).
[0018] The beneficial effects of this invention are as follows:
[0019] This invention provides a method for identifying lineage-specific amplified gene families. It utilizes orthologous analysis to screen and identify lineage-specific amplified genes appearing in different subfamilies. This method is simple to operate, requires minimal data, and the analysis software is easy to learn. Furthermore, the screened genes are highly representative, laying the foundation for future large-scale identification and functional analysis of lineage-specific genes. Attached Figure Description
[0020] Figure 1This is the original result of Orthogroup.tsv after the orthofinder run is complete (Note: In the figure, Vv, At, Gm, Mb, Prm, Pa, Ps, Py, Pom, Fn, Fv, Rc, Ro represent Vitis vinifera L (grape), Arabidopsis thaliana (Arabidopsis thaliana), Glycine max (soybean), Malus baccata (Mountain hawthorn), Prunus mume (plum blossom), Prunus avium (European sweet cherry), Prunus salicina (plum), Prunus yedoensis (Tokyo cherry), Potentilla micrantha (Potentilla micrantha), Fragaria nubicola (Tibetan strawberry), Fagaria vesca (strawberry), Rosa chinensis (rose), Rubus occidentalis (bright yellow dwarf tree), the same below);
[0021] Figure 2 The results are after a rough screening.
[0022] Figure 3 The final result for all groups that have been amplified in the Rosoidae subfamily after careful screening;
[0023] Figure 4After performing GO enrichment analysis on lineage-specific amplification genes selected in the Rosoidae subfamily, the major pathways enriched were as follows (Note: the top 10 enriched Goids are: GO: 0000209, protein polyubiquitination; GO: 0070647, protein modification by small protein conjugation or removal; GO: 0048759, xylem vessel membercell differentiation; GO: 0009808, lignin metabolic process; GO: 0006511, ubiquitin-dependent protein catabolic process; GO: 0019941, modification-dependent protein catabolic process; GO: 0043632, modification-dependent macromolecule catabolic process). process, modification-dependent macromolecular degradation process, GO: 0006508, proteolysis, protein degradation, GO: 0051603, proteolysis involved in cellular protein catabolic process, GO: 0044257, cellular protein catabolic process. Detailed Implementation
[0024] This invention provides a method for identifying lineage-specific amplified gene families, comprising the following steps:
[0025] (1) Obtain the genomic information of the target species and record the genomic data of the target species. Then, based on the known evolutionary information, sort the target species according to their evolutionary relationship.
[0026] (2) Orthologous genes were obtained by performing orthologous analysis on the whole proteins of the target species;
[0027] (3) Check whether the order of the evolutionary relationship of the orthologous genes is consistent with the evolutionary relationship of the target species that has been sorted in step (1). If they are inconsistent, sort them according to the evolutionary relationship of the target species that has been sorted in step (1).
[0028] (4) Using the subfamily of the target species as the experimental subfamily and the non-experimental subfamily as the control subfamily, the genomic data of the experimental and control subfamilies are compared pairwise, and preliminary screening and detailed screening are carried out in sequence.
[0029] The rough screening was based on the difference in the average number of genes compared in pairs between the experimental subfamily and the control subfamily as the first criterion for sorting. Gene data with an average difference ≤ 6 were removed, and the remaining data were carefully screened.
[0030] The careful screening refers to screening for genes that meet any of the following conditions:
[0031] Genes that are absent in the control subfamily but present in over 70% of species in the experimental subfamily;
[0032] Or genes present in less than 10% of species in the control subfamily, but present in more than 80% of species in the experimental subfamily;
[0033] Or it may be present in more than 80% of species in the experimental and control subfamilies, and the average number of genes in the experimental subfamilies is 2 to 3 times the average number of genes in the control subfamilies;
[0034] The genes that were carefully selected and retained above are those that have undergone lineage-specific amplification in the experimental subfamily.
[0035] In this invention, the careful screening described in step (4) is to be present in more than 80% of the species in the experimental subfamily and the control subfamily, and the average number of genes in the experimental subfamily is 2 to 3 times the average number of genes in the control subfamily, more preferably present in more than 80% of the species in the experimental subfamily and the control subfamily, and the average number of genes in the experimental subfamily is 2.5 times the average number of genes in the control subfamily.
[0036] In this invention, the number of species in the experimental subfamily and the non-experimental subfamily mentioned in step (4) is preferably the same.
[0037] In this invention, it is preferred to further include a step of performing GO enrichment analysis on the genes that have undergone lineage-specific amplification in the experimental subfamily obtained in step (4), so as to provide a functional description of the screened genes that have undergone lineage-specific amplification in the experimental subfamily.
[0038] The technical solutions provided by the present invention will be described in detail below with reference to the embodiments, but they should not be construed as limiting the scope of protection of the present invention.
[0039] Example 1
[0040] (1) Download Rosaceae plants from the following websites: GDR (https: / / www.rosaceae.org / ), ensemblplants (http: / / plants.ensembl.org / index.html), and phytozome (https: / / phytozome-next.jgi.doe.gov / ): Potentilla micrantha, Fragaria nubicola, Fagaria vesca, Rubus occidentalis, and Rosa chinese. Maloideae plants: Malus baccata, Prunus avium, Prunus mume, Prunus salicina, and Prunus yedoensis. In addition, the genome data of 13 species of non-Rosaceae plants, including grape (Vitis vinifera), Arabidopsis (Athaliana), and soybean (Glycine max), were collected. Then, based on known evolutionary information, these 13 species were sorted according to their evolutionary relationship.
[0041] (2) Use gffreader to extract the PEP and CDS sequences of the above 13 species, change all PEP sequence files to the Latin name of the species, and place them in the same folder. Run orthofinder to perform orthologous analysis on the whole proteins of the 13 species to obtain orthologous genes. The parameters used by Orthofinder during operation are: -t: number of threads used for sequence search: 16; -a: number of threads used for sequence analysis: 1; -M: gene tree inference method: dendroblast; -S: program used for sequence alignment: Blast.
[0042] (3) Locate the Orthogroups folder in the Orthoofinder output, find the Orthogroup.tsv file, create a heatmap to visualize it, and the result is as follows: Figure 1 As shown. By Figure 1It can be seen that the Rosoidae subfamily has a large number of expanded gene families compared to the Maloideae subfamily. Paste the data into a new Excel spreadsheet. Check whether the order of the evolutionary relationships of the obtained orthologous genes is consistent with the evolutionary relationships of the target species that have been sorted in step (1). If they are inconsistent, sort them according to the evolutionary relationships of the target species that have been sorted in step (1). Use the conditional filtering function of the Excel spreadsheet to fill in the colors according to the different values in each table;
[0043] (4) Using Rosoidae as the experimental subfamily and Maloideae as the control subfamily, the genomic data of Rosoidae and Maloideae were compared pairwise, and preliminary screening and detailed screening were carried out in sequence.
[0044] Genes exhibiting irregular amplification and contraction patterns were removed. Then, the gene counts were sorted using the difference in the average number of genes compared pairwise within the Rosoidae and Maloideae subfamilies as the first criterion. Genes with an average difference ≤ 6 were removed, yielding a rough screening result (e.g., ...). Figure 2 (as shown);
[0045] The remaining data is carefully sifted, meaning genes that meet any of the following criteria are selected:
[0046] This gene is absent in the Rosinoideae subfamily but present in over 70% of species in the Maloideae subfamily.
[0047] Or genes present in less than 10% of species in the Maloideae subfamily, but present in more than 80% of species in the Rosoidae subfamily;
[0048] Or it may be present in more than 80% of species in the Rosinoideae and Maloideae subfamilies, and the average number of genes in the Rosinoideae subfamily is 2.5 times the average number of genes in the Maloideae subfamily;
[0049] After the above careful screening, a total of 62 sets of orthologous genes were retained (such as...). Figure 3 As shown in the figure, this is a gene that has undergone lineage-specific amplification in the Rosine subfamily.
[0050] (5) In the Orthogroup_Sequences folder of the orthofinder results file, select the 62 orthologous gene groups that were finally screened out in the table. Figure 3 The corresponding groups were extracted from the folder and subjected to GO enrichment analysis. It was found that the functions enriched in these selected lineage-specific genes were manifested in lignin degradation and protein metabolism pathways (e.g., in the Rosoidae subfamily). Figure 4 (As shown).
[0051] The identification method described in this invention is simple to operate, requires little data, and the analysis software is easy to learn. Furthermore, the selected genes are highly representative, laying the foundation for the large-scale identification and functional analysis of lineage-specific genes in the future.
[0052] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for identifying lineage-specific amplified gene families, characterized in that, Includes the following steps: (1) Obtain the genomic information of the target species and record the genomic data of the target species. Then, based on the known evolutionary information, sort the target species according to their evolutionary relationship. (2) Orthologous genes were obtained by performing orthologous analysis on the whole proteins of the target species; (3) Check whether the order of the evolutionary relationship of the orthologous genes is consistent with the evolutionary relationship of the target species that has been sorted in step (1). If they are inconsistent, sort them according to the evolutionary relationship of the target species that has been sorted in step (1). (4) Using the subfamily of the target species as the experimental subfamily and the non-experimental subfamily as the control subfamily, the genomic data of the experimental and control subfamilies are compared pairwise, and preliminary screening and detailed screening are carried out in sequence. The rough screening was based on the difference in the average number of genes compared in pairs between the experimental subfamily and the control subfamily as the first criterion for sorting. Gene data with an average difference ≤ 6 were removed, and the remaining data were carefully screened. The careful screening refers to screening for genes that meet any of the following conditions: Genes that are absent in the control subfamily but present in over 70% of species in the experimental subfamily; Or genes present in less than 10% of species in the control subfamily, but present in more than 80% of species in the experimental subfamily; Or it may be present in more than 80% of species in the experimental and control subfamilies, and the average number of genes in the experimental subfamilies is 2 to 3 times that in the control subfamilies; The genes that were carefully selected and retained above are those that have undergone lineage-specific amplification in the experimental subfamily; The target species is a plant of the Rosaceae family.
2. The method for identifying lineage-specific amplified gene families according to claim 1, characterized in that, The number of species in the experimental subfamily and the non-experimental subfamily mentioned in step (4) is the same.
3. The method for identifying lineage-specific amplified gene families according to claim 1 or 2, characterized in that, It also includes a step of performing GO enrichment analysis on genes that have undergone lineage-specific amplification in the experimental subfamily obtained in step (4).