Method, device, equipment and medium for identifying key genes of biological complex characters
By performing orthologous gene group classification and phenotypic data analysis on biological reference genomic data, and using machine algorithm models to construct phenotypic prediction models, the problem of insufficient accuracy and comprehensiveness of the identification of key genes of complex traits in advanced organisms in existing technology, and efficient and comprehensive identification of key genes of complex traits is achieved.
Patent Information
- Application Number
- CN202510052860.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-03
AI Technical Summary
Existing cross-species genomic localization methods are insufficient in the identification of key genes of complex traits in higher organisms, especially when dealing with unknown functional genes and new key sequences, it is difficult to meet the needs of complex trait research.
By obtaining biological reference genomic data, orthologous gene group segments are divided into the encoded protein sequences, a characteristic value data set is constructed, and a phenotype matrix is constructed using machine algorithm models to train phenotype prediction models and determine the impact weight of subpopulation sequences on biological phenotypes, thereby identifying key genes that regulate biological complex traits.
Accurate, efficient and comprehensive identification of key genes in complex traits of prokaryotes and higher organisms has been achieved, and unidentified key gene sequences can be found to meet the needs of comprehensive trait research for full coverage of key genes.
Smart Images

Figure CN120089193A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics high-throughput data analysis, and particularly to a method, apparatus, device and medium for identifying key genes of complex biological traits. Background Art
[0002] Complex biological traits are formed by the functional interaction of multiple genes and environmental factors, and their regulatory mechanisms are complex, involving multiple interactions of gene-gene, gene-environment and epigenetic networks. Such traits include the growth and development of organisms, stress resistance, yield, and the formation of specific phenotypes, etc., which are one of the research focuses in the field of life science. However, the genetic basis of complex traits is often difficult to analyze because it involves the coordinated action of multiple genes, multi-level regulatory networks, and the combined effects of genetic and non-genetic factors. In recent years, the rapid development of genomics, transcriptomics, epigenetics and systems biology has provided strong technical support for the research of complex trait regulatory gene modules.
[0003] Through genome-wide association studies (GWAS), scientists can initially identify candidate gene regions associated with complex traits; subsequently, by integrating transcriptome data, weighted gene co-expression network analysis (WGCNA) can be used to further screen for functional genes. In addition, epigenetic regulation studies have revealed the important roles of DNA methylation, histone modification, and non-coding RNAs in the regulation of gene module functions. On this basis, using CRISPR / Cas9 gene editing technology, RNA interference, and the construction of transgenic models, the functions of key genes in the module can be verified, and their roles in the regulation of complex traits can be further explored. However, despite the increasing maturity of technology, the study of gene modules regulating complex traits still faces many challenges. Especially in higher organisms, due to their complex genomic structure, high gene function diversity, and more complex regulatory network hierarchies, the research difficulty has increased significantly. To address this problem, we have developed and applied for a new method for identifying bacterial functional genes based on cross-species genome mapping (Patent No.: ZL202310714738.3, Patent Name: Identification Method, Device, Equipment, and Medium for Candidate Genes Regulating Bacterial Shape) to identify the candidate gene composition associated with specific common phenotypes. The existing patented method quickly locates functional genes at the genomic level by integrating the characteristic information and functional analysis of protein domains. Compared with traditional methods, this technology has high accuracy and efficiency in analyzing the functions of microbial genes. However, the existing method also exposes some limitations. For example, this method relies on the third-party tool pfam_scan to analyze protein domains. When dealing with proteins with unknown functions, the pfam_scan database is difficult to provide complete classification or annotation, resulting in the direct loss of important unknown gene information for some traits of interest during gene identification. The existing method has weak ability to mine genes with unknown functions and new key sequences, and cannot meet the need for full coverage of key genes in complex trait research.
[0004] Although the functional genes of prokaryotes can be efficiently predicted through the compositional characteristic values of domains, due to the large number of members of the gene family-based characteristics in higher organisms such as plants, the identification efficiency of existing gene identification methods based on cross-species genomes is not ideal in higher organisms, and the accuracy and comprehensiveness of functional gene mapping in higher organisms need to be further improved technically. Therefore, there is an urgent need to develop a more accurate method that combines efficient mapping of known functional genes and comprehensive mining of unknown functional genes to break through the current technical bottleneck. Summary of the Invention
[0005] To solve the above problems, the present invention provides a method, device, equipment, and medium for identifying key genes of biological complex traits to achieve more accurate, efficient, and comprehensive identification of key genes of complex traits in prokaryotes and higher organisms.
[0006] The present invention realizes the above object through the following technical solutions:
[0007] In the first aspect of the present invention, a method for identifying key genes of biological complex traits is provided, and the method includes:
[0008] Obtain biological reference genome data, and perform orthologous gene group classification on the protein sequences encoded by the biological reference genome data;
[0009] According to the subgroup sequence numbers and the number of genes in the subgroup sequences obtained after classification, determine the eigenvalue data set;
[0010] Obtain each piece of the biological phenotype data;
[0011] Using a machine algorithm model, construct a phenotype matrix according to each piece of the biological phenotype data and the eigenvalue data set to train a phenotype prediction model, and determine the influence weight of each subgroup sequence on the biological phenotype according to the phenotype prediction model;
[0012] According to each influence weight, determine the key genes that regulate biological complex traits.
[0013] Further, the acquisition source of the biological reference genome data includes any one or more of the following databases: NCBI database, ORCAE database, GWH database, fernbase database, Phytozome database, Ensembl database or CuGenDB database.
[0014] Further, the performing orthologous gene group classification on the protein sequences encoded by the biological reference genome data includes:
[0015] According to the protein sequences encoded by the genomes of different organisms, divide the gene sequences into different orthologous gene groups, that is, subgroups.
[0016] Further, the determining the eigenvalue data set according to the subgroup sequence numbers and the number of genes in the subgroup sequences obtained after classification includes:
[0017] According to the subgroup numbers obtained by classifying the orthologous gene groups of each organism, construct a data set of the number of genes in each subgroup sequence as the eigenvalue data set.
[0018] Further, the machine algorithm model is constructed by a random forest algorithm.
[0019] In the second aspect of the present invention, a device for identifying key genes of biological complex traits is provided, and the device includes:
[0020] A first acquisition module, configured to acquire each biological reference genome data, and perform orthologous gene group classification on the protein sequences encoded by each biological reference genome data;
[0021] A first determination module, configured to determine a characteristic value data set according to the subgroup sequence numbers and the number of genes in the subgroup sequences obtained after classification;
[0022] A second acquisition module, configured to acquire each of the biological phenotype data;
[0023] A processing module, configured to use a machine algorithm model to construct a phenotype matrix according to each of the biological phenotype data and the characteristic value data set to train a phenotype prediction model, and determine the influence weight of each of the subgroup sequences on the biological phenotype according to the phenotype prediction model;
[0024] A second determination module, configured to determine key genes regulating biological complex traits according to each influence weight.
[0025] In a third aspect of the present invention, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the steps of the method according to any one of the above when executing the computer program.
[0026] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the method according to any one of the above when being executed by a processor.
[0027] The beneficial effects of the present invention are as follows:
[0028] The present invention provides a method for identifying key genes of biological complex traits. The method acquires biological reference genome data, performs orthologous gene group classification on the protein sequences encoded by each biological genome, constructs a gene number data set based on the subgroup sequences as characteristic values based on the subgroup numbers and the number of subgroup genes of the orthologous gene groups. Subsequently, in combination with the phenotype characteristic data and the orthologous gene group classification characteristic value data set, a machine learning prediction model is established, and the influence weight of different subgroup sequences on the phenotype is evaluated according to the model, so as to determine the composition of key genes regulating biological complex traits.
[0029] The present invention classifies orthologous genes in the gene sequences of different species genomes to construct a data set containing gene classification sequences and numbers. This data set serves as a characteristic value data set and combines phenotypic characteristics as markers of complex traits, and then constructs a prediction model based on the sequence numbers of each subgroup. This model can accurately predict the key gene compositions of prokaryotes and higher organisms for complex traits, achieving the purpose of quickly and accurately identifying key gene modules without relying on third-party gene annotation tools. The present invention constructs a phenotypic prediction model through cross-species gene sequence classification and sequence subgroup gene number matrix, and mines the key gene compositions of complex traits. In addition to being able to identify genes with known functions, the present invention can also discover key gene sequences that have not been recognized in the past, thus realizing a comprehensive and efficient identification of the gene composition of functional modules.
[0030] The development of the new method provides strong key technical support for accurate and efficient gene identification required in fields such as agricultural breeding, biological improvement, and medical research. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the technical solution roadmap provided by the present invention;
[0032] Figure 2 is the random forest prediction results of two models provided in Example 1 of the present invention;
[0033] Figure 3 is the random forest prediction results of two models provided in Example 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] The following further describes the present application in detail with reference to the drawings. It is necessary to point out here that the following specific embodiments are only used to further illustrate the present application and should not be construed as limiting the protection scope of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application according to the above application content.
[0035] Example 1
[0036] The symbiosis between plants and AM fungi depends on the coordinated action of multiple key genes, which are involved in important biological processes such as signal transduction, structural development, and nutrient exchange. In the signal transduction stage, the SYMRK and DMI gene families are involved in the recognition and transmission of mycorrhizal signals; during the formation of symbiotic structures, genes such as RAM1 and EXO70I regulate the development of arbuscules; in terms of nutrient exchange, the phosphate transporter PT4 and the lipid transporter STR / STR2 play key roles. In addition, CCaMK and IPD3, as signal regulators, are widely involved in signal conduction and regulation. The above known genes promote the establishment of the symbiosis between plants and AM fungi through their interactions.
[0037] Based on the above, this embodiment provides a method for analyzing plant regulatory genes for symbiosis between plant roots and AM fungi. The technical solution roadmap is as shown in Figure 1 shown, and the steps include:
[0038] (1) Obtain plant reference genome data from public databases. Perform orthologous gene group classification on the proteins encoded by the reference genomes of each plant species. For the coding genes of each plant species, write a program to remove duplicates from the multiple transcripts of each coding gene of each plant species, and only retain one transcript for each gene. Translate the retained transcript into a protein or directly extract the corresponding protein sequence.
[0039] Statistics of plant species and data sources are shown in Table 1.
[0040] Table 1. Statistics of plant species and data sources
[0041]
[0042]
[0043]
[0044]
[0045]
[0046] Note: Arbuscular Mycorrhizal Symbiosis (AMS); 0 (non-symbiotic), 1 (symbiotic).
[0047] Open-source software Orthofinder can be used for orthologous gene sequence analysis, or other software with the same function can be used for orthologous gene group classification. In this embodiment, the open-source software Orthofinder is used to divide the gene sequences into different orthologous gene groups according to the protein sequences encoded by the genomes of different plants.
[0048] Specifically, place the protein sequence files of each species in the Proteins_Data folder with the extensions required by the open-source software Orthofinder, and use the default parameters: "OrthoFinder / orthofinder - fOrthoFinder / Proteins_Data". In the case of a large number of species, first take the protein sequences of a small number of representative species and operate according to the above parameters. For the protein sequences of additional species, use the extended append parameters provided by the software (version 3.0 and above): "orthofinder--core Proteins_Data / OrthoFinder / Results_Core / --assign Proteins_Data / AdditionalSpecies" for the analysis of protein sequences of more species to improve the software running speed.
[0049] (2) Search for literature and determine the phenotypic data on whether each plant can symbiose with AM fungi. Use 1 to represent symbiosis and 0 to represent non - symbiosis.
[0050] (3) According to the numbered division of plant orthologous gene sequences and the number of genes in each subgroup, construct a dataset of the number of genes with the numbered division of each orthologous gene sequence as the eigenvalue dataset.
[0051] Refer to Table 2 for the data structure table in the eigenvalue dataset of the subgroup of the above - mentioned plant orthologous gene sequences in Example 1.
[0052] Table 2. Schematic table of data structure in Example 1
[0053]
[0054] As shown in Table 2 above, the above - mentioned eigenvalue dataset can include the numbers after the classification of each gene sequence class of each plant and the number of genes of the numbered classification of each gene sequence class in each plant.
[0055] (4) Construction of the symbiosis ability matrix. In this example, by combining the obtained plant dataset with the research results on whether it can symbiose with AM fungi, construct a phenotypic matrix of symbiosis or not.
[0056] (5) Make an experimental design table design.txt. First, divide the plant mycorrhizal symbiosis data into a training group (group1) and a test group (group2). Generally, adjust the ratio of the training group to the test group to 3:1 or 1:1 to achieve the best prediction model.
[0057] Refer to Table 3 for the grouping list.
[0058] Table 3. Grouping list in Example 1
[0059]
[0060] Specifically, the present application can divide each plant into two groups. In one possible design, each representative plant can be divided into a training group and a test group at a ratio of 1:1, that is, 50% as the training group and 50% as the test group. It should be noted that when grouping, the training group after grouping should cover the above two phenotypes, symbiotic and non-symbiotic as much as possible, and the test group after grouping should also cover the above two phenotypes, symbiotic and non-symbiotic phenotypes.
[0061] This embodiment can construct a machine learning model through the random forest algorithm in R language. The random forest algorithm is a machine learning method based on decision trees. It completes classification or regression tasks by constructing multiple decision trees and voting on them. The random forest algorithm can effectively overcome the problem of overfitting of a single decision tree and has good performance in various machine learning tasks. It is a commonly used machine learning algorithm. The Orthogroups model can be constructed based on the Orthogroups dataset using the random forest algorithm. The accuracy rates of both the training group and the test group reach more than 90%, and the Kappa coefficients are all higher than 0.80, showing a high degree of consistency( Figure 2 )
[0062] In Example 1 of the present invention, we further performed comparative annotation analysis on each key gene model predicted by different methods and found that although the Pfam model of the pfam_scan tool can annotate a few known symbiosis-related genes when parsing protein domains, the accuracy is relatively low (Table 4).
[0063] Table 4. Comparison and statistics of the top 50 important feature alignments of the two models in Example 1 with known symbiotic genes
[0064]
[0065] Among the top 50 important features, the Pfam model can only identify 2 known symbiosis-related genes and 1 downstream gene bound by the AP2 transcription factor (Tables 4 and 5).
[0066] Table 5. Alignment results of the top 50 important features of the Pfam model provided in Example 1 with known symbiotic genes
[0067]
[0068]
[0069] In contrast, the Orthogroups model constructed based on the number of orthologous genes of the present invention not only improves the accuracy and consistency, but also can annotate more symbiosis-related genes with known functions. Among the top 50 important features, the Orthogroups model identified 25 known symbiosis-related genes, and 5 downstream genes bound by AP2 transcription factors were identified (Tables 4 and 6). In addition, there are 20 unknown gene sequences to be further experimentally verified for their symbiosis-relatedness.
[0070] Table 6. Results of aligning the top 50 important features of the Orthogroups model provided in Example 1 with known symbiotic genes
[0071]
[0072]
[0073]
[0074] The present invention provides a new method for predicting functional genes, taking the symbiotic genes of plants and arbuscular mycorrhizal fungi (AMF) as an example. Compared with the existing methods, the method of the present invention does not rely on the annotation of a third-party functional gene database, but directly establishes the association between gene sequences and symbiotic phenotypes through an algorithm, thereby predicting the most likely symbiotic candidate gene sequences, and successfully achieving a comprehensive, accurate and efficient identification of the key functional genes of different plant common traits. The sequences obtained by prediction include both most of the symbiosis-related genes that have been experimentally verified and a part of the sequences that have not been experimentally verified. This new method can comprehensively, accurately and efficiently identify functional genes, providing strong support for the localization of plant functional genes and biological breeding, and is of great significance in solving the technical bottlenecks existing in traditional breeding methods and promoting the mining and application of functional genes.
[0075] Example 2
[0076] This example provides a method for efficiently identifying bacterial motility regulatory gene modules. The specific method is similar to the idea of Example 1, and the specific steps are as follows:
[0077] (1) Obtain the reference genome data related to bacterial motility and the encoded protein sequence data from a public database. The open-source software Orthofinder can be used for orthologous gene sequence analysis, or other software with the same function can be used for orthologous gene group classification. In this example, the open-source software Orthofinder is used to divide the gene sequences into different orthologous gene groups (Table 7) according to the protein sequences encoded by each genome of different bacteria. Place the protein sequence files of each species in the Proteins_Data folder with the extension required by the open-source software Orthofinder, and use the default parameters: "OrthoFinder / orthofinder -f OrthoFinder / Proteins_Data". In the case of a large number of species, first take the protein sequences of a small number of representative species and operate according to the above parameters. For the protein sequences of additional species, use the extended append parameters provided by the software (version 3.0 and above): "orthofinder --core Proteins_Data / OrthoFinder / Results_Core / --assign Proteins_Data / AdditionalSpecies" for the protein sequence analysis of more species to improve the software running speed.
[0078] Table 7. Schematic table of data structure in Example 2
[0079]
[0080]
[0081] (2) Search for literature materials to determine the ability of each bacterium to move as phenotypic data. The ability to move is represented by 1, and the inability to move is represented by 0, and a phenotypic matrix is constructed.
[0082] (3) Design the experiment table design.txt. First, divide the bacterial motility dataset into a training group (group1) and a test group (group2). Specifically, each bacterium can be divided into two groups. In one possible design, each representative bacterium can be divided into a training group and a test group at a ratio of 3:1, that is, 75% as the training group and 25% as the test group. The grouping list is shown in Table 8.
[0083] Table 8. Grouping list in Example 2
[0084]
[0085] In this embodiment, a machine learning model can be constructed through the random forest algorithm. The model constructed based on the number of orthologous genes is the Orthogroups model. The accuracy rates of both the training group and the test group reach over 80%, and the Kappa coefficients are all higher than 0.65, showing a high degree of consistency( Figure 3 ). In the existing method of using pfam for positioning, 35 known motility-related genes can be identified in the top 50 important domain rankings (Tables 9 and 10), while the Orthogroups model method of the present invention can identify 39 known motility-related genes (Tables 9 and 11). In addition, there are also a small number of sequences whose functions have not been experimentally studied (Table 11), which are very likely to be motility-related genes.
[0086] Table 9. Comparative statistics of the comparison results between the top 50 important features of the two models in Example 2 and known motility genes
[0087]
[0088] Table 10. Comparison results between the top 50 important features of the Pfam model provided in Example 2 and known motility genes
[0089]
[0090]
[0091]
[0092]
[0093] Table 11. Comparison results between the top 50 important features of the Orthogroups model provided in Example 2 and known motility genes
[0094]
[0095]
[0096] The results of Example 1 and Example 2 show that the method of the present invention not only greatly improves the identification accuracy of genes with complex traits in higher organisms, but also improves the identification accuracy of genes with complex traits in prokaryotes to a certain extent. In addition, the new method can identify some sequences whose functions have not been verified, and these sequences are very likely to be the key functional genes corresponding to the traits of interest, providing targeted research targets for the discovery of new genes. The present invention has good technical versatility and shows good application value for the identification of common trait genes in animals, plants, bacteria, fungi, protozoa, slime molds, myxobacteria, etc.
[0097] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.
Claims
1. A method for identifying key genes for complex biological traits, characterized in that: The method comprises: Acquiring biological reference genome data, and dividing the protein sequences encoded by the biological reference genome data into orthologous gene groups; Determine the characteristic value data set according to the subgroup sequence number obtained after the division and the number of genes in the subgroup sequence; Acquiring phenotypic data of each of the organisms; Using a machine algorithm model, a phenotype matrix is constructed according to each of the biological phenotype data and the eigenvalue data set to train a phenotype prediction model, and the influence weight of each of the subgroup sequences on the biological phenotype is determined according to the phenotype prediction model; Based on the weights of each impact, the key genes that regulate complex biological traits are determined.
2. The method according to claim 1, characterized in that: The biological reference genome data is obtained from any one or more of the following databases: NCBI database, ORCAE database, GWH database, fernbase database, Phytozome database, Ensembl database or CuGenDB database.
3. The method according to claim 1, characterized in that: The dividing of the protein sequence encoded by the biological reference genome data into orthologous gene groups comprises: According to the protein sequences encoded by the genomes of different organisms, gene sequences are divided into different orthologous gene groups, i.e. subgroups.
4. The method according to claim 1, characterized in that: The step of determining the characteristic value data set according to the subgroup sequence numbers obtained after the division and the number of genes in the subgroup sequences comprises: According to the subgroup numbers obtained by dividing the orthologous gene groups of each organism, a gene number dataset of each subgroup sequence was constructed as a feature value dataset.
5. The method according to claim 1, characterized in that: The machine algorithm model is constructed by a random forest algorithm.
6. A device for identifying key genes of biological complex traits, characterized in that: The device comprises: The first acquisition module is used to acquire the reference genome data of each organism and divide the protein sequence encoded by the reference genome data of each organism into orthologous gene groups; A first determination module is used to determine a feature value data set according to the subgroup sequence numbers and the number of genes in the subgroup sequences obtained after the division; A second acquisition module is used to acquire the phenotypic data of each organism; A processing module, for using a machine algorithm model to construct a phenotype matrix according to each of the biological phenotype data and the eigenvalue data set to train a phenotype prediction model, and determining the influence weight of each of the subgroup sequences on the biological phenotype according to the phenotype prediction model; The second determination module is used to determine the key genes regulating the complex traits of organisms according to the influence weights.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Method, device, apparatus and medium for identifying candidate genes regulating bacterial shape
CN116721695B
Cited By
Trait heritable variation site prediction method fusing transfer learning and interpretable mechanism
CN120895090A
A method for predicting trait genetic variation sites that integrates transfer learning and explainable mechanisms
CN120895090B
Method for identifying biological phenotype key functional genes based on cross-species gene sequence expression characteristics and application
CN122290727A