Method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics and application thereof

By using a cross-species gene sequence expression feature identification method, combined with machine learning models and transcriptome data, the problem of insufficient genome data for species has been solved. This enables the accurate and efficient identification of key functional genes in higher organisms, expanding the scope of application and making it suitable for scenarios involving interactions between plants, animals, and microorganisms.

CN122290727APending Publication Date: 2026-06-26ANHUI AGRICULTURAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-16
Publication Date
2026-06-26

Smart Images

  • Figure CN122290727A_ABST
    Figure CN122290727A_ABST
Patent Text Reader

Abstract

This invention discloses a method and its application for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics. The method includes: obtaining a reference genome of a small number of target organisms and classifying them into orthologous gene groups; obtaining raw transcriptome data of tissues of interest or specific phenotypes for each target species, cleaning and aligning them to the reference genome to obtain a gene expression matrix; constructing a feature value dataset by combining temporary sequence numbers and expression levels for each subgroup of transcriptome samples; obtaining phenotypic feature information for each corresponding sample and constructing a machine learning prediction model; and determining the influence weight of each subgroup of total sequence expression based on the optimal model, thereby locating and obtaining key functional genes for specific phenotypes. This invention achieves a wider range of applications with a small requirement for species genome size, and has promising application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biological high-throughput data analysis technology, and in particular to methods and applications for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics. Background Technology

[0002] There are many methods for identifying key functional genes. Traditional methods are gene identification methods based on single species, including functional gene identification methods based on mutants, functional gene identification methods based on human family genetic maps, and functional gene identification methods based on single species populations. Mutant screening methods involve inducing large-scale population mutagenesis in a species using chemical, physical, or biological mutagenesis techniques, resulting in the random distribution of mutation points. However, not all mutations are valuable for research; some may be ineffective, lethal, or irrelevant to the research objective. This makes identifying and screening meaningful mutants time-consuming and extremely laborious, further increasing the difficulty of discovering and identifying candidate genes. In addition, maintaining a mutant library requires preserving and propagating a large number of biological individuals. In contrast, pedigree-based functional gene identification methods have successfully identified many genetic genes, mainly for the identification of human disease genes, but discovering new functional genes in this field is becoming increasingly difficult. Genome-wide association studies (GWAS) require significant resources, including financial resources, effort, and time, to collect different strains or inbred lines of the same species, or to create recombinant inbred line populations. Subsequently, genome sequencing and resequencing are performed, and finally, candidate key superior alleles are identified through genome-wide association analysis. A single study typically yields only a very small number of key genes (usually 1-2).

[0003] Despite the increasing maturity of traditional functional gene identification techniques, the study of key functional genes for complex traits still faces many challenges. It is time-consuming, inefficient, and requires significant research to locate only a very small number of key genes. To address this problem, we previously developed and applied for a novel method for bacterial functional gene identification based on cross-species genome mapping (Patent No.: ZL202310714738.3, Patent Title: Method, Apparatus, Device, and Medium for Identifying Candidate Genes Regulating Bacterial Shape). This method is used to identify candidate gene composition with specific shared phenotypes, achieving a technological breakthrough. However, this method relies on third-party software tools. The pfam_scan tool is used for protein domain alignment and analysis. However, the method of locating functional genes based on protein domain feature values ​​can lead to the loss of some key genes, making it difficult to predict unknown gene sequences. Although domain composition feature values ​​can efficiently predict functional genes in prokaryotes, the large number of gene family members in higher organisms such as plants significantly increases the difficulty of identifying key functional genes, resulting in less than ideal identification efficiency of the protein domain feature value method in actual testing in higher organisms. Through technical improvements, we have submitted another method that utilizes cross-regional protein domains without relying on third-party tools. The genomic and phenotypic data of species (application number: 202510052860.8, patent title: Method, device, equipment and medium for identifying key genes of complex traits in organisms) enables accurate and efficient identification of key genes of complex traits in prokaryotes and higher organisms. The identified functional genes include those experimentally verified and those with completely unknown sequence functions. A single batch of studies can identify most key genes (generally several to dozens, depending on the complexity of the trait), representing a comprehensive technological advancement in functional gene identification methods. Although the improved method represents a major technological breakthrough in the field of functional gene identification, it has already been measured... The shared information from the genomes and phenotypes of a large number of species is sometimes insufficient to meet the requirements of models. Furthermore, some induced gene expression may not be practically applicable in specific real-world gene mapping scenarios, such as when plants are invaded by pathogenic microorganisms or in symbiotic relationships with probiotics. In such cases, the number of non-redundant higher organism genomes and phenotypes may not meet model requirements (generally, more than 60 species are needed to ensure reliable results for both the genome and the phenotype of interest). Existing functional gene identification methods based on cross-species genomes cannot be widely applied given the current limitations on the number of higher organism genomes. Therefore, there is an urgent need to develop a method that requires only a small number of species (just a few genomes), is accurate and efficient, and comprehensively identifies both known and unknown functional genes. This would overcome the current technical bottlenecks, expand the practical application of functional gene identification, and reduce economic costs. Summary of the Invention

[0004] Therefore, it is necessary to provide a method and application for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics to address the above-mentioned technical problems, so as to achieve a wider range of applications with a small demand for species genomes. This invention is not only a brand-new functional gene identification technology, but also a new technology for analyzing key functional genes in the interaction between animals, plants and microorganisms.

[0005] The present invention achieves the above objectives through the following technical solutions: As a first aspect of the present invention, a method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics is provided, the method comprising the following steps: S1. Obtain biological reference genome data and classify the proteins encoded by the biological reference genome data into orthologous gene groups; S2. Obtain transcriptome data of biological tissues or specific phenotypes, and analyze the expression levels of gene sequences of each biological tissue or specific phenotype. S3. Based on the subgroup numbers obtained after division and the expression levels of gene sequences in each biological tissue or specific phenotype, construct a temporary sequence expression total for each subgroup as a feature value dataset. S4. Obtain phenotypic information for each organism; S5. Train a machine learning prediction model based on the phenotypic information of each organism and the feature value dataset, and determine the influence weight of the total amount of temporary sequence expression of each subgroup sequence on the phenotypic based on the prediction model. S6. Based on the influence weights, identify the key genes that influence an organism's specific phenotype.

[0006] As a further optimization of the present invention, the sources for obtaining the biological reference genome data and the transcriptome data of biological tissues include, but are not limited to, NCBI database, ORCAE database, GWH database, fernbase database, Phytozome database, Ensembl database or CuGenDB database, or a large amount of transcriptome data obtained by oneself.

[0007] As a further optimization of the present invention, the protein sequences encoded by the biological reference genome data are classified into orthologous gene groups, including: Extract protein sequences from the biological reference genome data and remove duplicates from multiple transcripts of the same gene; Ortholog analysis was performed on protein sequences from all species; All genes are divided into different orthologous gene groups, i.e., subgroups; Obtain transcriptome data from biological tissues and analyze the expression levels of gene sequences in each tissue, including: The raw transcriptome data underwent quality control and data cleaning, and the cleaned data was split into two datasets: coding region fragments and rRNA fragments. The two datasets, the coding region fragment and the rRNA fragment, are aligned to the reference genome of their respective organisms, and the expression level of each gene is calculated to obtain the expression level of gene sequences of each biological tissue or specific phenotype.

[0008] As a further optimization of the present invention, the total temporary sequence expression of each subgroup sequence is constructed as a feature value dataset, including: Using the subgroup as the unit, the expression levels of all genes within the same subgroup are summed to construct a subgroup-sample total expression matrix, which serves as the feature value dataset.

[0009] As a further optimization of the present invention, the phenotypic information of each organism is obtained by directly observing the phenotype through experiments or by constructing a plant-microbe interaction phenotypic model.

[0010] As a further optimization of the present invention, a plant-microbe interaction phenotypic model is constructed, specifically as follows: For plant-microbe interaction scenarios, transcriptome data is split into coding region fragment datasets and rRNA fragment datasets; Determine whether there is accompanying target microbial information on the rRNA fragment dataset; The encoded region fragment dataset is aligned to the target microbial genome, and biological phenotypic information is obtained based on preset judgment rules.

[0011] As a further optimization of the present invention, S5-S6 specifically refers to: The phenotypic information of each organism and the feature value dataset are divided into training set and test set. The total expression feature of the subpopulation is used as input and the phenotypic label is used as output. The machine learning prediction model is trained and the parameters are tuned to obtain the optimal prediction model. Based on the optimal prediction model, the contribution weight of each subgroup to the phenotype is calculated, sorted from high to low, features are selected, and the high-weight subgroups are mapped back to the original genes to determine the key functional genes controlling specific phenotypes.

[0012] As a second aspect of the present invention, a system for identifying key functional genes of biological phenotypes is provided, characterized in that the system comprises: The first acquisition module is used to acquire reference genome data of each organism and classify the protein sequences encoded by the reference genome data of each organism into orthologous gene groups. The second acquisition module is used to acquire transcriptome data of biological tissues or specific phenotypes, and to parse and obtain the expression level information of gene sequences of each biological tissue or specific phenotype. The module is used to construct a temporary sequence expression total of each subgroup as a feature value dataset based on the subgroup number obtained after division and the gene sequence expression level of each biological tissue or specific phenotype. The third acquisition module is used to acquire the phenotypic data of each of the organisms; The processing module is used to train a machine learning prediction model based on the phenotypic information and feature value dataset of each organism, and to determine the influence weight of the total amount of temporary sequence expression of each organism on the phenotypic based on the prediction model. The determination module is used to identify key genes that influence an organism's specific phenotype based on the weights of each influence.

[0013] As a third aspect of the invention, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method as described in any of the preceding claims.

[0014] As a fourth aspect of the invention, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method as described in any of the preceding claims.

[0015] This invention categorizes gene sequences encoded by the genomes of organisms of interest into orthologous genes. Simultaneously, it combines gene expression level information to construct an expression level dataset for each gene classification sequence as a feature value dataset. Combined with target phenotypic features, a machine learning prediction model is built to assess the impact of expression levels of each subgroup sequence on the target phenotype. This invention provides a method that does not rely on third-party annotation tools. By introducing total gene sequence expression features, it effectively compensates for the information gaps caused by insufficient genome size in a species. It enables rapid and accurate identification of key functional genes for specific phenotypes in specific organisms with only a small number of genomes (e.g., 4 species), reducing the required genome size by more than 10 times compared to traditional methods. This invention classifies gene expression sequences across species transcriptomes and constructs phenotypic prediction models using expression level matrices. This allows for the identification of key functional genes for specific traits. It enables the identification of key functional sequences (genes) under most conditions with only a small number (or a few) of genomes, achieving the goal of efficiently and accurately identifying key genes that function under specific conditions using only a small number of species. This invention is not only a novel functional gene identification technology but also a new technique for analyzing key functional genes in interactions between plants, animals, and microorganisms, with broad prospects for practical application. Attached Figure Description

[0016] Figure 1 This is a technical roadmap of a method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics in one embodiment of the present invention; Figure 2 The following is a performance comparison of two models (Transcriptome and Genome) provided in one embodiment of the present invention; Note: Figure a shows the comparison of the training set, and Figure b shows the comparison of the test set. Figure 3 This is a structural block diagram of a system for identifying key functional genes of biological phenotypes in one embodiment of the present invention; Figure 4 This is an internal structural diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0017] The present application will now be described in further detail with reference to the accompanying drawings. It should be noted that the following specific embodiments are only used to further illustrate the present application and should not be construed as limiting the scope of protection of the present application. Those skilled in the art can make some non-essential improvements and adjustments to the present application based on the above application content.

[0018] In one embodiment, taking a key gene in plant-AM fungal symbiosis as an example, this embodiment provides a method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics. The technical solution roadmap is as follows: Figure 1 As shown, the steps include: (1) Obtain plant reference genome data from public databases, and classify the protein sequences encoded by each plant reference genome into orthologous gene groups, specifically: For the coding genes of various plant species, a self-developed program was used to remove duplicates from multiple transcripts of each coding gene in each plant species, so that only one transcript is retained for each coding gene, which is more accurate (multiple transcripts exist in the CDS sequence, which can also be left unremoved with similar efficiency), and the transcripts are translated into proteins or the corresponding proteins are directly extracted. There are many types of orthologous gene classification software. In this implementation, the open-source software Orthofinder is used for orthologous gene sequence analysis. Based on the protein sequences encoded by the genomes of different plants, the gene sequences are divided into different orthologous gene groups. The protein sequence files of each species are placed in the Proteins_Data folder according to the extension required by the open-source software Orthofinder, using the default parameter: "OrthoFinder / orthofinder -fOrthoFinder / Proteins_Data". Other orthologous software can also replace Orthofinder to obtain similar results. (2) Obtain the transcriptome of plant roots from public databases and analyze it to obtain gene expression data. Specifically: Raw transcriptome data of plant roots is obtained from public databases. After cleaning the raw data using open-source software FastP or other cleaning software, Sortmerna is used to split the transcriptome data into coding region fragment datasets and rRNA fragment datasets. The cleaned data or the coding region data sequences obtained by splitting using Sortmerna software are directly aligned to the reference genome of their respective plants to obtain gene expression information of each plant tissue sample. (3) Based on the division and numbering of plant orthologous gene sequences and the expression levels of plant tissue gene sequences, the expression levels of gene sequences belonging to the same orthologous sequence temporary number are summed to construct a matrix of total expression levels of each temporary sequence as a feature value dataset. Referring to Table 1, which is a table format of the data structure in the feature value dataset mentioned above in Example 1; Table 1. Data structure table format of the feature value dataset. ; Note: Expression levels are measured using TPM or FPKM.

[0019] As shown in Table 1 above, the feature value dataset can include the numbering of each gene sequence group after classification of each plant, and the total gene expression of the temporary numbered sequence in each plant tissue after classification of each gene sequence group. (4) Phenotypic matrix construction. Phenotyps can be obtained through direct experimental observations or through data analysis methods. In this example 1, the method involves mining the microbial sequences to be studied contained in the plant transcriptome data (due to the large number of false positives in the Kraken2 software alignment results, many results are not very reliable and need to be verified in conjunction with the genome of the microorganism to be studied). The rRNA fragment dataset is mined using Kraken2 software to obtain the accompanying microbial information. At the same time, the datasets of each coding region fragment are aligned to the target mycorrhizal fungi of interest in Example 1 using Hisat2. Glomus intraradices On the reference genome, a comprehensive assessment of whether a symbiotic relationship has formed between the plant and AM fungi is made based on a combination of microbial information and mycorrhizal fungi alignment rate information. Using microbial analysis results and transcriptome alignment to the corresponding microbial reference genomes (mapping results), transcriptomes with a hiast2 alignment rate of 0 to the mycorrhizal fungi genome are identified as non-symbiotic tissues, those with an alignment rate greater than 0.5% are considered symbiotic tissues, and samples with indeterminate alignment results are discarded. Symbiotic tissues are represented by 1, and non-symbiotic tissues by 0. A phenotypic matrix of symbiotic relationships is constructed using the alignment results. (5) Create an experimental design table design.txt. First, divide the plant mycorrhizal symbiosis data into a training set (group1) and a test set (group2). Adjust the model ratio to achieve the best prediction model. Please refer to Table 2 below, which is an example of the grouping list in Example 1.

[0020] Table 2 Grouping list of embodiments ; Specifically, this application can divide the plants into two groups. In one possible design, the two types of symbiotic relationships can be divided into a training group and a test group in a 3:1 ratio, i.e., 75% as the training group and 25% as the test group. It should be noted that when grouping, efforts should be made to ensure that the training group, after grouping, covers both phenotypes mentioned above, both symbiotic and non-symbiotic, and that the test group, after grouping, also covers both phenotypes mentioned above, both symbiotic and non-symbiotic.

[0021] The symbiosis between plants and AM fungi depends on the coordinated action of several key genes, which are involved in important biological processes such as signal transduction, symbiotic structure development, and nutrient exchange. During the signal transduction stage... SYMRK and DMI Gene families are involved in the recognition and transmission of mycorrhizal signals; in the formation of symbiotic structures, RAM1 and EXO70I Genes such as PT4 and STR / STR2 play key roles in nutrient exchange, regulating dendritic development. Furthermore, CCaMK and IPD3, as signaling regulators, are widely involved in signal transduction and regulation. These known genes interact to promote the establishment of plant-AM fungal symbionts.

[0022] In Example 1 of this application, we conducted a systematic comparative analysis of the performance of different prediction methods in key gene identification. Existing models rely solely on large-scale pure genomes and their shared phenotypic information (hereinafter referred to as the Orthofinder method). In the patent example of this application, modeling is required using 248 genome sample data from 52 species (application number: 202510052860.8, patent name: method, device, equipment and medium for identifying key genes of complex biological traits) (Table 3).

[0023] The Orthofinder-based transcriptome-based method (transcriptome model) proposed in this application exhibits low sample dependence (hereinafter referred to as the transcriptome-based method), requiring only 4 representative species. Daucus carota , Solanum tuberosum , Vitis vinifera , Zea maysModeling can be completed using the transcriptome expression data of 122 transcriptomes under different symbiotic states (Table 4). Compared with traditional methods, the transcriptome-integrated method in this application reduces the number of species required by more than 10 times (Table 5).

[0024] Table 3. Statistics on plant species symbiosis and data sources using the Orthofinder method. ; ; ; ; ; ; Note: 0, non-symbiotic; 1, symbiotic; Genome is a modeling method based on large-scale pure genomes and their shared phenotypic information, proposed in patent application number 202510052860.8, namely the Orthofinder method.

[0025] Table 4. Sources of genome data for representative plant species in this invention. ; Note: 1. It can form a symbiotic relationship with AM fungi; Transcriptome refers to the Orthofinder binding transcriptome method proposed in this patent application, i.e., the binding transcriptome method.

[0026] Table 5. Comparison of species counts between transcriptomics and Orthofinder methods. ; The performance metrics of the two construction models were evaluated, and the results are as follows: Figure 2 As shown, the transcriptome-based method significantly outperforms the Orthofinder method across all five performance metrics in the training set. In cross-sectional evaluations on the test set, the proposed method demonstrates superior generalization ability, leading in all performance metrics except Recall. This result proves that by introducing gene sequence expression totality features, information gaps caused by insufficient genome size in a species can be compensated for, thereby achieving accurate prediction of key genes for complex phenotypes. Further annotation analysis of the predicted high-importance feature gene sequences revealed that among the top 50 features predicted using the transcriptome-based method, receptor genes for plant recognition of mycorrhizal fungal signaling molecules were successfully identified. OsAM3 The study identified 30 known symbiotic-related genes, including the LysM domain; in contrast, the existing Orthofinder method identified 25 known genes and 5 others in the first 50 features. AP2 A total of 30 genes, including downstream genes (see Tables 6 and 7), were obtained using both methods, and the number of genes that have been experimentally verified is consistent.

[0027] Table 6. Statistics on the co-occurrence of species genome and transcriptome files in grouped files of Example 1. ; Table 7. Statistics on the coexistence of species genome files in grouped files of patent application number 202510052860.8 ; ; ; Table 8 shows the statistical comparison of the results of the alignment of the top 50 important features of Example 1 with known symbiotic genes and the results of the alignment with patent application number 202510052860.8.

[0028] Table 8. Comparison of the top 50 most important features of Example 1 with known symbiotic genes and results with patent application number 202510052860.8. ; The analysis of the top 50 important characteristic sequences of Example 1 and the comparison with the results of patent application number 202510052860.8 are shown in Table 9.

[0029] Table 9. Analysis of the top 50 most important characteristic sequences of known co-occurring genes in Example 1 and comparison with the results of patent application number 202510052860.8 ; ; ; ; ; Note: Bold underlined genes represent reported key symbiotic genes.

[0030] The results of Example 1 demonstrate that this invention solves the problem of the limited application of existing efficient functional gene identification methods due to the scarcity (or insufficiency) of reference genome resources and phenotypic matching in higher organisms, which necessitates an insufficient number of genomes required by previous machine learning models. Because genes exhibit tissue-specific and spatiotemporally specific expression, the previously applied Orthofinder method did not consider this gene expression characteristic, potentially leading to higher false positives in certain situations. This application introduces expression levels into a novel method for identifying functional genes using genomic and phenotypic characteristics of different species, achieving almost the same accuracy in identifying functional genes while significantly reducing the number of species. Given the extremely low cost of eukaryotic transcriptome sequencing (currently around 130-150 RMB / sample), the cost of sequencing over 100 plant and animal transcriptome samples is less than 20,000 RMB, reaching a level affordable for ordinary laboratories. This greatly expands the application scope of the original patented method and has strong application scenarios.

[0031] In one embodiment, such as Figure 3 As shown, a system for identifying key functional genes of biological phenotypes is provided. The system includes: The first acquisition module 11 is used to acquire reference genome data of each organism and classify the protein sequences encoded by the reference genome data of each organism into orthologous gene groups. The second acquisition module 12 is used to acquire transcriptome data of biological tissues or specific phenotypes, and to parse and obtain the expression level information of gene sequences of each biological tissue or specific phenotype. Module 13 is used to construct a temporary sequence expression total of each subgroup sequence as a feature value dataset based on the subgroup number obtained after division and the gene sequence expression level of each biological tissue or specific phenotype. The third acquisition module 14 is used to acquire the phenotypic data of each of the organisms; Processing module 15 is used to train a machine learning prediction model based on the phenotypic information and feature value dataset of each organism, and to determine the influence weight of the total amount of temporary sequence expression of each organism on the phenotypic based on the prediction model. Module 16 is used to determine the key genes that influence an organism's specific phenotype based on the influence weights of each factor.

[0032] Specific limitations regarding the identification system for key functional genes of biological phenotypes can be found in the limitations on identification methods for key functional genes of biological phenotypes mentioned above, and will not be repeated here. Each module in the aforementioned identification system for key functional genes of biological phenotypes can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0033] In one embodiment, a computer device is provided, which may be a terminal or a server, and its internal structure diagram may be as follows. Figure 4 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a method for identifying key functional genes of biological phenotypes.

[0034] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for identifying key functional genes of biological phenotypes in any of the above embodiments.

[0035] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for identifying key functional genes of biological phenotypes in any of the above embodiments.

[0036] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0037] The above-described embodiments are merely one implementation of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. The new method is also applicable to other species, such as animals, and lower organisms such as fungi, protozoa, and bacteria. It should be noted that those skilled in the art can make various modifications, software substitutions for individual similar functions within the steps, algorithm replacements, and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention.

Claims

1. A method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics, characterized in that, The method includes the following steps: S1. Obtain biological reference genome data and classify the proteins encoded by the biological reference genome data into orthologous gene groups; S2. Obtain transcriptome data of biological tissues or specific phenotypes, and analyze the expression levels of gene sequences of each biological tissue or specific phenotype. S3. Based on the subgroup numbers obtained after division and the expression levels of gene sequences in each biological tissue or specific phenotype, construct a temporary sequence expression total for each subgroup as a feature value dataset. S4. Obtain phenotypic information for each organism; S5. Train a machine learning prediction model based on the phenotypic information of each organism and the feature value dataset, and determine the influence weight of the total amount of temporary sequence expression of each subgroup sequence on the phenotypic based on the prediction model. S6. Based on the influence weights, identify the key genes that influence an organism's specific phenotype.

2. The method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics according to claim 1, characterized in that, The sources for obtaining the biological reference genome data and the transcriptome data of biological tissues or specific phenotypes include, but are not limited to, the NCBI database, ORCAE database, GWH database, fernbase database, Phytozome database, Ensembl database, or CuGenDB database.

3. The method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics according to claim 1, characterized in that, The protein sequences encoded by the biological reference genome data are classified into orthologous gene groups, including: Extract protein sequences from the biological reference genome data and remove duplicates from multiple transcripts of the same gene; Ortholog analysis was performed on protein sequences from all species; All genes are divided into different orthologous gene groups, i.e., subgroups; Obtain transcriptome data from biological tissues and analyze the expression levels of gene sequences in each tissue, including: The raw transcriptome data underwent quality control and data cleaning, and the cleaned data was split into two datasets: coding region fragments and rRNA fragments. The two datasets, the coding region fragment and the rRNA fragment, are aligned to the reference genome of their respective organisms, and the expression level of each gene is calculated to obtain the expression level of gene sequences of each biological tissue or specific phenotype.

4. The method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics according to claim 1, characterized in that, Construct a temporary sequence expression total for each subgroup of sequences as a feature value dataset, including: Using the subgroup as the unit, the expression levels of all genes within the same subgroup are summed to construct a subgroup-sample total expression matrix, which serves as the feature value dataset.

5. The method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics according to claim 1, characterized in that, The phenotypic information of each organism was obtained through direct experimental observation of phenotypes or by constructing plant-microbe interaction phenotypic models.

6. The method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics according to claim 5, characterized in that, Constructing a plant-microbe interaction phenotypic model, specifically: For plant-microbe interaction scenarios, transcriptome data is split into coding region fragment datasets and rRNA fragment datasets; Determine whether there is accompanying target microbial information on the rRNA fragment dataset; The encoded region fragment dataset is aligned to the target microbial genome, and biological phenotypic information is obtained based on preset judgment rules.

7. The method for identifying key functional genes of biological phenotypes based on cross-species gene sequence expression characteristics according to claim 1, characterized in that, Specifically, S5-S6 are: The phenotypic information of each organism and the feature value dataset are divided into training set and test set. The total expression feature of the subpopulation is used as input and the phenotypic label is used as output. The machine learning prediction model is trained and the parameters are tuned to obtain the optimal prediction model. Based on the optimal prediction model, the contribution weight of each subgroup to the phenotype is calculated, sorted from high to low, features are selected, and the high-weight subgroups are mapped back to the original genes to determine the key functional genes controlling specific phenotypes.

8. A system for identifying key functional genes in biological phenotypes, characterized in that, The system includes: The first acquisition module is used to acquire reference genome data of each organism and classify the protein sequences encoded by the reference genome data of each organism into orthologous gene groups. The second acquisition module is used to acquire transcriptome data of biological tissues or specific phenotypes, and to parse and obtain the expression level information of gene sequences of each biological tissue or specific phenotype. The module is used to construct a temporary sequence expression total of each subgroup as a feature value dataset based on the subgroup number obtained after division and the gene sequence expression level of each biological tissue or specific phenotype. The third acquisition module is used to acquire the phenotypic data of each of the organisms; The processing module is used to train a machine learning prediction model based on the phenotypic information and feature value dataset of each organism, and to determine the influence weight of the total amount of temporary sequence expression of each organism on the phenotypic based on the prediction model. The determination module is used to identify key genes that influence an organism's specific phenotype based on the weights of each influence.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes a computer program, it implements the steps of the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method, device, apparatus and medium for identifying candidate genes regulating bacterial shape

    CN116721695B

  • Method, device, equipment and medium for identifying key genes of biological complex characters

    CN120089193A