A method, apparatus and storage medium for inferring a time of origin of a gene

By statistically analyzing the distribution of orthologous genes in gene origin time inference and optimizing the judgment of positive groups using evolutionary timelines and ancestral group weights, the problem of large errors in gene origin time inference in existing technologies has been solved, achieving more accurate gene origin time inference.

CN117612604BActive Publication Date: 2026-07-21SHENZHEN RES INST THE CHINESE UNIV OF HONG KONG
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN RES INST THE CHINESE UNIV OF HONG KONG
Filing Date
2023-03-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies suffer from significant errors when inferring the time of gene origin due to the scarcity of reference species or sequencing errors in genome fragments, making it difficult to improve the accuracy of gene origin time inference.

Method used

By statistically analyzing the distribution of orthologous genes in a reference species, using the evolutionary timeline and ancestral group weights, setting a parameter p to identify positive groups, calculating confidence scores and continuity scores, optimizing the optimal evolutionary path, correcting low-weight groups, and selecting the evolutionary path with the highest confidence score as the optimal evolutionary path, the origin time of genes can be inferred.

Benefits of technology

It improves the accuracy of gene origin time estimation, avoids errors caused by gene variation in individual species, and achieves accurate and rapid gene origin time estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117612604B_ABST
    Figure CN117612604B_ABST
Patent Text Reader

Abstract

The application discloses a method, device and storage medium for inferring gene origin time. The method for inferring gene origin time comprises the following steps: a reference species statistics step, an evolutionary time axis acquisition step, an ancestor taxon weight setting step, an optimal evolutionary route analysis step and a gene origin time inference step. The gene origin time inference method provided by the application utilizes the characteristic that genes basically present continuity in an evolutionary time table. If a gene appears at an outlier time point in the time table, the time point is considered to be a false positive caused by technical error, and the judgment parameter of a positive taxon is optimized. From a possible evolutionary route formed by a subset of all positive taxons obtained from the optimal parameter, the optimal evolutionary route is screened by using a confidence score. Then, the origin time of a gene to be analyzed is accurately and rapidly inferred, and the error in gene origin time inference caused by the variation of the gene in individual species is avoided, and the accuracy of gene origin time inference is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gene origin analysis technology, and in particular to a method, apparatus and storage medium for inferring the time of gene origin. Background Technology

[0002] As bioinformatics methods are increasingly applied to research on biological evolution, more and more people are exploring the evolutionary process of genes by inferring their origin time, thereby sorting out the evolutionary history of gene function, and ultimately predicting the development direction of genes from different perspectives, and looking for mechanisms that are unfavorable to evolution and may lead to disease.

[0003] Traditional methods for inferring gene origin or formation time rely on the similarity of gene sequences or functions, searching for possible orthologous relationships in only a few other species. When orthologous relationships exist, it proves the gene's presence in a common ancestral group shared by the studied and reference species, thus demonstrating the gene's existence within that ancestral group. However, this method is only inefficient in inferring homologous relationships. Furthermore, an unavoidable limitation of this method for inferring gene origin time is that if the number of reference species is small, or if sequencing errors exist in the genome segments of some reference species, the inferred gene origin time will be questionable.

[0004] Therefore, how to reduce the errors caused by gene variation in individual species when estimating the time of gene origin and improve the accuracy of gene origin inference remains an important problem that needs to be solved in the field of gene origin research technology. Summary of the Invention

[0005] The purpose of this application is to provide a new method, apparatus, and storage medium for inferring the time of origin of genes.

[0006] To achieve the above objectives, this application adopts the following technical solution:

[0007] The first aspect of this application discloses a method for inferring the time of gene origin, comprising the following steps:

[0008] The reference species statistical steps include defining the species containing the gene to be analyzed as the study species, defining genes from other species that share the same ancestor as the gene to be analyzed as orthologous genes, using all species that are orthologous to the study species as reference species, and statistically analyzing the distribution of the gene to be analyzed and its orthologous genes in the reference species.

[0009] The steps for obtaining the evolutionary timeline include using species groups to depict the major events in the development of life on Earth to obtain the evolutionary timeline; among them, the taxa in the major events are the ancestral taxa;

[0010] The steps for setting ancestral group weights include classifying the reference species into the nearest ancestral group of the study species based on the distance between the ancestral groups of the reference species and the study species in the phylogenetic tree; deleting an ancestral group if no reference species are classified into it; and assigning weights to ancestral groups based on the number of reference species classified in the ancestral group, with higher weights for more reference species.

[0011] The optimal evolutionary path analysis steps include: a) setting a parameter p, and counting the number of reference species in each ancestral group containing the gene to be analyzed or its orthologous gene. If the percentage of this number to the total number of all reference species in that ancestral group exceeds p%, then the ancestral group is classified as a positive group; otherwise, it is classified as a negative group; b) constructing a preliminary evolutionary path of the gene to be analyzed in the studied species from the positive groups, traversing all possible evolutionary paths composed of subsets of all positive groups, calculating the confidence score of all possible evolutionary paths according to weights, and selecting the gene with the highest confidence score for each gene when using parameter p. c) Calculate the continuity score of the evolutionary path chosen by each gene when using parameter p, and sum the continuity scores of the evolutionary paths of all genes to obtain the overall continuity score of the evolutionary path obtained using parameter p; d) Adjust parameter p, calculate the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Using the optimal parameter, calculate the confidence score of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and take the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed.

[0012] The gene origin time inference step includes taking the age of the oldest ancestral group in the optimal evolutionary path as the origin time of the gene to be analyzed.

[0013] The gene origin time inference method of this application utilizes the characteristic that genes exhibit continuity in the evolutionary timeline to optimize the judgment parameters of positive groups; and from the possible evolutionary routes composed of subsets of all positive groups obtained from the optimal parameters, the optimal evolutionary route is selected by confidence score; thereby accurately and quickly inferring the origin time of the gene to be analyzed, avoiding the error in gene origin time inference caused by gene variation in individual species, and improving the accuracy of gene origin time inference.

[0014] In one implementation of this application, step b) of the optimal evolutionary path analysis step includes:

[0015] Data preparation involves extracting all possible subsets from all ancestral groups that are judged to be positive. Each ancestral group subset constitutes a sub-distribution of the gene to be analyzed in the evolutionary lineage, i.e., a sub-distribution of the group.

[0016] Extract and select a possible sub-distribution of the taxonomic group to form a possible evolutionary path;

[0017] Correction: If the low-weight class in the sub-distribution of the class is a negative class, and the two ancestral classes before and after this negative class in the evolutionary timeline are both positive classes, then correct this negative class to a positive class.

[0018] Calculate the confidence score of the corrected sub-distribution of the class group based on the weights, traverse the possible evolutionary paths consisting of subsets of all positive classes, calculate the confidence score of all possible evolutionary paths, and select the evolutionary path with the highest confidence score.

[0019] In one implementation of this application, the low-weight class group refers to the ancestor class group with a weight of less than 5.

[0020] In one implementation of this application, the weight of the ancestral group is determined by sorting the reference species classified within the ancestral group and using the sorted sequence number as the weight of the ancestral group.

[0021] It should be noted that in the improved scheme of this application, by correcting for the situation where the reference species in the low-weight group, that is, the ancestral group, is too scarce, the false negative problem caused by too few reference species is avoided, and the accuracy of gene origin time inference is further improved.

[0022] In one implementation of this application, the confidence score is calculated using Formula 1 based on the weights.

[0023] In one implementation of this application, the method for calculating the continuity score in step c) of the optimal evolutionary path analysis step includes assigning sorting numbers to the research species and its ancestral groups according to the chronological order of the evolutionary timeline, and calculating the continuity score of the positive groups using Formula 2.

[0024] The second aspect of this application discloses a device for inferring the time of gene origin, including a reference species statistics module, an evolutionary timeline acquisition module, an ancestral group weight setting module, an optimal evolutionary route analysis module, and a gene origin time inference module;

[0025] The reference species statistics module is used to define the species containing the gene to be analyzed as the study species, define genes from other species that share the same ancestor as the gene to be analyzed as orthologous genes, and use all species that are orthologous to the study species as reference species to count the distribution of the gene to be analyzed and its orthologous genes in the reference species.

[0026] The evolutionary timeline acquisition module is used to depict the major events in the life development of species on Earth using species groups to obtain an evolutionary timeline, where the taxa in the major events are the ancestral taxa.

[0027] The ancestral group weight setting module is used to classify the reference species into the nearest ancestral group of the study species based on the distance between the ancestral groups of the reference species and the ancestral groups of the study species in the evolutionary tree; if no reference species are classified into a certain ancestral group, the ancestral group is deleted; the weight of the ancestral group is set according to the number of reference species classified in the ancestral group, and the more reference species, the higher the weight.

[0028] The optimal evolutionary path analysis module is used to: a) set a parameter p, and, on a per-ancestral-group basis, count the number of reference species in each ancestral-group containing the gene to be analyzed or its orthologous gene. If this number accounts for more than p percent of the total number of reference species in that ancestral-group, then that ancestral-group is classified as a positive group; otherwise, it is classified as a negative group; b) construct a preliminary evolutionary path for the gene to be analyzed in the studied species from the positive groups, traverse all possible evolutionary paths composed of subsets of all positive groups, calculate the confidence score of all possible evolutionary paths according to weights, and, when using parameter p, select a confidence score for each gene. c) Calculate the continuity score of the evolutionary path chosen by each gene when using parameter p, and sum the continuity scores of all genes to obtain the overall continuity score of the evolutionary path obtained using parameter p; d) Adjust parameter p, calculate the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Using the optimal parameter, calculate the confidence score of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and take the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed.

[0029] The gene origin time inference module is used to determine the origin time of the gene to be analyzed by taking the age of the oldest ancestral group in the optimal evolutionary path as the origin time of the gene.

[0030] It should be noted that the device for inferring the time of gene origin in this application is actually implemented by various modules to realize the various steps of the method for inferring the time of gene origin in this application. Therefore, the specific implementation method or parameter conditions of each module in the device of this application can refer to the method of this application. For example, the calculation method and formula of confidence score, the calculation method and formula of continuity score, etc. can all refer to the method for inferring the time of gene origin in this application, and will not be elaborated here.

[0031] A third aspect of this application discloses an apparatus for inferring the time of gene origin, the apparatus comprising a memory and a processor; the memory for storing a program; and the processor for implementing the method of inferring the time of gene origin of this application by executing the program stored in the memory.

[0032] The fourth aspect of this application discloses a computer-readable storage medium including a program that can be executed by a processor to implement the method of inferring the time of gene origin of this application.

[0033] Due to the adoption of the above technical solutions, the beneficial effects of this application are as follows:

[0034] The method and apparatus for inferring the origin time of genes in this application can accurately and quickly infer the origin time of the gene to be analyzed through optimal parameter and optimal evolutionary route analysis, avoiding the error in inferring the origin time caused by gene variation in individual species, and improving the accuracy of origin time inference. Attached Figure Description

[0035] Figure 1 This is a flowchart of the method for inferring the time of gene origin in the embodiments of this application;

[0036] Figure 2 This is a structural block diagram of the device for inferring the time of gene origin in the embodiments of this application;

[0037] Figure 3 This is a schematic diagram illustrating the classification of reference species in the embodiments of this application;

[0038] Figure 4 This is a schematic diagram of the positive and negative class groups in the embodiments of this application;

[0039] Figure 5 This is a schematic diagram of the evolutionary path of positive groups in the embodiments of this application. Detailed Implementation

[0040] The present application will now be described in further detail with reference to specific embodiments and accompanying drawings. In the following embodiments, many details are described to facilitate a better understanding of the present application. However, those skilled in the art will readily recognize that some features may be omitted in different situations, or may be replaced by other elements, materials, or methods. In some cases, certain operations related to the present application are not shown or described in the specification to avoid obscuring the core parts of the application with excessive description. For those skilled in the art, detailed description of these related operations is not necessary; they can fully understand the related operations based on the description in the specification and general technical knowledge in the art.

[0041] With the reduction in sequencing costs, the genomes of many new species have been included in ENSEMBL, increasing the number of species from 26 to 180. Furthermore, the ENSEMBL database provides meticulously calculated orthologous relationships between genes, making it possible to improve the accuracy of gene origin time estimation through multi-species cross-validation. This application reasonably utilizes a large number of existing reference species to cross-validate orthologous relationships to reduce errors. Simultaneously, considering the common continuity of genes in ancestral lineages and the interference of paralogous genes on orthologous relationship estimation, a method for estimating gene origin time that comprehensively considers various factors is developed.

[0042] The method for inferring the time of gene origin in this application, such as Figure 1 As shown, the steps include: reference species statistics step 11, evolutionary timeline acquisition step 12, ancestral group weight setting step 13, optimal evolutionary route analysis step 14, and gene origin time inference step 15.

[0043] The reference species statistical step 11 includes defining the species containing the gene to be analyzed as the study species, defining genes from other species that share a common ancestor with the gene to be analyzed as orthologous genes, and using all species that are orthologous to the study species as reference species. The distribution of the gene to be analyzed and its orthologous genes in the reference species is then statistically analyzed. For example, if a gene G has its orthologous gene g in a reference species, it proves that gene G exists in both the study species and the reference species. Based on this rule, the distribution of the gene across all reference species is then examined.

[0044] Step 12, obtaining the evolutionary timeline, involves using species groups to depict the major events in the development of life on Earth, thus obtaining the evolutionary timeline, or evolutionary lineage. The groups involved in these major events are the ancestral groups. For example, the human evolutionary lineage consists of different groups, including apes, primates, mammals, eumezoans, and cellular organisms; these groups are the ancestral groups of humans. A common ancestral group for a reference species can be found in the human evolutionary lineage; for example, both humans and monkeys belong to the primate order.

[0045] Step 13, which sets ancestral group weights, includes classifying the reference species into the closest ancestral group of the study species based on the proximity of their ancestral groups in the phylogenetic tree. Figure 3 As shown; if no reference species are classified into a certain ancestral group, the ancestral group is deleted; based on the number of reference species classified into the ancestral group, a weight is assigned to the ancestral group, and the more reference species, the higher the weight.

[0046] Step 14 of the optimal evolutionary path analysis includes: a) setting a parameter p, and counting the number of reference species in each ancestral group that contain the gene to be analyzed or its orthologous gene. If the percentage of this number to the total number of all reference species in the ancestral group exceeds p%, the ancestral group is judged as a positive group; otherwise, it is a negative group; b) constructing a preliminary evolutionary path of the gene to be analyzed in the studied species from the positive groups, traversing the possible evolutionary paths composed of subsets of all positive groups, calculating the confidence score of all possible evolutionary paths according to the weights, and, when using parameter p, selecting the gene that makes the target gene more likely to be analyzed. c) Calculate the continuity score of the evolutionary path chosen by each gene when using parameter p, and sum the continuity scores of all genes to obtain the overall continuity score of the evolutionary path obtained using parameter p; d) Adjust parameter p, calculate the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Using the optimal parameter, calculate the confidence scores of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and take the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed.

[0047] Step b) specifically includes:

[0048] Data preparation involves extracting all possible subsets from all ancestral groups that are judged to be positive. Each ancestral group subset constitutes a sub-distribution of the gene to be analyzed in the evolutionary lineage, i.e., a sub-distribution of the group.

[0049] Extract and select a possible sub-distribution of the taxonomic group to form a possible evolutionary path;

[0050] Correction: If the low-weight class in the sub-distribution of the class is a negative class, and the two ancestral classes before and after this negative class in the evolutionary timeline are both positive classes, then correct this negative class to a positive class.

[0051] Calculate the confidence score of the corrected sub-distribution of the class group based on the weights, traverse the possible evolutionary paths consisting of subsets of all positive classes, calculate the confidence score of all possible evolutionary paths, and select the evolutionary path with the highest confidence score; specifically, use Formula 1 to calculate the confidence score.

[0052] Formula 1

[0053] In Formula 1, ConfidenceScore is the confidence score, x ij y is the weight of the j-th ancestral group in the i-th consecutive positive group. pq It is the weight of the q-th ancestral group of the p-th consecutive negative group, l i2 The number of ancestral groups that constitute the i-th contiguous group set is the square of l. p 2 This is the square of the number of ancestral groups that make up the p-th consecutive group set. For example... Figure 4 As shown, on the evolutionary timeline, white dots represent positive groups, and black dots represent negative groups. Consecutive positive or negative groups constitute set X. i Or Y p .

[0054] The specific method for calculating the continuity score in step c) includes assigning sorting numbers to the research species and their ancestral groups according to the chronological order of the evolutionary timeline, and calculating the continuity score of the positive groups using Formula 2.

[0055] Formula 2

[0056] In Formula 2, y i Let be the sorting index of the i-th positive group, and n be the number of positive groups.

[0057] Those skilled in the art will understand that all or part of the functions of the methods described above can be implemented in hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. Alternatively, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or portable hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.

[0058] Therefore, based on the method for inferring the time of gene origin in this application, this application proposes a device for inferring the time of gene origin, such as... Figure 2 As shown, it includes a reference species statistics module 21, an evolutionary timeline acquisition module 22, an ancestral group weight setting module 23, an optimal evolutionary route analysis module 24, and a gene origin time inference module 25.

[0059] The reference species statistics module 21 is used to define the species containing the gene to be analyzed as the study species, define genes from other species that share the same ancestor as the gene to be analyzed as orthologous genes, and use all species that have orthologous relationships with the study species as reference species to count the distribution of the gene to be analyzed and its orthologous genes in the reference species.

[0060] The evolutionary timeline acquisition module 22 is used to depict the major events in the life development process of the species on Earth using species groups to obtain the evolutionary timeline, where the taxa in the major events are the ancestral taxa.

[0061] The ancestral group weight setting module 23 is used to classify the reference species into the nearest ancestral group of the research species based on the distance between the ancestral groups of the reference species and the ancestral groups of the research species in the evolutionary tree; if no reference species are classified into a certain ancestral group, the ancestral group is deleted; the weight of the ancestral group is set according to the number of reference species classified in the ancestral group, and the more reference species there are, the higher the weight.

[0062] The optimal evolutionary path analysis module 24 is used to: a) set a parameter p, and, taking ancestral groups as units, count the number of reference species in each ancestral group that contain the gene to be analyzed or its orthologous gene. If the percentage of this number to the total number of all reference species in that ancestral group exceeds p percent, then the ancestral group is judged as a positive group; otherwise, it is a negative group; b) construct the preliminary evolutionary path of the gene to be analyzed in the studied species from the positive groups, traverse the possible evolutionary paths composed of subsets of all positive groups, calculate the confidence score of all possible evolutionary paths according to the weights, and, when using parameter p, select a weighted path for each gene. c) Calculate the continuity score of the evolutionary path chosen by each gene when using parameter p, and sum the continuity scores of all genes to obtain the overall continuity score of the evolutionary path obtained using parameter p; d) Adjust parameter p, calculate the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Using the optimal parameter, calculate the confidence scores of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and take the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed.

[0063] The gene origin time inference module 25 is used to determine the origin time of the gene to be analyzed by taking the age of the oldest ancestral group in the optimal evolutionary path as the origin time of the gene.

[0064] Another implementation of this application provides an apparatus for inferring the origin time of a gene, the apparatus including a memory and a processor; the memory includes a program for storing data; the processor includes a method for executing the program stored in the memory to implement the following steps: a reference species statistical step, including defining the species containing the gene to be analyzed as the study species, defining genes from other species that share a common ancestor with the gene to be analyzed as orthologous genes, using all species that are orthologous to the study species as reference species, and statistically analyzing the distribution of the gene to be analyzed and its orthologous genes in the reference species; and an evolutionary timeline acquisition step, including using species population delineation... The study investigates the major events in the evolution of life on Earth to obtain an evolutionary timeline; the taxa involved in these major events are the ancestral taxa; the steps for setting the weights of ancestral taxa include classifying reference species into the nearest ancestral taxa of the study species based on the proximity of the reference species to the ancestral taxa of the study species in the evolutionary tree; if no reference species are classified into a certain ancestral taxa, that ancestral taxa is deleted; the weights of ancestral taxa are assigned based on the number of reference species classified within them, with higher weights for larger numbers of reference species; the steps for analyzing the optimal evolutionary path include a) setting a parameter p, and statistically analyzing the ancestral taxa in each ancestral taxa. a) The number of reference species containing the gene to be analyzed or its orthologous genes. If this number accounts for more than p percent of the total number of all reference species in the ancestral group, the ancestral group is classified as a positive group; otherwise, it is classified as a negative group. b) The preliminary evolutionary path of the gene to be analyzed in the studied species is constructed from positive groups. All possible evolutionary paths composed of subsets of all positive groups are traversed, and the confidence scores of all possible evolutionary paths are calculated according to their weights. When using parameter p, the evolutionary path of the gene with the highest confidence score is selected for each gene. c) The continuity score of the selected evolutionary path for each gene is calculated when using parameter p, and the scores are summed. The steps include: d) obtaining the continuity score of the evolutionary path of the gene, i.e., obtaining the overall continuity score of the evolutionary path using parameter p; d) adjusting parameter p, calculating the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and taking the parameter p with the largest overall continuity score as the optimal parameter; e) using the optimal parameter, calculating the confidence scores of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and taking the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed; and the gene origin time inference step, including taking the age of the oldest ancestral group in the optimal evolutionary path as the origin time of the gene to be analyzed.

[0065] Another implementation of this application also provides a computer-readable storage medium including a program executable by a processor to implement the following method: a reference species statistical step, including defining the species containing the gene to be analyzed as the study species, defining genes from other species that share a common ancestor with the gene to be analyzed as orthologous genes, using all species that are orthologous to the study species as reference species, and statistically analyzing the distribution of the gene to be analyzed and its orthologous genes in the reference species; and an evolutionary timeline acquisition step, including using species populations to depict the major events in the development of life on Earth for the study species. Obtain the evolutionary timeline; the taxa in the major events are the ancestral taxa; the ancestral taxa weight setting steps include: classifying the reference species into the nearest ancestral taxa of the study species based on the distance between the reference species and the ancestral taxa of the study species in the phylogenetic tree; if no reference species are classified into a certain ancestral taxa, then the ancestral taxa is deleted; assigning weights to the ancestral taxa based on the number of reference species classified in the ancestral taxa, with higher weights for more reference species; the optimal evolutionary path analysis steps include: a) setting the parameter p, and counting the presence of the gene to be analyzed or its ortholog in each ancestral taxa on a per-taxonomic basis. a) The number of reference species for a gene is used. If this number accounts for more than p percent of the total number of reference species in the ancestral group, the ancestral group is classified as a positive group; otherwise, it is classified as a negative group. b) The positive groups form the initial evolutionary path of the gene under study in the species under study. All possible evolutionary paths consisting of subsets of all positive groups are traversed, and the confidence scores of all possible evolutionary paths are calculated according to their weights. When using parameter p, the gene with the highest confidence score is selected for each gene. c) The continuity score of the selected evolutionary path for each gene is calculated when using parameter p, and the evolution scores of all genes are summed. The continuity score of the route is obtained by using parameter p to obtain the overall continuity score of the evolutionary route; d) Adjust parameter p, calculate the overall continuity score of the evolutionary route with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Use the optimal parameter, calculate the confidence score of all possible evolutionary routes corresponding to the optimal parameter according to steps a) and b), and take the evolutionary route with the largest confidence score as the optimal evolutionary route of the gene to be analyzed; The gene origin time inference step includes taking the age of the oldest ancestral group in the optimal evolutionary route as the origin time of the gene to be analyzed.

[0066] The following is a definition of some of the terms used in this application:

[0067] Study species: The species in which the gene to be studied, i.e. the gene to be analyzed, belongs.

[0068] Reference species: Species used to infer gene age. For example, in the embodiments of this application, all other species that are directly related to the study species and are included in the ENSEMBL database are used as reference species.

[0069] Evolutionary timeline: also known as evolutionary lineage, uses species groups to depict the major events in the development of life on Earth.

[0070] Ancestor groups: The groups that constitute the main events in the study of the evolutionary lineage of species are called ancestral groups. For example, apes, primates, mammals, eumetazoids, and cellular organisms are the ancestral groups of humans. We can find the common ancestral groups of reference species in the human evolutionary lineage. For example, humans and monkeys both belong to the primate class.

[0071] Orthologous genes: If a gene originates from the same ancestor as genes in other species, then the genes in other species are called orthologous genes.

[0072] Positive group: When more than p% of the reference species in an ancestral group contain orthologous genes, the ancestral group is considered to contain the gene and is called a positive group.

[0073] Negative group: When less than p% of the reference species in an ancestral group contain orthologous genes, the ancestral group is considered not to contain the gene and is called a negative group.

[0074] Genetic evolutionary lineage: The distribution of genes, consisting of positive or negative ancestral groups, along the evolutionary timeline of a species under study.

[0075] Evolutionary trajectory: the evolutionary timeline composed of positive gene groups, such as... Figure 5 As shown.

[0076] Expected distribution: The evolutionary path with the highest confidence score.

[0077] Ancestor group weights are ranked according to the number of reference species included in the ancestral group. The more reference species included, the higher the weight. The higher the weight, the higher the confidence score. The weights of ancestral groups are ranked by the number of reference species.

[0078] Low-weight class group: Ancestor class group with a weight of less than 5.

[0079] The following specific experiments will further illustrate this application in detail. These experiments are merely illustrative and should not be construed as limiting the scope of this application.

[0080] Example

[0081] This example uses human genes as the genes to be analyzed and humans as the research species to infer the time of gene origin. Because this method requires first traversing all human genes to obtain preliminary results, and then estimating these results, the optimal parameters are obtained after multiple trials. This method performs batch estimations on all genes; the specific experimental details below only list a few genes.

[0082] Data processing: A batch study of all human protein-coding genes was conducted. Humans were used as the research species, and 186 other species provided by the ENSEMBL database were used as reference species. Orthologous relationship data between humans and other species were obtained. Evolutionary phylogenetic trees of all species and TAXON IDs of each taxa were obtained from the NCBI database. Some results are shown in Table 1.

[0083] Table 1. Evolutionary lineage consisting of taxonomy IDs.

[0084]

[0085]

[0086] Reference species were categorized, and by comparing their evolutionary timelines with those of humans, the most recent common ancestral group between the reference and study species was identified. Partial results are shown in Table 1. The common ancestral groups of the reference and study species are indicated by the underlined sections in the table, with the most recent ancestral group being Euarchontoglires (TAXON ID: 314146). Based on the position of this common ancestral group within the human evolutionary spectrum, the reference species were categorized as follows: Figure 3 As shown in Table 2, some of the classification results are presented.

[0087] Table 2 Reference Species Classification

[0088]

[0089]

[0090]

[0091] After classification, human ancestral groups that do not contain the reference species are removed.

[0092] Weighting of human ancestral groups: First, the number of reference species contained in the ancestral group is calculated; then, the ancestral groups are ranked according to the number of reference species contained, with a higher ranking for a larger number of reference species; the results are shown in Table 3.

[0093] Table 3. Ancestor Group Ranking

[0094] Ancestor Groups Evolutionary lineage position Weighted ranking [ranking by the number of reference species] Homo sapiens (humans) 1 1 Homininae (subfamily Homininae) 2 3 Hominidae (subfamily Homininae) 3 1 Hominoidea (superfamily Hominoidea) 4 1 Catarrhini (narrow-nosed type) 5 5 Simiiformes (inferior order of hominids) 6 4 Haplorrhini (Suborder Rhinoidea) 7 1 Primates 8 4 Euarchontoglires (Primates) 9 6 Boreoeutheria (northern mammals) 10 8 Eutheria (lower class of mammals) 11 1 Theria (Subclass Theria) 12 4 Mammalia (Class Mammalia) 13 1 Amniota (amniotic membrane type) 14 7 Tetrapoda (quadrupedal) 15 2 Sarcopterygii (Lobe-finned fish) 16 1 Euteleostomi (bony vertebrates) 17 9 Gnathostomata (jawed animals) 18 1 Vertebrata (phylum Vertebrata) 19 2 Chordata (chordates) 20 2 Bilateria (bilaterally symmetrical animals) 21 2 Opisthokonta (a posterior flagellated organism) 22 1

[0095] The more reference species an ancestral group uses for cross-validation, the lower the error in inference and the higher the accuracy of the group's results. Therefore, the ranking of the number of reference species is the weight of the ancestral group; when the weight is less than 5, the group is considered a low-weight group.

[0096] The method for determining the distribution of human genes in a reference species is as follows: Look for highly reliable orthologous genes in the reference species. If present, it indicates that the human gene also has a ancestor in that species (positive); conversely, if absent, it indicates that no ancestor exists in the reference species (negative). The obtained negative or positive distribution of genes in each reference species represents the desired distribution, and some results are shown in Table 4.

[0097] Table 4. Negative or positive gene distribution in the reference species.

[0098]

[0099]

[0100]

[0101] Table 4 shows the negative or positive distribution of the gene BRD3OS in the reference species, where "1" indicates positive and "0" indicates negative.

[0102] Data processing

[0103] Back-engineering the optimal parameter p:

[0104] To reduce errors caused by gene variation in individual species when inferring gene evolutionary routes, this example utilizes the characteristic that most human ancestral groups are classified into a large number of reference species to cross-validate the existence of genes in ancestral groups.

[0105] Set the parameter p, which is 10 in this example; if more than 10 percent of the reference species in an ancestral group contain a gene that is orthologous to a gene, then the gene is considered to exist in this ancestral group, and the ancestral group is defined as a positive group; otherwise, it is a negative group.

[0106] This allows us to estimate the preliminary positive distribution of genes in the evolutionary lineage, i.e., the preliminary evolutionary route. Some results are shown in Table 5.

[0107] When certain low-weight ancestral groups are negative, to avoid false negatives due to insufficient reference species or the absence of genes in such groups (e.g., 1. the gene actually exists in the ancestral group, but is missing due to variation in the reference species, leading to the erroneous assumption that the ancestral group as a whole lacks the gene due to the scarcity of the reference species), this example assumes the low-weight negative group between two positive groups is a false negative; corrects it to positive, and obtains the human evolutionary path of the completed gene. Some correction results are shown in Table 5.

[0108] Table 5 Preliminary evolutionary pathways of genes and their correction results

[0109]

[0110] Table 5 shows the evolutionary path of the BRD3OS gene, where the path before correction (i.e., before correction of the low-weight ancestral groups) and the path after correction (i.e., after correction of the low-weight ancestral groups) are shown.

[0111] Subsequently, to further reduce the impact of false positives, for each gene-completed evolutionary path, all possible evolutionary paths consisting of subsets of all positive groups were traversed, and the confidence score of all evolutionary paths was calculated using Formula 1. The evolutionary path that maximizes the confidence score is the human evolutionary path of the gene with the lowest error rate. For example, the evolutionary path with the highest confidence score obtained by BRD3OS using parameter 10 is: [13,12,11,10,8,7,6,5,4,3,2,1]. Confidence score = 1498.

[0112] Formula 1

[0113] In Formula 1, ConfidenceScore is the confidence score, x ij y is the weight of the j-th ancestral group in the i-th consecutive positive group. pq It is the weight of the q-th ancestral group of the p-th consecutive negative group, l i 2 The number of ancestral groups that constitute the i-th contiguous group set is the square of l. p 2 It is the square of the number of ancestral groups that constitute the p-th continuous group set.

[0114] Subsequently, the continuity score of the evolutionary distribution of each gene is calculated using Formula 2 with parameter p. The overall continuity of the evolutionary path obtained using parameter p is estimated by summing the continuity scores of all genes. For example, the continuity score of the BRD3OS evolutionary path is 0.931. The mean of the continuity scores of all genes is the overall continuity score of the evolutionary path obtained with parameter p = 10, which is 0.969295.

[0115] Formula 2

[0116] In Formula 2, y i Let be the sorting index of the i-th positive group, and n be the number of positive groups.

[0117] Repeat the above steps and assign new p values, selecting the p value that maximizes the overall continuous score as the optimal parameter. Specifically, this example conducted experiments with multiple point values ​​ranging from p values ​​of 5 to 100, and the p values ​​and their overall continuous score results are shown in Table 6.

[0118] Table 6. p-values ​​and their overall continuity scores

[0119]

[0120]

[0121] The results in Table 6 show that the maximum overall continuous score in this example is 0.9713681. The corresponding p value is taken as the optimal parameter, that is, the optimal parameter in this example is 25.

[0122] Estimate the evolutionary path of genes:

[0123] Using the optimal parameter p, i.e. 25, the evolutionary path of the gene was re-inferred. The time period of the emergence of the earliest positive ancestral group in the evolutionary path is the origin time of the gene. Some results are shown in Table 7.

[0124] Table 7 Origin Time of Gene BRD3OS [ENSG00000235106]

[0125] Parameter p Evolutionary path Initial date 10 [13,12,11,10,8,7,6,5,4,3,2,1] 177 million years ago 25 [8,7,6,5,4,3,2,1] 740,000 years ago

[0126] Table 7 shows the inferred evolutionary path and gene origin time of the BRD3OS gene under different parameter p values; the results show that the evolutionary path and gene origin time inferred with the optimal parameter 25 are more accurate and are basically consistent with existing studies or reports.

[0127] The above description, in conjunction with specific embodiments, provides a further detailed explanation of this application and should not be construed as limiting the specific implementation of this application to these descriptions. Those skilled in the art to which this application pertains can make several simple deductions or substitutions without departing from the concept of this application.

Claims

1. A method for inferring the time of gene origin, characterized in that: Includes the following steps, The reference species statistical steps include defining the species containing the gene to be analyzed as the study species, defining genes from other species that share the same ancestor as the gene to be analyzed as orthologous genes, and using all species that have orthologous relationships with the study species as reference species, and statistically analyzing the distribution of the gene to be analyzed and its orthologous genes in the reference species. The steps for obtaining the evolutionary timeline include using species groups to depict the major events in the life development process of the studied species on Earth to obtain the evolutionary timeline, wherein the taxa in the major events are the ancestral taxa. The ancestral group weight setting step includes classifying the reference species into the nearest ancestral group of the study species based on the distance between the ancestral group of the reference species and the ancestral group of the study species in the phylogenetic tree; deleting the ancestral group if no reference species are classified into it; and assigning a weight to the ancestral group based on the number of reference species classified in the ancestral group, with a higher weight for a larger number of reference species. The optimal evolutionary path analysis steps include: a) setting a parameter p, and counting the number of reference species in each ancestral group containing the gene to be analyzed or its orthologous gene. If the percentage of this number to the total number of reference species in that ancestral group exceeds p%, then the ancestral group is classified as a positive group; otherwise, it is classified as a negative group; b) constructing a preliminary evolutionary path of the gene to be analyzed in the studied species from the positive groups, traversing all possible evolutionary paths composed of subsets of all positive groups, and calculating the confidence score of all possible evolutionary paths according to their weights; when using parameter p, for each gene, a confidence score is selected. c) Calculate the continuity score of the evolutionary path chosen by each gene when using parameter p, and sum the continuity scores of all genes to obtain the overall continuity score of the evolutionary path obtained using parameter p; d) Adjust parameter p, calculate the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Using the optimal parameter, calculate the confidence score of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and take the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed. The gene origin time inference step includes taking the age of the oldest ancestral group in the optimal evolutionary path as the origin time of the gene to be analyzed.

2. The method according to claim 1, characterized in that: In the optimal evolutionary path analysis step, step b) includes, Data preparation involves extracting all possible subsets from all ancestral groups that are judged to be positive. Each ancestral group subset constitutes a sub-distribution of the gene to be analyzed in the evolutionary lineage, i.e., a sub-distribution of the group. Extract and select a possible sub-distribution of the taxonomic group to form a possible evolutionary path; Correction: If the low-weight class in the sub-distribution of the class is a negative class, and the two ancestral classes before and after this negative class in the evolutionary timeline are both positive classes, then correct this negative class to a positive class. Calculate the confidence score of the corrected sub-distribution of the class group based on the weights, traverse the possible evolutionary paths consisting of subsets of all positive classes, calculate the confidence score of all possible evolutionary paths, and select the evolutionary path with the highest confidence score.

3. The method according to claim 2, characterized in that: Assigning weights to ancestral groups includes sorting them according to the number of reference species classified within the ancestral group, and using the sorted sequence number as the weight of the ancestral group.

4. The method according to claim 2, characterized in that: The low-weight class group refers to the ancestral class group with a weight of less than 5.

5. The method according to claim 2, characterized in that: The confidence score is calculated using Formula 1 based on the weights. Formula 1 ; In Formula 1, ConfidenceScore is the confidence score, x ij y is the weight of the j-th ancestral group in the i-th consecutive positive group. pq It is the weight of the q-th ancestral group of the p-th consecutive negative group, l i 2 The number of ancestral groups that constitute the i-th contiguous group set is the square of l. p 2 It is the square of the number of ancestral groups that constitute the p-th continuous group set.

6. The method according to any one of claims 1-5, characterized in that: In the optimal evolutionary path analysis step, step c) continuity score includes assigning sorting numbers to the research species and their ancestral groups according to the chronological order of the evolutionary timeline, and calculating the continuity score of positive groups using Formula 2. Formula 2 ; In Formula 2, y i Let be the sorting index of the i-th positive group, and n be the number of positive groups.

7. An apparatus for inferring the time of gene origin, characterized in that: It includes a reference species statistics module, an evolutionary timeline acquisition module, an ancestral group weight setting module, an optimal evolutionary route analysis module, and a gene origin time inference module; The reference species statistics module is used to define the species containing the gene to be analyzed as the study species, define genes from other species that share the same ancestor as the gene to be analyzed as orthologous genes, and use all species that have orthologous relationships with the study species as reference species to count the distribution of the gene to be analyzed and its orthologous genes in the reference species. The evolutionary timeline acquisition module is used to depict the major events in the life development process of the studied species on Earth using species groups to obtain the evolutionary timeline, wherein the taxa in the major events are the ancestral taxa; The ancestral group weight setting module is used to classify the reference species into the nearest ancestral group of the research species based on the distance between the ancestral group of the reference species and the ancestral group of the research species in the phylogenetic tree; if no reference species are classified into a certain ancestral group, the ancestral group is deleted; and a weight is set for the ancestral group based on the number of reference species classified in the ancestral group, with a higher weight for a larger number of reference species. The optimal evolutionary path analysis module is used to: a) set a parameter p, and, on a per-ancestral-group basis, count the number of reference species in each ancestral-group containing the gene to be analyzed or its orthologous gene. If this number accounts for more than p percent of the total number of reference species in that ancestral-group, then that ancestral-group is classified as a positive group; otherwise, it is classified as a negative group; b) construct a preliminary evolutionary path for the gene to be analyzed in the studied species from the positive groups, traverse all possible evolutionary paths composed of subsets of all positive groups, and calculate the confidence score of all possible evolutionary paths according to their weights; when using parameter p, for each gene, select a confidence score. c) Calculate the continuity score of the evolutionary path chosen by each gene when using parameter p, and sum the continuity scores of all genes to obtain the overall continuity score of the evolutionary path obtained using parameter p; d) Adjust parameter p, calculate the overall continuity score of the evolutionary path with adjusted parameter p according to steps a) to c), and take the parameter p with the largest overall continuity score as the optimal parameter; e) Using the optimal parameter, calculate the confidence score of all possible evolutionary paths corresponding to the optimal parameter according to steps a) and b), and take the evolutionary path with the largest confidence score as the optimal evolutionary path of the gene to be analyzed. The gene origin time inference module is used to determine the origin time of the gene to be analyzed by taking the age of the oldest ancestral group in the optimal evolutionary path as the origin time of the gene.

8. The apparatus according to claim 7, characterized in that: In the optimal evolutionary path analysis module, step b) includes: Data preparation involves extracting all possible subsets from all ancestral groups that are judged to be positive. Each ancestral group subset constitutes a sub-distribution of the gene to be analyzed in the evolutionary lineage, i.e., a sub-distribution of the group. Extract and select a possible sub-distribution of the taxonomic group to form a possible evolutionary path; Correction: If the low-weight class in the sub-distribution of the class is a negative class, and the two ancestral classes before and after this negative class in the evolutionary timeline are both positive classes, then correct this negative class to a positive class. Calculate the confidence score of the corrected sub-distribution of the class group based on the weights, traverse the possible evolutionary paths consisting of subsets of all positive classes, calculate the confidence score of all possible evolutionary paths, and select the evolutionary path with the highest confidence score.

9. The apparatus according to claim 8, characterized in that: Assigning weights to ancestral groups includes sorting them according to the number of reference species classified within the ancestral group, and using the sorted sequence number as the weight of the ancestral group.

10. The apparatus according to claim 8, characterized in that: The low-weight class group refers to the ancestral class group with a weight of less than 5.

11. The apparatus according to claim 8, characterized in that: The confidence score is calculated using Formula 1 based on the weights. Formula 1 ; In Formula 1, ConfidenceScore is the confidence score, x ij y is the weight of the j-th ancestral group in the i-th consecutive positive group. pq It is the weight of the q-th ancestral group of the p-th consecutive negative group, l i 2 The number of ancestral groups that constitute the i-th contiguous group set is the square of l. p 2 It is the square of the number of ancestral groups that constitute the p-th continuous group set.

12. The apparatus according to any one of claims 7-11, characterized in that: In the optimal evolutionary path analysis module, step c) of calculating the continuity score includes assigning sorting numbers to the research species and their ancestral groups according to the chronological order of the evolutionary timeline, and calculating the continuity score of the positive groups using Formula 2. Formula 2 ; In Formula 2, y i Let be the sorting index of the i-th positive group, and n be the number of positive groups.

13. An apparatus for inferring the time of gene origin, characterized in that: The device includes a memory and a processor; The memory includes a storage device for storing programs; The processor includes a method for implementing the method of any one of claims 1-6 by executing a program stored in the memory.

14. A computer-readable storage medium, characterized in that: The storage medium includes a program that can be executed by a processor to implement the method of any one of claims 1-6.