The application discloses a method and device for identifying orthologous
gene groups in a pan-genetic family, and belongs to the field of
bioinformatics. The method first acquires a
phylogenetic tree, a sequence set and a target species number of genes to be grouped, and calculates a
similarity matrix of
gene pairs; then, a data-driven 5th percentile (P5) is used to automatically infer a split threshold, and an initial
gene grouping is performed in combination with the topological structure of the
phylogenetic tree; subsequently, a post-
processing pipeline including isolated leaf node redistribution, micro-group topological absorption, adjacent group merging and super-large gene group targeted re-splitting is executed; finally, based on a multi-dimensional
quality score function and Latin
hypercube sampling, parameter optimization iteration is performed, and the best orthologous gene group (OGG) is output. The application breaks the limitation of traditional clustering which only depends on sequence similarity, effectively eliminates over-merging, fragmentation and multi-lineage
genome errors, significantly improves the accuracy and evolutionary rationality of gene grouping of complex
polyploid species, and has important significance for dividing core and non-core genes in pan-
genome analysis. The specific
flowchart is shown in FIG. 1.